Skip to content

tp: pass the SQL a pipeline reads to it as dataframes - #7641

Draft
LalitMaganti wants to merge 1 commit into
dev/lalitm/pipeline-prepfrom
dev/lalitm/pipeline-sql-inputs
Draft

LalitMaganti wants to merge 1 commit into
dev/lalitm/pipeline-prepfrom
dev/lalitm/pipeline-sql-inputs

Conversation

@LalitMaganti

@LalitMaganti LalitMaganti commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

A pipeline ran its SQL sources itself, as statements of their own. So they could see nothing of the statement the pipeline was written in.

Now SQLite evaluates each relation a pipeline reads where the pipeline is written, and builds it into a dataframe, exactly as a PERFETTO TABLE is built. The dataframes are gathered into one list, passed to the pipeline after its plan, and the pipeline reads each like any other dataframe. A pipeline never runs SQL itself, and a plan never holds SQL.

-- FROM (SELECT id, parent_id, self FROM tree) |> TREE ACCUMULATE UP SUM(self) AS total
-- becomes:
SELECT c0 AS "id", ... FROM __intrinsic_pipeline(
  X'<plan>',
  __intrinsic_dataframes(
    (SELECT __intrinsic_dataframe_agg('id,parent_id,self', "id", "parent_id", "self")
     FROM (SELECT id, parent_id, self FROM tree))))

How:

  • __intrinsic_dataframe_agg builds a dataframe of a relation's rows with the same builder CREATE PERFETTO TABLE uses, and passes it to SQL as a "TABLE" pointer, like other functions returning dataframes. It skips the builder's analysis (integer downcasting, sort detection, distinct-count statistics), as a pipeline only scans the result.
  • __intrinsic_dataframes gathers them into one list argument, so the table function has a single hidden column for them however many relations a pipeline reads.
  • A plan reads its i-th SQL source as dataframe argument i. When the plan is loaded, each argument is bound to the dataframe passed, taking its column types from it, so lowering only ever sees dataframes: no new source or storage type.
  • The columns of a relation come from semantic analysis rather than from preparing its SQL: dataframes and views as before, and other tables SQLite knows (including table functions) by their schema. Anything else, like VALUES, fails clearly for now.
  • Anyone can write a plan and its arguments into SQL, so binding checks each argument has the columns the plan reads.

Removes: SqlScan (and all of perfetto_sql/exec), DescribeQuery, positional pruning of SQL sources, and the string pool lowering needed.

Behaviour changes, on purpose:

  • As in a PERFETTO TABLE, a column read from SQL must hold one type throughout.
  • Duplicate or unnamed columns are reported from the source as written, not as SQLite renamed them.

This is what lets pipelines become subqueries later in the stack: their SQL then simply runs in the statement around them.

@LalitMaganti
LalitMaganti added this pull request to stack #7645 September 27, 2026 20:08
@LalitMaganti
LalitMaganti removed this pull request from stack #7645 September 27, 2026 20:08
@github-actions

github-actions Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

@LalitMaganti
LalitMaganti force-pushed the dev/lalitm/pipeline-sql-inputs branch 2 times, most recently from b9e909b to cd9005e Compare September 27, 2026 21:13
@LalitMaganti LalitMaganti changed the title tp: let pipelines read SQL through their arguments tp: pass the SQL a pipeline reads to it as dataframes Sep 27, 2026
@LalitMaganti
LalitMaganti force-pushed the dev/lalitm/pipeline-prep branch from ded4510 to 584b139 Compare September 27, 2026 22:55
@LalitMaganti
LalitMaganti force-pushed the dev/lalitm/pipeline-sql-inputs branch 2 times, most recently from 790f2ac to dd2fbbb Compare September 27, 2026 23:22
@LalitMaganti
LalitMaganti force-pushed the dev/lalitm/pipeline-prep branch from 584b139 to 4d7a3ba Compare September 27, 2026 23:22
@LalitMaganti
LalitMaganti force-pushed the dev/lalitm/pipeline-sql-inputs branch 2 times, most recently from 520d311 to 8c8772f Compare September 28, 2026 03:53
@LalitMaganti
LalitMaganti force-pushed the dev/lalitm/pipeline-prep branch from 4d7a3ba to c7fdc78 Compare September 28, 2026 04:21
@LalitMaganti
LalitMaganti force-pushed the dev/lalitm/pipeline-sql-inputs branch from 8c8772f to 0a66bea Compare September 28, 2026 04:21
@LalitMaganti
LalitMaganti force-pushed the dev/lalitm/pipeline-sql-inputs branch from 0a66bea to fbd9f11 Compare September 28, 2026 05:28
A pipeline ran its SQL sources itself, as statements of their own. So
they could see nothing of the statement the pipeline was written in.

Now SQLite evaluates each relation a pipeline reads where the pipeline
is written, and builds it into a dataframe, as a PERFETTO TABLE is
built. __intrinsic_dataframes gathers them into one list, passed to
__intrinsic_pipeline after the plan, and the pipeline reads each like
any other dataframe. A pipeline never runs SQL itself, and a plan never
holds SQL.

- __intrinsic_dataframe_agg builds the dataframe. A plan reads its i-th
  as dataframe argument i, which is bound to the dataframe when the
  plan is loaded, before it is lowered, so lowering only ever sees
  dataframes.
- The columns of the relation come from semantic analysis rather than
  from preparing its SQL. Relations analysis cannot describe yet, such
  as VALUES, fail clearly. Tables SQLite knows, including table
  functions, are described from its schema.
- As in a PERFETTO TABLE, a column must hold one type throughout. The
  dataframe is built without analysis, keeping integers as Int64 and
  estimating no statistics, as a pipeline only scans it.

This removes SqlScan, DescribeQuery and positional pruning of SQL
sources, and lowering no longer needs a string pool.
@LalitMaganti
LalitMaganti force-pushed the dev/lalitm/pipeline-prep branch from c7fdc78 to 4dd9866 Compare September 28, 2026 06:05
@LalitMaganti
LalitMaganti force-pushed the dev/lalitm/pipeline-sql-inputs branch from fbd9f11 to 552c503 Compare September 28, 2026 06:06

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant