Skip to content

tp: let pipelines read SQL through their arguments - #7637

Closed
LalitMaganti wants to merge 1 commit into
dev/lalitm/pipeline-rereadfrom
dev/lalitm/pipeline-inputs
Closed

LalitMaganti wants to merge 1 commit into
dev/lalitm/pipeline-rereadfrom
dev/lalitm/pipeline-inputs

Conversation

@LalitMaganti

Copy link
Copy Markdown
Member

A pipeline ran its SQL sources itself, as statements of their own. So they could see nothing of the statement the pipeline was written in: not its CTEs, not the arguments of the function it was in, not an outer query's columns.

Now SQLite evaluates each relation a pipeline reads where the pipeline is written, and hands it over. The pipeline never runs SQL itself.

-- FROM (SELECT id, parent_id, self * $k AS self FROM tree) |> TREE ACCUMULATE UP SUM(self) AS total
-- is replaced with:
SELECT c0 AS "id", ... FROM __intrinsic_pipeline(
  X'<plan>',
  (SELECT __intrinsic_rows('v:id,v:parent_id,v:self', "id", "parent_id", "self")
   FROM (SELECT id, parent_id, self * $k AS self FROM tree)))

How:

  • Collecting: __intrinsic_rows is an aggregate that collects a relation's rows into batches (CollectedRows). The table function takes one collection per input, after the plan.

  • Plans hold no SQL: a plan reads its inputs by index. SQL sources are moved out into inputs when the plan is written, and only the columns the pipeline uses are collected.

  • Columns from semantic analysis, never from preparing SQL:

    • dataframes and views, as before;
    • other relations SQLite knows (plain tables, table functions), by their schema (pragma_table_xinfo), untyped;
    • other pipelines in the same statement, by the plan they compiled to, typed.

    Anything else fails clearly. CTEs come next.

  • Safety: anyone can write a plan and its inputs into SQL, so a run checks each input has the columns the plan reads from it.

What this fixes and removes:

  • Function arguments now just work, since SQLite evaluates them in the collecting subquery.
  • SqlScan, DescribeQuery and positional pruning of SQL sources are gone.
  • It's also faster: collecting takes about 95 ms against 220 ms for stepping, for 1M rows of three columns.

Behaviour changes, on purpose:

  • A relation read again, as on the inner side of a join, is collected once, so random() in a source gives the same rows to every read.
  • Duplicate or unnamed columns are reported from the source as written.

Needs syntaqlite 0.12 (#7635), which expands pipelines once their statement is parsed, so the analyzer sees the whole statement.

Tests: trace processor unittests (2391, including new CollectedRows tests and function arguments) and diff tests pass, apart from ProfilingLlvmSymbolizer:stack_profile_symbols, which also fails without this. The pipeline tests also pass under ASan.

@LalitMaganti
LalitMaganti requested a review from a team as a code owner September 27, 2026 19:23
@LalitMaganti
LalitMaganti added this pull request to stack #7636 September 27, 2026 19:30
@LalitMaganti
LalitMaganti force-pushed the dev/lalitm/pipeline-inputs branch from 22f8f37 to 38c7270 Compare September 27, 2026 19:30
@LalitMaganti
LalitMaganti marked this pull request as draft September 27, 2026 19:36
@LalitMaganti
LalitMaganti removed this pull request from stack #7636 September 27, 2026 19:36
@LalitMaganti
LalitMaganti added this pull request to stack #7639 September 27, 2026 19:37
A pipeline ran its SQL sources itself, as statements of their own. So
they could see nothing of the statement the pipeline was written in: not
its CTEs, not the arguments of the function it was in, not the columns
of an outer query. And the pipeline paid SQLite's row-at-a-time boundary
twice, stepping each statement from outside.

Now SQLite evaluates each relation a pipeline reads where the pipeline
is written, and hands it over. The SQL a pipeline is replaced with
collects each SQL source with an aggregate, `__intrinsic_rows`, into
batches, and passes them to `__intrinsic_pipeline` after the plan:

  __intrinsic_pipeline(X'<plan>',
                       (SELECT __intrinsic_rows('<columns>', a, b) FROM src))

The plan reads them as inputs by index, and never holds SQL: a SQL
source is moved out into an input when the plan is written. Only the
columns the pipeline uses are collected. Collecting is also cheaper than
stepping: about 95 ms against 220 ms for 1M rows of three columns.

A SQL source's columns now come from semantic analysis, never from
preparing it, since only where the pipeline is written is everything it
reads in scope:
- dataframes and views, as before;
- other tables SQLite knows, such as plain tables and table functions,
  by their schema (pragma_table_xinfo), untyped;
- pipelines of the same statement, by the plan they compiled to, typed.
A relation the analyzer cannot describe fails clearly. Leaf relations
now own their strings, and the analyzer names columns as SQLite sees
them, after macro expansion.

Anyone can write a plan and its inputs into SQL, so a run checks that
each input has the columns the plan reads from it.

SqlScan, DescribeQuery and positional pruning of SQL sources are gone.
Function arguments work with no special handling, as SQLite evaluates
them in the collecting subquery. Needs syntaqlite 0.12, which expands
pipelines once their statement is parsed.
@LalitMaganti
LalitMaganti force-pushed the dev/lalitm/pipeline-reread branch from 57efc33 to a8afe5b Compare September 27, 2026 19:41
@LalitMaganti
LalitMaganti force-pushed the dev/lalitm/pipeline-inputs branch from 38c7270 to ee626da Compare September 27, 2026 19:42
@LalitMaganti
LalitMaganti deleted the branch dev/lalitm/pipeline-reread September 27, 2026 20:06
@LalitMaganti

Copy link
Copy Markdown
Member Author

Superseded by #7641 (reading SQL through arguments), split out ahead of subqueries, with its refactors in #7640.

@LalitMaganti
LalitMaganti deleted the dev/lalitm/pipeline-inputs branch September 27, 2026 20:06
@github-actions

Copy link
Copy Markdown

🎨 Perfetto UI Builds & Tests

@LalitMaganti
LalitMaganti removed this pull request from stack #7639 September 27, 2026 20:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant