tp: let pipelines read SQL through their arguments - #7637
Closed
LalitMaganti wants to merge 1 commit into
Closed
LalitMaganti wants to merge 1 commit into
LalitMaganti wants to merge 1 commit into
Conversation
LalitMaganti
added this pull request to stack #7636
September 27, 2026 19:30
LalitMaganti
force-pushed
the
dev/lalitm/pipeline-inputs
branch
from
September 27, 2026 19:30
22f8f37 to
38c7270
Compare
LalitMaganti
marked this pull request as draft
September 27, 2026 19:36
LalitMaganti
removed this pull request from stack #7636
September 27, 2026 19:36
LalitMaganti
added this pull request to stack #7639
September 27, 2026 19:37
A pipeline ran its SQL sources itself, as statements of their own. So
they could see nothing of the statement the pipeline was written in: not
its CTEs, not the arguments of the function it was in, not the columns
of an outer query. And the pipeline paid SQLite's row-at-a-time boundary
twice, stepping each statement from outside.
Now SQLite evaluates each relation a pipeline reads where the pipeline
is written, and hands it over. The SQL a pipeline is replaced with
collects each SQL source with an aggregate, `__intrinsic_rows`, into
batches, and passes them to `__intrinsic_pipeline` after the plan:
__intrinsic_pipeline(X'<plan>',
(SELECT __intrinsic_rows('<columns>', a, b) FROM src))
The plan reads them as inputs by index, and never holds SQL: a SQL
source is moved out into an input when the plan is written. Only the
columns the pipeline uses are collected. Collecting is also cheaper than
stepping: about 95 ms against 220 ms for 1M rows of three columns.
A SQL source's columns now come from semantic analysis, never from
preparing it, since only where the pipeline is written is everything it
reads in scope:
- dataframes and views, as before;
- other tables SQLite knows, such as plain tables and table functions,
by their schema (pragma_table_xinfo), untyped;
- pipelines of the same statement, by the plan they compiled to, typed.
A relation the analyzer cannot describe fails clearly. Leaf relations
now own their strings, and the analyzer names columns as SQLite sees
them, after macro expansion.
Anyone can write a plan and its inputs into SQL, so a run checks that
each input has the columns the plan reads from it.
SqlScan, DescribeQuery and positional pruning of SQL sources are gone.
Function arguments work with no special handling, as SQLite evaluates
them in the collecting subquery. Needs syntaqlite 0.12, which expands
pipelines once their statement is parsed.
LalitMaganti
force-pushed
the
dev/lalitm/pipeline-reread
branch
from
September 27, 2026 19:41
57efc33 to
a8afe5b
Compare
LalitMaganti
force-pushed
the
dev/lalitm/pipeline-inputs
branch
from
September 27, 2026 19:42
38c7270 to
ee626da
Compare
Member
Author
🎨 Perfetto UI Builds & Tests
|
LalitMaganti
removed this pull request from stack #7639
September 27, 2026 20:08
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A pipeline ran its SQL sources itself, as statements of their own. So they could see nothing of the statement the pipeline was written in: not its CTEs, not the arguments of the function it was in, not an outer query's columns.
Now SQLite evaluates each relation a pipeline reads where the pipeline is written, and hands it over. The pipeline never runs SQL itself.
How:
Collecting:
__intrinsic_rowsis an aggregate that collects a relation's rows into batches (CollectedRows). The table function takes one collection per input, after the plan.Plans hold no SQL: a plan reads its inputs by index. SQL sources are moved out into inputs when the plan is written, and only the columns the pipeline uses are collected.
Columns from semantic analysis, never from preparing SQL:
pragma_table_xinfo), untyped;Anything else fails clearly. CTEs come next.
Safety: anyone can write a plan and its inputs into SQL, so a run checks each input has the columns the plan reads from it.
What this fixes and removes:
SqlScan,DescribeQueryand positional pruning of SQL sources are gone.Behaviour changes, on purpose:
random()in a source gives the same rows to every read.Needs syntaqlite 0.12 (#7635), which expands pipelines once their statement is parsed, so the analyzer sees the whole statement.
Tests: trace processor unittests (2391, including new
CollectedRowstests and function arguments) and diff tests pass, apart fromProfilingLlvmSymbolizer:stack_profile_symbols, which also fails without this. The pipeline tests also pass under ASan.