tp: pass the SQL a pipeline reads to it as dataframes - #7641
Draft
LalitMaganti wants to merge 1 commit into
Draft
LalitMaganti wants to merge 1 commit into
LalitMaganti wants to merge 1 commit into
Conversation
LalitMaganti
added this pull request to stack #7645
September 27, 2026 20:08
LalitMaganti
removed this pull request from stack #7645
September 27, 2026 20:08
LalitMaganti
force-pushed
the
dev/lalitm/pipeline-sql-inputs
branch
2 times, most recently
from
September 27, 2026 21:13
b9e909b to
cd9005e
Compare
LalitMaganti
force-pushed
the
dev/lalitm/pipeline-prep
branch
from
September 27, 2026 22:55
ded4510 to
584b139
Compare
LalitMaganti
force-pushed
the
dev/lalitm/pipeline-sql-inputs
branch
2 times, most recently
from
September 27, 2026 23:22
790f2ac to
dd2fbbb
Compare
LalitMaganti
force-pushed
the
dev/lalitm/pipeline-prep
branch
from
September 27, 2026 23:22
584b139 to
4d7a3ba
Compare
LalitMaganti
force-pushed
the
dev/lalitm/pipeline-sql-inputs
branch
2 times, most recently
from
September 28, 2026 03:53
520d311 to
8c8772f
Compare
LalitMaganti
force-pushed
the
dev/lalitm/pipeline-prep
branch
from
September 28, 2026 04:21
4d7a3ba to
c7fdc78
Compare
LalitMaganti
force-pushed
the
dev/lalitm/pipeline-sql-inputs
branch
from
September 28, 2026 04:21
8c8772f to
0a66bea
Compare
LalitMaganti
force-pushed
the
dev/lalitm/pipeline-sql-inputs
branch
from
September 28, 2026 05:28
0a66bea to
fbd9f11
Compare
A pipeline ran its SQL sources itself, as statements of their own. So they could see nothing of the statement the pipeline was written in. Now SQLite evaluates each relation a pipeline reads where the pipeline is written, and builds it into a dataframe, as a PERFETTO TABLE is built. __intrinsic_dataframes gathers them into one list, passed to __intrinsic_pipeline after the plan, and the pipeline reads each like any other dataframe. A pipeline never runs SQL itself, and a plan never holds SQL. - __intrinsic_dataframe_agg builds the dataframe. A plan reads its i-th as dataframe argument i, which is bound to the dataframe when the plan is loaded, before it is lowered, so lowering only ever sees dataframes. - The columns of the relation come from semantic analysis rather than from preparing its SQL. Relations analysis cannot describe yet, such as VALUES, fail clearly. Tables SQLite knows, including table functions, are described from its schema. - As in a PERFETTO TABLE, a column must hold one type throughout. The dataframe is built without analysis, keeping integers as Int64 and estimating no statistics, as a pipeline only scans it. This removes SqlScan, DescribeQuery and positional pruning of SQL sources, and lowering no longer needs a string pool.
LalitMaganti
force-pushed
the
dev/lalitm/pipeline-prep
branch
from
September 28, 2026 06:05
c7fdc78 to
4dd9866
Compare
LalitMaganti
force-pushed
the
dev/lalitm/pipeline-sql-inputs
branch
from
September 28, 2026 06:06
fbd9f11 to
552c503
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A pipeline ran its SQL sources itself, as statements of their own. So they could see nothing of the statement the pipeline was written in.
Now SQLite evaluates each relation a pipeline reads where the pipeline is written, and builds it into a dataframe, exactly as a
PERFETTO TABLEis built. The dataframes are gathered into one list, passed to the pipeline after its plan, and the pipeline reads each like any other dataframe. A pipeline never runs SQL itself, and a plan never holds SQL.How:
__intrinsic_dataframe_aggbuilds a dataframe of a relation's rows with the same builderCREATE PERFETTO TABLEuses, and passes it to SQL as a"TABLE"pointer, like other functions returning dataframes. It skips the builder's analysis (integer downcasting, sort detection, distinct-count statistics), as a pipeline only scans the result.__intrinsic_dataframesgathers them into one list argument, so the table function has a single hidden column for them however many relations a pipeline reads.VALUES, fails clearly for now.Removes:
SqlScan(and all ofperfetto_sql/exec),DescribeQuery, positional pruning of SQL sources, and the string pool lowering needed.Behaviour changes, on purpose:
PERFETTO TABLE, a column read from SQL must hold one type throughout.This is what lets pipelines become subqueries later in the stack: their SQL then simply runs in the statement around them.