diff --git a/.claude-plugin/plugin.json b/.claude-plugin/plugin.json index 63af4557..87ff732d 100644 --- a/.claude-plugin/plugin.json +++ b/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "name": "dex", "description": "Analytics engineering for Claude Code and any agent: data warehouse exploration, dbt transformation and semantic modeling, and schema-drift maintenance on dbt.", - "version": "1.12.2", + "version": "1.12.3", "author": { "name": "Exmergo, Inc.", "email": "support@exmergo.com", diff --git a/AGENTS.md b/AGENTS.md index bf35d749..f9f20fc6 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -76,11 +76,11 @@ credentials and no network. | `demo [path]` | generates a seeded local DuckDB warehouse (7 tables, 29,512 rows) plus a `.dex/config.yml` beside it, so a first run needs no warehouse, no credentials, and no network; both artifacts are listed in `data.created` and `data.next_steps` names the commands worth running next. The path is positional and resolves against the working directory, defaulting to `dex_demo.duckdb`; `--path` is refused here rather than honored, since everywhere else it names the warehouse dex *reads*. Create-only, with no `--confirm` that can talk past it: an existing file at the target is a refusal (`reason: guard`), a missing parent directory is a refusal (`reason: request`), no directories are ever created, and a `.dex/config.yml` at or above the target is left untouched with a warning rather than shadowed by a second one. The data is generated from a pinned seed, so the counts quoted in the docs are the counts a user sees, and it is deliberately flawed: a key that lost uniqueness to a double-loaded batch, a key mixing two id schemes, a join whose columns share a name and none of their values, an empty table, two columns whose declared type contradicts their content, and personal data alongside two designed false positives. Needs the `[duckdb]` extra and says so by name when it is absent (`reason: prerequisite`) | | `connect test` | capabilities, dialect, `read_only: true`; DuckDB takes `--path`, every warehouse connector takes repeatable `--scope` (a bare database on ClickHouse, whose identifiers are two-part `database.table`) (BigQuery also accepts its older `--project`/`--dataset`), never written to config; Snowflake and Databricks report the pinned warehouse and its credit or DBU rate; ClickHouse Cloud reports live replica memory and its derived compute-unit rate | | `explore inventory [--rank] [--limit N] [--all]` | ranked object summary (counts, sizes; no rows). `--rank` caps at 30 objects by default so the shortlist stays a shortlist on a large warehouse; `--limit` widens it, `--all` lifts the cap; both are no-ops without `--rank`, since the unranked list carries no order to cut from | -| `explore profile [--columns all]` | column profiles + PII flags (column, category, confidence) + candidate keys, grain, data-quality warnings; the verdict fields (grain, keys, data quality, row count) lead the serialized payload and `columns` trails, so a truncating harness cuts schema rather than the answer. By default each dataset's `columns` is summarized to the ones carrying a finding (PII, a non-zero null fraction, key membership, a reported value domain, or a mention in a `data_quality` note), with the rest counted in `elided_column_count`; `--columns all` restores every column. A per-column field that is null on every column shown is dropped from each of them and named once in `suppressed_fields` (always present, so an empty list means the columns shown carry their full shape), and each `value_domain` carries its `profile_value_domain_cap` most frequent values (`.dex/config.yml`, default 25) with the rest counted in its `elided`; both reduce the payload only, never the cached profile. `--use-project` lets a semantic model's declared primary entity override the heuristic grain (disagreements noted) | +| `explore profile [--columns all]` | column profiles + PII flags (column, category, confidence) + ranked candidate keys, grain, `key_evidence`, data-quality warnings; the verdict fields (grain, keys, key evidence, data quality, row count) lead the serialized payload and `columns` trails, so a truncating harness cuts schema rather than the answer. `candidate_keys` is ranked, tightest proven key first, and `key_evidence` says why for each one: a combination that is unique only because one of its members is unique on almost every row, or because a continuous measure completes it, is suppressed rather than reported, with its reason kept, since a caller who cannot tell the real key from the filler is worse served by several candidates than by one named defect and none. Where a column is unique on almost every row, `data_quality` says so with the ratio, the counts, and the exact number of rows that would have to be removed for it to be unique, because the finding is duplicates in the source rather than an absent key. By default each dataset's `columns` is summarized to the ones carrying a finding (PII, a non-zero null fraction, key membership, a reported value domain, or a mention in a `data_quality` note), with the rest counted in `elided_column_count`; `--columns all` restores every column. A per-column field that is null on every column shown is dropped from each of them and named once in `suppressed_fields` (always present, so an empty list means the columns shown carry their full shape), and each `value_domain` carries its `profile_value_domain_cap` most frequent values (`.dex/config.yml`, default 25) with the rest counted in its `elided`; both reduce the payload only, never the cached profile. `--use-project` lets a semantic model's declared primary entity override the heuristic grain (disagreements noted) | | `explore relationships [--verify] [--use-project] [--use-hosted-semantic-layer]` | inferred joins with confidences, plus notes on what inference examined; `--verify` measures each join with an aggregate overlap probe, declared and inferred alike, including the full ordered tuple of a composite key. `--use-project` folds local project and semantic-layer declarations in at confidence 1.0; native composite relationships retain every column pair and provenance. `--use-hosted-semantic-layer` independently authorizes a configured hosted catalog read, but a backend with no physical relations cannot add warehouse edges or exposure annotations. A measurement never revises a declared join's confidence, which stays at the 1.0 the project asserts; a declared join whose probe finds the parent largely missing is reported as a finding instead | -| `explore map [--detail] [--verify] [--use-project]` | writes/updates the `.dex/` map and returns it: the counts as before, plus `data.objects` (per top-ranked object: row count, detected grain, candidate key, the notable columns with the role that earned each one a place, PII flags as category and confidence, and data-quality findings) and `data.edges` (the join edges, shaped exactly as `explore relationships` returns them). Budgeted like `explore diagram`: 25 objects by rank, 12 columns per object, 40 edges, 5 findings per object, every cap binding in every mode and every elision counted in `notes` and in an `elided_*` field, so a truncated answer never reads as a complete one. `--detail` widens the selection to every column and to objects that were inventoried but never profiled, and lifts no cap; it is not `--full`, which decides how much gets scanned and therefore what the run costs. No column value ever appears: the cache holds min/max and value domains and this command does not read them. `--use-project` additionally applies declared grain, ranks metric-backing models higher, folds the semantic layer's declared entity graph into `data.edges`, and marks each object with `semantic_models`, the semantic models that sit on that relation. Empty there is an answer: a relation nothing in the layer reads is a different object from one several metrics are built on, and row counts and PII flags cannot tell them apart. Every object in view is rewritten whenever the layer was read, so a model dropped from the layer clears rather than leaving a stale claim, and a project with no compiled semantic layer contributes nothing here rather than erroring | +| `explore map [--detail] [--verify] [--use-project]` | writes/updates the `.dex/` map and returns it: the counts as before, plus `data.objects` (per top-ranked object: row count, detected grain, the best-ranked candidate key, the notable columns with the role that earned each one a place, PII flags as category and confidence, and data-quality findings) and `data.edges` (the join edges, shaped exactly as `explore relationships` returns them). Budgeted like `explore diagram`: 25 objects by rank, 12 columns per object, 40 edges, 5 findings per object, every cap binding in every mode and every elision counted in `notes` and in an `elided_*` field, so a truncated answer never reads as a complete one. `--detail` widens the selection to every column and to objects that were inventoried but never profiled, and lifts no cap; it is not `--full`, which decides how much gets scanned and therefore what the run costs. No column value ever appears: the cache holds min/max and value domains and this command does not read them. `--use-project` additionally applies declared grain, ranks metric-backing models higher, folds the semantic layer's declared entity graph into `data.edges`, and marks each object with `semantic_models`, the semantic models that sit on that relation. Empty there is an answer: a relation nothing in the layer reads is a different object from one several metrics are built on, and row counts and PII flags cannot tell them apart. Every object in view is rewritten whenever the layer was read, so a model dropped from the layer clears rather than leaving a stale claim, and a project with no compiled semantic layer contributes nothing here rather than erroring | | `explore diagram [--full]` | the `.dex/` map serialized as a Mermaid `erDiagram` under `data.mermaid`, plus an `entities` legend mapping each entity name back to its fully-qualified identifier. Free and connectionless: it reads the cache and never opens the warehouse, so it needs no credential and cannot spend. Declared joins are solid, inferred joins dotted, and a cardinality is drawn only where the cache proved it (an unverified inference never claims "exactly one"). A solid edge whose label names a semantic entity is a join the semantic layer declares; the cardinality rule is unchanged for it, so a primary entity is the layer's claim and still buys no "exactly one" the cache has not proven. The default draws profiled objects that participate in a join, with their grain, key, join, and PII-flagged columns; `--full` widens to every eligible object and column. An entity cap always binds and every elision is counted in `notes`. No column value ever appears; PII renders as category and confidence. dex writes no file: reproduce the string in a fenced ```mermaid block, or save it yourself | -| `explore query "" [more...]` | runs agent-authored SELECTs through the query firewall: columnar, capped results; `data.shape` is `columnar`, making the `columns`/`types`/`cells` layout discoverable from the envelope itself. Values only come from profiled columns whose PII flag is absent or below the blocking threshold (sub-threshold projections warn in the envelope; `references/pii-policy.md`); the FROM clause may unnest JSON/array columns in the connector's native idiom (UNNEST, LATERAL FLATTEN, LATERAL VIEW EXPLODE, set-returning functions, PartiQL) when the unnested value derives from a queried table's column, with the outputs inheriting that column's flags. The positional is variadic, so a chain of questions is one call: each argument is one statement, adjudicated, executed, and ledgered on its own, and `--sql-file ` reads a larger batch from a file (one statement per line, or semicolon-separated). Several statements in one string is still refused, so batching never widens what a call may do. One statement returns the envelope described here; two or more return `data.results`, one entry per statement carrying its own `shape` discriminator and this same `columns`/`types`/`cells`/`row_count`/`truncated` layout plus its own `status` (`ok`, `refused`, `failed`, `skipped`) and `error`, so a refusal on the third does not discard the first two, and the envelope's own status is `error` whenever any statement failed. `query.max_payload_bytes` is the budget for the whole call rather than for one statement, and `query.max_statements` (default 10) refuses an oversized batch. An object a statement names that the connection has but the cache cannot adjudicate (never profiled, inventoried without column detail, or profiled against a column signature the warehouse has since changed) is profiled first and the statement then runs, with a warning naming what was profiled and `data.profiled_on_demand` listing it; that profile is a full one, so the flags governing the query are the flags a deliberate `explore profile` would have produced. On a metered connector it is priced, not implied: one handshake covers the profiles and every statement together, itemized per table and per statement, and the objects a whole batch needs are scanned once rather than once per statement. An object the connection does not have refuses only the statements that named it, naming the connection rather than the cache. `--no-auto-profile` (or `auto_profile: false` in `.dex/config.yml`) restores the strict prerequisite, and on that path nothing opens a connection before the firewall has spoken | | `explore cluster [--features a,b] [-k N]` | k-means over a bounded, column-pruned, dialect-sampled scan of numeric columns; returns cluster sizes + centroids (feature means) + silhouette, never rows; auto-selects non-PII, non-key numeric features (or takes `--features`; a named PII column is opt-in, mean only); needs the `[cluster]` extra; profiles the named object on demand exactly as `explore query` does, including `--no-auto-profile`, and reports it under `data.profiled_on_demand`; billed connectors take the cost handshake, and where a profile was needed the sample scan is priced after it (the feature columns come out of that profile), so a budget too small for the sample comes back as `needs_confirmation` with the profile already saved rather than as a refusal | | `explore semantic list [--local\|--api] [--metric ] [--for-dimension ] [--search ] [--full]` | discover the semantic layer's objects in one shape from either backend: semantic models (the layer's organizing unit, with the transformation model each sits on and its default time dimension), metrics (type, the tokens each can be grouped by, the measures it reads, a ratio's two sides, any filter, the queryable grains, and `time_axis`: the physical time column(s) `metric_time` resolves to for that metric, where more than one means its measures aggregate over different timestamps and a time grouping buckets parts of the number differently), dimensions (the token a query groups by, plus the bare definition, owning model, and queryable grains behind it), entities (one declaration per semantic model, each with its own join key and caveats, and a derived `type` that is primary wherever any declaration is), and measures (the aggregation and expression the number is actually made of). Every label and description is the project's own words where it declared one, and an unset field is omitted rather than null. Costs no warehouse query: one GraphQL round trip hosted, one compiled-artifact read locally, through the project seam rather than a dbt-specific parser. `--metric` narrows the catalog to those metrics and what they reach, and names the scope in `scoped_to` so a subset is never mistaken for the layer. `--for-dimension` (repeatable, comma-separated) asks the reverse question, returning the metrics groupable by **all** the named tokens and narrowing the catalog to them, which is also the cheapest way to find the metrics that can go on one chart against one axis; it is an inversion of the `dimensions` list each metric already carries rather than a second call, so it costs nothing, answers for a metric's own time token, and refuses an unknown token by name instead of returning the empty list a caller would read as a fact about the layer. `--search` is for a caller who knows a word rather than a name: it matches case-insensitively against every element's name and against the project's own label and description, resolves to the metrics that word touches, and names a term that matched nothing in a note rather than refusing, since a substring matching nothing is an honest answer where a misspelled metric name is not. The three compose, applied in that order so `--metric x --search y` reads as "within x, the parts about y", and a named metric that cannot be grouped that way is dropped with a note naming it. Budgeted like `explore map`: 50 semantic models, 60 metrics, 150 dimension rows, 50 entities, 60 measures and 40 groupable tokens per metric, with every cut counted in `elided` and named in `notes`, `elided` present with its zeros so a complete catalog states that it is complete, and `--full` to lift the caps. The defaults leave a layer of a dozen models and a few dozen metrics uncut, so a cap only bites one that was already too large to read in one payload; the narrowing flags are the better answer either way, because they decide which part comes back rather than letting a cap decide. Two legitimate backend differences are declared in the payload rather than left to be inferred: `dimension_scope` says whether a dimension row is one declaration or one groupable path (which is why the two backends can report different dimension counts for one layer; `--local` resolves the join graph through MetricFlow where the `[semantic]` extra is installed, and says `declarations` plus a note where it could not), and `unavailable` names the fields a backend structurally cannot supply (the dbt Cloud API exposes no entity label, no semantic model metadata beyond a name, and no words on a measure). The catalog also resolves the layer onto the warehouse: a semantic model carries the `relation` it sits on, and each dimension, entity declaration and measure carries the `column` behind it, so "which table is behind this metric" is the metric's `semantic_models` followed to their relations and `explore profile` is the next call. The relation is carried once per model rather than once per element, and an element defined as a computed expression carries no column rather than a guessed one, because the PII gate resolves a dimension to a column and reads that column's evidence. `relation` is in `unavailable` on the hosted backend, which exposes columns but no relations at all. Distinct from the top-level `semantic` group, which *authors* the layer; `explore semantic` *queries* it. The catalog names which layer answered on four fields: `backend`, `vendor`, `deployment`, and `execution`. `semantic.vendor: ossie` reads native Apache Ossie documents from the repository rather than a dbt project, with no MetricFlow in the path and needing the `[ossie]` extra; it is catalog-first because Ossie specifies interchange metadata and no portable query runtime, so `query`, `values` and `--for-dimension` refuse by name, and the catalog declares in `unavailable` what the format structurally cannot carry (no measures, no entities, no metric groupability) rather than returning empty fields a caller would read as facts about the layer | | `explore semantic values [--metric ] [--local\|--api]` | one semantic dimension's value domain: what a filter on it may be filtered to, capped and columnar like `explore query`. The precondition for writing a filter, and on a hosted layer the only dex command that can reach it at all, since dbt Cloud is not a connector and `explore profile` cannot see a semantic dimension. Takes exactly one dimension and accepts a grain suffix (`user__created_at__month`), split against the grains the layer reports. The token is resolved against the catalog first, so a misspelling is refused by name. `scoped_to` says how the values were reached and changes what they mean: empty is the domain of the column behind the dimension, and a metric name means the values present for that metric. A dimension reached through a join has no other answer (neither layer will run a distinct-values query with no measure to join from), so dex renders the cheap form first, escalates once to a metric that reaches it, and names that metric and `--metric` in a note rather than narrowing silently. PII is screened harder than on a query: the whole output is values, so a flagged dimension refuses the command, and the refusal names the durable ways to clear one reviewed as not PII. dex reports the values that came back and never claims an exact cardinality, which would cost a second scan; a large domain comes back capped, `truncated`, and saying so. Local renders through MetricFlow and takes the full cost handshake (needs the `[semantic]` extra, unlike `list`); hosted is executed by dbt Cloud and carries the same cost-guard-unavailable warning as a hosted query. `semantic.vendor: ossie` refuses this command by name: Ossie defines no distinct-values API and no query runtime, so the answer is the physical route it names, following the dimension's `semantic_model` to its `relation` and reading that relation with `explore profile` and `explore query` under the firewall and the cost guard | @@ -109,7 +109,7 @@ credentials and no network. | `maintain check` | sweep every drift axis vs the snapshot; ranked drift report (read-only); two-phase on billed connectors: the free axes complete and return `ok`, with one estimate for the scanning axes under `data.offer`. A repository whose only layer is a semantic one still sweeps: the axes that need a transformation project report what they could not check rather than reporting no drift | | `maintain schema []` | structural drift: columns/tables added, dropped, retyped, renamed; nullability; dangling sources; a model added, removed, or content-changed since the baseline (free) | | `maintain volume []` | freshness drift: row counts that collapsed, emptied, or spiked (free metadata). Free metadata carries no count for an object the warehouse does not maintain one for (a view anywhere, an external table on BigQuery), so those are named in `warnings` as not compared rather than returning no finding | -| `maintain grain []` | cardinality/identity drift: lost key uniqueness, changed grain, join fanout, plus the grains the project itself declares (model-level `unique_combination_of_columns`) re-verified against the data (scans; gated on billed connectors). Two codes come out of the uniqueness checks and the difference is the baseline: `key_lost_uniqueness` is a key measurement proved unique and no longer is, `declared_grain_not_unique` is a declared combination that never held, which is a declaration to fix rather than drift to absorb. Below `maintain.grain_min_rows` rows (default 100, set in `.dex/config.yml`), a uniqueness-regression finding is damped to `low` rather than `high`: on a handful of rows, losing uniqueness means the least, and a 4-row table's boolean column "loses" a uniqueness it never meaningfully had. Damped, never dropped, and the damping is named in the finding's own `data` (`severity_floor_applied`, `grain_min_rows`) and prose. The declared grains come from whichever layers declare one: a dbt model-level `unique_combination_of_columns`, and a semantic layer's own key declarations, including a multi-column one, which is measured as a complete composite and never one column at a time. A declaration that never held is `declared_grain_not_unique`, a finding to fix rather than drift to absorb, and it is never an automatic rewrite | +| `maintain grain []` | cardinality/identity drift: lost key uniqueness, changed grain, join fanout, plus the grains the project itself declares (model-level `unique_combination_of_columns`) re-verified against the data (scans; gated on billed connectors). The composites it re-probes are the ranked, artifact-suppressed set `explore profile` reported, so a table whose only combinations were filler for a near-unique column costs nothing to re-check on a metered connector. Two codes come out of the uniqueness checks and the difference is the baseline: `key_lost_uniqueness` is a key measurement proved unique and no longer is, `declared_grain_not_unique` is a declared combination that never held, which is a declaration to fix rather than drift to absorb. Below `maintain.grain_min_rows` rows (default 100, set in `.dex/config.yml`), a uniqueness-regression finding is damped to `low` rather than `high`: on a handful of rows, losing uniqueness means the least, and a 4-row table's boolean column "loses" a uniqueness it never meaningfully had. Damped, never dropped, and the damping is named in the finding's own `data` (`severity_floor_applied`, `grain_min_rows`) and prose. The declared grains come from whichever layers declare one: a dbt model-level `unique_combination_of_columns`, and a semantic layer's own key declarations, including a multi-column one, which is measured as a complete composite and never one column at a time. A declaration that never held is `declared_grain_not_unique`, a finding to fix rather than drift to absorb, and it is never an automatic rewrite | | `maintain semantic []` | definition drift and dangling refs (free, and returned as `ok`) plus categorical dimension cardinality change (scans; offered under `data.offer` and gated on billed connectors). Definition drift is a definition added, removed, or changed; a source relation that no longer exists; a dimension, entity, measure, or declared key naming a column that is gone; and a relationship whose endpoint or column pairs no longer resolve, which is reported at high severity because a join nothing can resolve is a broken layer rather than a stale one. The cardinality half needs a semantic model that names a transformation model, so on a layer that names none this axis is free and this command never offers a scan | | `maintain reconcile []` | propose the edits that reconcile detected drift, as a stored plan of diffs tagged mechanical or advisory (never applied; apply with `transform apply `). It composes every layer's declarations first, so a grain a semantic layer already declares is not proposed as though nothing declared it. Where no editable project is configured there is nothing for it to author, so every proposal is advisory and no plan is stored; authoring into a native semantic layer is `semantic ossie` rather than this command | | `maintain verify []` | is the project correct *right now*, with no `.dex/snapshot.json` baseline required, unlike every other `maintain` subcommand above. Two finding classes. **Build status**: nodes that failed to build, nodes skipped because a parent failed (naming it, walking back through a chain of transitively-skipped parents when the immediate parent was itself only skipped), nodes that warned rather than failed (`node_warned`, ranked low: a project running relationship tests at `severity: warn` over documented gaps has warnings by design, and what it needs is a list to compare against last run's), and models the project declares that have no relation in the warehouse. **Row population**: each model's compiled SQL is read for its *driving parent*, the relation in the FROM clause as distinct from anything joined, followed through the CTE chain a dbt model compiles to; `row_loss` reports a model holding materially fewer rows than that parent when nothing in its SQL (a `WHERE`, `HAVING`, `QUALIFY`, `GROUP BY`, `DISTINCT`, `LIMIT`, a semi or anti join, a set operation) accounts for the shortfall, naming the join that can explain it, and `row_fanout` reports one holding materially more, naming the join and its key columns. Both state the two counts so a reader can judge the threshold, and both are deliberately conservative: a model with any reason to hold fewer rows is never reported for loss, and an incremental model is skipped outright because it holds what previous runs loaded rather than a function of its parent. Free wherever the answer is free: the manifest read opens no connection, the relation check and the row counts read cheap object metadata, and on a connector with no cost gate every count is made exact because doing so bills nothing. A relation the warehouse keeps no count for (any view, which is dbt's default materialization) is the only part that costs anything: those counts are batched into one aggregate-only statement, priced, and returned as an `offer` beside the findings that are already final, never taken. A project that fails to compile is reported first and suppresses every other check here, since a finding computed from a manifest a broken project could not have produced honestly is not a finding at all; `data.suppressed` names every finding class that did not run and why, so an empty `data.findings` from a run that skipped everything is never mistaken for a clean project, and a model whose driving parent could not be identified is named in `warnings` rather than passing as checked | @@ -125,12 +125,10 @@ is not free: `schema`, `volume`, and the reference half of `semantic` are metada scan and go through the `--confirm --budget` handshake on billed connectors. The engine does not care which skill fronts a subcommand. -A command whose free half completed reports `ok` and puts the price of the -scanning half in `data.offer`, rather than gating the whole answer behind a -confirmation. `needs_confirmation` means dex is waiting on you for work you asked -for; an offer is work you did not ask for, and ignoring it is a valid choice. -Read `data.axes_run` for what completed and `data.offer.axes` for what the -estimate would add, since with an `ok` status those are no longer implied. +A command whose free half completed reports `ok` and offers the scanning half +(see the cost section below). Read `data.axes_run` for what completed and +`data.offer.axes` for what the estimate would add, since with an `ok` status +those are no longer implied. Authored content reaches the engine through `--edits-file ` (or `-` for stdin): a JSON payload of `{"edits": [{"path", "kind", "op", "content"}, ...]}` @@ -211,65 +209,37 @@ operations and last-resort credential discovery; it does not select dex's connector or override `--connector`/`--path` or `.dex/config.yml`. Cost is a preflight estimate surfaced **before** any spend. Any command that -would spend requires an explicit `--confirm` and a session budget: on a -metered connector (BigQuery, Snowflake, Databricks, Redshift, Postgres, and -ClickHouse) -the first call returns `needs_confirmation` with a free estimate, and the -same command is re-issued with `--confirm --budget ` once the user -has agreed to the spend. +would spend requires an explicit `--confirm` and a session budget: on a metered +connector the first call returns `needs_confirmation` with a free estimate, and +the same command is re-issued with `--confirm --budget ` once the +user has agreed to the spend. One exception to the status, not to the rule: a +command that finished free work the caller did want, and can offer paid work +they did not ask for, returns `ok` with the estimate under `data.offer` instead. +That is `maintain check` and `maintain semantic`. Reserve `needs_confirmation` +for reading "dex is waiting on me", and an offer for "there is more available if +I want it". -One exception to the status, not to the rule: a command that finished free work -the caller did want, and can offer paid work they did not ask for, returns `ok` -with the estimate under `data.offer` instead. That is `maintain check` and -`maintain semantic`. The re-issue is identical (`--confirm --budget`), nothing -runs until it arrives, and `cost.estimate` stays empty so an `ok` never reads as -though it spent. Reserve `needs_confirmation` for reading "dex is waiting on -me", and an offer for "there is more available if I want it". +The magnitude is paradigm-relative: -The first billed command in a project that has never -decided whether the *day's* total is bounded also carries a -`suggested_session_ceiling`, and adding `--session-ceiling ` (or -`--no-session-ceiling`) to that same re-issue answers both asks at once and is -recorded in `.dex/config.yml`; skip it and the confirmed run stops once to ask. -The magnitude is paradigm-relative: **bytes** on -BigQuery (an exact free dry-run figure), **warehouse-seconds** on Snowflake -(a heuristic labeled `estimate_quality: "heuristic"`, with a credit -translation alongside) and on Databricks (a floor labeled -`estimate_quality: "low"` that sharpens itself inside the confirmed budget, -with a DBU translation alongside), **compute-seconds** on Redshift (a -heuristic with an RPU-hour translation alongside; Serverless estimates carry -the 60-second wake minimum once), **database-seconds** on Postgres (nothing -is billed in dollars; the guarded quantity is load on the operational -database, estimated free via EXPLAIN) and on ClickHouse (database-seconds when -self-hosted; compute-seconds in Cloud with live-capacity CU-hours and optional -USD alongside; estimated free via the non-executing EXPLAIN ESTIMATE, which -prices after primary-key pruning). On every time paradigm the budget -still binds exactly via a server-side statement timeout, except on ClickHouse, -where it binds via `max_execution_time` **and** `max_bytes_to_read`, because -time alone is checked only at block boundaries there. Actual spend comes -back under `data.spend` (`bytes_billed` or `seconds_billed`) and accumulates -in the `.dex/spend.jsonl` ledger per connector. That is the only place spend is -reported: every command that can bill carries the unit key whatever it settled -at, zero included, and no command puts a billed magnitude anywhere else in -`data`, because a key present on one command and absent on another reads as a -spend of zero rather than as a key to look for elsewhere. That ledger gates billing and -nothing else: it is read where work is admitted, not where a connection is -assembled, so a command that cannot spend does not depend on it, and a ledger -that cannot be read refuses billed work by name while reporting the day's total -as `null` on the two surfaces that quote it. Every row in it declares what -it is, so the file is readable by anything that wants to audit spend: one JSON -object per line, always the same keys, an `entry` of `reservation`, `settlement` -or `release`, and settled spend is the `entry == "settlement"` filter summed over -one connector's unit. A reservation and its release cancel, so that filter and -the day's total agree whenever nothing is still running, and where they differ -the difference is headroom a live command is holding. The ledger also records each -command's estimate beside what it settled at, which is what lets a refusal over -the ceiling end with this connector's own observed ratio ("the last 8 settled -bigquery commands billed a median 69% of estimate, range 61%-88%") instead of -leaving the next budget to a guess; with too little history it says so rather -than quoting a ratio. The refusal itself is unchanged and still cannot be -confirmed through, and the ceiling is checked against the estimate, so a budget -set at that fraction of the estimate is refused again. Credentials never appear in +| Connector | Magnitude | Estimate | +|---|---|---| +| BigQuery | bytes | exact, from a free dry run | +| Snowflake | warehouse-seconds | heuristic, with a credit translation | +| Databricks | warehouse-seconds | a floor that sharpens inside the confirmed budget, with a DBU translation | +| Redshift | compute-seconds | heuristic, with an RPU-hour translation | +| Postgres | database-seconds | free, via EXPLAIN | +| ClickHouse | database-seconds, or compute-seconds in Cloud | free, via EXPLAIN ESTIMATE | +| DuckDB | free and local | nothing to confirm | + +On every time paradigm the budget binds via a server-side statement timeout, +except on ClickHouse, where it binds via `max_execution_time` **and** +`max_bytes_to_read`, because time alone is checked only at block boundaries +there. Actual spend comes back under `data.spend` and accumulates in the +`.dex/spend.jsonl` ledger. The daily cumulative ceiling, the ledger's row shape, +the one-time ceiling ask, and what a refusal quotes back from spend history are +in `references/cost-controls.md`. + +Credentials never appear in `data` (BigQuery authenticates via discovered Application Default Credentials, Snowflake via a discovered `connections.toml` entry, environment, or dbt profile, Databricks via the SDK's unified chain, Redshift @@ -300,38 +270,21 @@ them. 4. Cost-aware by connector. Nothing dex runs touches the warehouse without a ceiling. The source allowlist in `.dex/config.yml` is a committed cost boundary: `--scope` narrows it for one command and can never widen it, and a - scope that names nothing is refused rather than dropped. The one place dex - cannot enforce a ceiling is the hosted dbt Cloud Semantic Layer - (`explore semantic query --api`): dbt Cloud owns the warehouse connection and - executes the query server-side, so no dry-run estimate and no `maximum_bytes_billed` - are possible from dex. That backend therefore runs without a `--confirm` - handshake and instead states, explicitly and on every result, that the cost - guard is unavailable and spend is governed by the dbt Cloud environment, not - by dex. The local backend (`--local`) executes through dex's own connector and - keeps the full cost-before-spend handshake. `budget.session_ceiling` binds - across commands that overlap in time as well as across commands that follow - one another: an admitted command books its estimate against the day's - headroom before it runs, so issuing several billed commands at once cannot - spend the same budget twice. If a cache backend cannot serialize that, every - billed command says so, and if the ledger it binds against cannot be read, - billed admission refuses rather than deciding a ceiling from nothing. And a - project is asked for that daily cap once rather than warned about it forever: - the first billed command in a project that has never decided returns - `needs_confirmation` with a `suggested_session_ceiling`, answered by - `--session-ceiling ` or `--no-session-ceiling` and recorded in - `.dex/config.yml`, so an unbounded day is a decision somebody made rather - than the default nobody noticed. + scope that names nothing is refused rather than dropped. One surface cannot + be ceilinged at all, the hosted dbt Cloud Semantic Layer, and it says so on + every result rather than implying a guard it does not have. The daily + cumulative ceiling, the spend ledger, the one-time ceiling ask, and that one + unguarded surface: `references/cost-controls.md`. 5. Nothing reaches agent context except through the sanitized envelope. Credentials never; data values only from profiled, PII-cleared columns, bounded and capped. 6. PII is flagged (column, category, confidence), never surfaced, and a flag is never removed by evidence: value-shape statistics computed in the profiling - scan only move its confidence, in both directions and fail-closed. The query - firewall enforces the policy on agent SQL: any expression that would carry - values from a column flagged at confidence 0.5 or above is refused (the - threshold is a hard-coded engine constant); a projection of a lower-confidence - flag runs with an envelope warning. A human clears a reviewed column durably - with a `pii_overrides` entry in `.dex/config.yml`, never by editing the cache. + scan only move its confidence, in both directions and fail-closed. Only a + human clears a flag, durably and in committed config, never by editing the + cache. The policy in full, including which surfaces gate on a flag's presence + and which on its confidence, the blocking threshold, and both clearing + routes: `references/pii-policy.md`. 7. Persistence is git, not a service. The repository is the source of truth, on two axes: the transformation project and the semantic layer. The `.dex/` directory is a non-canonical cache (exploration artifacts and the reconcile @@ -345,6 +298,11 @@ them. - DexEngine: `packages/dex-core/` (PyPI: `exmergo-dex-core`, Apache-2.0). - Connector and methodology notes: `references/`. - The contract in full: `references/command-contract.md`. +- The PII policy: what a flag means, which surfaces gate on presence and which + on the threshold, and how a human clears one: `references/pii-policy.md`. +- The cost guard: the confirm handshake, the spend ledger, the cumulative daily + ceiling, and the one surface where no ceiling is possible: + `references/cost-controls.md`. - The source of truth and the `.dex/` cache: `references/canonical-model.md`. - Where `.dex/` state lives, selecting a backend, and writing one: `references/storage.md`. diff --git a/CHANGELOG.md b/CHANGELOG.md index f8a9aa38..391d5546 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,6 +9,140 @@ tag releases both in lockstep, so entries below are keyed by the engine version. ## [Unreleased] +### Fixed + +- **`explore profile` no longer offers composite keys that are artifacts of a + near-unique column or of a continuous measure, ranks the ones it does report, + and says why for each** ([#292]). On a 2,037-row orders table where `order_id` + held 1,927 distinct values, the profile returned five candidate keys and + elected a grain out of them: `order_id, customer_id`, and then `CREATED_AT`, + `SUBTOTAL`, `GRAND_TOTAL` and `UPDATED_AT` each paired with `order_id`. Four + of the five were one fact wearing four hats, that `order_id` is unique on all + but 110 rows so any wider column completes it. `(SUBTOTAL, order_id)` is not + a grain: it is the observation that two rows sharing an order id happened to + differ in their subtotal. The list carried no order, so a caller could not + tell the real key from the filler, and the one thing worth saying, that this + table has duplicate order ids, was the one thing the profile did not say. + + The same bug shipped in the demo warehouse, where `order_items` reported a + grain of `(unit_price, order_id)`: a `DECIMAL(10,2)` money column paired with + a foreign key, on a table whose actual story is 1,000 duplicate + `order_item_id` rows from a double-loaded batch. + + Two exclusions, both applied before a pair is priced. A pair whose member is a + **continuous measure** is dropped, because a measurement's cardinality grows + with the table and it completes a partner by arithmetic rather than by + meaning. A pair **anchored on a column already unique on almost every row** is + dropped too, unless its partner has a domain bounded by the fanout rather than + by the table: a few duplicate order ids separated by a three-value line number + really is that grain, where any wider partner separates them by accident. A + near-unique timestamp gets no such exception, since a per-row event time is + not an entity whose rows a position column enumerates. Neither rule applies + below a fixed row floor, where every column looks near-unique and every + measure looks continuous, so a small table has everything asked as before. + + The measure test reads a declared type with an explicit nonzero scale, and a + measure-name vocabulary above a modest cardinality bar. The vocabulary is not + decoration: Snowflake's `SHOW COLUMNS` renders every `NUMBER` as the bare + token `FIXED` with the scale dropped, and BigQuery `NUMERIC` carries no scale + either, so on those connectors the type test goes silent by design and the + name is the only signal left. A name never decides alone, because `quantity`, + `amount` and `total` are legitimately low-cardinality members of real fact + grains. + + Where nothing survives, `grain` is `null` and the profile names the near-unique + column with its counts, including how many rows would have to be removed for + it to be unique (`order_id is not unique: 1927 distinct over 2037 rows (110 + rows would have to be removed for it to be unique, so it is unique for 94.6% + of rows)`). That figure needs no new measurement: it is the non-null row count + less the distinct count. It is phrased as rows to remove rather than as "110 + duplicate order ids" because the arithmetic counts surplus rows and not values + that repeat, and the two differ. Two long-standing imprecisions in that + sentence are fixed with it: the surplus was computed against the total row + count rather than the non-null count, overstating it on a nullable column, and + it carried a `~` even when derived from two exact numbers. The marker now + follows the arithmetic, and a percentage that would round to `100.0%` prints + `>99.9%` rather than contradicting the sentence it sits in. + + Suppression is not silence. One new field, `key_evidence`, carries an entry + per combination the profile considered, each with its `columns`, a `status` of + `reported` or `suppressed`, and the `reason` in the profile's own words; the + reported entries are in the same order as `candidate_keys`, which is the + invariant that keeps the two from drifting. `data_quality` states the + suppression too, and deliberately names only the anchor: a column named in a + note is kept in the serialized `columns`, so naming the filler would drag four + columns back into a payload that already elides them correctly. + + Field-visible beyond `explore profile`, because five things read + `candidate_keys`. `explore diagram` stopped marking a money column `PK`, and + `order_id` on the demo's `order_items` reads `FK` rather than `PK`. + `explore map`'s best-ranked `candidate_key` and its column roles no longer + point at filler, while a column named by a suppressed entry stays in the + notable set with no role, so a map that reports duplicates in a column still + lists it. `transform plan --scaffold` reads `candidate_keys[0]` to decide + which columns get tests, so a scaffolded `schema.yml` no longer puts + `not_null` on a money column. And `maintain grain` re-probes every composite + in `candidate_keys` on every run, billed on metered connectors, so four + artifacts per table stop being a recurring charge. On the demo the pairs + actually probed fall from five to two. A declared grain still overrides all of + it, and both declared-grain notes now state their origin even when there was + no heuristic grain to disagree with, so a declaration that fills a gap says so + rather than applying silently. + + `CACHE_SCHEMA_VERSION` moves to 4. An older engine reads a version-4 cache + fine, since an unknown key is ignored; the direction that breaks is a current + engine reading a version-3 one, where a suppressed combination still reads as + a ranked candidate and an empty `key_evidence` is indistinguishable from a run + that suppressed nothing. Rather than warn about it, the profile freshness gate + now treats a pre-4 profile as stale, so the first `explore profile` or + `explore map` after upgrading re-scans and the cache heals itself. The + `explore query` degradation on an older cache is unchanged and still refuses + nothing. + + Deliberately unchanged: no flag. The issue's position is that a spurious key is + worse than nothing, so there is no opt-out to talk past it, and none is needed + since `key_evidence` hands back every suppressed combination and its reason in + the same payload. No `probed` boolean either, because the probe's existing + budget notes already say when it did not run, and duplicating them into a + field would be growth for nothing. `key_evidence` is not in `explore map`'s + payload, which is budgeted per object; the full ranking belongs to the command + whose subject is one relation in full. The demo warehouse's data is untouched, + since its determinism is a contract and it is the reproduction rather than + something to fix. + +### Changed + +- **The PII policy and the cost guard each have one document, and every other + document links to it.** Both guardrails cut across every connector, command and + backend, so neither had an owner: the PII blocking threshold was stated in five + places in four wordings, one paragraph about auto-profile pricing was copied + verbatim into five connector references, the `--confirm` handshake was fully + restated in roughly thirteen files, and the hosted dbt Cloud exception appeared + ten times across seven. `references/pii-policy.md` and + `references/cost-controls.md` now own the policy, the constants and the + end-to-end flow. Connector references keep their own cost models, which are + genuinely per-warehouse, and `references/storage.md` keeps the store protocol a + backend implements, which is a different reader's question. `AGENTS.md` + guardrails 4 and 6 state the invariant and point. + + Two behaviors that lived only in this changelog are now documented: the + `pii_overrides` mismatch warning that `explore profile` and `explore map` both + emit, and the `column_name` plus `scope` pattern form of an override entry. + + The three skills keep their copies, because `npx skills add exmergo/dex` + installs each one standalone with no engine repository to link into. What they + no longer carry is the constant: a skill states the consequence ("below the + blocking threshold it projects with a warning") and names the file that states + the number, so a skill can go stale on wording but not on the threshold. + + `packages/dex-core/tests/test_docs_policy.py` holds this in place. It asserts + that any document stating the threshold states the engine's + `PII_BLOCK_CONFIDENCE`, that the set of files allowed to state it has not grown, + that the shared cost prose lives in one file, that every connector's + `session_ceiling` example keeps one wording, and that relative links between + documents resolve. `packages/dex-core/README.md` is deliberately untouched: the + PyPI package ships without `references/`, so that file stays self-contained. + ### Added - **`transform test --mutate ` measures whether a model's tests would diff --git a/README.md b/README.md index 23a4ac0a..ceb192f9 100644 --- a/README.md +++ b/README.md @@ -109,7 +109,7 @@ from a pinned seed, so what you see is what is written here. It is seeded to be **realistically broken**, because a first run that reports a clean bill of health teaches you nothing. `explore map` flags 6 columns as personal data, -infers 5 joins, and reports 5 data-quality findings. Then: +infers 5 joins, and reports 6 data-quality findings. Then: ``` dex explore profile order_items products @@ -117,8 +117,13 @@ dex explore relationships --verify dex explore query "select email from customers" ``` -- **A broken grain.** `order_item_id is not unique: 13000 distinct over 14000 rows`, - because a batch was loaded twice. Any join on it silently fans out. +- **A broken grain, reported as duplicates rather than as a missing key.** + `order_item_id is not unique: 13000 distinct over 14000 rows (1000 rows would have + to be removed for it to be unique, so it is unique for 92.9% of rows)`, because a + batch was loaded twice. Any join on it silently fans out. The grain comes back + unknown rather than as one of the several column pairs that are technically unique + here only because `order_item_id` almost is; those are in `key_evidence` with the + reason each was suppressed. - **A key that mixes id schemes.** `sku` is `90% numeric, 10% 32-character hexadecimal (md5-shaped)`, from a merged catalogue. Cast it to a number and you drop 10% of your rows without an error. @@ -335,6 +340,9 @@ More info in the package's [`README.md`](packages/dex-core/README.md) command contract, the source of truth and the `.dex/` cache, the semantic layer and Ossie compatibility, the project, storage and host-integration seams, methodology, and evaluation. +- The two guardrails that cut across all of it: + [`references/pii-policy.md`](references/pii-policy.md) and + [`references/cost-controls.md`](references/cost-controls.md). ## Contributing diff --git a/packages/dex-core/README.md b/packages/dex-core/README.md index 2f6e8a7e..63c42a6a 100644 --- a/packages/dex-core/README.md +++ b/packages/dex-core/README.md @@ -74,7 +74,7 @@ uniqueness to a double-loaded batch, a key mixing two id schemes from a merged catalogue, a join whose columns share a name and none of their values, a table an interrupted load left empty, two columns whose declared type contradicts their content, and personal data alongside two deliberate false positives. `explore map` -finds 6 PII columns, 5 joins, and 5 data-quality findings; `explore query "select +finds 6 PII columns, 5 joins, and 6 data-quality findings; `explore query "select email from customers"` is refused, and the same count over the same column is not. The generation is create-only: it writes a new file and refuses rather than replace @@ -246,7 +246,8 @@ on its own path, never through a connector, which is what keeps the read-only ru true everywhere else. `explore`: ranks what matters in an unfamiliar warehouse, profiles columns -selectively, flags PII, surfaces grain and data-quality warnings, infers joins +selectively, flags PII, surfaces grain and data-quality warnings with the ranked +keys and the reasoning behind each, infers joins and verifies them with overlap probes (`--verify`), and executes agent-authored ad-hoc SELECTs behind a PII-aware query firewall (`explore query`, which takes several statements per call, or a `--sql-file`, and adjudicates each on its own), diff --git a/packages/dex-core/src/exmergo_dex_core/cache.py b/packages/dex-core/src/exmergo_dex_core/cache.py index 77b35b7c..62e45df7 100644 --- a/packages/dex-core/src/exmergo_dex_core/cache.py +++ b/packages/dex-core/src/exmergo_dex_core/cache.py @@ -15,11 +15,21 @@ import re from enum import Enum +from typing import Literal from pydantic import BaseModel, Field # Bump when the stored cache shape changes in a way old readers cannot handle. -CACHE_SCHEMA_VERSION = 3 +# +# 4 added `Dataset.key_evidence` and, with it, changed what `candidate_keys` +# means: it was every combination measured unique, and it is now every one that +# survived artifact suppression. An older reader handed a version-4 cache is +# fine, since an unknown key is ignored. The direction that breaks is a current +# reader handed a version-3 one, where a suppressed combination still reads as a +# ranked candidate and an empty `key_evidence` is indistinguishable from "this +# run suppressed nothing". So a pre-4 profile is treated as stale rather than +# reused (see `_split_fresh_stale`), and the cache heals on the next profile. +CACHE_SCHEMA_VERSION = 4 class PIICategory(str, Enum): @@ -70,6 +80,28 @@ class ValueDomain(BaseModel): elided: int = 0 +class KeyEvidence(BaseModel): + """Why one column combination is, or is not, reported as a key. + + ``candidate_keys`` is ranked but says nothing about why, and a caller who + cannot tell a real key from filler is worse served by several candidates + than by one named defect and none. This is where the reasoning lives. + + ``status`` is the only discriminator a consumer needs. ``reason`` is prose + rather than a code, because the set of causes is open and a new one should + not need a contract change and an exhaustive match in every consumer; the + sentences are engine-authored templates over column names and counts, never + a customer string and never a column value. + + The reported entries appear in the same order as ``candidate_keys``, which + is what keeps the two fields from drifting; the suppressed ones follow. + """ + + columns: list[str] + status: Literal["reported", "suppressed"] + reason: str + + class ColumnProfile(BaseModel): """Aggregate-derived understanding of one column, built from SQL aggregates and never from raw rows in context.""" @@ -124,6 +156,11 @@ class Dataset(BaseModel): #: annotation pass recomputes from column stats) so the proof survives #: re-annotation and its provenance stays distinct from derived signals. composite_keys: list[list[str]] = Field(default_factory=list) + #: Why each combination is or is not a key, reported entries first and in + #: ``candidate_keys`` order. Empty on a profile written before this existed; + #: the cache schema version is what tells those apart from a run that + #: suppressed nothing. + key_evidence: list[KeyEvidence] = Field(default_factory=list) rank_score: float | None = None data_quality: list[str] = Field(default_factory=list) profiled_at: str | None = None @@ -152,9 +189,19 @@ def notable_columns( reporting it is the enumeration dex exists not to do. Returns each kept column paired with its role (``"grain"``, ``"key"``, - ``"join"``, or ``None`` for a column kept only because it is flagged) and - the number dropped, so a caller can always say how much it did not show. - ``everything`` keeps every column and still assigns the roles. + ``"join"``, or ``None`` for a column kept only because it is flagged or + because it almost keys the table) and the number dropped, so a caller + can always say how much it did not show. ``everything`` keeps every + column and still assigns the roles. + + A column named by a **suppressed** ``key_evidence`` entry is kept, with + no role, because it is the subject of the grain verdict on exactly the + tables that have no grain: a map reporting "order_item_id is not unique" + while omitting ``order_item_id`` from the columns would be answering + past the question. It gets no role because it is not a key; claiming one + is the thing the suppression exists to stop. Only suppressed entries are + read, and each names a single anchor column rather than the partners it + was paired with, so this cannot pull filler columns in. ``join_columns`` is supplied rather than derived, because which joins are in view is the caller's question: a diagram marks FK against the edges it @@ -168,6 +215,12 @@ def notable_columns( keyed |= {c.lower() for group in self.composite_keys for c in group} grain = {c.lower() for c in (self.grain or [])} joins = {c.lower() for c in (join_columns or ())} + near_keys = { + c.lower() + for entry in self.key_evidence + if entry.status == "suppressed" + for c in entry.columns + } kept: list[tuple[ColumnProfile, str | None]] = [] dropped = 0 @@ -181,7 +234,12 @@ def notable_columns( role = "key" else: role = None - if role is None and column.pii is None and not everything: + if ( + role is None + and column.pii is None + and lowered not in near_keys + and not everything + ): dropped += 1 continue kept.append((column, role)) @@ -214,6 +272,14 @@ def columns_with_findings( predicate exists so a real finding is never the reason it gets truncated away, not to be a precise finding-to-column index. ``everything`` keeps every column. + + Needs no ``key_evidence`` check of its own: a suppressed entry's + anchor is a column with duplicates, so it is already named in one of + this dataset's ``data_quality`` sentences and kept by the last rule + below. That is also why the suppression prose names only the anchor and + never the partners it was paired with; naming those would pull four + columns back in through the same rule, which is exactly the payload + bloat the finding summary removed. """ keyed = {c.lower() for group in self.candidate_keys for c in group} diff --git a/packages/dex-core/src/exmergo_dex_core/explore/commands.py b/packages/dex-core/src/exmergo_dex_core/explore/commands.py index 09c20536..693e829d 100644 --- a/packages/dex-core/src/exmergo_dex_core/explore/commands.py +++ b/packages/dex-core/src/exmergo_dex_core/explore/commands.py @@ -39,6 +39,7 @@ # same-named row carrier is only a type hint on one shaping helper. from ..adapters.base import QueryResult as AdapterQueryResult from ..cache import ( + CACHE_SCHEMA_VERSION, ColumnProfile, Dataset, DexCache, @@ -3923,6 +3924,9 @@ def _annotate_grain( for ds in datasets: ds.candidate_keys = rel_mod.candidate_keys(ds) ds.grain = rel_mod.detect_grain(ds) + # After candidate_keys, because the reported half has to come back in + # exactly its order; the probe's suppressed entries ride through. + ds.key_evidence = rel_mod.key_evidence(ds) ds.data_quality.extend(rel_mod.data_quality_notes(ds)) if orphaned and ds.identifier in orphaned: ds.data_quality.append( @@ -3939,10 +3943,21 @@ def _annotate_grain( ) if profiled is not None: declared = [profiled.name] - if ds.grain and ds.grain != declared: + if ds.grain != declared: + # Says where the grain came from either way. "Measurement + # found none" is as much worth stating as a disagreement: + # it is the case where the declaration is carrying the + # answer on its own, and a reader who cannot tell that + # apart from a confirmed guess does not know how much the + # grain rests on. + origin = ( + f"heuristic suggested {', '.join(ds.grain)}" + if ds.grain + else "measurement found no key of its own" + ) ds.data_quality.append( f"grain {profiled.name} comes from the project's declared " - f"primary entity (heuristic suggested {', '.join(ds.grain)})" + f"primary entity ({origin})" ) ds.grain = declared for col in ds.columns: @@ -3999,11 +4014,15 @@ def _annotate_grain( f"table; using {', '.join(chosen)} (also declared: " f"{others})" ) - if ds.grain and ds.grain != chosen: + if ds.grain != chosen: + origin = ( + f"heuristic suggested {', '.join(ds.grain)}" + if ds.grain + else "measurement found no key of its own" + ) ds.data_quality.append( f"grain {', '.join(chosen)} comes from the project's " - "declared composite key (heuristic suggested " - f"{', '.join(ds.grain)})" + f"declared composite key ({origin})" ) ds.grain = chosen @@ -4193,7 +4212,13 @@ def _split_fresh_stale( mismatched-connector or absent prior — mirroring ``cmd_map``'s reuse gate. Freshness is fail-closed: a missing or unparseable ``profiled_at``, or any - doubt, re-profiles rather than trusting a stale scan. + doubt, re-profiles rather than trusting a stale scan. A cache written before + the current ``CACHE_SCHEMA_VERSION`` is stale for the same reason: an older + profile can carry verdicts this version would no longer reach (a composite + key later recognized as an artifact of a near-unique column, say), and + reusing one would keep a superseded answer alive and keep charging a metered + connector to re-check it. One scan per table per upgrade, and no caller has + to know to pass ``--refresh``. This is the gate for a *deliberate* profile, which is why the age window belongs in it. The on-demand path asks a narrower question (see @@ -4207,7 +4232,12 @@ def _split_fresh_stale( common, not on ``ObjectMeta``. """ - if refresh or prior is None or prior.provenance.connector != connector: + if ( + refresh + or prior is None + or prior.provenance.connector != connector + or prior.schema_version < CACHE_SCHEMA_VERSION + ): return list(identifiers), {} prior_by_id = {d.identifier: d for d in prior.datasets if d.columns} diff --git a/packages/dex-core/src/exmergo_dex_core/explore/profile.py b/packages/dex-core/src/exmergo_dex_core/explore/profile.py index 98bcf54a..fc0eda17 100644 --- a/packages/dex-core/src/exmergo_dex_core/explore/profile.py +++ b/packages/dex-core/src/exmergo_dex_core/explore/profile.py @@ -12,7 +12,7 @@ import re from collections.abc import Callable -from dataclasses import replace +from dataclasses import dataclass, replace from datetime import UTC, date, datetime from ..adapters.base import ( @@ -29,6 +29,7 @@ from ..cache import ( ColumnProfile, Dataset, + KeyEvidence, PIICategory, PIIFlag, ValueCount, @@ -68,6 +69,112 @@ # cap has room, because a discarded candidate cannot be the grain. _COMPOSITE_REDUNDANCY_RATIO = 3.0 +# A column whose EXACT distinct count reaches this fraction of the rows is an +# identifier that almost holds. The handful of rows it fails to separate get +# separated by accident by any partner carrying more than a value or two, so a +# pair anchored on such a column restates its near-uniqueness rather than +# discovering a grain: it is unique for a reason that has nothing to do with the +# table's identity. Deliberately NOT `NEAR_UNIQUE_RATIO`, and well clear of it. +# That one means "an approximation cannot tell whether this column is unique", +# which is a statement about HLL noise and says nothing about structure. This is +# a structural judgement, so it is only ever applied to a count that was proven, +# and a column still resting on an approximation is left alone. +_ANCHOR_NEAR_UNIQUE_RATIO = 0.90 + +# A key member that enumerates positions within a parent (a line number, a +# version, a fiscal period, a status, a snapshot date) has a domain bounded by +# the fanout rather than by the table; a continuous measurement has the opposite +# property, its distinct count grows with the row count. Below this fraction of +# the rows a column is the former whatever its type or name suggests, which is +# what keeps a legitimate numeric grain member out of the measure exclusion: a +# Snowflake `FIXED` line number, a `NUMERIC(6,0)` fiscal period, a `DECIMAL` +# latitude snapped to a grid. It is also what lets a near-unique anchor keep the +# one partner that genuinely refines it. +# +# Deliberately not `VALUE_DOMAIN_MAX_FRACTION`, whose value this happens to +# match: that one decides whether values are worth printing in a payload and is +# tuned to a payload budget. Coupling them would let a payload change silently +# move grain detection. +_KEY_MEMBER_ENUM_FRACTION = 0.10 + +# Below this row count, cardinality-versus-rows reasoning says nothing: on a +# 20-row table every column looks near-unique and every measure looks +# continuous, so the two exclusions below would suppress grains rather than +# junk. The probes are nearly free at this size too, so the honest default is to +# ask everything and let the exact combination count decide. A fixed floor +# rather than a config knob, because it changes what gets measured rather than +# how loudly a finding reads. +_COMPOSITE_MIN_ROWS = 100 + +# Declared types that are unambiguously fractional, so a column carrying one is +# a measurement and never an identifier, whatever its cardinality. +# +# The scale pattern is the load-bearing half, and it is narrow on purpose. A +# blanket "numeric but not integer" test would be wrong on a whole connector: +# `adapters.base.is_integer_type` documents that Snowflake's `SHOW COLUMNS` +# renders `NUMBER(38,0)` and `NUMBER(10,2)` identically as `FIXED`, so on +# Snowflake every non-id numeric column would read as fractional and a real +# `LINE_NUMBER` grain member would be excluded. Requiring an explicit nonzero +# scale covers DuckDB `DECIMAL(10,2)`, Postgres and Redshift `numeric(10,2)` and +# Databricks `decimal(10,2)`, and goes silent on BigQuery `NUMERIC` and +# Snowflake `FIXED`, which carry no scale at all. Going silent is the house's +# under-report posture, and it is what `_MEASURE_TOKENS` exists to cover. +_FLOAT_HINTS = ("DOUBLE", "FLOAT", "REAL") +_NONZERO_SCALE = re.compile(r"\(\s*\d+\s*,\s*[1-9]\d*\s*\)") + +# Measure tokens, matched as whole tokens on the snake-normalized name (so +# `order_amount`, `amountUsd` and `grand_total` hit and `am` does not). Two +# jobs: they catch a measure declared as an integer (money in minor units, a +# duration in milliseconds, a byte count), and they are the only signal left on +# a connector whose type strings cannot separate `NUMBER(38,0)` from +# `NUMBER(10,2)`, where the fractional-type test above goes silent by design. +# +# A token never decides alone. It only matters above `_KEY_MEMBER_ENUM_FRACTION`, +# because a column with a domain that small is a bounded enumeration whatever it +# is called, and several of these names sit on legitimate low-cardinality key +# members. +# +# Deliberately excluded, and do not "complete" this list: +# `qty` / `quantity`: an integer quantity is a bounded count of units, not a +# continuous measurement, and a fractional one (a weight, a rate) is already +# caught by the type test. Both spellings are load-bearing key members in +# this probe's own fixtures. +# `rate` / `score` / `weight`: fractional in practice, so the type test has +# them; as integers they read as ordinals. +# `value` / `number` / `num`: too generic, and frequently identifiers +# (`account_number`). +# `version` / `seq` / `period` / `line` / `rank` / `position` / `index` / +# `year` / `month` / `day`: ordinals and dimensions that legitimately +# complete a grain, which is the failure mode this rule must not cause. +_MEASURE_TOKENS = ( + "amount", + "amt", + "total", + "subtotal", + "revenue", + "price", + "cost", + "fee", + "tax", + "discount", + "balance", + "spend", + "margin", + "profit", + "duration", + "elapsed", + "latency", + "bytes", +) +_MEASURE_TOKEN_PATTERN = re.compile(r"(^|_)(" + "|".join(_MEASURE_TOKENS) + r")(_|$)") + +# The four parts a column can play in a composite key, returned by +# :func:`key_member_verdict`. +KEY_MEMBER_ELIGIBLE = "eligible" +KEY_MEMBER_ENUMERATION = "bounded_enumeration" +KEY_MEMBER_MEASURE = "continuous_measure" +KEY_MEMBER_ANCHOR = "near_unique_anchor" + # Name patterns mapped to a PII category and a base confidence. Matched on the # snake-normalized column name (camelCase is split first, so "firstName" matches # the same as "first_name") with word-ish boundaries so "email" hits but @@ -236,6 +343,77 @@ def is_numeric_type(data_type: str) -> bool: return any(h in upper for h in _NUMERIC_HINTS) +def key_member_verdict( + name: str, + data_type: str, + distinct_count: int | None, + *, + distinct_count_exact: bool, + row_count: int | None, +) -> str: + """What part a column can play in a composite key: ``KEY_MEMBER_ELIGIBLE``, + ``KEY_MEMBER_ENUMERATION``, ``KEY_MEMBER_MEASURE`` or ``KEY_MEMBER_ANCHOR``. + + Composite-key candidates are drawn from this, and so is the sentence a + profile writes when the strongest thing it found was a column that almost + keys the table. Both read the same verdict from the same constants, so the + prose can never disagree with the pruning. + + **The order of the tests is the substance**, so it is stated rather than + implied: + + 1. A column at or below ``_KEY_MEMBER_ENUM_FRACTION`` of the rows is a + bounded enumeration and nothing else. It is checked first because it is + the escape hatch: a fractional latitude, a Snowflake ``FIXED`` line + number and a low-cardinality snapshot date are all legitimate key + members, and a domain that small cannot be continuous whatever the + declared type or the name says. + 2. A measure-shaped column above that bar is a continuous measure. Checked + before the anchor test deliberately: a money column at 93% distinct is a + measurement, not an identifier that almost holds, and calling it an + anchor would put it in front of a reader as the closest thing to a key. + 3. A column whose **exact** distinct count reaches + ``_ANCHOR_NEAR_UNIQUE_RATIO`` of the rows is a near-unique anchor. Exact + only: a key claim is withheld on the strength of this verdict, and an + approximation inside the HLL band cannot carry that weight. A column + left approximate by the escalation cap, or by an adapter with no + ``exact_distinct_counts``, is simply eligible and behaves as it always + did. + + An id-shaped name never reads as a measure, and below + ``_COMPOSITE_MIN_ROWS`` rows every column is eligible, because none of the + three tests means anything at that size. + """ + + if not row_count or row_count < _COMPOSITE_MIN_ROWS or distinct_count is None: + return KEY_MEMBER_ELIGIBLE + + # Imported here rather than at module scope: relationships imports + # NEAR_UNIQUE_RATIO from this module, so a top-level import would cycle. + from .relationships import is_id_shaped + + if distinct_count <= row_count * _KEY_MEMBER_ENUM_FRACTION: + return KEY_MEMBER_ENUMERATION + + if not is_id_shaped(name) and is_numeric_type(data_type): + normalized = _normalize(name) + upper = data_type.upper() + fractional = any(h in upper for h in _FLOAT_HINTS) or bool( + _NONZERO_SCALE.search(data_type) + ) + if ( + fractional + or _MEASURE_TOKEN_PATTERN.search(normalized) is not None + or normalized.endswith(_AGGREGATE_SUFFIXES) + ): + return KEY_MEMBER_MEASURE + + if distinct_count_exact and distinct_count >= row_count * _ANCHOR_NEAR_UNIQUE_RATIO: + return KEY_MEMBER_ANCHOR + + return KEY_MEMBER_ELIGIBLE + + # Per-category structural type gate. Type evidence is known before any scan, so # an impossible pairing (an EMAIL on an integer) suppresses the flag outright at # classification time rather than being re-scored later. The gate is @@ -811,9 +989,14 @@ def profile( aggregates = _escalate_near_unique( adapter, identifier, meta.row_count, aggregates ) - composite_keys = _probe_composite_keys( - adapter, identifier, meta.row_count, aggregates + composite_probe = _probe_composite_keys( + adapter, + identifier, + meta.row_count, + aggregates, + {c.name: c.data_type for c in columns}, ) + composite_keys = composite_probe.keys value_domains = _probe_value_domains( adapter, identifier, @@ -901,6 +1084,7 @@ def profile( byte_size=meta.byte_size, columns=profiles, composite_keys=composite_keys, + key_evidence=composite_probe.evidence, data_quality=data_quality, profiled_at=datetime.now(UTC).isoformat(), ) @@ -967,15 +1151,173 @@ def _escalate_near_unique( return escalated +def uniqueness_shortfall( + distinct_count: int | None, + null_fraction: float | None, + row_count: int | None, +) -> tuple[int, float, bool] | None: + """How far a column is from being unique: ``(surplus_rows, fraction, + exact)``, or ``None`` when the numbers cannot support the arithmetic. + + ``surplus_rows`` is how many rows would have to be removed for the column + to be unique, which is the non-null row count less the distinct count. That + is deliberately **not** "how many values repeat": those differ, and only + the first is derivable from counts already in hand. One id appearing 111 + times is 110 surplus rows and exactly one repeated value, so a sentence + claiming 110 duplicate ids would be wrong. The removal figure is exactly + true in every case, needs no extra scan, and is the form that tells a + caller what to fix. + + ``exact`` is true only when the distinct count was proven **and** the + column has no nulls. With nulls the non-null count is derived from a + fraction, so everything downstream of it is approximate even though the + distinct count is not, and the ``~`` marker has to follow the derivation + rather than the distinct count alone. + """ + + if not row_count or distinct_count is None: + return None + exact = bool(null_fraction in (0.0, None)) + non_null = ( + row_count + if null_fraction in (0.0, None) + else round((1 - null_fraction) * row_count) + ) + if non_null <= 0: + return None + return non_null - distinct_count, distinct_count / non_null, exact + + +def format_uniqueness_fraction(fraction: float) -> str: + """A uniqueness ratio as a percentage, one decimal place. + + Never prints ``100.0%``: a column with two duplicates in a million rows + rounds there, and "unique for 100.0% of rows" in a sentence that then + reports duplicates contradicts itself. Such a column reads ``>99.9%``. + """ + + percent = fraction * 100 + if percent >= 99.95: + return ">99.9%" + return f"{percent:.1f}%" + + +def _near_unique_clause(name: str, agg: ColumnAggregate, row_count: int) -> str: + """``order_id is already unique for 94.6% of rows``, the phrase both the + suppression reason and the suppression note are built from.""" + + shortfall = uniqueness_shortfall(agg.distinct_count, agg.null_fraction, row_count) + if shortfall is None: + return f"{name} is already near-unique" + _surplus, fraction, exact = shortfall + marker = "" if (exact and agg.distinct_count_exact) else "~" + return ( + f"{name} is already unique for {marker}" + f"{format_uniqueness_fraction(fraction)} of rows" + ) + + +def _proven_composite_reason( + key: tuple[str, ...] | list[str], keys: list[list[str]], row_count: int +) -> str: + """Why a proven combination is reported, and where it sits in the ranking.""" + + columns = ", ".join(key) + reason = ( + f"{columns} is unique as a combination on all {row_count} rows, proven " + "by an exact distinct-combination probe" + ) + if list(key) == keys[0]: + return f"{reason}; it is the smallest proven combination, so it is the grain" + best = ", ".join(keys[0]) + return ( + f"{reason}, and ranks behind {best}, which is proven at a smaller cardinality" + ) + + +def _suppressed_evidence( + anchors: set[str], aggregates: dict[str, ColumnAggregate], row_count: int +) -> list[KeyEvidence]: + """One entry per near-unique column whose pairs were dropped, so a + suppression is never silent. Keyed on the anchor rather than on each + dropped pair: the pairs are interchangeable filler, the anchor is the + finding, and naming every partner would put four column names into the + payload that nothing else needs (see ``Dataset.columns_with_findings``).""" + + entries = [] + for name in sorted(anchors): + agg = aggregates.get(name) + if agg is None: + continue + entries.append( + KeyEvidence( + columns=[name], + status="suppressed", + reason=( + f"combinations pairing {name} with another column were not " + f"probed because {_near_unique_clause(name, agg, row_count)}, " + "so any partner completes it by accident rather than keying " + "the table; the duplicates in " + f"{name} are the finding, not a composite key" + ), + ) + ) + return entries + + +@dataclass(frozen=True) +class CompositeKeyProbe: + """What the composite-key probe proved, and the reasoning it owes a reader. + + ``keys`` is what :class:`~..cache.Dataset` persists as ``composite_keys``, + best candidate first. ``evidence`` is one :class:`~..cache.KeyEvidence` per + combination the probe considered, reported or suppressed. + + Both come back from the probe rather than being recomputed later because + only the probe knows what it excluded and what it measured: a caller + re-deriving the suppression would have to reimplement the pool rules, and + the two would drift. The prose a suppression owes a reader is *not* here, + though. It belongs beside the "grain unknown" verdict it explains, and + :func:`~.relationships.data_quality_notes` builds it there from these + entries, so one sentence carries both facts instead of two carrying one + each. + """ + + keys: list[list[str]] + evidence: list[KeyEvidence] + + def _probe_composite_keys( adapter: Adapter, identifier: str, row_count: int | None, aggregates: dict[str, ColumnAggregate], -) -> list[list[str]]: + data_types: dict[str, str], +) -> CompositeKeyProbe: """Prove 2-column keys on tables where no single column is one: the shape of a fact table, whose grain is exactly what a profile must answer. + **Unique is necessary for a key and not sufficient to be one**, and two + shapes of pair are unique for reasons that have nothing to do with the + table's identity. Both are excluded from the pool before anything is + priced, because the exact answer could not move a decision either way and + a pair reported and then disclaimed is worse than a pair never offered. + Pruning early also frees cap slots for pairs that might be the real grain, + which is the same argument the redundancy rule below makes for keeping a + demoted candidate in the running. + + - A member that is a continuous measure (``KEY_MEMBER_MEASURE``) leaves the + pool outright. A key made of a measurement is almost never the intended + grain, and its cardinality grows with the table, so it completes a + partner by arithmetic rather than by meaning. + - A pair **anchored** on a near-unique column (``KEY_MEMBER_ANCHOR``) is + dropped unless its partner is a bounded enumeration. When a handful of + duplicate order ids are separated by a three-value line number, the data + really does have that grain; any wider partner separates them by + accident, and the pair only restates the anchor's near-uniqueness. That + near-uniqueness is what the profile reports instead, with the counts, and + it names a defect in the source rather than a key. + A pair can only be a key if the product of its members' distinct counts reaches the row count, so pairs are pruned on that necessary condition (relaxed per approximate member to absorb HLL undershoot) and ranked: @@ -1005,30 +1347,65 @@ def _probe_composite_keys( pairs it dropped simply unproven. """ + empty = CompositeKeyProbe(keys=[], evidence=[]) if not row_count: - return [] + return empty combo_counts = getattr(adapter, "distinct_combination_counts", None) if combo_counts is None: - return [] + return empty for agg in aggregates.values(): if agg.is_unique and agg.null_fraction in (0.0, None): - return [] # a proven single-column key makes the probe waste - + return empty # a proven single-column key makes the probe waste + + verdicts = { + agg.name: key_member_verdict( + agg.name, + data_types.get(agg.name, ""), + agg.distinct_count, + distinct_count_exact=agg.distinct_count_exact, + row_count=row_count, + ) + for agg in aggregates.values() + } pool = [ agg for agg in aggregates.values() - if agg.distinct_count and not agg.is_unique and agg.null_fraction in (0.0, None) + if agg.distinct_count + and not agg.is_unique + and agg.null_fraction in (0.0, None) + and verdicts[agg.name] != KEY_MEMBER_MEASURE ] if len(pool) < 2: - return [] + return empty # Imported here: relationships imports NEAR_UNIQUE_RATIO from this module, # so a module-level import would be circular. from .relationships import is_id_shaped ranked: list[tuple[int, int, tuple[str, str]]] = [] + suppressed_anchors: set[str] = set() for i, a in enumerate(pool): for b in pool[i + 1 :]: + anchored = [m for m in (a, b) if verdicts[m.name] == KEY_MEMBER_ANCHOR] + if anchored and not any( + verdicts[m.name] == KEY_MEMBER_ENUMERATION + # The enumeration escape holds only for an anchor that could + # plausibly be a parent identifier. A near-unique *timestamp* + # is a per-row event time, not an entity whose rows a position + # column enumerates, so pairing it with a low-cardinality + # dimension separates rows by accident exactly the way a wider + # partner would: `(status, created_at)` is unique on a table + # whose created_at is 98% unique, and it is not a grain. A + # near-unique identifier paired with a line number is. + and not any( + is_temporal_type(data_types.get(m.name, "")) for m in anchored + ) + for m in (a, b) + ): + # Unique only because one member almost is: the pair restates + # that, and the near-uniqueness note says it properly. + suppressed_anchors.update(m.name for m in anchored) + continue product = a.distinct_count * b.distinct_count n_approx = sum(1 for m in (a, b) if not m.distinct_count_exact) if product < row_count * NEAR_UNIQUE_RATIO**n_approx: @@ -1042,7 +1419,10 @@ def _probe_composite_keys( ) ranked.append((-id_shaped, product, (members[0], members[1]))) if not ranked: - return [] + return CompositeKeyProbe( + keys=[], + evidence=_suppressed_evidence(suppressed_anchors, aggregates, row_count), + ) ranked.sort() preferred: list[tuple[int, int, tuple[str, str]]] = [] @@ -1068,7 +1448,17 @@ def _probe_composite_keys( for _ids, product, pair in selected if exact.get(pair) == row_count ) - return [list(pair) for _product, pair in proven] + keys = [list(pair) for _product, pair in proven] + evidence = [ + KeyEvidence( + columns=list(key), + status="reported", + reason=_proven_composite_reason(key, keys, row_count), + ) + for key in keys + ] + evidence.extend(_suppressed_evidence(suppressed_anchors, aggregates, row_count)) + return CompositeKeyProbe(keys=keys, evidence=evidence) def _probe_value_domains( diff --git a/packages/dex-core/src/exmergo_dex_core/explore/relationships.py b/packages/dex-core/src/exmergo_dex_core/explore/relationships.py index e6cba845..0f9c9d99 100644 --- a/packages/dex-core/src/exmergo_dex_core/explore/relationships.py +++ b/packages/dex-core/src/exmergo_dex_core/explore/relationships.py @@ -22,10 +22,11 @@ import re from typing import NamedTuple -from ..adapters.base import Adapter +from ..adapters.base import Adapter, is_temporal_type from ..cache import ( ColumnProfile, Dataset, + KeyEvidence, Relationship, RelationshipKind, match_identifier, @@ -34,7 +35,13 @@ from ..dbt_project import ProjectDefinitions from ..progress import ProgressReporter from ..semantic_catalog import EntityJoin -from .profile import NEAR_UNIQUE_RATIO +from .profile import ( + KEY_MEMBER_ANCHOR, + NEAR_UNIQUE_RATIO, + format_uniqueness_fraction, + key_member_verdict, + uniqueness_shortfall, +) # Warehouse-layer prefixes stripped from a table name before entity matching, so # RAW_HOSTS, stg_races, and dim_customers all match FKs named after the bare entity. @@ -427,6 +434,7 @@ def data_quality_notes(dataset: Dataset) -> list[str]: notes: list[str] = [] if not dataset.row_count: return notes + counted: set[str] = set() entity = entity_of(dataset.identifier.rsplit(".", 1)[-1]) for col in dataset.columns: @@ -445,20 +453,165 @@ def data_quality_notes(dataset: Dataset) -> list[str]: # noise (an approx 500 distinct over 1,125 rows) still warns. continue if col.distinct_count < dataset.row_count: - duplicates = dataset.row_count - col.distinct_count - # An unescalated count is honest about being approximate. - marker = "" if col.distinct_count_exact else "~" - notes.append( - f"{col.name} is not unique: {marker}{col.distinct_count} distinct " - f"over {dataset.row_count} rows (~{duplicates} duplicate rows); " - "joins on it will fan out" - ) + notes.append(_not_unique_note(col, dataset.row_count)) + counted.add(col.name) if not candidate_keys(dataset): - notes.append("no candidate key detected; grain unknown") + notes.append(_grain_unknown_note(dataset, counted)) return notes +def _not_unique_note(col: ColumnProfile, row_count: int) -> str: + """How far a column is from keying its table, in the terms a caller acts on. + + Three numbers: the distinct count, the row count, and how many rows would + have to be removed for the column to be unique. The last is what a caller + needs in order to decide, and deriving it from an approximate distinct + count was the thing they previously had to do by hand. + + The `~` marker follows the arithmetic rather than the distinct count alone. + A surplus derived from two exact numbers is itself exact, so the marker + comes off there; with nulls the non-null count is derived from a fraction, + so it goes back on even though the distinct count is proven. + """ + + shortfall = uniqueness_shortfall(col.distinct_count, col.null_fraction, row_count) + distinct_marker = "" if col.distinct_count_exact else "~" + if shortfall is None: + return ( + f"{col.name} is not unique: {distinct_marker}{col.distinct_count} " + f"distinct over {row_count} rows; joins on it will fan out" + ) + surplus, fraction, exact = shortfall + marker = "" if (exact and col.distinct_count_exact) else "~" + return ( + f"{col.name} is not unique: {distinct_marker}{col.distinct_count} distinct " + f"over {row_count} rows ({marker}{surplus} rows would have to be removed " + f"for it to be unique, so it is unique for " + f"{marker}{format_uniqueness_fraction(fraction)} of rows); joins on it " + "will fan out" + ) + + +def _grain_unknown_note(dataset: Dataset, counted: set[str]) -> str: + """ "Grain unknown", and where the reader should look instead. + + A bare "no candidate key detected" leaves the most useful fact unsaid: on + the shape this exists for, one column very nearly keys the table and the + real defect is the duplicates in it. Naming that column turns the note from + an absence into an instruction. + + Bounded to one column, the highest-cardinality anchor, so a wide table does + not get five of these. Temporal anchors are skipped as uninteresting to + report even though they still prune pairs: a per-row timestamp being + near-unique is not news, and it is never the key anyone meant. + + ``counted`` is the columns that already have a non-uniqueness note above, + carrying the distinct count, the row count and the surplus. For one of + those this note points at them rather than restating them: several notes + repeating one column's arithmetic reads as padding and spends a budget + ``explore map`` caps per object. + + Where the profile-time probe suppressed combinations, this sentence says so + too rather than leaving that to a note of its own. The two always co-occur + (a suppressed-everything probe is exactly a probe that proved no composite, + and the probe only runs when no single column is a key), so they are one + finding and belong in one sentence. + """ + + bare = "no candidate key detected; grain unknown" + suppressed = [e for e in dataset.key_evidence if e.status == "suppressed"] + probe = ( + "; the composite-key probe found combinations that were unique and " + "suppressed every one of them as an artifact of that, and key_evidence " + "carries each with its reason" + if suppressed + else "" + ) + if not dataset.row_count: + return bare + anchors = [ + col + for col in dataset.columns + if col.pii is None + and col.null_fraction in (0.0, None) + and not is_temporal_type(col.data_type) + and key_member_verdict( + col.name, + col.data_type, + col.distinct_count, + distinct_count_exact=col.distinct_count_exact, + row_count=dataset.row_count, + ) + == KEY_MEMBER_ANCHOR + ] + if not anchors: + return bare + probe + best = max(anchors, key=lambda c: c.distinct_count or 0) + shortfall = uniqueness_shortfall( + best.distinct_count, best.null_fraction, dataset.row_count + ) + if shortfall is None: + return bare + probe + surplus, fraction, _exact = shortfall + if best.name in counted: + return ( + f"{bare}: {best.name} is the closest thing to one, and the " + f"duplicates in it noted above are the reason{probe}" + ) + return ( + f"{bare}: {best.name} is the closest thing to one at " + f"{format_uniqueness_fraction(fraction)} unique ({best.distinct_count} " + f"distinct over {dataset.row_count} rows, {surplus} rows would have to " + "be removed), and a combination pairing it with any wider column would " + f"prove unique without describing the grain{probe}" + ) + + +def key_evidence(dataset: Dataset) -> list[KeyEvidence]: + """Why each of this dataset's keys is or is not reported, ranked. + + The reported entries are in exactly ``candidate_keys`` order, which is the + invariant that keeps the two fields from drifting; the suppressed entries + the profile-time probe recorded follow. Composite reasons come from the + probe (only it measured the combination); single-column reasons are derived + here from the same column statistics ``candidate_keys`` reads. + """ + + from_probe = { + tuple(entry.columns): entry + for entry in dataset.key_evidence + if entry.status == "reported" + } + suppressed = [e for e in dataset.key_evidence if e.status == "suppressed"] + by_name = {col.name: col for col in dataset.columns} + + reported: list[KeyEvidence] = [] + for key in candidate_keys(dataset): + existing = from_probe.get(tuple(key)) + if existing is not None: + reported.append(existing) + continue + col = by_name.get(key[0]) + if col is None or len(key) != 1: + continue + rows = dataset.row_count + proof = ( + "proven by an exact distinct count" + if col.distinct_count_exact + else "but its distinct count is still approximate, so this is a " + "signal rather than a proof" + ) + reported.append( + KeyEvidence( + columns=list(key), + status="reported", + reason=f"{col.name} is unique and non-null on all {rows} rows, {proof}", + ) + ) + return reported + suppressed + + # Below this, a high orphan rate is still just weaker evidence for the # inferred join (verify_relationships already demotes confidence starting at # 0.2). At or above it, the two columns are effectively disjoint: a shared diff --git a/packages/dex-core/src/exmergo_dex_core/explore/results.py b/packages/dex-core/src/exmergo_dex_core/explore/results.py index 1b39a77d..7b1e93ad 100644 --- a/packages/dex-core/src/exmergo_dex_core/explore/results.py +++ b/packages/dex-core/src/exmergo_dex_core/explore/results.py @@ -120,6 +120,9 @@ def _profile_dataset_payload( "byte_size": dataset.byte_size, "candidate_keys": dataset.candidate_keys, "grain": dataset.grain, + # Beside the two fields it explains, because a reason separated from + # them is meaningless and has to survive the same truncation. + "key_evidence": [e.model_dump(mode="json") for e in dataset.key_evidence], "composite_keys": dataset.composite_keys, "rank_score": dataset.rank_score, "data_quality": dataset.data_quality, diff --git a/packages/dex-core/tests/demo/test_commands.py b/packages/dex-core/tests/demo/test_commands.py index 9c59ff02..4286941a 100644 --- a/packages/dex-core/tests/demo/test_commands.py +++ b/packages/dex-core/tests/demo/test_commands.py @@ -185,6 +185,23 @@ def test_the_generated_warehouse_drives_the_whole_explore_tour(tmp_path: Path, c assert any("order_item_id is not unique" in n for n in notes) assert any("mixes value shapes" in n for n in notes) + # The shipped-artifact regression for the double-loaded batch: this table's + # story is 1,000 duplicate order_item_ids, and it used to be told as a + # composite grain of (unit_price, order_id), a money column paired with a + # foreign key. Nothing is reported as a key now, and the duplicates are. + items = next( + d + for d in profiled["data"]["datasets"] + if d["identifier"].endswith("order_items") + ) + assert items["grain"] is None + assert items["candidate_keys"] == [] + assert not any("unit_price" in e["reason"] for e in items["key_evidence"]) + assert any( + "1000 rows would have to be removed for it to be unique" in n + for n in items["data_quality"] + ) + verified = _run(["explore", "relationships", "--verify"], capsys)["data"] edges = { (r["from_dataset"].split(".")[-1], r["from_columns"][0]): r diff --git a/packages/dex-core/tests/explore/conftest.py b/packages/dex-core/tests/explore/conftest.py index 9bfeee6e..532dd1ed 100644 --- a/packages/dex-core/tests/explore/conftest.py +++ b/packages/dex-core/tests/explore/conftest.py @@ -281,6 +281,45 @@ def composite_grain_duckdb(tmp_path: Path) -> Path: return path +@pytest.fixture +def near_unique_key_duckdb(tmp_path: Path) -> Path: + """The field failure of issue #292: an orders table whose own key is unique + on all but 110 of 2,037 rows, so every high-cardinality column completes it. + + Before the artifact rules, this table reported five candidate keys and + elected a grain out of them: `(order_id, customer_id)`, and `created_at`, + `subtotal`, `grand_total` and `updated_at` each paired with `order_id`. All + five are genuinely unique over these rows, so the fixture is a real + regression test rather than a mock of one: rows 1928..2037 repeat ids + 1..110, and the twin rows differ in every one of those five columns. + + No column is unique on its own, so the composite probe really runs. + `status` is functionally determined by `order_id` on purpose: it is a + bounded enumeration, so it is the one partner still asked, and it honestly + fails to prove. + """ + + duckdb = pytest.importorskip("duckdb") + path = tmp_path / "near_unique_key.duckdb" + conn = duckdb.connect(str(path)) + conn.execute( + "CREATE TABLE orders AS SELECT " + " (CASE WHEN r <= 1927 THEN r ELSE r - 1927 END)::BIGINT AS order_id, " + " ((r % 640) + 1)::INTEGER AS customer_id, " + " (['new', 'paid', 'shipped', 'closed'])" + " [(CASE WHEN r <= 1927 THEN r ELSE r - 1927 END) % 4 + 1] AS status, " + " (((r % 1900) * 7 + 13) / 100.0)::DECIMAL(12,2) AS subtotal, " + " (((r % 1850) * 11 + 29) / 100.0)::DECIMAL(12,2) AS grand_total, " + " (TIMESTAMP '2026-01-01 00:00:00' + INTERVAL (r % 2000) SECOND) " + " AS created_at, " + " (TIMESTAMP '2026-01-01 00:00:00' + INTERVAL (r % 1950) MINUTE) " + " AS updated_at " + "FROM (SELECT UNNEST(range(1, 2038)) AS r)" + ) + conn.close() + return path + + @pytest.fixture def parent_line_grain_duckdb(tmp_path: Path) -> Path: """The fact-table shape whose grain the pair probe used to discard: a diff --git a/packages/dex-core/tests/explore/test_diagram.py b/packages/dex-core/tests/explore/test_diagram.py index 7ad1b428..371e1ada 100644 --- a/packages/dex-core/tests/explore/test_diagram.py +++ b/packages/dex-core/tests/explore/test_diagram.py @@ -740,3 +740,42 @@ def test_the_legend_names_both_sources_of_a_solid_line(): legend = next(line for line in mermaid.splitlines() if "solid lines" in line) assert "relationships test" in legend and "semantic-layer entity" in legend + + +def test_a_near_unique_anchor_is_drawn_without_claiming_a_key(): + """A table whose only "keys" were artifacts of a near-unique column has no + key to draw, but the column carrying that finding must still appear: a + diagram that omitted it would answer past the question. It is drawn with no + key mark, because it is not a key, which is the whole point.""" + + from exmergo_dex_core.cache import ColumnProfile, Dataset, DexCache, KeyEvidence + from exmergo_dex_core.explore.diagram import render_er_mermaid + + items = Dataset( + identifier="shop.main.order_items", + row_count=14000, + columns=[ + ColumnProfile( + name="order_item_id", + data_type="BIGINT", + distinct_count=13000, + distinct_count_exact=True, + is_unique=False, + null_fraction=0.0, + ), + ColumnProfile(name="filler", data_type="VARCHAR"), + ], + key_evidence=[ + KeyEvidence( + columns=["order_item_id"], + status="suppressed", + reason="order_item_id is already unique for 92.9% of rows", + ) + ], + ) + mermaid = render_er_mermaid(DexCache(datasets=[items]), full=True).mermaid + + assert "order_item_id" in mermaid + line = next(row for row in mermaid.splitlines() if "order_item_id" in row) + for mark in (" PK", " UK", " FK"): + assert mark not in line, line diff --git a/packages/dex-core/tests/explore/test_explore.py b/packages/dex-core/tests/explore/test_explore.py index 4e4d0818..9ddbc976 100644 --- a/packages/dex-core/tests/explore/test_explore.py +++ b/packages/dex-core/tests/explore/test_explore.py @@ -238,8 +238,54 @@ def test_profile_leads_with_the_verdict_not_columns(duckdb_file: Path, capsys): ) keys = list(payload["data"]["datasets"][0].keys()) assert keys.index("columns") == len(keys) - 2 # elided_column_count trails it - for verdict_field in ("grain", "candidate_keys", "data_quality", "row_count"): + for verdict_field in ( + "grain", + "candidate_keys", + "key_evidence", + "data_quality", + "row_count", + ): assert keys.index(verdict_field) < keys.index("columns") + # `key_evidence` explains `grain` and `candidate_keys`, and a reason + # separated from them says nothing, so it sits directly beside them rather + # than merely somewhere ahead of `columns`. + assert keys.index("key_evidence") == keys.index("grain") + 1 + + +def test_key_evidence_reported_half_matches_candidate_keys_in_order( + duckdb_file: Path, capsys +): + """The invariant that keeps the two fields from drifting. A caller reads the + ranking off `candidate_keys` and the reasoning off `key_evidence`, so the + reported entries have to be the same list in the same order.""" + + payload = _run( + ["explore", "profile", "customers,orders", "--path", str(duckdb_file)], capsys + ) + for dataset in payload["data"]["datasets"]: + reported = [ + e["columns"] for e in dataset["key_evidence"] if e["status"] == "reported" + ] + assert reported == dataset["candidate_keys"], dataset["identifier"] + for entry in dataset["key_evidence"]: + assert entry["columns"], entry + assert entry["status"] in {"reported", "suppressed"} + assert entry["reason"] + + +def test_key_evidence_is_a_profile_field_and_not_a_map_field(duckdb_file: Path, capsys): + """`explore map` is budgeted at a handful of findings per object across many + objects; the full key ranking belongs to the command whose subject is one + relation in full. The same split already puts summarized columns in `map` + and every column in `profile`.""" + + from exmergo_dex_core.explore.summary import MapObject + + assert "key_evidence" not in MapObject.model_fields + + payload = _run(["explore", "map", "--path", str(duckdb_file)], capsys) + for obj in payload["data"]["objects"]: + assert "key_evidence" not in obj def test_profile_columns_default_summarizes_to_findings( diff --git a/packages/dex-core/tests/explore/test_profile.py b/packages/dex-core/tests/explore/test_profile.py index 280c9090..da0a2f3c 100644 --- a/packages/dex-core/tests/explore/test_profile.py +++ b/packages/dex-core/tests/explore/test_profile.py @@ -1634,6 +1634,33 @@ def test_profile_reuses_fresh_cached_profile( assert refreshed["data"]["cache_hit_count"] == 0 +def test_a_cache_written_before_the_current_schema_version_is_not_reused( + airbnb_duckdb: Path, tmp_path: Path, capsys +): + """A profile written by an older cache schema can carry verdicts this + version would no longer reach (a composite key later recognized as an + artifact of a near-unique column). Reusing one would keep a superseded + answer alive and keep a metered connector paying to re-check it, so the + bump heals itself on the next profile instead of needing `--refresh`.""" + + from exmergo_dex_core.cache import CACHE_SCHEMA_VERSION + + repo = tmp_path / "repo" + repo.mkdir() + _map(airbnb_duckdb, repo, capsys) + store = FilesystemStore(repo) + + stale = store.load_cache() + stale.schema_version = CACHE_SCHEMA_VERSION - 1 + store.save_cache(stale) + + payload = _profile(["RAW_HOSTS"], airbnb_duckdb, repo, capsys) + assert payload["data"]["profiled_count"] == 1 + assert payload["data"]["cache_hit_count"] == 0 + # And the re-scan writes the current version back, so it happens once. + assert store.load_cache().schema_version == CACHE_SCHEMA_VERSION + + def test_profile_freshness_zero_disables_reuse( airbnb_duckdb: Path, tmp_path: Path, capsys ): @@ -1840,7 +1867,9 @@ class _StubAdapter: """Metadata-only double: crafted approximate aggregates, recorded escalations and composite probes. ``combos`` maps a column tuple to its exact distinct combination count; unlisted tuples come back just below the row count, so - they read as probed-but-not-unique.""" + they read as probed-but-not-unique. ``types`` overrides the declared type of + named columns (everything else is INTEGER), which is what the measure and + ordinal rules read.""" name = "stub" dialect = "duckdb" @@ -1852,12 +1881,14 @@ def __init__( nulls: dict[str, float] | None = None, combos: dict[tuple[str, ...], int] | None = None, domains: dict[str, object] | None = None, + types: dict[str, str] | None = None, ): self.rows = rows self.approx = approx self.nulls = nulls or {} self.combos = combos or {} self.domains = domains or {} + self.types = types or {} self.calls: list[list[str]] = [] self.combo_calls: list[list[list[str]]] = [] self.domain_calls: list[list[str]] = [] @@ -1875,7 +1906,12 @@ def table_metadata(self, identifier): column_count=len(self.approx), ) columns = [ - ColumnMeta(name=n, data_type="INTEGER", nullable=True, ordinal=i) + ColumnMeta( + name=n, + data_type=self.types.get(n, "INTEGER"), + nullable=True, + ordinal=i, + ) for i, n in enumerate(self.approx) ] return meta, columns @@ -2137,14 +2173,19 @@ def test_composite_probe_reports_the_minimal_grain_when_two_pairs_prove(): """Filling the cap makes two proven composites reachable on one table, and consumers read the first entry as the grain. ``(order_id, product_id)`` is two id-shaped columns so it wins the probe ranking, but it is a superkey of - the parent-plus-line grain, so the tighter pair has to come back first.""" + the parent-plus-line grain, so the tighter pair has to come back first. + + ``product_id`` is kept well below the near-unique anchor bar on purpose. A + decoy at 90%+ of the rows would be pruned as an anchor before it could be + probed, and this test is about ordering two proven pairs rather than about + that prune.""" from exmergo_dex_core.explore import profile as profile_mod from exmergo_dex_core.explore import relationships as rel_mod adapter = _StubAdapter( rows=1000, - approx={"order_id": 250, "line_number": 4, "product_id": 900}, + approx={"order_id": 250, "line_number": 4, "product_id": 300}, combos={("order_id", "line_number"): 1000, ("product_id", "order_id"): 1000}, ) datasets = profile_mod.profile(adapter, ["db.s.order_items"]) @@ -2178,6 +2219,270 @@ def test_parent_line_grain_detected_end_to_end(parent_line_grain_duckdb: Path, c assert not any("grain unknown" in n for n in ds["data_quality"]) +# --- artifact suppression in composite keys (#292) ------------------------------ + + +def test_continuous_measure_columns_never_enter_a_composite_pair(): + """A money column pairs with anything high-cardinality and means nothing, so + it leaves the pool before a pair is priced. Nothing here is near-unique, so + this isolates the measure rule from the anchor rule.""" + + from exmergo_dex_core.explore import profile as profile_mod + + adapter = _StubAdapter( + rows=1000, + approx={"order_id": 400, "grand_total": 700, "subtotal": 650, "status": 4}, + types={"grand_total": "DECIMAL(10,2)", "subtotal": "DECIMAL(10,2)"}, + # The junk pair would prove if it were ever asked. + combos={("grand_total", "order_id"): 1000}, + ) + datasets = profile_mod.profile(adapter, ["db.s.orders"]) + + probed = adapter.combo_calls[0] + assert not any("grand_total" in pair or "subtotal" in pair for pair in probed) + assert ["order_id", "status"] in probed, probed + assert datasets[0].composite_keys == [] + + +def test_a_measure_named_in_integer_minor_units_is_still_a_measure(): + """Money stored as an integer count of cents survives the fractional-type + test, so the name is what catches it. `qty` sits in the same table at low + cardinality and must stay a key member, which is why a name never decides + on its own.""" + + from exmergo_dex_core.explore import profile as profile_mod + + adapter = _StubAdapter( + rows=1000, + approx={"order_id": 250, "amount_cents": 700, "qty": 30}, + combos={("amount_cents", "order_id"): 1000, ("order_id", "qty"): 1000}, + ) + datasets = profile_mod.profile(adapter, ["db.s.orders"]) + + probed = adapter.combo_calls[0] + assert not any("amount_cents" in pair for pair in probed) + assert probed == [["order_id", "qty"]] + assert datasets[0].composite_keys == [["order_id", "qty"]] + + +def test_a_snowflake_style_numeric_ordinal_still_completes_the_grain(): + """Snowflake reports every NUMBER as the bare token FIXED with the scale + dropped, so a type-only measure rule would delete `NUMBER(38,0)` key members + on that whole connector. A low-cardinality ordinal is a bounded enumeration + whatever its type string says.""" + + from exmergo_dex_core.explore import profile as profile_mod + + adapter = _StubAdapter( + rows=1000, + approx={"order_id": 250, "fiscal_period": 24}, + types={"order_id": "FIXED", "fiscal_period": "FIXED"}, + combos={("order_id", "fiscal_period"): 1000}, + ) + datasets = profile_mod.profile(adapter, ["db.s.order_facts"]) + + assert datasets[0].composite_keys == [["order_id", "fiscal_period"]] + + +def test_a_pair_anchored_on_a_near_unique_column_is_never_probed(): + """The issue's shape at stub scale: `order_id` is unique on all but a + handful of rows, so every high-cardinality column completes it. No pair is + priced at all, the grain comes back unknown rather than wrong, and the note + names the one column worth looking at.""" + + from exmergo_dex_core.explore import profile as profile_mod + from exmergo_dex_core.explore import relationships as rel_mod + + adapter = _StubAdapter( + rows=1000, + # Every one of these escalates to 990 exact, so all three are anchors. + approx={ + "order_id": 950, + "created_at": 980, + "grand_total": 940, + "customer_id": 400, + }, + types={"created_at": "TIMESTAMP", "grand_total": "DECIMAL(10,2)"}, + # The artifact would prove if it were ever asked. + combos={("order_id", "customer_id"): 1000}, + ) + datasets = profile_mod.profile(adapter, ["db.s.orders"]) + ds = datasets[0] + + assert adapter.combo_calls == [], "no combination statement is issued at all" + assert ds.composite_keys == [] + assert rel_mod.candidate_keys(ds) == [] + assert rel_mod.detect_grain(ds) is None + + notes = rel_mod.data_quality_notes(ds) + closest = [n for n in notes if "closest thing to one" in n] + assert len(closest) == 1, notes + assert "order_id" in closest[0] + # A near-unique timestamp is not news, and a measure is never the key + # anyone meant, so neither is offered as the closest thing to a key. + assert "created_at" not in closest[0] and "grand_total" not in closest[0] + + # The evidence is complete where the prose is bounded: every anchor whose + # pairs were dropped is recorded, including the timestamp the note leaves + # out, so a host reading key_evidence sees the whole suppression and a + # human reading the notes gets the one column worth acting on. + # `grand_total` never became an anchor because it left the pool as a + # measure first, which is the documented test order. + suppressed = [e for e in ds.key_evidence if e.status == "suppressed"] + assert [e.columns for e in suppressed] == [["created_at"], ["order_id"]] + assert all("already unique for 99.0% of rows" in e.reason for e in suppressed) + + +def test_a_near_unique_anchor_still_pairs_with_a_bounded_enumeration(): + """The over-reach guard, and why the anchor rule is a pair test rather than + a member test. On a parent-line table where most orders have a single line, + `order_id` is near-unique and `(order_id, line_number)` is still the real + grain: a three-value partner separates the duplicates on purpose, where a + wider one would separate them by accident.""" + + from exmergo_dex_core.explore import profile as profile_mod + + adapter = _StubAdapter( + rows=1000, + approx={"order_id": 950, "line_number": 3}, + combos={("order_id", "line_number"): 1000}, + ) + datasets = profile_mod.profile(adapter, ["db.s.order_items"]) + + assert adapter.combo_calls[0] == [["order_id", "line_number"]] + assert datasets[0].composite_keys == [["order_id", "line_number"]] + + +def test_the_measure_and_anchor_rules_do_not_apply_below_the_minimum_row_count(): + """Cardinality-versus-rows reasoning says nothing on a tiny table: every + column looks near-unique and every measure looks continuous, so both rules + would suppress grains rather than junk. The probe is nearly free at this + size, so everything is asked and the exact count decides.""" + + from exmergo_dex_core.explore import profile as profile_mod + + below = _StubAdapter( + rows=profile_mod._COMPOSITE_MIN_ROWS - 1, + approx={"order_id": 90, "amount": 95}, + types={"amount": "DOUBLE"}, + combos={("amount", "order_id"): profile_mod._COMPOSITE_MIN_ROWS - 1}, + ) + datasets = profile_mod.profile(below, ["db.s.t"]) + assert below.combo_calls[0] == [["amount", "order_id"]] + assert datasets[0].composite_keys == [["amount", "order_id"]] + + at_bar = _StubAdapter( + rows=profile_mod._COMPOSITE_MIN_ROWS, + approx={"order_id": 90, "amount": 95}, + types={"amount": "DOUBLE"}, + ) + profile_mod.profile(at_bar, ["db.s.t"]) + # `amount` leaves the pool and one column cannot make a pair. + assert at_bar.combo_calls == [] + + +def test_key_evidence_reports_every_candidate_in_candidate_keys_order(): + """The invariant that keeps the two fields from drifting: the reported half + of key_evidence is exactly candidate_keys, in order.""" + + from exmergo_dex_core.explore import commands as cmd_mod + from exmergo_dex_core.explore import profile as profile_mod + + adapter = _StubAdapter( + rows=1000, + approx={"order_id": 250, "line_number": 4, "product_id": 300}, + combos={("order_id", "line_number"): 1000, ("product_id", "order_id"): 1000}, + ) + datasets = profile_mod.profile(adapter, ["db.s.order_items"]) + cmd_mod._annotate_grain(datasets, None) + ds = datasets[0] + + assert ds.candidate_keys == [ + ["order_id", "line_number"], + ["product_id", "order_id"], + ] + reported = [e.columns for e in ds.key_evidence if e.status == "reported"] + assert reported == ds.candidate_keys + assert all(e.reason for e in ds.key_evidence) + assert "it is the grain" in ds.key_evidence[0].reason + assert "ranks behind order_id, line_number" in ds.key_evidence[1].reason + + +def test_a_near_unique_timestamp_does_not_get_the_enumeration_escape(): + """The escape exists for a parent identifier whose rows a position column + enumerates. A near-unique timestamp is a per-row event time, so a + low-cardinality dimension beside it separates rows by accident just as a + wider partner would, and `(status, created_at)` is not a grain.""" + + from exmergo_dex_core.explore import profile as profile_mod + + adapter = _StubAdapter( + rows=1000, + approx={"created_at": 980, "status": 4, "customer_id": 300}, + types={"created_at": "TIMESTAMP"}, + combos={("created_at", "status"): 1000}, + ) + datasets = profile_mod.profile(adapter, ["db.s.events"]) + + probed = (adapter.combo_calls or [[]])[0] + assert not any("created_at" in pair for pair in probed), probed + assert datasets[0].composite_keys == [] + + +def test_a_near_unique_key_is_reported_instead_of_composites_built_on_it( + near_unique_key_duckdb: Path, capsys +): + """The issue's acceptance case, end to end against a real warehouse. Five + combinations are genuinely unique on this table and every one of them is an + artifact of `order_id`'s 110 duplicate rows, so none is reported: the grain + comes back unknown and the duplicates come back exactly.""" + + payload = _run( + ["explore", "profile", "orders", "--path", str(near_unique_key_duckdb)], + capsys, + ) + ds = payload["data"]["datasets"][0] + + assert ds["row_count"] == 2037 + assert ds["candidate_keys"] == [] + assert ds["grain"] is None + assert ds["composite_keys"] == [] + + notes = " ".join(ds["data_quality"]) + assert ( + "order_id is not unique: 1927 distinct over 2037 rows (110 rows would " + "have to be removed for it to be unique, so it is unique for 94.6% of " + "rows)" + ) in notes + # The suppression is stated in prose as well as in the field, and it says + # a probe ran, which is what separates this from a probe that never did. + assert "suppressed every one of them as an artifact" in notes + assert "order_id is the closest thing to one" in notes + # Two notes, not four: one column's arithmetic is stated once. + assert len(ds["data_quality"]) == 2, ds["data_quality"] + # Exact numbers throughout: the distinct count was escalated and the column + # has no nulls, so nothing here is marked approximate. + assert "~" not in notes + + evidence = ds["key_evidence"] + assert [e["columns"] for e in evidence if e["status"] == "reported"] == [] + # The two money columns never reach the evidence at all: they left the pool + # as measures before the anchor test ran, which is the documented order. + reasons = " ".join(e["reason"] for e in evidence) + assert "subtotal" not in reasons and "grand_total" not in reasons + # The near-unique timestamps do appear, but as anchors whose pairs were + # dropped, never as a member of a key that was reported. + suppressed = sorted(tuple(e["columns"]) for e in evidence) + assert suppressed == [("created_at",), ("order_id",), ("updated_at",)] + + # The suppression names the anchor and never the filler it was paired with, + # so `columns_with_findings` does not drag four extra columns back into the + # serialized `columns` that the finding summary correctly elides. + assert "customer_id" not in notes + assert "customer_id" not in {c["name"] for c in ds["columns"]} + assert ds["elided_column_count"] >= 1 + + # --- value-domain reporting (#203) ----------------------------------------------- diff --git a/packages/dex-core/tests/explore/test_relationships.py b/packages/dex-core/tests/explore/test_relationships.py index a7d0feae..8a67f1f7 100644 --- a/packages/dex-core/tests/explore/test_relationships.py +++ b/packages/dex-core/tests/explore/test_relationships.py @@ -742,7 +742,10 @@ def test_own_key_duplicates_produce_fan_out_warning(): notes = data_quality_notes(hosts) warning = next(n for n in notes if "not unique" in n) assert "ID is not unique: ~9590 distinct over 14111 rows" in warning - assert "4521 duplicate rows" in warning + # The surplus is stated as what it is, a number of rows to remove, and not + # as "duplicate ids": those differ, and only this one follows from the counts. + assert "~4521 rows would have to be removed for it to be unique" in warning + assert "unique for ~68.0% of rows" in warning assert "fan out" in warning assert any("grain unknown" in n for n in notes) diff --git a/packages/dex-core/tests/test_docs_policy.py b/packages/dex-core/tests/test_docs_policy.py new file mode 100644 index 00000000..f1a42194 --- /dev/null +++ b/packages/dex-core/tests/test_docs_policy.py @@ -0,0 +1,284 @@ +"""The two policy documents are the source of truth, and this is what makes that true. + +`references/pii-policy.md` and `references/cost-controls.md` exist to end a +restatement problem: the PII blocking threshold was stated in five places in four +wordings, and one paragraph about auto-profile pricing was copied verbatim into +five connector documents. Adding two documents only helps if the copies leave, +and only stays helped if new copies cannot arrive. + +So three things are held here. The number every document states is the number the +engine holds. The set of documents allowed to state it is frozen. The prose that +was copied lives in one file. + +When this fails, fix the document. Editing an allowlist is the fix only when a new +file has a real reason to carry the policy, and that reason belongs next to the +entry, not in a commit message. +""" + +from __future__ import annotations + +import re +import shutil +import subprocess +from pathlib import Path + +import pytest + +from exmergo_dex_core.guards import PII_BLOCK_CONFIDENCE + +ROOT = Path(__file__).resolve().parents[3] + +# Resolved once, and absolute: the corpus below is whatever git reports, so a +# `git` picked up from a mutable PATH would decide what this suite checks. +GIT = shutil.which("git") + +# The suite runs from a checkout in CI and in development. An installed wheel has +# no repository around it, and nothing checked here is a property of the wheel, so +# skip rather than fail. `tests/ossie/test_corpus.py` makes the same assumption +# about `references/`; this only states it out loud. A checkout with no `git` on +# PATH skips for the same reason rather than failing on a missing executable. +pytestmark = pytest.mark.skipif( + not (ROOT / "references").is_dir() or GIT is None, + reason="needs a repository checkout with git available", +) + +PII_DOC = "references/pii-policy.md" +COST_DOC = "references/cost-controls.md" + +# Append-only history. Every past wording of every policy is in here on purpose, +# and rewriting history to satisfy a lint is worse than the drift it would catch. +EXEMPT = frozenset({"CHANGELOG.md"}) + +# Where the PII blocking threshold may be stated, and why. Everything else links +# to references/pii-policy.md. +# +# The rule behind the list, so a reviewer can apply it without reading the list: +# a file may carry the number only when it is the file that owns the policy. +# Every other document links, including the three skills, which ship standalone +# through `npx skills add exmergo/dex` and cannot link to `references/`: they +# carry the policy's consequence ("below the blocking threshold it projects with +# a warning") and never its constant, so a skill can go stale on wording but not +# on the number. +PII_RESTATEMENT_ALLOWED = frozenset({PII_DOC}) + +# "Confidence" alone is too broad to key on: join inference reports a declared +# join at confidence 1.0, and a walkthrough transcript quotes the flag on one +# column at 0.95. Neither states the policy. What identifies a restatement of the +# constant is the threshold concept itself, so key on that and let example +# confidences and join confidences through. +THRESHOLD_SUBJECT = re.compile(r"(?i)\bthreshold\b|blocks projection") +DECIMAL = re.compile(r"(? dict[str, str]: + """Every committed markdown file, keyed by repository-relative path. + + `git ls-files` rather than a glob, for the same reason the em-dash CI job uses + it: `unversioned-docs/` is present in a working tree and absent in CI, and a + check that sees different files on the two machines is worse than no check. + """ + + # Tracked files, plus files that are new but not ignored. Without the second + # call a policy document added in the working tree is invisible here, so the + # census would pass on a branch that has not committed yet and fail the + # moment it does. `--exclude-standard` keeps `unversioned-docs/` out, which + # is the whole reason for asking git rather than globbing. + paths: list[str] = [] + for extra in ([], ["--others", "--exclude-standard"]): + listed = subprocess.run( # noqa: S603 (a fixed argv, no shell) + [GIT, "ls-files", "-z", *extra, "*.md", "*.mdc"], + cwd=ROOT, + capture_output=True, + text=True, + check=True, + ).stdout + paths.extend(p for p in listed.split("\0") if p) + return { + path: (ROOT / path).read_text(encoding="utf-8") + for path in sorted(set(paths)) + if path not in EXEMPT and (ROOT / path).is_file() + } + + +def prose_sentences(text: str): + """Yield (line number, sentence) for prose only. + + Fenced blocks are dropped whole and inline code spans blanked, on the same + reasoning as `scripts/check_no_em_dashes.py`: a YAML fragment showing a + connector's own ceiling is configuration, not a restatement of policy, and a + check that cannot tell them apart gets silenced rather than satisfied. + """ + + in_fence = False + block: list[str] = [] + block_start = 1 + + def flush(): + if not block: + return + joined = " ".join(block) + joined = re.sub(r"`[^`]*`", lambda m: " " * len(m.group()), joined) + joined = re.sub(r"https?://\S+", lambda m: " " * len(m.group()), joined) + for sentence in re.split(r"(?<=[.;:])\s+", joined): + if sentence.strip(): + yield block_start, sentence + + for line_no, line in enumerate(text.splitlines(), start=1): + if line.lstrip().startswith("```"): + yield from flush() + block = [] + in_fence = not in_fence + continue + if in_fence or not line.strip(): + yield from flush() + block = [] + continue + # A table row is its own paragraph. Joining a backtick-dense table into + # one block makes the inline-code masking pair backticks across rows and + # blank out prose that is not code at all, which is how a restatement + # hides in a command table. + if line.lstrip().startswith("|"): + yield from flush() + block = [line.strip()] + block_start = line_no + yield from flush() + block = [] + continue + # These documents hard-wrap, so a sentence routinely straddles a line + # break. Scanning line by line would miss every claim whose number and + # subject land on different lines, which is most of them. Paragraphs are + # the smallest unit that keeps a sentence whole; the reported line is + # where the paragraph starts. + if not block: + block_start = line_no + block.append(line.strip()) + yield from flush() + + +def threshold_carriers() -> dict[str, list[str]]: + """Files whose prose states a decimal in a sentence about the PII threshold.""" + + carriers: dict[str, list[str]] = {} + for path, text in versioned_markdown().items(): + for line_no, sentence in prose_sentences(text): + if THRESHOLD_SUBJECT.search(sentence) and DECIMAL.search(sentence): + carriers.setdefault(path, []).append(f"{path}:{line_no}") + return carriers + + +def test_every_document_that_states_the_pii_threshold_states_the_engine_s_number(): + expected = f"{PII_BLOCK_CONFIDENCE:g}" + wrong = [] + for path, text in versioned_markdown().items(): + for line_no, sentence in prose_sentences(text): + if not THRESHOLD_SUBJECT.search(sentence): + continue + wrong.extend( + f"{path}:{line_no}: states {hit.group()}, engine holds " + f"{expected}: {sentence.strip()[:120]}" + for hit in DECIMAL.finditer(sentence) + if hit.group() != expected + ) + assert not wrong, ( + "the PII blocking threshold moved in the engine and these documents still " + f"carry the old number. PII_BLOCK_CONFIDENCE is {expected} " + "(packages/dex-core/src/exmergo_dex_core/guards/__init__.py).\n " + + "\n ".join(wrong) + ) + + +def test_the_canonical_pii_document_states_the_threshold_at_all(): + """Deleting the sentence is the other way a source of truth stops being one.""" + + text = (ROOT / PII_DOC).read_text(encoding="utf-8") + assert any( + THRESHOLD_SUBJECT.search(sentence) and f"{PII_BLOCK_CONFIDENCE:g}" in sentence + for _, sentence in prose_sentences(text) + ), f"{PII_DOC} names no blocking threshold, so nothing else can point at it" + + +def test_only_the_canonical_document_restates_the_threshold(): + carriers = threshold_carriers() + added = sorted(set(carriers) - PII_RESTATEMENT_ALLOWED) + gone = sorted(PII_RESTATEMENT_ALLOWED - set(carriers)) + assert not added, ( + "a new restatement of the PII blocking threshold appeared. Link to " + f"{PII_DOC} instead of repeating the number, or, if this file genuinely " + "cannot link, add it to PII_RESTATEMENT_ALLOWED with the reason.\n " + + "\n ".join(f"{p}: {', '.join(carriers[p])}" for p in added) + ) + assert not gone, ( + "these files no longer state the threshold, which is the good direction. " + "Delete them from PII_RESTATEMENT_ALLOWED so the list keeps meaning what " + "it says.\n " + "\n ".join(gone) + ) + + +# Prose that was copied verbatim across connector documents before the refactor. +# Matched whitespace-normalized so a reflow does not read as a rewrite. +SHARED_PROSE = { + "auto-profile pricing": ( + "That scan is billed, and it is priced into the same handshake as the " + "statements rather than added afterward, so the estimate you confirm is " + "the whole cost." + ), +} +SHARED_PROSE_HOME = {"auto-profile pricing": COST_DOC} + + +def test_the_shared_cost_prose_lives_in_exactly_one_document(): + corpus = { + path: " ".join(text.split()) for path, text in versioned_markdown().items() + } + for label, paragraph in SHARED_PROSE.items(): + needle = " ".join(paragraph.split()) + carriers = sorted(path for path, text in corpus.items() if needle in text) + home = SHARED_PROSE_HOME[label] + assert carriers == [home], ( + f"the {label} prose is the kind that was copied into five connector " + f"documents and drifted. It belongs in {home} and nowhere else; the " + f"other files link. Found in: {carriers}" + ) + + +CEILING_FRAGMENT = re.compile( + r"session_ceiling:\s+\d+\s+# cumulative (bytes|seconds) per UTC day" +) + + +def test_every_connector_config_example_uses_the_same_ceiling_fragment(): + """The value is per-connector and stays so. The comment around it is policy. + + Six connector documents show this fragment with three different magnitudes and + two units, all correct. What must not drift is the sentence, because a reader + copying one fragment reads the comment as the definition of the key. + """ + + offenders = [] + for path, text in versioned_markdown().items(): + for line_no, line in enumerate(text.splitlines(), start=1): + if "session_ceiling:" in line and not CEILING_FRAGMENT.search(line): + offenders.append(f"{path}:{line_no}: {line.strip()}") + assert not offenders, ( + "a connector document invented its own wording for the cumulative ceiling. " + "Use `session_ceiling: # cumulative per UTC day`, and see " + f"{COST_DOC} for what it means.\n " + "\n ".join(offenders) + ) + + +LINK = re.compile(r"\]\((?!https?:|#|mailto:)([^)\s]+)\)") + + +def test_relative_links_between_documents_resolve(): + """In a refactor that replaces prose with pointers, a broken pointer deletes + documentation rather than centralizing it.""" + + broken = [] + for path, text in versioned_markdown().items(): + base = (ROOT / path).parent + for line_no, line in enumerate(text.splitlines(), start=1): + for target in LINK.findall(line): + file_part = target.partition("#")[0] + if file_part and not (base / file_part).exists(): + broken.append(f"{path}:{line_no}: -> {target}") + assert not broken, "a relative link points at nothing.\n " + "\n ".join(broken) diff --git a/packages/dex-core/tests/test_safety_spine.py b/packages/dex-core/tests/test_safety_spine.py index 9bc5d801..46ae7501 100644 --- a/packages/dex-core/tests/test_safety_spine.py +++ b/packages/dex-core/tests/test_safety_spine.py @@ -1958,6 +1958,7 @@ def test_the_er_diagram_marks_pii_and_carries_no_column_value(): from exmergo_dex_core.cache import ( Dataset, DexCache, + KeyEvidence, PIICategory, Relationship, ) @@ -1988,6 +1989,18 @@ def test_the_er_diagram_marks_pii_and_carries_no_column_value(): ], candidate_keys=[["customer_id"]], grain=["customer_id"], + # A key reason is prose, so it is the one new place a value could get + # spliced in by a later change. The diagram must not read this field at + # all, and asserting it here is what confines it to `explore profile` + # by test rather than by accident: `notable_columns` is shared with the + # renderer and is one line away from consulting it. + key_evidence=[ + KeyEvidence( + columns=["customer_id"], + status="reported", + reason="customer_id is unique and non-null, aaron@example.com", + ) + ], ) orders = Dataset( identifier="shop.main.orders", @@ -2018,6 +2031,78 @@ def test_the_er_diagram_marks_pii_and_carries_no_column_value(): assert "pii:government_id 0.90" in mermaid +def test_profile_key_evidence_and_notes_carry_counts_not_values(): + """Where the line between a measurement and a value falls, stated in a test + rather than left to a reviewer's judgement. + + A value is a datum read out of a row and rendered as itself: a min, a max, + a value domain entry, an address, an id. A count, a distinct count, a ratio + and the number of rows that would have to go for a column to be unique are + measurements *over* rows, and they are the currency this whole guardrail is + denominated in. So the test forbids the first and **requires** the second: + a version that reported no numbers would pass a forbid-only assertion while + being useless. + + Scoped to `data_quality` and `key_evidence` deliberately, not to the whole + payload. A profile legitimately carries min/max and a value domain for safe + columns, so a blanket "no sentinel anywhere" assertion would be wrong here + in a way it is not wrong for the diagram and the map. + """ + + from exmergo_dex_core.cache import ( + Dataset, + KeyEvidence, + ValueCount, + ValueDomain, + ) + from exmergo_dex_core.explore.relationships import data_quality_notes, key_evidence + from exmergo_dex_core.explore.results import _profile_dataset_payload + + # Nothing here could hold a value, and the pinned field set is what makes + # that a fact rather than a claim about today's code. + assert set(KeyEvidence.model_fields) == {"columns", "status", "reason"} + + orders = Dataset( + identifier="shop.main.orders", + row_count=2037, + columns=[ + ColumnProfile( + name="order_id", + data_type="BIGINT", + distinct_count=1927, + distinct_count_exact=True, + is_unique=False, + null_fraction=0.0, + min_value=1, + max_value=987654, + ), + ColumnProfile( + name="tier", + data_type="VARCHAR", + distinct_count=2, + value_domain=ValueDomain( + values=[ValueCount(value="platinum", count=3)] + ), + ), + ], + ) + orders.key_evidence = key_evidence(orders) + orders.data_quality = data_quality_notes(orders) + + payload = _profile_dataset_payload(orders, show_all_columns=True) + notes = " ".join(payload["data_quality"]) + reasons = " ".join(e["reason"] for e in payload["key_evidence"]) + + for value in ("987654", "platinum"): + assert value not in notes, "a column value reached a data-quality note" + assert value not in reasons, "a column value reached a key reason" + + # The positive half: the measurements a caller acts on are all present. + assert "1927 distinct over 2037 rows" in notes + assert "110 rows would have to be removed" in notes + assert "94.6% of rows" in notes + + def test_the_map_payload_marks_pii_and_carries_no_column_value(): """`explore map` returns findings rather than a receipt (issue #202), which puts profile content into the envelope for the first time. The rule the @@ -2033,6 +2118,7 @@ def test_the_map_payload_marks_pii_and_carries_no_column_value(): from exmergo_dex_core.cache import ( Dataset, DexCache, + KeyEvidence, PIICategory, Relationship, ValueCount, @@ -2067,6 +2153,16 @@ def test_the_map_payload_marks_pii_and_carries_no_column_value(): ], candidate_keys=[["customer_id"]], grain=["customer_id"], + # As in the diagram test: a key reason is prose, and this payload must + # not reach for it. `columns_with_findings` and `notable_columns` are + # both shared with this command and both one line from consulting it. + key_evidence=[ + KeyEvidence( + columns=["customer_id"], + status="reported", + reason="customer_id is unique and non-null, aaron@example.com", + ) + ], ) orders = Dataset( identifier="shop.main.orders", diff --git a/references/bigquery.md b/references/bigquery.md index e0103bce..aa9dc83e 100644 --- a/references/bigquery.md +++ b/references/bigquery.md @@ -53,29 +53,17 @@ budget: - **Billed:** profiling aggregates, `explore query`, relationship verification probes, and `transform build`. -`explore query` and `explore cluster` profile an object they name that this connection has but the `.dex/` cache cannot adjudicate. That scan is billed, and it is priced into the same handshake as the statements rather than added afterward, so the estimate you confirm is the whole cost. A call carrying several statements is quoted once for all of them, itemized per statement, and an object two of them share is scanned once rather than twice. Resolving which objects need it stays free: it is object listing and column metadata, the same reads the inventory uses. Pass `--no-auto-profile` (or set `auto_profile: false` in `.dex/config.yml`) to be refused instead. +`explore query` and `explore cluster` bill an auto-profile of an object this connection has that the `.dex/` cache cannot adjudicate, priced into the same handshake as the statements: see [`cost-controls.md`](cost-controls.md). Pass `--no-auto-profile` (or set `auto_profile: false` in `.dex/config.yml`) to be refused instead. -Every billed command is estimated first with free dry-runs. Without -`--confirm` it returns a `needs_confirmation` envelope carrying the byte -estimate (per table where relevant); re-issue with `--confirm` and -`--budget `. Nothing executes unconfirmed or without a ceiling, and an estimate -over the ceiling is refused outright (confirmation cannot override it). +Every billed command is estimated first with free dry-runs, which is why the +estimate here is `exact` rather than modelled. Budgets are bytes; the handshake +itself is in [`cost-controls.md`](cost-controls.md). On the confirmed run, every statement is dry-run again and charged against the budget, and every job carries a server-side `maximum_bytes_billed` cap, so a -drifting estimate cannot overrun the budget. Billed bytes are appended to -`.dex/spend.jsonl` (byte counts, job ids, and statement hashes; never SQL text -or values), and `budget.session_ceiling` binds cumulatively against that -ledger per UTC day. - -The ledger gates billing and nothing else. A gate is built whenever a BigQuery -connection is assembled, free commands included, but the day's spend is read only -where it is needed: billed admission reads it and refuses if it cannot (a named -`reason: guard` refusal saying nothing ran), settlement tolerates a failure, and a -free command never reaches it. So a store keeping the ledger somewhere that can be -unreachable does not put `explore inventory` behind it. `connect test` is the one -free exception, because reporting the budget is its job: it takes one guarded read -and reports `budget.session_spent_today: null` when the ledger cannot be reached. +drifting estimate cannot overrun the budget. Where the billed bytes land +afterwards, and when a ledger that cannot be read refuses a command, are in +[`cost-controls.md`](cost-controls.md). BigQuery bills a 10 MB minimum per query; a remaining budget below that is refused with the math rather than letting the job fail server-side. Query-cache @@ -94,6 +82,15 @@ rows, a table of nested or repeated columns only (no approximate distinct, which every probe starts from), a table too small for a value domain, or one with too few countable columns to form a composite pair. +The reserve cannot mirror every reason a composite probe ends up issuing nothing. +A pair built on a continuous measure, or one that merely completes a column +already unique on almost every row, is excluded before the probe runs, and both +verdicts come from distinct counts that do not exist at estimate time. So a table +whose every candidate pair is excluded that way still carries its reserve and +then spends nothing. That is the loose direction and it is the safe one: +reserving for a query that does not run costs a caller headroom, while failing to +reserve for one that does is an overrun. + An object BigQuery keeps no row count for, meaning every view and every external table, reserves all three. Unknown is not empty: the count arrives inside the aggregate scan, so every probe can run, and at estimate time there is no number @@ -199,16 +196,10 @@ and sums the result (downstream nodes whose dev inputs are not built yet cannot be dry-run, so on a cold target the total is a partial floor). It still requires `--confirm` and a `--budget`, and its billed bytes land in the spend ledger. -`transform build --verify` prices its row counts into the same estimate as the -build itself, as a `(row counts)` entry in the per-table breakdown, so one -`--budget` covers both phases. Only a relation the warehouse keeps no row count -for costs anything, which is any view (dbt's default materialization); a table's -count is free metadata and its verdict is reported `exact: false` to say so. On -a cold dev target the counts cannot be dry-run priced before the build has -written the relations, so a note says so and they are priced again afterwards as -a phase drawn against the reservation the build is already holding. A phase that -does not fit returns `ok` with the counts in `data.offer`, never -`needs_confirmation` for a build that has already run and billed. +`transform build --verify` costs only where the warehouse keeps no row +count, so a table's count is free metadata; how the counts are priced into +the build's own estimate is in +[`cost-controls.md`](cost-controls.md). `--verify` also folds `bigquery.dev_dataset` into its read scope for the length of that one command, because dbt writes the relations it is judging there and that namespace diff --git a/references/canonical-model.md b/references/canonical-model.md index 8f4dea72..5f6be3d0 100644 --- a/references/canonical-model.md +++ b/references/canonical-model.md @@ -54,10 +54,14 @@ informs proposals, never the source of truth: .dex/ config.yml non-secret config: connector + dbt target, budgets, ranking hints cache.json exploration artifacts (DexCache): profiles, PII flags, relationships, - candidate keys, grain, rankings, data-quality observations. A column - profile records whether its distinct count is exact or approximate - (distinct_count_exact); each dataset carries the time it was profiled - (profiled_at) so carried-forward profiles stay attributable + ranked candidate keys with the evidence behind each (key_evidence, + including the combinations suppressed as artifacts and why), grain, + rankings, data-quality observations. A column profile records whether + its distinct count is exact or approximate (distinct_count_exact); each + dataset carries the time it was profiled (profiled_at) so + carried-forward profiles stay attributable, and the file carries a + schema version, so a profile written before the artifact exclusions + existed is re-profiled rather than read as if they had applied snapshot.json the maintain baseline: a frozen fingerprint of the warehouse schema, the transformation project's state, the semantic layer's definitions with the relationships and keys they declare, and declared grain assumptions. The diff --git a/references/clickhouse.md b/references/clickhouse.md index 1fa1691d..1f333190 100644 --- a/references/clickhouse.md +++ b/references/clickhouse.md @@ -313,17 +313,9 @@ fine; what it cannot do is be capped. dex **warns** in that case rather than refusing, and the build result says the run was uncapped rather than claiming a cap it never injected. -`transform build --verify` prices its row counts into the same estimate as the -build itself, as a `(row counts)` entry in the per-table breakdown, so one -`--budget` covers both phases. Only a relation the warehouse keeps no row count -for costs anything, which is any view (dbt's default materialization); a table's -count is free `system.tables` metadata, and a verdict resting on it is reported -`exact: false` to say so. On a cold dev target the counts cannot be priced -before the build has written the relations, so a note says so and they are -priced again afterwards as a phase drawn against the reservation the build is -already holding. A phase that does not fit returns `ok` with the counts in -`data.offer`, never `needs_confirmation` for a build that has already run and -billed. +`transform build --verify` costs only where the warehouse keeps no row count, so a +table's count is free `system.tables` metadata; how the counts are priced into the +build's own estimate is in [`cost-controls.md`](cost-controls.md). `transform test --mutate` prices its whole batch as one number and confirms it once, then runs one dbt invocation per mutant. Nothing is materialized: a mutant diff --git a/references/command-contract.md b/references/command-contract.md index 5e458ffc..6af9964d 100644 --- a/references/command-contract.md +++ b/references/command-contract.md @@ -466,11 +466,12 @@ in `data.cap.elided`, per class. `seed_csv` is the first kind that puts **values**, not logic, into a reviewable diff, and a diff goes into git and stays there. So a seed's header is checked both against the PII detector `explore` profiles warehouse columns with and -against the flags already in the `.dex/` cache. A column at or above the block -threshold is refused, and the refusal names the `pii_overrides` entry that would -clear it (never a value). The standing limit is worth knowing: dex detects PII -from names and types and never from values, everywhere, so a seed column named -`code` full of email addresses passes this gate. +against the flags already in the `.dex/` cache. A column at or above the +blocking threshold is refused, and the refusal names the `pii_overrides` entry +that would clear it, never a value ([`pii-policy.md`](pii-policy.md)). The +standing limit is worth knowing: dex detects PII from names and types and never +from values, everywhere, so a seed column named `code` full of email addresses +passes this gate. `project_yml` and `profiles_yml` bring the two project-root config files into the same plan/diff/apply flow; because they carry @@ -832,18 +833,13 @@ the warehouse and take the `--confirm --budget` handshake on billed connectors; `check` runs the free axes first and returns one combined estimate for the scanning axes. -That estimate arrives as an **offer on a complete answer**, not as a pending -charge. `check` and `semantic` finish their free axes on every call and return -`ok`, with the price of the scanning axes under `data.offer` (`estimated_bytes` -or the connector's own estimate shape, `per_table_bytes`, `axes` naming what the -estimate would add, and the `--confirm --budget` hint). `data.axes_run` names -what completed, and reading both is how a caller tells "grain found nothing" -from "grain did not run", which the status used to imply. Nothing scans until -the confirmed re-issue arrives, and `cost.estimate` stays unset on the offer so -an `ok` never carries a number that reads as spend. The reason for the split is -that `needs_confirmation` is a request for a decision dex is blocked on, and -spending it on work the caller never asked for teaches them to confirm -reflexively, which is the one habit the handshake cannot survive. +That estimate arrives as an **offer on a complete answer** +([`cost-controls.md`](cost-controls.md)), not as a pending charge. `check` and +`semantic` finish their free axes on every call and return `ok`, with the price +of the scanning axes under `data.offer`, which carries `axes` naming what the +estimate would add. `data.axes_run` names what completed, and reading both is +how a caller tells "grain found nothing" from "grain did not run", which the +status used to imply. **A baseline reports its own coverage, and the axes only compare what it covers.** `maintain snapshot` pins the exploration cache, and a cache is thin @@ -945,10 +941,35 @@ leaving a stale claim. A project with no compiled semantic layer contributes neither direction and is not an error on this path; `explore semantic list` is the command whose subject is the layer, and it is the one that refuses by name. +**A key is ranked, and the reason travels with it.** `candidate_keys` is ordered, +tightest proven key first, so the first entry is the one `grain` elects and the +rest are alternatives rather than an unordered set a caller has to guess through. +`key_evidence` carries one entry per combination the profile considered, each with +its `columns`, a `status` of `reported` or `suppressed`, and a `reason` in the +profile's own words. The reported entries are in the same order as +`candidate_keys`, which is the invariant that keeps the two fields from drifting; +the suppressed ones follow. + +Suppression exists because unique is not the same as identifying. Where one column +is unique on all but a handful of rows, every wider column in the table completes +it, and the resulting combinations are the same fact restated: the column has +duplicates. Reporting them as keys buries the real key among filler and puts a +test in scaffolded dbt on a tuple nobody meant. So they are suppressed from +`candidate_keys`, kept in `key_evidence` with the reason, and stated once in +`data_quality` alongside the counts. A caller learns from `key_evidence` that a +probe ran and found only artifacts; that a probe never ran, or was narrowed by the +budget, is what the probe's own notes say, and the two are not the same answer. + +`key_evidence` is a `profile` field. `explore map` reports the best-ranked key as +`candidate_key` and is budgeted per object, so the full ranking belongs to the +command whose subject is one relation in full. + **`explore map` returns the map, not a receipt for it.** Alongside the counts, `data.objects` carries each top-ranked object's row count, detected grain, -candidate key, notable columns (each with the role that earned it a place: -`grain`, `key`, `join`, or a PII flag) and data-quality findings, and `data.edges` +best-ranked candidate key (the full ranking and the reason behind each key are +`explore profile`'s `key_evidence`, which this payload deliberately does not +carry), notable columns (each with the role that earned it a place: `grain`, +`key`, `join`, or a PII flag) and data-quality findings, and `data.edges` carries the join edges in exactly the shape `explore relationships` returns them. It is budgeted the way `explore diagram` is budgeted: at most 25 objects kept by rank, 12 columns per object, 40 edges, and 5 data-quality findings per object. @@ -1020,15 +1041,11 @@ connector reads it in its own namespace vocabulary: a `dataset` on BigQuery, a `catalog.schema` on Databricks, a `schema` on Postgres. It is never written back to config, so `connect test --scope X` works before a connector block exists. -Two rules make it a cost control rather than a hint: - -- **Scope narrows, never widens.** When `.dex/config.yml` commits a source - allowlist, that allowlist is a cost boundary and every `--scope` entry must - resolve inside it. A scope that reaches outside is refused. -- **A scope is honored or named in an error, never dropped.** An entry that - names nothing refuses and lists what exists. `--project` and `--dataset` are - BigQuery vocabulary and error on any other connector; DuckDB has no namespace - to scope and refuses all three (its target is `--path`). +Two rules make it a cost control rather than a hint, both in +[`cost-controls.md`](cost-controls.md): it narrows and never widens, and it is +honored or named in an error, never dropped. `--project` and `--dataset` are +BigQuery vocabulary and error on any other connector; DuckDB has no namespace to +scope and refuses all three (its target is `--path`). ## The query firewall @@ -1086,22 +1103,12 @@ The gate, in order: sample passes a mid-command gate afterward: a budget too small for the sample returns `needs_confirmation` with the profile already saved, rather than refusing and discarding it. -3. **Classify the projection.** Output may carry values only from profiled - columns whose PII flag is absent or below the blocking threshold. A flag at - confidence 0.5 or above blocks projection; the threshold is a hard-coded - engine constant, uniform across categories, deliberately not configurable. - Every value path from a blocking column must pass through a measuring - aggregate (COUNT, APPROX_COUNT_DISTINCT, AVG, SUM, STDDEV, ...). - Value-carrying aggregates (MIN, MAX, ANY_VALUE, STRING_AGG, ...) do not - qualify, unknown functions fail closed, and `SELECT *` is refused when the - expansion includes a blocking column. Projecting a column whose flag sits - below the threshold (de-rated by value-shape evidence at profile time) runs, - with an envelope warning naming the column, category, and confidence. - Filters, join conditions, GROUP BY and ORDER BY are unrestricted: values - flow in, not out. A column a human has reviewed as not PII is cleared by a - `pii_overrides` entry in `.dex/config.yml` (fully qualified column, optional - reason), which unblocks querying immediately and suppresses the flag durably - on every later profile. +3. **Classify the projection.** Output may carry values only from columns the + PII policy clears, which is what the profiles gathered in step 2 are for: + the firewall cannot judge a column whose flags it does not have. Which + aggregates carry a value out, what a sub-threshold flag does, and how a human + clears one are in [`pii-policy.md`](pii-policy.md). Filters, join conditions, + GROUP BY and ORDER BY are unrestricted: values flow in, not out. 4. **Bound the result.** LIMIT is clamped (default 50 rows), long cells are cut (default 256 chars), the payload is byte-capped (default 16 KiB), at most 10 statements ride in one call, and every cut is announced in `notes`. A watchdog @@ -1165,11 +1172,9 @@ Every command prints one object of this shape (`exmergo_dex_core.envelope`): Rules the envelope enforces, all of them Tier-2 eval targets: -- **Cost before spend.** `cost` is a preflight estimate. Any command that would - spend returns `needs_confirmation` unless given `--confirm` (and a `--budget` - on billed connectors; DuckDB is free, so the confirm handshake alone gates it). - An estimate over the ceiling is refused outright; confirmation cannot override - it. +- **Cost before spend.** `cost` is a preflight estimate, and any command that + would spend returns `needs_confirmation` until it is confirmed and budgeted. + The handshake in full is in [`cost-controls.md`](cost-controls.md). - **An estimate says how much it is worth.** `estimate_quality` is `exact` on BigQuery, where a dry run is what the job will bill, and `approximate` on every connector that models a run instead. `unknown` means dex tried to price the @@ -1188,11 +1193,6 @@ Rules the envelope enforces, all of them Tier-2 eval targets: `ok` as settled preflight for work that ran. `needs_confirmation` stays reserved for work the caller asked for and has not authorized, which is the only case where dex is genuinely blocked. - On billed connectors the estimate comes from free dry-runs, the confirmed - run re-checks every statement against the budget with a server-side cap as - backstop, actual spend is reported under `data.spend`, and every billed byte - is appended to the `.dex/spend.jsonl` ledger, against which the optional - `budget.session_ceiling` binds cumulatively per UTC day. - **`data.spend` reports settled spend for every billed command**, including `transform build`, and always matches what the same command appended to the ledger. It carries the connector's unit (`bytes_billed` or `seconds_billed`) @@ -1218,107 +1218,12 @@ Rules the envelope enforces, all of them Tier-2 eval targets: binding ceiling and ledger unit, and adds approximate compute-unit-hours from live per-replica memory plus optional USD from the configured price. Missing or partial capacity refuses before billed work. -- **The ledger is a readable artifact, and every row says what it is.** - `.dex/spend.jsonl` is one JSON object per line, appended and never rewritten. - Every row carries the same keys, `null` where one does not apply, so the file - parses to a stable schema and no reader has to interpret an absent key: - - | key | what it holds | - |---|---| - | `at` | UTC ISO-8601 stamp of the write | - | `connector` | the connector that billed. Several can share one file, so filter on it before summing | - | `command` | the command that wrote the row | - | `entry` | the kind: `reservation`, `settlement` or `release`, and nothing else | - | `reservation_id` | ties one command's rows together. `null` on a `transform build` settlement, which settles outside any gate because dbt runs the statements | - | `billed_bytes` or `billed_seconds` | the magnitude, in the connector's unit. Signed | - | `estimate` | the whole-command preflight figure the settlement was admitted on. `null` on the other two kinds, and on a settlement whose pricing degraded to no estimate at all | - | `job_id` | the warehouse's identifier for the job, where it has one | - | `statement_sha256` | a hash of the statement. Never its text, never a value | - - **Settled spend is the `entry == "settlement"` filter**, summed over one - connector's unit. That is not the same number as `session_spent_today` - whenever a command is still in flight: a reservation is positive and its - release is the same magnitude negative, so the three kinds net to actual spend - once a command has settled, and until then the day's total legitimately reads - higher by the headroom being held. A process killed outright leaves its - reservation standing until the UTC rollover. So the two figures agree exactly - when nothing is running, and where they differ the difference is held - headroom rather than an accounting error. Sum what you are given; clamping the - negative would leave a release uncancelled. - - **A ledger holding settlements alone is not a ledger missing rows.** A - reservation exists to be seen by a concurrent command settling against - `budget.session_ceiling`, so a project that has never set a daily cap has - nothing for one to protect and writes none, and no release either. Both kinds - appear from the first billed command after a ceiling is set. - - **A row with no `entry` at all was written by a dex older than v1.5.1**, and it - is a settlement. The ledger is an append-only audit trail and dex does not - rewrite it, so a project that has been running since before that release has - both shapes in one file. Filter on `entry` being `"settlement"` or absent to - include them. dex's own reads already sum them correctly, because the day's - total does not branch on kind. - -- **The spend ledger is a dependency of billing, not of every command.** A gate - is built at every connection assembly on a billed connector, free commands - included, but nothing reads the ledger until something needs the day's total. - Billed admission reads it and fails closed if it cannot: the refusal is named - (`reason: guard`), carries the cost, and says nothing ran, so re-issuing the - same command is safe. Settlement tolerates a read failure instead, so a backend - that goes away mid-command does not turn a command that already ran into a - refusal. A command that cannot spend never reaches the ledger, so a store - keeping it on a network can be unreachable without taking down a cache-served - answer. Two fields can therefore report `null`: `data.spend.session_spent_today` - when the ledger failed at settlement (what that command billed is still exact, - because the warehouse said so; only the day's total is unavailable), and - `connect test`'s `budget.session_spent_today`, which takes one guarded read - because reporting the budget is that command's job. -- **The cumulative ceiling binds across commands that overlap in time**, not - only across commands that follow one another. An admitted command books its - estimate against the day's headroom before it runs and releases the unspent - part when it settles, so a second command issued while the first is still - running is measured against what is genuinely left, and the server-side cap - each statement carries is bounded by that booking rather than by the whole - ceiling. Three consequences a caller can see: - - `cost.ceiling` on a refusal reflects headroom another command is holding, so - two runs of the same command can be refused against different numbers. - - `session_spent_today` counts headroom held by commands still in flight, so - while another billed command is running it reads higher than settled spend - and can briefly exceed `session_ceiling` without anything having overspent. - Run commands one at a time and it is exactly settled spend. - - A command killed outright leaves its estimate booked until the UTC rollover. - Every softer exit, including an interrupt, releases. This errs conservative - on purpose: the alternative is a hold that expires while its command is - still spending. -- **A billed command with no cumulative ceiling warns.** `budget.ceiling` is - *refused* when missing, because nothing runs unbudgeted; - `budget.session_ceiling` is only warned about, because refusing would break - every project that never set one. Without the warning the two are - indistinguishable from outside, and an unset daily cap reads as one that - bound. Config is read from `/.dex/config.yml` and does not inherit, - so a second repo root has its own budget or none, and the warning says so. -- **And a project is asked for one, once.** The warning above is accurate and it - is also the default state of every new project, so it repeated on every billed - command, which is the condition under which warnings stop being read: several - billed commands can run bound by their per-command caps alone, each carrying - the same sentence, with the aggregate bounded by nothing. So the first billed - command in a project with no recorded decision returns `needs_confirmation` - naming a `suggested_session_ceiling` (five times that command's own estimate, - in the connector's unit, as a starting point) and a `session_ceiling_hint` - spelling out both answers, under its own key because a two-phase command's - findings payload owns `hint`. Answer it with `--session-ceiling ` to - set one or `--no-session-ceiling` to record that the project runs unbounded; - either answer is written to `.dex/config.yml` and nothing asks again in that - project. The ask is the *last* check before - spend, so an unanswered one has run nothing, booked no headroom, and reached - the ledger not at all, and the unconfirmed cost ask that precedes it carries - the suggestion in `notes` so one re-run can answer both. Three cases are never - asked: a project that already set `budget.session_ceiling` (nothing changes - for it), one that recorded a decline, and a config-free ad-hoc read, which has - no committed file to record an answer in and keeps the warning alone. A - decline records a decision and loosens nothing: the warning still fires on - every billed command, and now names the decline so a reader can tell a settled - choice from a project that was never asked. +- **Spend, the ledger, and the daily ceiling are governed by + [`cost-controls.md`](cost-controls.md).** That file owns the + `.dex/spend.jsonl` row shape, how reservations and settlements net, the + cumulative `budget.session_ceiling` and its one-time ask, and what a ledger + that cannot be read does to a billed command. This section states only what + the envelope carries. - **`cost.paradigm` names the connector the command ran against**, not what the command happened to cost. A free metadata command on BigQuery reports `bytes_scanned` with a null estimate, so a caller learns what a billed command diff --git a/references/cost-controls.md b/references/cost-controls.md new file mode 100644 index 00000000..36def390 --- /dev/null +++ b/references/cost-controls.md @@ -0,0 +1,195 @@ +# Cost controls + +Nothing dex runs touches the warehouse without a ceiling. This file is the whole +cost guard: the handshake that surfaces a price before spending, the ledger that +records what was spent, the cumulative ceiling that binds a day's work, and the +one surface where no ceiling is possible. Every other document links here rather +than restating it. What a scan costs on a particular warehouse, and how that +figure is derived, belongs to each connector's own file; this file owns the +mechanism that is the same everywhere. + +## What "budget" means here + +Money, or load on a machine somebody pays for: `--budget`, `budget.ceiling` and +`budget.session_ceiling`. Two unrelated things are also called a budget and are +not governed here, because they protect agent context rather than spend: +`query.max_payload_bytes` and the semantic catalog's caps. `--confirm` likewise +carries a second meaning on the project commands, where it overwrites a stored +plan or resolves an apply conflict; that use is not a cost gate, and it lives in +[`dbt-project.md`](dbt-project.md). + +## The handshake + +`cost` is a preflight estimate, and it comes from free dry-runs. Any command +that would spend returns `needs_confirmation` unless given `--confirm`, plus a +`--budget` on billed connectors; DuckDB is free, so the confirm handshake alone +gates it. Re-issue the same command with `--confirm` and `--budget ` +in the connector's unit, and the confirmed run re-checks every statement against +the budget with a server-side cap as backstop. An estimate over the ceiling is +refused outright; confirmation cannot override it. + +`transform build --verify` prices its row counts into the same estimate as the +build itself, as a `(row counts)` entry in the per-table breakdown, so one +`--budget` covers both phases. Only a relation the warehouse keeps no row count +for costs anything, which is any view; a table's count is free metadata, and a +verdict resting on it is reported `exact: false` to say so. On a cold dev target +the counts cannot be priced before the build has written the relations, so a +note says so and they are priced again afterwards as a phase drawn against the +reservation the build is already holding. + +**A priced phase the caller did not request is an offer, not a refusal.** When a +command's free half is a complete answer in its own right, the envelope is `ok` +and the price of the optional half sits in `data.offer`, carrying the same +estimate, breakdown and hint a refusal would. `maintain check` and +`maintain semantic` are the two. `cost.estimate` stays unset on an offer, so a +populated `cost.estimate` on an `ok` still means settled preflight for work that +ran, and `needs_confirmation` stays reserved for work the caller asked for and +has not authorized. + +## Paradigms + +`cost.paradigm` names the connector the command ran against, not what the +command happened to cost, so a free metadata command still reports the paradigm +a billed one would bill in. `free_local` is a positive claim that the connector +bills nothing; `null` means no connector was resolved. + +| Connector | Paradigm | Binding unit | Cost model | +|---|---|---|---| +| BigQuery | `bytes_scanned` | bytes | [`bigquery.md`](bigquery.md) | +| Snowflake | `compute_time` | warehouse-seconds | [`snowflake.md`](snowflake.md) | +| Databricks | `compute_time` | warehouse-seconds | [`databricks.md`](databricks.md) | +| Redshift | `compute_time` | compute-seconds | [`redshift.md`](redshift.md) | +| Postgres | `db_load` | database-seconds | [`postgres.md`](postgres.md) | +| ClickHouse | `db_load`, or `compute_time` on Cloud | database-seconds | [`clickhouse.md`](clickhouse.md) | +| DuckDB | `free_local` | nothing to confirm | [`duckdb.md`](duckdb.md) | +| dbt Cloud Semantic Layer | `hosted` | no ceiling possible | see below | + +What the envelope carries alongside the paradigm, including `estimate_quality` +and the `data.spend` key parity rule, is in +[`command-contract.md`](command-contract.md). + +## The per-command ceiling and `--scope` + +`--budget` sets the ceiling for one command and is refused when missing, because +nothing runs unbudgeted. `--scope` narrows the committed source allowlist for +one command, in each connector's own namespace vocabulary. Two rules make it a +cost control rather than a hint: a committed allowlist is a cost boundary, so +every `--scope` entry must resolve inside it and one reaching outside is +refused; and a scope is honored or named in an error, never dropped, so an entry +that names nothing refuses and lists what exists. It is never written back to +config. The per-connector vocabulary is in +[`command-contract.md`](command-contract.md). + +## Pricing an auto-profile + +`explore query` and `explore cluster` profile an object they name that the +connection has but the `.dex/` cache cannot adjudicate, because the firewall +cannot judge a column whose flags it does not have. That scan is billed, and it +is priced into the same handshake as the statements rather than added +afterward, so the estimate you confirm is the whole cost. A call carrying +several statements is quoted once for all of them, itemized per statement, and +an object two of them share is scanned once rather than twice. Resolving which +objects need it stays free: it is object listing and column metadata, the same +reads the inventory uses. `explore cluster` prices its profile first and gates +its sample mid-command, so a budget too small for the sample returns +`needs_confirmation` with the profile already saved rather than discarding it. + +## The spend ledger + +`.dex/spend.jsonl` is one JSON object per line, appended and never rewritten. +Every row carries the same keys, `null` where one does not apply, so the file +parses to a stable schema and no reader has to interpret an absent key. + +| key | what it holds | +|---|---| +| `at` | UTC ISO-8601 stamp of the write | +| `connector` | the connector that billed. Several can share one file, so filter on it before summing | +| `command` | the command that wrote the row | +| `entry` | the kind: `reservation`, `settlement` or `release`, and nothing else | +| `reservation_id` | ties one command's rows together. `null` on a `transform build` settlement, which settles outside any gate because dbt runs the statements | +| `billed_bytes` or `billed_seconds` | the magnitude, in the connector's unit. Signed | +| `estimate` | the whole-command preflight figure the settlement was admitted on. `null` on the other two kinds, and on a settlement whose pricing degraded to no estimate at all | +| `job_id` | the warehouse's identifier for the job, where it has one | +| `statement_sha256` | a hash of the statement. Never its text, never a value | + +**Settled spend is the `entry == "settlement"` filter**, summed over one +connector's unit. That is not the same number as `session_spent_today` whenever +a command is in flight: a reservation is positive and its release is the same +magnitude negative, so the three kinds net to actual spend once a command has +settled, and until then the day's total legitimately reads higher by the +headroom being held. Sum what you are given; clamping the negative would leave a +release uncancelled. + +**A ledger holding settlements alone is not a ledger missing rows.** A +reservation exists to be seen by a concurrent command settling against +`budget.session_ceiling`, so a project that never set a daily cap has nothing to +protect and writes none. **A row with no `entry` at all** was written by a dex +older than v1.5.1 and is a settlement; filter on `entry` being `"settlement"` or +absent to include them. + +**The ledger is a dependency of billing, not of every command.** Billed +admission reads it and fails closed if it cannot, with a named refusal +(`reason: guard`) that says nothing ran, so re-issuing is safe. Settlement +tolerates a read failure instead, so a backend that goes away mid-command does +not turn a command that already ran into a refusal. Writing a store that serves +this contract is [`storage.md`](storage.md). + +## The cumulative ceiling + +`budget.session_ceiling` binds per UTC day and across commands that overlap in +time, not only across commands that follow one another. An admitted command +books its estimate against the day's headroom before it runs and releases the +unspent part when it settles, so a second command issued while the first is +running is measured against what is genuinely left. Three consequences a caller +can see: + +- `cost.ceiling` on a refusal reflects headroom another command is holding, so + two runs of the same command can be refused against different numbers. +- `session_spent_today` counts held headroom, so it reads higher than settled + spend while another billed command runs and can briefly exceed the ceiling + without anything having overspent. Run commands one at a time and it is + exactly settled spend. +- A command killed outright leaves its estimate booked until the UTC rollover. + Every softer exit, including an interrupt, releases. + +Unlike `budget.ceiling`, a missing `budget.session_ceiling` only warns, because +refusing would break every project that never set one. Config is read from +`/.dex/config.yml` and does not inherit, so a second repo root has its +own budget or none. + +## The one-time ask + +That warning is accurate and it is also the default state of every new project, +so on its own it would repeat on every billed command, which is the condition +under which warnings stop being read. Instead, the first billed command in a +project with no recorded decision returns `needs_confirmation` naming a +`suggested_session_ceiling` (five times that command's own estimate, in the +connector's unit, as a starting point) and a `session_ceiling_hint` spelling out +both answers. Answer with `--session-ceiling ` to set one or +`--no-session-ceiling` to record that the project runs unbounded. Either answer +is written to `.dex/config.yml` and nothing asks again in that project. + +The ask is the last check before spend, so an unanswered one has run nothing, +booked no headroom, and never reached the ledger. The unconfirmed cost ask that +precedes it carries the suggestion in `notes`, so one re-run can answer both. +Three cases are never asked: a project that already set the ceiling, one that +recorded a decline, and a config-free ad-hoc read, which has no committed file +to record an answer in. A decline loosens nothing: the warning still fires and +now names the decline, so a reader can tell a settled choice from a project that +was never asked. + +## The one surface with no cost guard + +The hosted dbt Cloud Semantic Layer (`explore semantic query --api`) is the one +place dex cannot enforce a ceiling. dbt Cloud owns the warehouse connection and +executes the query server-side, so no dry-run estimate and no server-side cap +are possible from dex. That backend therefore does not ask for `--confirm`, +because a confirmation dex could not back with a ceiling would be dishonest. It +reports `cost.paradigm: hosted` with no estimate and no ceiling, and states on +every result that the cost guard is unavailable and spend is governed by the dbt +Cloud environment. Do not present a hosted result as cost-guarded, and do not +read the absence of an estimate as "it was free". + +The local backend (`--local`) executes through dex's own connector and keeps the +full handshake. Which axis decides this is in +[`semantic-layer.md`](semantic-layer.md). diff --git a/references/databricks.md b/references/databricks.md index 8e3e9594..93ae7550 100644 --- a/references/databricks.md +++ b/references/databricks.md @@ -73,7 +73,7 @@ paths provably free: `explore query`, relationship verification probes, DESCRIBE DETAIL size probes, and `transform build`. -`explore query` and `explore cluster` profile an object they name that this connection has but the `.dex/` cache cannot adjudicate. That scan is billed, and it is priced into the same handshake as the statements rather than added afterward, so the estimate you confirm is the whole cost. A call carrying several statements is quoted once for all of them, itemized per statement, and an object two of them share is scanned once rather than twice. Resolving which objects need it stays free: it is object listing and column metadata, the same reads the inventory uses. Pass `--no-auto-profile` (or set `auto_profile: false` in `.dex/config.yml`) to be refused instead. +`explore query` and `explore cluster` bill an auto-profile of an object this connection has that the `.dex/` cache cannot adjudicate, priced into the same handshake as the statements: see [`cost-controls.md`](cost-controls.md). Pass `--no-auto-profile` (or set `auto_profile: false` in `.dex/config.yml`) to be refused instead. Budgets (`budget.ceiling`, `--budget`, `budget.session_ceiling`) are warehouse-seconds: the number you budget is the number the server enforces. @@ -96,16 +96,10 @@ prices `transform build`: each compiled model, snapshot, and test is estimated and summed into the build's upfront cost, labeled `estimate_quality: "low"` like any other Databricks estimate. -`transform build --verify` prices its row counts into the same estimate as the -build itself, as a `(row counts)` entry in the per-table breakdown, so one -`--budget` covers both phases. Only a relation the warehouse keeps no row count -for costs anything, which is any view (dbt's default materialization); a table's -count is free metadata, and a verdict resting on it is reported `exact: false` -to say so. On a cold dev target the counts cannot be priced before the build has -written the relations, so a note says so and they are priced again afterwards as -a phase drawn against the reservation the build is already holding. A phase that -does not fit returns `ok` with the counts in `data.offer`, never -`needs_confirmation` for a build that has already run and billed. +`transform build --verify` costs only where the warehouse keeps no row +count, so a table's count is free metadata; how the counts are priced into +the build's own estimate is in +[`cost-controls.md`](cost-controls.md). `transform test --mutate` prices its whole batch as one number and confirms it once, then runs one dbt invocation per mutant. Nothing is materialized: a mutant @@ -125,15 +119,11 @@ billed statement the session's `STATEMENT_TIMEOUT` is set to the remaining budget, so a weak floor cannot overrun the ceiling: Databricks kills the statement and dex reports the over-ceiling refusal. Actual spend is wall-clock seconds per statement (including any wake the statement caused), -recorded to `.dex/spend.jsonl` as `billed_seconds` and summed into the daily -session ceiling. Connections identify themselves with a `dex` user-agent for +recorded as `billed_seconds` in the ledger. Connections identify themselves with a `dex` user-agent for attribution. -The handshake is the same strict two-step as every billed connector: a -scanning command without `--confirm` returns `needs_confirmation` carrying -the seconds estimate (per table where relevant) and its DBU translation; -re-issue with `--confirm --budget `. Nothing executes unconfirmed or -without a ceiling, and an estimate over the ceiling is refused outright. +The estimate carries its DBU translation; the handshake itself is in +[`cost-controls.md`](cost-controls.md). ## Read-only, enforced in depth diff --git a/references/methodology.md b/references/methodology.md index bbae2f3b..d4c164de 100644 --- a/references/methodology.md +++ b/references/methodology.md @@ -73,12 +73,47 @@ that reuses a column from a better-ranked pair at a similar product is the same hypothesis with different filler, so it ranks behind every pair that is not, but it is never removed from the running: the cap is the spend guard, and while it has room a discarded candidate costs a grain and saves nothing. A pair whose -combination count equals the row count is a proven composite key: it enters the -candidate keys and, absent any single-column key, becomes the reported grain, -smallest product first so the tightest proven key is the one reported. On metered -connectors the probe spends only inside the already-confirmed budget, narrowing to -the best-ranked pairs the budget covers, and skipping with an explanatory note -only when it cannot cover one. +combination count equals the row count is unique, which is necessary for a key and +not sufficient to be one. On metered connectors the probe spends only inside the +already-confirmed budget, narrowing to the best-ranked pairs the budget covers, and +skipping with an explanatory note only when it cannot cover one. + +Two shapes of pair are unique for reasons that have nothing to do with the table's +identity, and both are excluded from the pool before anything is priced. A pair +whose member is a **continuous measure** is one: a measurement's cardinality grows +with the table, so it completes a partner by arithmetic rather than by meaning, and +a key made of money is almost never the grain anybody intended. The other is a pair +**anchored on a column that is already unique on almost every row**. At that ratio +the handful of rows the column fails to separate get separated by accident by any +partner carrying more than a value or two, so the pair restates the column's +near-uniqueness instead of discovering a grain. The exception, and the reason this +is a test on the pair rather than on the column, is a partner whose domain is +bounded by the fanout rather than by the table: when a few duplicate order ids are +separated by a three-value line number, the data really does have that grain. A +near-unique *timestamp* gets no such exception, because a per-row event time is not +an entity whose rows a position column enumerates. + +Excluding before probing is deliberate on three counts: the exact answer could not +move a decision either way, the pair would have to be reported and immediately +disclaimed, and dropping it frees a capped slot for a pair that might be the real +grain. None of it applies to a table too small for cardinality-versus-rows +reasoning to mean anything, where every column looks near-unique and every measure +looks continuous: below a fixed row floor everything is asked and the exact count +decides. + +What survives enters the candidate keys, smallest proven cardinality first, so the +tightest proven key is the first entry and, absent any single-column key, the +reported grain. Where nothing survives, the grain is reported unknown and the +near-unique column is named with its counts, because the finding there is +duplicates in the source rather than an absent key. + +Suppression is never silent. Every combination the probe considered comes back in +`key_evidence`, one entry per combination, carrying its columns, whether it was +reported or suppressed, and the reason in the profile's own words. The reported +entries appear in the same order as the candidate keys, so the ranking is readable +rather than implied, and the suppressed ones follow. That is also what makes a +probe that ran and found only artifacts distinguishable from one that never ran, +which the probe's own budget notes say instead. Two safety rules are enforced at the source, in the SQL that is generated: @@ -113,15 +148,10 @@ When the evidence is missing or ambiguous, the name-derived confidence stands: absence of evidence never weakens a flag. The flag itself is never removed by evidence. What a weak flag means is the -consumer's decision: the query firewall blocks projection at confidence 0.5 and -above (a hard-coded engine constant) and allows lower-confidence columns with an -envelope warning, while min/max suppression and dbt `meta` stamping remain -presence-based at any confidence. The only way to clear a flag entirely is a -human decision recorded as a `pii_overrides` entry in `.dex/config.yml` (fully -qualified column plus an optional reason). An override is re-applied on every -profile, so it survives re-profiling, takes effect at query time immediately, -and leaves an audit trail in the cache recording which category the detector had -matched. +consumer's decision, and the consumers differ: some gate on a flag's presence at +any confidence, others compare it against a blocking threshold. That split, the +threshold, and the only way a human clears a flag are in +[`pii-policy.md`](pii-policy.md). ## Relationships: joins from metadata, not scans @@ -137,9 +167,14 @@ A join is proposed when a foreign-key-shaped column name matches a parent object corresponding column is a candidate key and the types are compatible; confidence reflects how strong the name and key signals are. Candidate keys and the most likely grain come from the uniqueness signals: single columns proven unique, plus -the composite keys proven at profile time; a single-column key is always -preferred as the grain, and a member of a composite key is never treated as -unique on its own. Declared joins come +the composite keys proven at profile time that survived the artifact exclusions +above; a single-column key is always preferred as the grain, and a member of a +composite key is never treated as unique on its own. Where the strongest signal is +a column unique on almost every row but not all of them, that is what the profile +says, naming the ratio, the counts, and how many rows would have to be removed for +the column to be unique. It reports no key at all rather than a combination built +on the shortfall, because a caller who cannot tell a real key from filler is worse +served by several candidates than by one named defect and none. Declared joins come from whichever channels declare one: the dbt project when one is present, and the semantic layer when one is configured, which need not be dbt's. Absent both, declared joins are simply empty, which is expected because explore is designed to diff --git a/references/ossie-compatibility.md b/references/ossie-compatibility.md index f557ec86..93a12a11 100644 --- a/references/ossie-compatibility.md +++ b/references/ossie-compatibility.md @@ -131,7 +131,8 @@ many screens the wrong column and reports the verdict as evidence-backed. PII linkage follows exactly this table. A dimension dex cannot resolve to a column is screened by the name heuristic and says so, rather than being screened -against a column that is not behind it. +against a column that is not behind it. The policy that linkage feeds is in +[`pii-policy.md`](pii-policy.md). ## Keys, relationships, snapshots, and authoring diff --git a/references/ossie-walkthrough.md b/references/ossie-walkthrough.md index 25542e3f..7cc7c795 100644 --- a/references/ossie-walkthrough.md +++ b/references/ossie-walkthrough.md @@ -424,8 +424,9 @@ dex explore query "select country_code, count(*) as customers ``` This route is governed, not a way around the guards. It runs through the query -firewall and, on a metered warehouse, through the cost handshake. Profiling -flagged `customers.email` at confidence 0.95, so: +firewall ([`pii-policy.md`](pii-policy.md)) and, on a metered warehouse, through +the cost handshake ([`cost-controls.md`](cost-controls.md)). Profiling flagged +`customers.email` at confidence 0.95, so: ``` dex explore query "select email from dex_demo.main.customers limit 5" @@ -455,11 +456,19 @@ on it, and empty is an answer: `dex_demo.main.products` and which is what separates a load-bearing table from a merely large one. **The declared grain.** `order_items` has a composite `primary_key` in the -document, and it overrides the heuristic, saying so rather than replacing it -silently: +document, and it supplies a grain measurement could not prove, saying where the +grain came from rather than leaving that silent: > grain order_id, product_id comes from the project's declared composite key -> (heuristic suggested unit_price, order_id) +> (measurement found no key of its own) + +That is the case the declaration earns its keep in. Nothing keys this table: +`order_item_id` is unique on 13,000 of its 14,000 rows because a batch was +loaded twice, and the column pairs that are unique here are unique only because +of that shortfall, so profiling reports none of them and says so in +`key_evidence`. A declaration cannot fix duplicates, and it does not claim to. +It states what the grain is meant to be, `maintain grain` is where that gets +verified against the data, and the measurement keeps the finding. **The declared joins**, at confidence 1.0, with the composite kept whole: diff --git a/references/pii-policy.md b/references/pii-policy.md new file mode 100644 index 00000000..9929d31d --- /dev/null +++ b/references/pii-policy.md @@ -0,0 +1,144 @@ +# The PII policy + +dex flags personal data and never surfaces it. This file is the whole policy: +what a flag is, what each surface does with one, and how a human clears one. +Every other document links here rather than restating it. + +How a flag is *produced* (the name-pattern table, the categories, the +value-shape statistics) belongs to [`methodology.md`](methodology.md). This file +owns what a flag *means* and who can clear it. + +## What a flag is + +A flag is recorded strictly as `(column, category, confidence)` with no example +value, and that triple is what propagates downstream into emitted dbt. Detection +runs on column names and aggregate shape, never by inspecting values, so a +column named `code` full of email addresses is not flagged anywhere in dex. + +A flag is never removed by evidence. Value-shape statistics computed during +profiling move its confidence in both directions, and they fail closed: when the +evidence is missing or ambiguous the name-derived confidence stands, because +absence of evidence never weakens a flag. Only a human decision removes a flag, +and only through the config entries below. + +## Presence versus threshold + +This is the distinction that decides every surface's behavior, and it is why two +commands can disagree about the same column without either being wrong. Some +consumers act on a flag's *presence* at any confidence; others compare its +confidence against the blocking threshold. + +| Surface | Gate | Effect | +|---|---|---| +| min and max suppression | presence, any confidence | the extremes are never computed, so no raw value leaves the engine | +| dbt `meta` stamping | presence, any confidence | the flag is stamped into model and column `meta` | +| cluster feature selection | presence, any confidence | flagged columns are excluded; naming one is opt-in and mean only | +| the query firewall | threshold | at or above, projection is refused; below, it runs with a warning | +| the seed header gate | threshold | a seed column at or above the threshold is refused | +| the semantic request gate | presence, any confidence | a flagged dimension refuses the whole command | + +A weak flag therefore still suppresses min and max and still stamps `meta`, +while allowing a query that projects the column. What a weak flag means is the +consumer's decision, and the consumers differ on purpose. + +## The threshold + +A flag at confidence **0.5** or above blocks projection. The threshold is a +hard-coded engine constant, uniform across categories, and deliberately not +configurable: a configurable threshold would let a one-line config edit quietly +widen the PII boundary. At today's base confidences everything blocks, so only a +flag de-rated by value-shape evidence at profile time falls below. + +This section is the only place in the repository that states the number. + +## The firewall rule + +Output may carry values only from profiled columns whose flag is absent or below +the threshold. Every value path from a blocking column must pass through a +measuring aggregate (COUNT, APPROX_COUNT_DISTINCT, AVG, SUM, STDDEV, and the +like). Value-carrying aggregates (MIN, MAX, ANY_VALUE, STRING_AGG) do not +qualify, unknown functions fail closed, and `SELECT *` is refused when the +expansion includes a blocking column. + +Filters, join conditions, GROUP BY and ORDER BY are unrestricted: values flow +in, not out. A count projects no column, which is why a filter over a flagged +column is still attributable. + +Projecting a column whose flag sits below the threshold runs, with an envelope +warning naming the column, category, and confidence. Treat that warning as +information for the user, not an error to fix. + +Unnested JSON and array outputs inherit the source column's flags. + +## The stricter surface + +Where a command's result *is* the values rather than an aggregate those values +slice, the gate moves from threshold to presence and from dropping a column to +refusing the command. + +A metric query returns aggregates that a dimension merely groups, so a flagged +dimension can be dropped from the grouping and the query still answers +something. Listing a dimension's values cannot degrade that way, so **any flag +refuses the command outright**. The refusal names both clearing routes. + +Evidence rules run in both directions: a dimension whose name reads innocuous is +refused when its resolved column is flagged, and a profiled, cleared column is +not re-blocked by a PII-shaped name. Where a dimension resolves to no physical +column, or to a relation that was never profiled, the name heuristic is the +fail-closed floor, so silence never clears. A result that ran on the floor alone +says so. + +Which evidence each backend can reach, and why a computed expression resolves to +no column, is in [`semantic-layer.md`](semantic-layer.md) and +[`ossie-compatibility.md`](ossie-compatibility.md). + +## Clearing a flag + +Two durable routes, both reviewable in git and both re-applied on every profile +so they survive re-profiling: + +1. A `pii_overrides` entry in `.dex/config.yml`. +2. `meta: {pii: false}` on the dimension in the project, for a semantic layer. + +An override takes effect at query time immediately, without re-profiling, and +leaves an audit trail in the cache recording which category the detector had +matched. **Never hand-edit `.dex/cache.json` to clear a flag**, and never +suggest weakening detection: neither survives the next profile, and neither is +reviewable. + +A `pii_overrides` entry takes one of two mutually exclusive shapes, enforced at +load: + +```yaml +pii_overrides: + # Exact: one reviewed column, fully qualified with the cache's + # connector-normalized identifier, so it can never silently widen to a + # same-named column elsewhere. + - column: MY_DB.PUBLIC.REGION.R_NAME + reason: region label, not a person + + # Pattern: one reviewed decision about a structurally identical column that + # exists by construction on many tables, where the exact form would cost one + # entry per table per environment. Glob scope, case-insensitive. + - column_name: document_name + scope: analytics.cdc_*.* + reason: resource path in this CDC export, not a person name +``` + +Both forms are typo-guarded at profile time. An exact entry naming a column its +table does not have warns, and so does a pattern entry whose scope matches +profiled tables when none of them carries the named column. A scope matching no +table yet stays silent, since new entities landing later under the same scope is +the point of the pattern form. Both `explore profile` and `explore map` emit +this warning with the same text, so a scheduled `map` is enough to catch a +rename that left an override pointing at nothing. + +## Where the policy is enforced + +[`command-contract.md`](command-contract.md) carries the surfaces that gate a +statement: the query firewall, the auto-profile that makes a column adjudicable +at all, the seed header gate on values entering git, and row-attribution counts. +Profiling's own suppression and the diagram renderer are in +[`methodology.md`](methodology.md), `meta` stamping on emitted models is in +[`project.md`](project.md), and the per-backend request gate is in +[`semantic-layer.md`](semantic-layer.md). diff --git a/references/postgres.md b/references/postgres.md index 72fd2ab7..f1753066 100644 --- a/references/postgres.md +++ b/references/postgres.md @@ -63,7 +63,7 @@ metered connector. - **Metered:** profiling aggregates, `explore query`, relationship verification probes, distinct-count escalations, and `transform build`. -`explore query` and `explore cluster` profile an object they name that this connection has but the `.dex/` cache cannot adjudicate. That scan is billed, and it is priced into the same handshake as the statements rather than added afterward, so the estimate you confirm is the whole cost. A call carrying several statements is quoted once for all of them, itemized per statement, and an object two of them share is scanned once rather than twice. Resolving which objects need it stays free: it is object listing and column metadata, the same reads the inventory uses. Pass `--no-auto-profile` (or set `auto_profile: false` in `.dex/config.yml`) to be refused instead. +`explore query` and `explore cluster` bill an auto-profile of an object this connection has that the `.dex/` cache cannot adjudicate, priced into the same handshake as the statements: see [`cost-controls.md`](cost-controls.md). Pass `--no-auto-profile` (or set `auto_profile: false` in `.dex/config.yml`) to be refused instead. Budgets (`budget.ceiling`, `--budget`, `budget.session_ceiling`) are database-seconds: the number you budget is the number the server enforces. @@ -80,15 +80,11 @@ metered statement the session's `statement_timeout` is set to the remaining budget, so a wrong heuristic cannot overrun the ceiling: Postgres kills the statement and dex reports the over-ceiling refusal. Actual spend is wall-clock seconds per statement (a killed statement still bills what ran), -recorded to `.dex/spend.jsonl` as `billed_seconds` and summed into the daily -session ceiling. Every session connects as `application_name = 'dex'` for +recorded as `billed_seconds` in the ledger. Every session connects as `application_name = 'dex'` for attribution in `pg_stat_activity`. -The handshake is the same strict two-step as every metered connector: a -scanning command without `--confirm` returns `needs_confirmation` carrying -the seconds estimate (per table where relevant); re-issue with -`--confirm --budget `. Nothing executes unconfirmed or without a -ceiling, and an estimate over the ceiling is refused outright. +Budgets here are database-seconds; the handshake itself is in +[`cost-controls.md`](cost-controls.md). ## Read-only, enforced in depth diff --git a/references/redshift.md b/references/redshift.md index cda48efb..5d9246aa 100644 --- a/references/redshift.md +++ b/references/redshift.md @@ -88,7 +88,7 @@ rather than hides: **Metered:** profiling aggregates, `explore query`, relationship verification probes, distinct-count escalations, and `transform build`. -`explore query` and `explore cluster` profile an object they name that this connection has but the `.dex/` cache cannot adjudicate. That scan is billed, and it is priced into the same handshake as the statements rather than added afterward, so the estimate you confirm is the whole cost. A call carrying several statements is quoted once for all of them, itemized per statement, and an object two of them share is scanned once rather than twice. Resolving which objects need it stays free: it is object listing and column metadata, the same reads the inventory uses. Pass `--no-auto-profile` (or set `auto_profile: false` in `.dex/config.yml`) to be refused instead. +`explore query` and `explore cluster` bill an auto-profile of an object this connection has that the `.dex/` cache cannot adjudicate, priced into the same handshake as the statements: see [`cost-controls.md`](cost-controls.md). Pass `--no-auto-profile` (or set `auto_profile: false` in `.dex/config.yml`) to be refused instead. Budgets (`budget.ceiling`, `--budget`, `budget.session_ceiling`) are compute-seconds: the number you budget is the number the server enforces. @@ -106,16 +106,10 @@ conservative capacity-scaled scan rate; every handshake payload carries each compiled model, snapshot, and test is estimated and summed into the build's upfront cost. -`transform build --verify` prices its row counts into the same estimate as the -build itself, as a `(row counts)` entry in the per-table breakdown, so one -`--budget` covers both phases. Only a relation the warehouse keeps no row count -for costs anything, which is any view (dbt's default materialization); a table's -count is free metadata, and a verdict resting on it is reported `exact: false` -to say so. On a cold dev target the counts cannot be priced before the build has -written the relations, so a note says so and they are priced again afterwards as -a phase drawn against the reservation the build is already holding. A phase that -does not fit returns `ok` with the counts in `data.offer`, never -`needs_confirmation` for a build that has already run and billed. +`transform build --verify` costs only where the warehouse keeps no row +count, so a table's count is free metadata; how the counts are priced into +the build's own estimate is in +[`cost-controls.md`](cost-controls.md). `transform test --mutate` prices its whole batch as one number and confirms it once, then runs one dbt invocation per mutant. Nothing is materialized: a mutant @@ -135,15 +129,11 @@ metered statement the session's `statement_timeout` is set to the remaining budget, so a wrong heuristic cannot overrun the ceiling: Redshift kills the statement and dex reports the over-ceiling refusal. Actual spend is wall-clock seconds per statement (a killed statement still bills what ran), -recorded to `.dex/spend.jsonl` as `billed_seconds` and summed into the daily -session ceiling. Every session connects as `application_name = 'dex'` +recorded as `billed_seconds` in the ledger. Every session connects as `application_name = 'dex'` (`SYS_CONNECTION_LOG`) and sets `query_group = 'dex'` for attribution. -The handshake is the same strict two-step as every metered connector: a -scanning command without `--confirm` returns `needs_confirmation` carrying -the seconds estimate (per table where relevant) and its RPU translation; -re-issue with `--confirm --budget `. Nothing executes unconfirmed -or without a ceiling, and an estimate over the ceiling is refused outright. +The estimate carries its RPU translation; the handshake itself is in +[`cost-controls.md`](cost-controls.md). ## Read-only, enforced in depth diff --git a/references/semantic-layer.md b/references/semantic-layer.md index 175634f0..e4d5036b 100644 --- a/references/semantic-layer.md +++ b/references/semantic-layer.md @@ -355,12 +355,11 @@ request), so the second attempt is free. The metric it settled on is in `scoped_ and named in a note along with the alternatives and the `--metric` flag that overrides the choice, because the narrowing must never be silent. -**PII is screened harder here than on a metric query.** A metric query returns -aggregates that a dimension merely slices, so a flagged dimension can be dropped -from the grouping and the query still answers something. Here the result *is* the -values, so a flagged dimension refuses the command, and the refusal names the -durable ways to clear a dimension reviewed as not PII (a `pii_overrides` entry in -`.dex/config.yml`, or `meta: {pii: false}` in the project). The evidence is the same +**PII is screened harder here than on a metric query**, because the result *is* +the values rather than aggregates a dimension slices, so a flagged dimension +refuses the command instead of being dropped from the grouping. That rule and +the two durable ways to clear a dimension are in +[`pii-policy.md`](pii-policy.md). The evidence is the same on each backend as it is for a query: the `.dex/` cache's flag on the resolved physical column locally, the layer's own `config.meta` hosted, fetched one metric at a time and unioned across every metric that reaches the dimension, with the name @@ -424,13 +423,9 @@ open a connection or see a credential, and dex then runs that SQL through its ow spine, in order: 1. **PII request-gate.** Each grouped or filtered dimension is resolved through the - manifest to its physical column, and that column's `.dex/` cache flag decides - (with `pii_overrides` from `.dex/config.yml` applied). Evidence rules in both - directions: a dimension whose name reads innocuous is refused when its column is - flagged, and a profiled, cleared column is not re-blocked by a PII-shaped name. - When the cache cannot speak to a dimension (never profiled, or a computed - expression rather than a bare column), the name heuristic is the fail-closed - floor, so silence never clears. + manifest to its physical column, and that column's `.dex/` cache flag decides. + The evidence rules, in both directions, and the name heuristic that floors them + are in [`pii-policy.md`](pii-policy.md). 2. **SELECT-only assertion.** Before anything else touches the statement or the connection, the rendered SQL is proven read-only. 3. **Relation pre-check.** The rendered SQL bakes in `relation_name` from the @@ -529,14 +524,10 @@ passes `--api`. Leaving the default in place there is refused with that fix name rather than failing further in on a missing project. **The cost guard is unavailable on this backend, and dex says so on every -result.** dbt Cloud owns the warehouse connection and executes the query -server-side under its own credential, so dex cannot dry-run to estimate cost and -cannot set a byte or credit ceiling. The hosted backend therefore does not ask for -a `--confirm` (a confirmation dex could not back with a ceiling would be -dishonest); it runs, and it attaches a warning to every result stating that dbt -Cloud, not dex, governs the spend, with the cost paradigm reported as `hosted` and -no estimate or ceiling. Spend is bounded only by the dbt Cloud environment's own -limits. +result.** It asks for no `--confirm`, reports the paradigm as `hosted` with no +estimate and no ceiling, and warns on every result that dbt Cloud governs the +spend. Why no ceiling is possible here, and what a caller must not conclude from +a missing estimate, are in [`cost-controls.md`](cost-controls.md). Both backends screen the group-by tokens and the dimensions a filter clause names. Reading a clause is the **backend's** job, not the shared gate's, because the @@ -547,11 +538,10 @@ no note is emitted, because nothing was found to adjudicate). A backend that can read its own filter dialect therefore refuses filtered queries instead of passing them. -PII is still screened before the query is sent: a dimension the layer's own -metadata marks as PII is refused, and a name heuristic (the same detector the -profiler uses) is the fail-closed floor for a layer that carries no such metadata. -Grouping or filtering by a PII-shaped dimension (`user__email`) is refused with a -recovery hint before anything reaches dbt Cloud. +PII is still screened before the query is sent, and the evidence here is the +layer's own metadata rather than the `.dex/` cache. Grouping or filtering by a +PII-shaped dimension (`user__email`) is refused with a recovery hint before +anything reaches dbt Cloud. That metadata is fetched one metric at a time and unioned, in a single request that carries one aliased field per metric. The API's `dimensions(metrics:)` field @@ -1015,9 +1005,7 @@ backend is bound to the same runtime contract in | Renders the SQL | dex, via MetricFlow `explain()` | dbt Cloud | nothing to render: the format defines no query runtime | | Executes the SQL | dex, through the active connector | dbt Cloud, server-side | nothing executes: `query` and `values` refuse | | Needs a local dbt project | yes | no | no, and it never reads one | -| Cost surfaced before spend | yes, the full handshake | no: cost guard unavailable, warns on every result | not applicable: reading the layer spends nothing | -| Ceiling enforced by dex | yes (`maximum_bytes_billed` / timeout) | no: the dbt Cloud environment's own limits | not applicable: nothing runs | -| `--confirm` required | yes, on billed connectors | no (nothing dex can gate) | no: the catalog is free on every connector | +| Cost guard ([`cost-controls.md`](cost-controls.md)) | the full handshake, ceiling enforced by dex | unavailable: dbt Cloud's own limits, warned on every result | not applicable: reading the layer spends nothing | | PII gate | `.dex/` cache flags on the resolved physical column, name heuristic as the floor | layer metadata, fetched per metric and unioned, plus a name heuristic | `.dex/` cache flags on the directly linked column, name heuristic where a field carries none | | When only the floor ran | disclosed on the result, naming the unprofiled relations | disclosed on the result, naming the dimensions the layer said nothing about | disclosed per field, naming which of the four no-column cases applies | | Namespace mismatch | refused before spend, against the connection's own inventory | dbt Cloud resolves its own relations | a source dex cannot address as one whole relation is opaque and links nothing | diff --git a/references/snowflake.md b/references/snowflake.md index 4b7f640f..e206a02c 100644 --- a/references/snowflake.md +++ b/references/snowflake.md @@ -118,7 +118,7 @@ no warehouse), while any data scan costs warehouse runtime. So dex guards - **Billed:** profiling aggregates, `explore query`, relationship verification probes, and `transform build`. -`explore query` and `explore cluster` profile an object they name that this connection has but the `.dex/` cache cannot adjudicate. That scan is billed, and it is priced into the same handshake as the statements rather than added afterward, so the estimate you confirm is the whole cost. A call carrying several statements is quoted once for all of them, itemized per statement, and an object two of them share is scanned once rather than twice. Resolving which objects need it stays free: it is object listing and column metadata, the same reads the inventory uses. Pass `--no-auto-profile` (or set `auto_profile: false` in `.dex/config.yml`) to be refused instead. +`explore query` and `explore cluster` bill an auto-profile of an object this connection has that the `.dex/` cache cannot adjudicate, priced into the same handshake as the statements: see [`cost-controls.md`](cost-controls.md). Pass `--no-auto-profile` (or set `auto_profile: false` in `.dex/config.yml`) to be refused instead. Budgets (`budget.ceiling`, `--budget`, `budget.session_ceiling`) are warehouse-seconds: the number you budget is the number the server enforces. @@ -137,16 +137,10 @@ probe against a cold warehouse is quoted at what the account will actually see. The same estimator prices `transform build`: each compiled model, snapshot, and test is estimated and summed into the build's upfront cost. -`transform build --verify` prices its row counts into the same estimate as the -build itself, as a `(row counts)` entry in the per-table breakdown, so one -`--budget` covers both phases. Only a relation the warehouse keeps no row count -for costs anything, which is any view (dbt's default materialization); a table's -count is free metadata, and a verdict resting on it is reported `exact: false` -to say so. On a cold dev target the counts cannot be priced before the build has -written the relations, so a note says so and they are priced again afterwards as -a phase drawn against the reservation the build is already holding. A phase that -does not fit returns `ok` with the counts in `data.offer`, never -`needs_confirmation` for a build that has already run and billed. +`transform build --verify` costs only where the warehouse keeps no row +count, so a table's count is free metadata; how the counts are priced into +the build's own estimate is in +[`cost-controls.md`](cost-controls.md). `transform test --mutate` prices the batch in warehouse-seconds, from the same heuristic the rest of this connector uses rather than from a dry run, so the @@ -170,14 +164,10 @@ billed statement the session's `STATEMENT_TIMEOUT_IN_SECONDS` is set to the remaining budget, so a wrong heuristic cannot overrun the ceiling: Snowflake kills the statement and dex reports the over-ceiling refusal. Actual spend is wall-clock seconds per statement (including any resume the statement caused), -recorded to `.dex/spend.jsonl` as `billed_seconds` and summed into the daily -session ceiling. Every session is tagged `QUERY_TAG = 'dex'` for attribution. - -The handshake is the same strict two-step as every billed connector: a -scanning command without `--confirm` returns `needs_confirmation` carrying the -seconds estimate (per table where relevant) and its credit translation; -re-issue with `--confirm --budget `. Nothing executes unconfirmed or -without a ceiling, and an estimate over the ceiling is refused outright. +recorded as `billed_seconds` in the ledger. Every session is tagged `QUERY_TAG = 'dex'` for attribution. + +The estimate carries its credit translation; the handshake itself is in +[`cost-controls.md`](cost-controls.md). ## Read-only, enforced in depth diff --git a/references/storage.md b/references/storage.md index dd9c5637..38576727 100644 --- a/references/storage.md +++ b/references/storage.md @@ -358,37 +358,19 @@ principals. A host that splits stores for some other reason, one per repo root say, has split the budget too and will not be told. **The ledger read-then-write has to be atomic for the cumulative ceiling to -bind**, and `SpendLock` below is how a backend provides it. The cost gate reads -`spend_since`, decides whether the command fits under `budget.session_ceiling`, -and appends, and two commands running that sequence at once read the same total -and both decide yes. Implement the lock and dex serializes the sequence through -it. Without one dex still books the headroom, which narrows the window from the -length of a warehouse query to the microseconds around the read, and warns on -every billed command that the ceiling is advisory on this backend. - -**Entries are stored, not interpreted.** Each carries an `entry` kind -(`reservation`, `settlement`, `release`) and a `reservation_id` tying the three -together, because the ceiling has to hold headroom for a command that has been -admitted and has not finished paying. A release carries a **negative** -magnitude, and that is the one thing to know here: a backend that clamps or -filters on sign would leak held headroom for the rest of the UTC day. Sum what -you are given. - -**Every entry has the same keys**, with `null` where one does not apply, because -the ledger is an artifact other tooling reads and an absent key there is a claim -of its own. So a reservation and a release carry `estimate`, `job_id` and -`statement_sha256` as nulls, and a `transform build` settlement carries a null -`reservation_id` rather than omitting it, which is how it says it settled outside -any gate. A backend needs to know none of this, and that is the point: store the -dict you were handed, whole, and the shape stays whatever dex wrote. A backend -that drops keys it does not recognize breaks a reader joining settlements to -reservations, and one that projects the entry onto columns of its own diverges -the first time a key is added. - -`estimate` is the whole-command preflight figure the command was admitted on, so -the ledger holds both halves of every "estimated this, billed that" pair rather -than only the half a budget is measured against. `SpendHistory` is what reads it -back. +bind**, and `SpendLock` is how a backend provides it. See +[Serializing the spend admission](#serializing-the-spend-admission) above, which +states the obligation and what dex does for a backend that provides no lock. + +**Entries are stored, not interpreted.** A backend stores the dict it was +handed, whole, and three properties of that dict are the whole obligation: a +release carries a **negative** magnitude, so clamping or filtering on sign leaks +held headroom for the rest of the UTC day; every entry carries the same keys with +`null` where one does not apply, so dropping unrecognized keys breaks a reader +joining settlements to reservations; and projecting the entry onto columns of +your own diverges the first time dex adds a key. Sum what you are given. What +each key holds, and what the three kinds mean to a reader, is in +[`cost-controls.md`](cost-controls.md). Two properties follow, and neither required a change to any backend written before reservations existed: diff --git a/skills/explore/SKILL.md b/skills/explore/SKILL.md index 709bb162..afc3e179 100644 --- a/skills/explore/SKILL.md +++ b/skills/explore/SKILL.md @@ -51,8 +51,17 @@ Subcommands, in the usual order: never rows). 3. `explore profile ` (space- or comma-separated) returns column profiles, PII flags recorded as (column, category, confidence) and never - example values, plus candidate keys, the likely grain, and data-quality - warnings (e.g. a non-unique id that will fan out on joins). A generic + example values, plus ranked candidate keys, the likely grain, `key_evidence`, + and data-quality warnings (e.g. an id unique on all but 110 rows, which will + fan out on joins). `candidate_keys` is ordered, tightest proven key first, + and `key_evidence` gives one entry per combination considered with its + `status` (`reported` or `suppressed`) and the reason. Read it before you + trust a composite: a combination unique only because one member is unique on + almost every row, or because a money column completes it, is suppressed + rather than reported. Where a near-unique column is the real story the + warning says so with the ratio, the counts, and how many rows would have to + be removed for it to be unique. That last number is the one to act on: it + names a source defect to fix rather than a key to work around. A generic `*_name` flag's confidence is refined by value-shape evidence from the same scan, in both directions: person-shaped values corroborate it, a closed reference vocabulary or long labels de-rate it below the firewall's blocking @@ -61,7 +70,9 @@ Subcommands, in the usual order: are approximate for scale, but any column that looks unique within approximation noise is escalated to an exact COUNT(DISTINCT) (`distinct_count_exact: true`), so uniqueness and grain verdicts rest on - proof; a `~` prefix in a warning marks a count that is still approximate. + proof; a `~` prefix marks a number that is still approximate, on a count and + on a percentage alike, so a figure quoted without one is exact arithmetic + over an exact distinct count on a column with no nulls. A requested object whose cached profile is still fresh (same connector, schema unchanged, within `profile_freshness_hours`, default 24) is served from the cache (`cache_hit_count`) instead of re-scanned, so profiling a @@ -79,8 +90,8 @@ Subcommands, in the usual order: columns that share no name at all. 5. `explore map` writes or updates the `.dex/` cache and returns the map (`--verify` works here too). Alongside the counts, `data.objects` gives each - top-ranked object its row count, detected grain, candidate key, notable - columns (each carrying the role that earned it a place: `grain`, `key`, + top-ranked object its row count, detected grain, best-ranked candidate key, + notable columns (each carrying the role that earned it a place: `grain`, `key`, `join`, or a PII flag) and data-quality findings, and `data.edges` gives the join edges in the same shape `explore relationships` returns. With `--use-project` each object also carries `semantic_models`, the semantic models @@ -276,14 +287,15 @@ DuckDB `t, UNNEST(json_keys(doc)) AS u(k)`, ClickHouse join; ARRAY JOIN is the expansion). The unnested value must come from a column of a table in the query (bare, or through a JSON/array function); unnesting a subquery, another table, a literal, or a generator is refused, -and the unnest's outputs inherit the source column's PII flags. A column whose flag was de-rated below the 0.5 -blocking threshold projects normally, with an envelope warning naming it; treat -the warning as information for the user, not an error to fix. If the user says a -refused column is not personal data, recommend a `pii_overrides` entry in -`.dex/config.yml` (fully qualified column, optional reason): it unblocks -querying immediately, survives re-profiles, and is reviewable in git. Never -hand-edit `.dex/cache.json` to clear a flag. Never fall back to raw Python or a -database CLI to run SQL; the firewall path is the only sanctioned one. +and the unnest's outputs inherit the source column's PII flags. A column whose +flag was de-rated below the blocking threshold projects normally, with an +envelope warning naming it; treat the warning as information for the user, not +an error to fix. If the user says a refused column is not personal data, +recommend a `pii_overrides` entry in `.dex/config.yml` (fully qualified column, +optional reason): it unblocks querying immediately, survives re-profiles, and is +reviewable in git. Never hand-edit `.dex/cache.json` to clear a flag. Never fall +back to raw Python or a database CLI to run SQL; the firewall path is the only +sanctioned one. ## Cloud and database targets (BigQuery, Snowflake, Databricks, Postgres, Redshift, ClickHouse) @@ -304,25 +316,17 @@ entry or `SNOWFLAKE_*` env; for Databricks `databricks auth login` or paste a key, token, or password. On a metered connector, scanning commands (`profile`, `map`, `relationships`, -`query`) run a two-step handshake. The first call returns -`needs_confirmation` with an estimate in `cost.estimate` (and a per-table -breakdown where relevant): an exact dry-run byte figure on BigQuery, a -heuristic labeled `estimate_quality: "heuristic"` in warehouse-seconds on -Snowflake (credits alongside), a floor labeled `estimate_quality: "low"` in -warehouse-seconds on Databricks (DBUs alongside; it sharpens itself inside -the confirmed budget), a heuristic in compute-seconds on Redshift (RPU-hours -alongside; Serverless estimates carry the 60-second wake minimum once), and -database-seconds on Postgres (no dollars; the guarded quantity is load on -the operational database) and on ClickHouse (self-hosted, also no dollars; -estimated free by the non-executing `EXPLAIN ESTIMATE`, which prices after -primary-key pruning, and reporting `estimate_basis` so you can tell a pruned -plan estimate from a whole-relation fallback). Surface the -estimate to the user in human units, get an explicit budget from them, and -re-issue the same command with `--confirm` and `--budget ` in the -paradigm's unit. Never invent a budget the user did not agree to, and never -retry with a raised budget on an over-ceiling refusal without asking. -Metadata is free (`connect test`, `inventory` run immediately), and OK -envelopes report actual spend under `data.spend`. +`query`) run a two-step handshake. The first call returns `needs_confirmation` +with an estimate in `cost.estimate`, a per-table breakdown where relevant, and +the unit it is counted in: bytes on BigQuery, warehouse-seconds on Snowflake +(credits alongside) and Databricks (DBUs), compute-seconds on Redshift +(RPU-hours), database-seconds on Postgres and ClickHouse (no dollars; the +guarded quantity is load). Surface the estimate to the user in human units, get +an explicit budget from them, and re-issue the same command with `--confirm` and +`--budget ` in that unit. Never invent a budget the user did not +agree to, and never retry with a raised budget on an over-ceiling refusal +without asking. Metadata is free (`connect test`, `inventory` run immediately), +and OK envelopes report actual spend under `data.spend`. An over-ceiling refusal now carries a calibration line drawn from `.dex/spend.jsonl`: what this connector's last few settled commands actually @@ -375,3 +379,5 @@ thing to reach for on a warehouse whose full map would be expensive. cross the envelope only from profiled columns whose flag is absent or below the blocking threshold, bounded and capped. Only a human's `pii_overrides` entry clears a flag entirely; never suggest weakening the detection. +- The two policies in full, in the engine repository: + `references/pii-policy.md` and `references/cost-controls.md`. diff --git a/skills/explore/references/probe-playbook.md b/skills/explore/references/probe-playbook.md index 026be48a..c0108a8a 100644 --- a/skills/explore/references/probe-playbook.md +++ b/skills/explore/references/probe-playbook.md @@ -5,7 +5,8 @@ firewall guarantees safety; this playbook is about effectiveness: asking the question in a shape that returns a small, decisive answer instead of a wall of rows. Map first (`explore map`) when you are getting your bearings: the profile usually already holds the answer (null fractions, distinct counts, min/max, -candidate keys), and probes exist for the questions it does not. But you do not +ranked candidate keys and the reason behind each), and probes exist for the +questions it does not. But you do not have to map before you can probe. A table the engine has not profiled, including a model you built moments ago, is profiled as part of answering, so a probe against something new costs one call. @@ -53,8 +54,11 @@ to this recipe for you: the joins that share a child are measured in one statement, so it costs what the relations cost rather than what the join count costs. -**2. Duplicate / grain check.** How badly is a key broken, and what does the -duplication look like? +**2. Duplicate distribution.** How badly a key is broken is already in the +profile: it reports the distinct count, the row count, and how many rows would +have to be removed for the column to be unique, exactly, whenever the distinct +count was escalated and the column has no nulls. Probe when you need the *shape* +of the duplication rather than its size. ```sql SELECT COUNT(*) AS rows, @@ -64,6 +68,10 @@ SELECT COUNT(*) AS rows, FROM (SELECT id, COUNT(*) AS cnt FROM t GROUP BY id) ``` +`worst_repeat` is the column the profile cannot give you, and it is the one that +separates a double-loaded batch (every repeat is 2) from a single id that +swallowed the table. + **3. Top-K categorical distribution.** What values dominate a (non-flagged) column, and how concentrated is it? diff --git a/skills/explore/scripts/run.py b/skills/explore/scripts/run.py index 67ef2c84..619f5473 100644 --- a/skills/explore/scripts/run.py +++ b/skills/explore/scripts/run.py @@ -65,7 +65,7 @@ # Rewritten by scripts/prepare_release.sh to the tagged version. The connector # extra is deliberately NOT part of this pin: it is chosen at runtime (see # _resolve_connector), so a release artifact is connector-neutral. -DEX_CORE_VERSION = "1.12.2" +DEX_CORE_VERSION = "1.12.3" # Connector id -> packaging extra. The engine's connector ids and the pyproject # extras share names, so this is the identity set today. An unknown or unset diff --git a/skills/maintain/SKILL.md b/skills/maintain/SKILL.md index 372b3871..c9ecfb88 100644 --- a/skills/maintain/SKILL.md +++ b/skills/maintain/SKILL.md @@ -270,3 +270,5 @@ since detection surfaces as a conflict, never a silent overwrite. rather than overwriting. - The repository is the source of truth, on both axes; the `.dex/` snapshot is a non-canonical fingerprint used only to detect change. +- The cost guard behind the scanning axes, in full, in the engine repository: + `references/cost-controls.md`. diff --git a/skills/maintain/scripts/run.py b/skills/maintain/scripts/run.py index 67ef2c84..619f5473 100644 --- a/skills/maintain/scripts/run.py +++ b/skills/maintain/scripts/run.py @@ -65,7 +65,7 @@ # Rewritten by scripts/prepare_release.sh to the tagged version. The connector # extra is deliberately NOT part of this pin: it is chosen at runtime (see # _resolve_connector), so a release artifact is connector-neutral. -DEX_CORE_VERSION = "1.12.2" +DEX_CORE_VERSION = "1.12.3" # Connector id -> packaging extra. The engine's connector ids and the pyproject # extras share names, so this is the identity set today. An unknown or unset diff --git a/skills/transform/SKILL.md b/skills/transform/SKILL.md index f19c5c7e..52c3e3c2 100644 --- a/skills/transform/SKILL.md +++ b/skills/transform/SKILL.md @@ -571,7 +571,9 @@ database, which is fine for model-only builds. never writes to source warehouse data. - Dev-target only. Prod-target execution is never initiated by dex. - Cost surfaced before any spend. A build that would spend requires explicit - confirmation and a session budget. + confirmation and a session budget. The cost guard in full, in the engine + repository: `references/cost-controls.md`; the PII policy that governs what + a seed may carry and what gets stamped into `meta`: `references/pii-policy.md`. - Propose, don't impose. Human edits to the project (SQL and semantic YAML) and to a native semantic document are authoritative; on conflict the engine surfaces a diff and asks rather than overwriting. diff --git a/skills/transform/scripts/run.py b/skills/transform/scripts/run.py index 67ef2c84..619f5473 100644 --- a/skills/transform/scripts/run.py +++ b/skills/transform/scripts/run.py @@ -65,7 +65,7 @@ # Rewritten by scripts/prepare_release.sh to the tagged version. The connector # extra is deliberately NOT part of this pin: it is chosen at runtime (see # _resolve_connector), so a release artifact is connector-neutral. -DEX_CORE_VERSION = "1.12.2" +DEX_CORE_VERSION = "1.12.3" # Connector id -> packaging extra. The engine's connector ids and the pyproject # extras share names, so this is the identity set today. An unknown or unset