From 4c45adca6726b061171cd4f4d690fdc8b41f4cd3 Mon Sep 17 00:00:00 2001 From: Marco Ciavarella Date: Wed, 16 Sep 2026 20:02:35 +0200 Subject: [PATCH 1/3] `explore profile` no longer offers composite keys that are artifacts of a near-unique column or of a continuous measure --- AGENTS.md | 6 +- CHANGELOG.md | 101 +++++ README.md | 11 +- packages/dex-core/README.md | 5 +- .../dex-core/src/exmergo_dex_core/cache.py | 76 +++- .../src/exmergo_dex_core/explore/commands.py | 44 +- .../src/exmergo_dex_core/explore/profile.py | 414 +++++++++++++++++- .../exmergo_dex_core/explore/relationships.py | 175 +++++++- .../src/exmergo_dex_core/explore/results.py | 3 + packages/dex-core/tests/demo/test_commands.py | 17 + packages/dex-core/tests/explore/conftest.py | 39 ++ .../dex-core/tests/explore/test_diagram.py | 39 ++ .../dex-core/tests/explore/test_explore.py | 48 +- .../dex-core/tests/explore/test_profile.py | 313 ++++++++++++- .../tests/explore/test_relationships.py | 5 +- packages/dex-core/tests/test_safety_spine.py | 96 ++++ references/bigquery.md | 9 + references/canonical-model.md | 12 +- references/command-contract.md | 29 +- references/methodology.md | 58 ++- references/ossie-walkthrough.md | 14 +- skills/explore/SKILL.md | 21 +- skills/explore/references/probe-playbook.md | 14 +- 23 files changed, 1474 insertions(+), 75 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 9b4142a7..4b9c83a3 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -76,9 +76,9 @@ credentials and no network. | `demo [path]` | generates a seeded local DuckDB warehouse (7 tables, 29,512 rows) plus a `.dex/config.yml` beside it, so a first run needs no warehouse, no credentials, and no network; both artifacts are listed in `data.created` and `data.next_steps` names the commands worth running next. The path is positional and resolves against the working directory, defaulting to `dex_demo.duckdb`; `--path` is refused here rather than honored, since everywhere else it names the warehouse dex *reads*. Create-only, with no `--confirm` that can talk past it: an existing file at the target is a refusal (`reason: guard`), a missing parent directory is a refusal (`reason: request`), no directories are ever created, and a `.dex/config.yml` at or above the target is left untouched with a warning rather than shadowed by a second one. The data is generated from a pinned seed, so the counts quoted in the docs are the counts a user sees, and it is deliberately flawed: a key that lost uniqueness to a double-loaded batch, a key mixing two id schemes, a join whose columns share a name and none of their values, an empty table, two columns whose declared type contradicts their content, and personal data alongside two designed false positives. Needs the `[duckdb]` extra and says so by name when it is absent (`reason: prerequisite`) | | `connect test` | capabilities, dialect, `read_only: true`; DuckDB takes `--path`, every warehouse connector takes repeatable `--scope` (a bare database on ClickHouse, whose identifiers are two-part `database.table`) (BigQuery also accepts its older `--project`/`--dataset`), never written to config; Snowflake and Databricks report the pinned warehouse and its credit or DBU rate; ClickHouse Cloud reports live replica memory and its derived compute-unit rate | | `explore inventory [--rank] [--limit N] [--all]` | ranked object summary (counts, sizes; no rows). `--rank` caps at 30 objects by default so the shortlist stays a shortlist on a large warehouse; `--limit` widens it, `--all` lifts the cap; both are no-ops without `--rank`, since the unranked list carries no order to cut from | -| `explore profile [--columns all]` | column profiles + PII flags (column, category, confidence) + candidate keys, grain, data-quality warnings; the verdict fields (grain, keys, data quality, row count) lead the serialized payload and `columns` trails, so a truncating harness cuts schema rather than the answer. By default each dataset's `columns` is summarized to the ones carrying a finding (PII, a non-zero null fraction, key membership, a reported value domain, or a mention in a `data_quality` note), with the rest counted in `elided_column_count`; `--columns all` restores every column. A per-column field that is null on every column shown is dropped from each of them and named once in `suppressed_fields` (always present, so an empty list means the columns shown carry their full shape), and each `value_domain` carries its `profile_value_domain_cap` most frequent values (`.dex/config.yml`, default 25) with the rest counted in its `elided`; both reduce the payload only, never the cached profile. `--use-project` lets a semantic model's declared primary entity override the heuristic grain (disagreements noted) | +| `explore profile [--columns all]` | column profiles + PII flags (column, category, confidence) + ranked candidate keys, grain, `key_evidence`, data-quality warnings; the verdict fields (grain, keys, key evidence, data quality, row count) lead the serialized payload and `columns` trails, so a truncating harness cuts schema rather than the answer. `candidate_keys` is ranked, tightest proven key first, and `key_evidence` says why for each one: a combination that is unique only because one of its members is unique on almost every row, or because a continuous measure completes it, is suppressed rather than reported, with its reason kept, since a caller who cannot tell the real key from the filler is worse served by several candidates than by one named defect and none. Where a column is unique on almost every row, `data_quality` says so with the ratio, the counts, and the exact number of rows that would have to be removed for it to be unique, because the finding is duplicates in the source rather than an absent key. By default each dataset's `columns` is summarized to the ones carrying a finding (PII, a non-zero null fraction, key membership, a reported value domain, or a mention in a `data_quality` note), with the rest counted in `elided_column_count`; `--columns all` restores every column. A per-column field that is null on every column shown is dropped from each of them and named once in `suppressed_fields` (always present, so an empty list means the columns shown carry their full shape), and each `value_domain` carries its `profile_value_domain_cap` most frequent values (`.dex/config.yml`, default 25) with the rest counted in its `elided`; both reduce the payload only, never the cached profile. `--use-project` lets a semantic model's declared primary entity override the heuristic grain (disagreements noted) | | `explore relationships [--verify] [--use-project] [--use-hosted-semantic-layer]` | inferred joins with confidences, plus notes on what inference examined; `--verify` measures each join with an aggregate overlap probe, declared and inferred alike, including the full ordered tuple of a composite key. `--use-project` folds local project and semantic-layer declarations in at confidence 1.0; native composite relationships retain every column pair and provenance. `--use-hosted-semantic-layer` independently authorizes a configured hosted catalog read, but a backend with no physical relations cannot add warehouse edges or exposure annotations. A measurement never revises a declared join's confidence, which stays at the 1.0 the project asserts; a declared join whose probe finds the parent largely missing is reported as a finding instead | -| `explore map [--detail] [--verify] [--use-project]` | writes/updates the `.dex/` map and returns it: the counts as before, plus `data.objects` (per top-ranked object: row count, detected grain, candidate key, the notable columns with the role that earned each one a place, PII flags as category and confidence, and data-quality findings) and `data.edges` (the join edges, shaped exactly as `explore relationships` returns them). Budgeted like `explore diagram`: 25 objects by rank, 12 columns per object, 40 edges, 5 findings per object, every cap binding in every mode and every elision counted in `notes` and in an `elided_*` field, so a truncated answer never reads as a complete one. `--detail` widens the selection to every column and to objects that were inventoried but never profiled, and lifts no cap; it is not `--full`, which decides how much gets scanned and therefore what the run costs. No column value ever appears: the cache holds min/max and value domains and this command does not read them. `--use-project` additionally applies declared grain, ranks metric-backing models higher, folds the semantic layer's declared entity graph into `data.edges`, and marks each object with `semantic_models`, the semantic models that sit on that relation. Empty there is an answer: a relation nothing in the layer reads is a different object from one several metrics are built on, and row counts and PII flags cannot tell them apart. Every object in view is rewritten whenever the layer was read, so a model dropped from the layer clears rather than leaving a stale claim, and a project with no compiled semantic layer contributes nothing here rather than erroring | +| `explore map [--detail] [--verify] [--use-project]` | writes/updates the `.dex/` map and returns it: the counts as before, plus `data.objects` (per top-ranked object: row count, detected grain, the best-ranked candidate key, the notable columns with the role that earned each one a place, PII flags as category and confidence, and data-quality findings) and `data.edges` (the join edges, shaped exactly as `explore relationships` returns them). Budgeted like `explore diagram`: 25 objects by rank, 12 columns per object, 40 edges, 5 findings per object, every cap binding in every mode and every elision counted in `notes` and in an `elided_*` field, so a truncated answer never reads as a complete one. `--detail` widens the selection to every column and to objects that were inventoried but never profiled, and lifts no cap; it is not `--full`, which decides how much gets scanned and therefore what the run costs. No column value ever appears: the cache holds min/max and value domains and this command does not read them. `--use-project` additionally applies declared grain, ranks metric-backing models higher, folds the semantic layer's declared entity graph into `data.edges`, and marks each object with `semantic_models`, the semantic models that sit on that relation. Empty there is an answer: a relation nothing in the layer reads is a different object from one several metrics are built on, and row counts and PII flags cannot tell them apart. Every object in view is rewritten whenever the layer was read, so a model dropped from the layer clears rather than leaving a stale claim, and a project with no compiled semantic layer contributes nothing here rather than erroring | | `explore diagram [--full]` | the `.dex/` map serialized as a Mermaid `erDiagram` under `data.mermaid`, plus an `entities` legend mapping each entity name back to its fully-qualified identifier. Free and connectionless: it reads the cache and never opens the warehouse, so it needs no credential and cannot spend. Declared joins are solid, inferred joins dotted, and a cardinality is drawn only where the cache proved it (an unverified inference never claims "exactly one"). A solid edge whose label names a semantic entity is a join the semantic layer declares; the cardinality rule is unchanged for it, so a primary entity is the layer's claim and still buys no "exactly one" the cache has not proven. The default draws profiled objects that participate in a join, with their grain, key, join, and PII-flagged columns; `--full` widens to every eligible object and column. An entity cap always binds and every elision is counted in `notes`. No column value ever appears; PII renders as category and confidence. dex writes no file: reproduce the string in a fenced ```mermaid block, or save it yourself | | `explore query "" [more...]` | runs agent-authored SELECTs through the query firewall: columnar, capped results; `data.shape` is `columnar`, making the `columns`/`types`/`cells` layout discoverable from the envelope itself. Values only come from profiled columns whose PII flag is absent or below the 0.5 blocking threshold (sub-threshold projections warn in the envelope); the FROM clause may unnest JSON/array columns in the connector's native idiom (UNNEST, LATERAL FLATTEN, LATERAL VIEW EXPLODE, set-returning functions, PartiQL) when the unnested value derives from a queried table's column, with the outputs inheriting that column's flags. The positional is variadic, so a chain of questions is one call: each argument is one statement, adjudicated, executed, and ledgered on its own, and `--sql-file ` reads a larger batch from a file (one statement per line, or semicolon-separated). Several statements in one string is still refused, so batching never widens what a call may do. One statement returns the envelope described here; two or more return `data.results`, one entry per statement carrying its own `shape` discriminator and this same `columns`/`types`/`cells`/`row_count`/`truncated` layout plus its own `status` (`ok`, `refused`, `failed`, `skipped`) and `error`, so a refusal on the third does not discard the first two, and the envelope's own status is `error` whenever any statement failed. `query.max_payload_bytes` is the budget for the whole call rather than for one statement, and `query.max_statements` (default 10) refuses an oversized batch. An object a statement names that the connection has but the cache cannot adjudicate (never profiled, inventoried without column detail, or profiled against a column signature the warehouse has since changed) is profiled first and the statement then runs, with a warning naming what was profiled and `data.profiled_on_demand` listing it; that profile is a full one, so the flags governing the query are the flags a deliberate `explore profile` would have produced. On a metered connector it is priced, not implied: one handshake covers the profiles and every statement together, itemized per table and per statement, and the objects a whole batch needs are scanned once rather than once per statement. An object the connection does not have refuses only the statements that named it, naming the connection rather than the cache. `--no-auto-profile` (or `auto_profile: false` in `.dex/config.yml`) restores the strict prerequisite, and on that path nothing opens a connection before the firewall has spoken | +| `explore query "