Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"name": "dex",
"description": "Analytics engineering for Claude Code and any agent: data warehouse exploration, dbt transformation and semantic modeling, and schema-drift maintenance on dbt.",
"version": "1.12.2",
"version": "1.12.3",
"author": {
"name": "Exmergo, Inc.",
"email": "support@exmergo.com",
Expand Down
146 changes: 52 additions & 94 deletions AGENTS.md

Large diffs are not rendered by default.

134 changes: 134 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,140 @@ tag releases both in lockstep, so entries below are keyed by the engine version.

## [Unreleased]

### Fixed

- **`explore profile` no longer offers composite keys that are artifacts of a
near-unique column or of a continuous measure, ranks the ones it does report,
and says why for each** ([#292]). On a 2,037-row orders table where `order_id`
held 1,927 distinct values, the profile returned five candidate keys and
elected a grain out of them: `order_id, customer_id`, and then `CREATED_AT`,
`SUBTOTAL`, `GRAND_TOTAL` and `UPDATED_AT` each paired with `order_id`. Four
of the five were one fact wearing four hats, that `order_id` is unique on all
but 110 rows so any wider column completes it. `(SUBTOTAL, order_id)` is not
a grain: it is the observation that two rows sharing an order id happened to
differ in their subtotal. The list carried no order, so a caller could not
tell the real key from the filler, and the one thing worth saying, that this
table has duplicate order ids, was the one thing the profile did not say.

The same bug shipped in the demo warehouse, where `order_items` reported a
grain of `(unit_price, order_id)`: a `DECIMAL(10,2)` money column paired with
a foreign key, on a table whose actual story is 1,000 duplicate
`order_item_id` rows from a double-loaded batch.

Two exclusions, both applied before a pair is priced. A pair whose member is a
**continuous measure** is dropped, because a measurement's cardinality grows
with the table and it completes a partner by arithmetic rather than by
meaning. A pair **anchored on a column already unique on almost every row** is
dropped too, unless its partner has a domain bounded by the fanout rather than
by the table: a few duplicate order ids separated by a three-value line number
really is that grain, where any wider partner separates them by accident. A
near-unique timestamp gets no such exception, since a per-row event time is
not an entity whose rows a position column enumerates. Neither rule applies
below a fixed row floor, where every column looks near-unique and every
measure looks continuous, so a small table has everything asked as before.

The measure test reads a declared type with an explicit nonzero scale, and a
measure-name vocabulary above a modest cardinality bar. The vocabulary is not
decoration: Snowflake's `SHOW COLUMNS` renders every `NUMBER` as the bare
token `FIXED` with the scale dropped, and BigQuery `NUMERIC` carries no scale
either, so on those connectors the type test goes silent by design and the
name is the only signal left. A name never decides alone, because `quantity`,
`amount` and `total` are legitimately low-cardinality members of real fact
grains.

Where nothing survives, `grain` is `null` and the profile names the near-unique
column with its counts, including how many rows would have to be removed for
it to be unique (`order_id is not unique: 1927 distinct over 2037 rows (110
rows would have to be removed for it to be unique, so it is unique for 94.6%
of rows)`). That figure needs no new measurement: it is the non-null row count
less the distinct count. It is phrased as rows to remove rather than as "110
duplicate order ids" because the arithmetic counts surplus rows and not values
that repeat, and the two differ. Two long-standing imprecisions in that
sentence are fixed with it: the surplus was computed against the total row
count rather than the non-null count, overstating it on a nullable column, and
it carried a `~` even when derived from two exact numbers. The marker now
follows the arithmetic, and a percentage that would round to `100.0%` prints
`>99.9%` rather than contradicting the sentence it sits in.

Suppression is not silence. One new field, `key_evidence`, carries an entry
per combination the profile considered, each with its `columns`, a `status` of
`reported` or `suppressed`, and the `reason` in the profile's own words; the
reported entries are in the same order as `candidate_keys`, which is the
invariant that keeps the two from drifting. `data_quality` states the
suppression too, and deliberately names only the anchor: a column named in a
note is kept in the serialized `columns`, so naming the filler would drag four
columns back into a payload that already elides them correctly.

Field-visible beyond `explore profile`, because five things read
`candidate_keys`. `explore diagram` stopped marking a money column `PK`, and
`order_id` on the demo's `order_items` reads `FK` rather than `PK`.
`explore map`'s best-ranked `candidate_key` and its column roles no longer
point at filler, while a column named by a suppressed entry stays in the
notable set with no role, so a map that reports duplicates in a column still
lists it. `transform plan --scaffold` reads `candidate_keys[0]` to decide
which columns get tests, so a scaffolded `schema.yml` no longer puts
`not_null` on a money column. And `maintain grain` re-probes every composite
in `candidate_keys` on every run, billed on metered connectors, so four
artifacts per table stop being a recurring charge. On the demo the pairs
actually probed fall from five to two. A declared grain still overrides all of
it, and both declared-grain notes now state their origin even when there was
no heuristic grain to disagree with, so a declaration that fills a gap says so
rather than applying silently.

`CACHE_SCHEMA_VERSION` moves to 4. An older engine reads a version-4 cache
fine, since an unknown key is ignored; the direction that breaks is a current
engine reading a version-3 one, where a suppressed combination still reads as
a ranked candidate and an empty `key_evidence` is indistinguishable from a run
that suppressed nothing. Rather than warn about it, the profile freshness gate
now treats a pre-4 profile as stale, so the first `explore profile` or
`explore map` after upgrading re-scans and the cache heals itself. The
`explore query` degradation on an older cache is unchanged and still refuses
nothing.

Deliberately unchanged: no flag. The issue's position is that a spurious key is
worse than nothing, so there is no opt-out to talk past it, and none is needed
since `key_evidence` hands back every suppressed combination and its reason in
the same payload. No `probed` boolean either, because the probe's existing
budget notes already say when it did not run, and duplicating them into a
field would be growth for nothing. `key_evidence` is not in `explore map`'s
payload, which is budgeted per object; the full ranking belongs to the command
whose subject is one relation in full. The demo warehouse's data is untouched,
since its determinism is a contract and it is the reproduction rather than
something to fix.

### Changed

- **The PII policy and the cost guard each have one document, and every other
document links to it.** Both guardrails cut across every connector, command and
backend, so neither had an owner: the PII blocking threshold was stated in five
places in four wordings, one paragraph about auto-profile pricing was copied
verbatim into five connector references, the `--confirm` handshake was fully
restated in roughly thirteen files, and the hosted dbt Cloud exception appeared
ten times across seven. `references/pii-policy.md` and
`references/cost-controls.md` now own the policy, the constants and the
end-to-end flow. Connector references keep their own cost models, which are
genuinely per-warehouse, and `references/storage.md` keeps the store protocol a
backend implements, which is a different reader's question. `AGENTS.md`
guardrails 4 and 6 state the invariant and point.

Two behaviors that lived only in this changelog are now documented: the
`pii_overrides` mismatch warning that `explore profile` and `explore map` both
emit, and the `column_name` plus `scope` pattern form of an override entry.

The three skills keep their copies, because `npx skills add exmergo/dex`
installs each one standalone with no engine repository to link into. What they
no longer carry is the constant: a skill states the consequence ("below the
blocking threshold it projects with a warning") and names the file that states
the number, so a skill can go stale on wording but not on the threshold.

`packages/dex-core/tests/test_docs_policy.py` holds this in place. It asserts
that any document stating the threshold states the engine's
`PII_BLOCK_CONFIDENCE`, that the set of files allowed to state it has not grown,
that the shared cost prose lives in one file, that every connector's
`session_ceiling` example keeps one wording, and that relative links between
documents resolve. `packages/dex-core/README.md` is deliberately untouched: the
PyPI package ships without `references/`, so that file stays self-contained.

### Added

- **`transform test --mutate <model>` measures whether a model's tests would
Expand Down
14 changes: 11 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -109,16 +109,21 @@ from a pinned seed, so what you see is what is written here.

It is seeded to be **realistically broken**, because a first run that reports a clean
bill of health teaches you nothing. `explore map` flags 6 columns as personal data,
infers 5 joins, and reports 5 data-quality findings. Then:
infers 5 joins, and reports 6 data-quality findings. Then:

```
dex explore profile order_items products
dex explore relationships --verify
dex explore query "select email from customers"
```

- **A broken grain.** `order_item_id is not unique: 13000 distinct over 14000 rows`,
because a batch was loaded twice. Any join on it silently fans out.
- **A broken grain, reported as duplicates rather than as a missing key.**
`order_item_id is not unique: 13000 distinct over 14000 rows (1000 rows would have
to be removed for it to be unique, so it is unique for 92.9% of rows)`, because a
batch was loaded twice. Any join on it silently fans out. The grain comes back
unknown rather than as one of the several column pairs that are technically unique
here only because `order_item_id` almost is; those are in `key_evidence` with the
reason each was suppressed.
- **A key that mixes id schemes.** `sku` is `90% numeric, 10% 32-character
hexadecimal (md5-shaped)`, from a merged catalogue. Cast it to a number and you
drop 10% of your rows without an error.
Expand Down Expand Up @@ -335,6 +340,9 @@ More info in the package's [`README.md`](packages/dex-core/README.md)
command contract, the source of truth and the `.dex/` cache, the semantic layer
and Ossie compatibility, the project, storage and host-integration seams,
methodology, and evaluation.
- The two guardrails that cut across all of it:
[`references/pii-policy.md`](references/pii-policy.md) and
[`references/cost-controls.md`](references/cost-controls.md).

## Contributing

Expand Down
5 changes: 3 additions & 2 deletions packages/dex-core/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -74,7 +74,7 @@ uniqueness to a double-loaded batch, a key mixing two id schemes from a merged
catalogue, a join whose columns share a name and none of their values, a table an
interrupted load left empty, two columns whose declared type contradicts their
content, and personal data alongside two deliberate false positives. `explore map`
finds 6 PII columns, 5 joins, and 5 data-quality findings; `explore query "select
finds 6 PII columns, 5 joins, and 6 data-quality findings; `explore query "select
email from customers"` is refused, and the same count over the same column is not.

The generation is create-only: it writes a new file and refuses rather than replace
Expand Down Expand Up @@ -246,7 +246,8 @@ on its own path, never through a connector, which is what keeps the read-only ru
true everywhere else.

`explore`: ranks what matters in an unfamiliar warehouse, profiles columns
selectively, flags PII, surfaces grain and data-quality warnings, infers joins
selectively, flags PII, surfaces grain and data-quality warnings with the ranked
keys and the reasoning behind each, infers joins
and verifies them with overlap probes (`--verify`), and executes agent-authored
ad-hoc SELECTs behind a PII-aware query firewall (`explore query`, which takes
several statements per call, or a `--sql-file`, and adjudicates each on its own),
Expand Down
76 changes: 71 additions & 5 deletions packages/dex-core/src/exmergo_dex_core/cache.py
Original file line number Diff line number Diff line change
Expand Up @@ -15,11 +15,21 @@

import re
from enum import Enum
from typing import Literal

from pydantic import BaseModel, Field

# Bump when the stored cache shape changes in a way old readers cannot handle.
CACHE_SCHEMA_VERSION = 3
#
# 4 added `Dataset.key_evidence` and, with it, changed what `candidate_keys`
# means: it was every combination measured unique, and it is now every one that
# survived artifact suppression. An older reader handed a version-4 cache is
# fine, since an unknown key is ignored. The direction that breaks is a current
# reader handed a version-3 one, where a suppressed combination still reads as a
# ranked candidate and an empty `key_evidence` is indistinguishable from "this
# run suppressed nothing". So a pre-4 profile is treated as stale rather than
# reused (see `_split_fresh_stale`), and the cache heals on the next profile.
CACHE_SCHEMA_VERSION = 4


class PIICategory(str, Enum):
Expand Down Expand Up @@ -70,6 +80,28 @@ class ValueDomain(BaseModel):
elided: int = 0


class KeyEvidence(BaseModel):
"""Why one column combination is, or is not, reported as a key.

``candidate_keys`` is ranked but says nothing about why, and a caller who
cannot tell a real key from filler is worse served by several candidates
than by one named defect and none. This is where the reasoning lives.

``status`` is the only discriminator a consumer needs. ``reason`` is prose
rather than a code, because the set of causes is open and a new one should
not need a contract change and an exhaustive match in every consumer; the
sentences are engine-authored templates over column names and counts, never
a customer string and never a column value.

The reported entries appear in the same order as ``candidate_keys``, which
is what keeps the two fields from drifting; the suppressed ones follow.
"""

columns: list[str]
status: Literal["reported", "suppressed"]
reason: str


class ColumnProfile(BaseModel):
"""Aggregate-derived understanding of one column, built from SQL aggregates and
never from raw rows in context."""
Expand Down Expand Up @@ -124,6 +156,11 @@ class Dataset(BaseModel):
#: annotation pass recomputes from column stats) so the proof survives
#: re-annotation and its provenance stays distinct from derived signals.
composite_keys: list[list[str]] = Field(default_factory=list)
#: Why each combination is or is not a key, reported entries first and in
#: ``candidate_keys`` order. Empty on a profile written before this existed;
#: the cache schema version is what tells those apart from a run that
#: suppressed nothing.
key_evidence: list[KeyEvidence] = Field(default_factory=list)
rank_score: float | None = None
data_quality: list[str] = Field(default_factory=list)
profiled_at: str | None = None
Expand Down Expand Up @@ -152,9 +189,19 @@ def notable_columns(
reporting it is the enumeration dex exists not to do.

Returns each kept column paired with its role (``"grain"``, ``"key"``,
``"join"``, or ``None`` for a column kept only because it is flagged) and
the number dropped, so a caller can always say how much it did not show.
``everything`` keeps every column and still assigns the roles.
``"join"``, or ``None`` for a column kept only because it is flagged or
because it almost keys the table) and the number dropped, so a caller
can always say how much it did not show. ``everything`` keeps every
column and still assigns the roles.

A column named by a **suppressed** ``key_evidence`` entry is kept, with
no role, because it is the subject of the grain verdict on exactly the
tables that have no grain: a map reporting "order_item_id is not unique"
while omitting ``order_item_id`` from the columns would be answering
past the question. It gets no role because it is not a key; claiming one
is the thing the suppression exists to stop. Only suppressed entries are
read, and each names a single anchor column rather than the partners it
was paired with, so this cannot pull filler columns in.

``join_columns`` is supplied rather than derived, because which joins are
in view is the caller's question: a diagram marks FK against the edges it
Expand All @@ -168,6 +215,12 @@ def notable_columns(
keyed |= {c.lower() for group in self.composite_keys for c in group}
grain = {c.lower() for c in (self.grain or [])}
joins = {c.lower() for c in (join_columns or ())}
near_keys = {
c.lower()
for entry in self.key_evidence
if entry.status == "suppressed"
for c in entry.columns
}

kept: list[tuple[ColumnProfile, str | None]] = []
dropped = 0
Expand All @@ -181,7 +234,12 @@ def notable_columns(
role = "key"
else:
role = None
if role is None and column.pii is None and not everything:
if (
role is None
and column.pii is None
and lowered not in near_keys
and not everything
):
dropped += 1
continue
kept.append((column, role))
Expand Down Expand Up @@ -214,6 +272,14 @@ def columns_with_findings(
predicate exists so a real finding is never the reason it gets
truncated away, not to be a precise finding-to-column index.
``everything`` keeps every column.

Needs no ``key_evidence`` check of its own: a suppressed entry's
anchor is a column with duplicates, so it is already named in one of
this dataset's ``data_quality`` sentences and kept by the last rule
below. That is also why the suppression prose names only the anchor and
never the partners it was paired with; naming those would pull four
columns back in through the same rule, which is exactly the payload
bloat the finding summary removed.
"""

keyed = {c.lower() for group in self.candidate_keys for c in group}
Expand Down
Loading
Loading