Skip to content

feat(vintage): disagg/TFF canonicalisers, frozen-year tripwire, flow decomposition - #83

Merged
mspinola merged 7 commits into
mainfrom
claude/cot-flow-decomposition
Jul 31, 2026
Merged

mspinola merged 7 commits into
mainfrom
claude/cot-flow-decomposition

Conversation

@mspinola

@mspinola mspinola commented Jul 31, 2026 •

Copy link
Copy Markdown
Owner

Adversarial review, then six fixes (9a34750). Reviewed by an agent given the spec
and diff and no prior context, because every earlier review was by the author and
shared its blind spots. It found six confirmed defects, four in code written the same
day, all reproduced before fixing and each now covered by a regression test. It also
confirmed hash purity empirically rather than by reading, which is the property the whole
design rests on. Full write-up in docs/design/cot_vintage.md §9.

# Defect Consequence
1 A fetch failure poisoned the dedupe comparison One network blip produced a false restatement alert and re-retained megabytes already on disk
2 An absolute local_path was re-rooted under the store On any replica using COTDATA_VINTAGE_ROOT (where it is mandatory), one run drained the whole backlog to failed with no way to reset
3 A bare ingest replayed every snapshot with observed_at = now Forged reversed revisions and inflated age_days, corrupting the primary artifact
4 schedule backfill bypassed the write lock Dropped a concurrent ingest's rows, leaving a revision asserting a change to a value no longer in observations/
5 Every flat week was labelled long_liquidation 11.1% of real transitions; made it the modal state with a third of its bucket being weeks where nothing happened
6 validate() accepted duplicate natural keys and had no null-rate band Append order decided stored values; a changed value format would coerce to nulls and be read as revisions

Not fixed, recorded as a real gap: the announced release-date tier is unreachable in
production and acceptance criterion 5 is unmet. Nothing writes source="announced"
rows, so the Oct-Dec 2025 backlog weeks resolve only to derived. The plumbing and
precedence are correct; what is missing is an extractor pulling a
(report_date, release_date) pair out of free-text announcement prose, and writing that
against no corpus of real announcements would be guessing. A release date resolved by a
guess is worse than a derived one, because it carries a flag claiming it was announced.


Four commits, in the order they were asked for. Everything below was measured against the real store, and several measurements changed the plan.

Commit What
1d8e1c9 frozen-year tripwire, flow decomposition (§6.4), corrected OI validation
38d6f90 step 2 proposal (contract master + normalisation), no step-2 code
b1ade42 Disaggregated and TFF canonicalisers

Disaggregated and TFF canonicalisers

The step 2 proposal identified this as the real prerequisite rather than step 2 itself. Every engine in the module spec keys on Managed Money and Leveraged Funds, which exist only in these two reports, and ingest wired Legacy only. asof() now returns both point-in-time.

Verified against the real first production capture (2026-07-31 01:15Z), not fixtures:

Report Canonical rows, 2026 Weeks Re-ingest
Legacy 31,041 10,347 0 rows
Disaggregated 39,235 7,847 0 rows
TFF 12,500 2,500 0 rows

All three ingest in about 5 seconds, and re-ingesting writes nothing, which is the strongest available confirmation that the permanent row_sha256 artifact is stable across the new code path.

These two reports populate three fields Legacy never does: per-category spreading, per-category trader counts, and CR4/CR8 net concentration. So the zero-sum identity closes completely for Disaggregated:

Report Exact Within tolerance oi_gap
Legacy 10,321 / 10,347 10,347 / 10,347 never zero (the uncaptured spreading)
Disaggregated 7,847 / 7,847 7,847 / 7,847 zero everywhere
TFF 2,463 / 2,500 2,500 / 2,500 zero on 98.6%

Disaggregated closing exactly, with a zero gap, is the confirming counterpart to the Legacy finding: that gap really is the missing spreading column and nothing else.

The rounding tolerance is derived, not fitted

Every off-by-one-or-two row in both Legacy and TFF falls in exactly three markets:

Code Name Legacy TFF Worst
13874+ S&P 500 Consolidated 12 17 2
20974+ NASDAQ-100 Consolidated 11 15 1
12460+ DJIA Consolidated 3 5 1

The + suffix is CFTC's own marker for a Consolidated contract, which aggregates several contract sizes onto a common unit and so involves a division. Summing n independently rounded category figures admits at most n contracts of error, which is what rounding_tolerance() returns. Without it, 48 off-by-one warnings fire on the 2026 files alone, the same cry-wolf rate the OI check was corrected for in the first commit.

Four implementation points

  • CFTC's header has a typo, and it is load-bearing. Swap__Positions_Short_All and _Spread_All carry a double underscore; _Long_All does not. Both spellings resolve, so the day CFTC fixes it is not the day Swap Dealer positions ingest as nulls.
  • An unresolvable column raises. Silent nulls would be written as real observations, and the next genuine value recorded as a revision that never happened, polluting the exact artifact this subsystem exists to produce.
  • Trader counts arrive as strings, because CFTC writes . for a suppressed count (3,578 of 7,847 Managed Money long counts on the 2026 file, so "null" is routine, not an error). Every value field is coerced so a marker never reaches row_sha256 as a literal string.
  • combined is now read from the file's FutOnly_or_Combined column, and a file mixing the two is refused. Constant-FutOnly today, so nothing changes, but adding the combined files becomes purely a fetch-list change.

canonicalize_legacy is deliberately not routed through the new shared helper. Its output feeds row_sha256 over rows already stored in production, and rewriting that path to save a little duplication risks registering every stored row as revised at once.

Fixed in passing: raw paths were unreadable off the producer

snapshots.json records local_path as written by the capturing machine, and the producer is Windows, so the real store carries vintage\raw\annual_zip\2026\....zip. On macOS or Linux that is one filename containing backslashes, so ingest on either replica failed with "no such file" on a file plainly sitting there. Normalised on read, so every snapshot already recorded on the producer stays readable.


Three things, all measured against the real store rather than assumed.

1. The frozen-year tripwire is now automatic

The rolling two-year regeneration window (measured in #78) means CFTC re-serves the prior year every week, byte-identical. That is the one place a content check on closed data is free, and it is the only automated detector for the failure mode this subsystem exists to guard against (July 2008: reports restated back to 2007-07-03).

The standing decision was "run --all monthly or quarterly". A tripwire that only fires when somebody remembers to run it by hand is an intention, not a detector, so the prior year moves into the default fetch set: one ~7 MB transfer per week (the other six days 304) and zero bytes on disk, since the content dedupes.

Regime Files Expected weekly outcome Violation
churn current year new bytes nothing, that is the data arriving
frozen_in_window prior year unchanged bytes (deduped) new sha (restatement), or 304/failure (check went blind)
frozen_out_of_window 2+ years back 304 not-modified a 200 with new bytes

Two alert shapes, deliberately triggered differently:

  • Content changed on a frozen year: always alerts. A discrete, diffable event; two restatements in consecutive weeks are two things worth knowing. January says so explicitly, since year-end finalisation is the one benign way it fires.
  • The detector went blind (frozen-in-window 304 or failure): alerts on the TRANSITION only. A standing condition, so level-triggering would fire every day forever, and an alert that never clears is one that gets ignored.

fetch still exits zero on purpose: the Windows wrapper aborts the run on a non-zero fetch, which would skip the ingest that turns a restatement into readable revision rows. ingest re-raises it and already writes the REVISIONS_<date>.txt marker.

2. Flow decomposition (module spec section 6.4)

vintage_flow.py plus cotdata-vintage flow. Weekly dLong vs dShort per market/category into new_longs / short_covering / new_shorts / long_liquidation. No prices, no contract master, so it is a clean smoke test of the canonical schema.

Classification is by dominant leg, because the spec's "~0" leg never happens in real data. Parameter-free by default, with an optional min_frac_oi dead zone that stays off, since any value for it is a judgement in the same class as the fragility weights.

Read-side only: writes no store domain, changes no producer/consumer contract. ADR-0007 narrows cotdata by instrument domain, not derived-vs-raw, and ADR-0008 settles that COT-derived-from-COT stays inside that boundary. The positioning engine (prices, contract master, configured weights) stays in crowdmon.

Worked example, gold non-commercial, week ending 2026-06-02: dNet +21,760, which reads as heavy fresh buying. Decomposed it is dShort -16,368 against dLong +5,392 with OI down 27,437. Short covering: a rally with a finite fuel supply. That distinction is the whole point of section 6.4, and it showed up on the first real week anyone looked at.

Swept over the whole store

95 Legacy markets, 1986 to 2026-07-21:

Measure Result
Zero-sum identity (sum of long across categories == sum of short) 149,412 / 149,412 weeks balanced
Flow rows 447,951
Exceptions 0
validate() warnings 0 (after the fix below)

Finding: non-commercial SPREADING is not captured, and never was

NonComm_Positions_Spread_All is absent from providers/cftc.py's TARGET_COLS, so every side total falls short of OI by exactly the spreading. Gold on 2026-07-21: OI 383,368, both sides 351,385, gap 31,983 (8%). 64 of 95 markets have at least one week where the gap is zero, which is the confirming case.

Net positioning is unaffected (spreading is long and short in equal measure). Anything denominated as a share of open interest is not, and that is section 5.2 step 2, the next build step.

Not fixed here: adding the column changes providers/cftc.py output, which breaks the current/ byte-identical guarantee and the golden baseline that exists to protect it. It is its own change, not a drive-by.

Finding: COT was FORTNIGHTLY until 1992-10-13

415,908 of 447,951 intervals are 7 days; the 15-day (9,057) and 14-day (5,775) intervals are almost all pre-October-1992. days_elapsed is emitted as a column rather than assumed, so callers filter on it instead of discovering it in a result. This is section 5.3's holiday hazard showing up in a second place: not only in the release date, but in the differencing interval.

3. Corrected validate()'s open-interest warning

It compared one category's long + short + spread against total OI, which is not a bound: the two sides of a market are counted separately, so a category holding 80k long and 294k short in a 383k-OI market sums to 374k with nothing wrong. Measured, it fired on 811 of 5,778 gold rows (14%), and a soft warning at that rate is one nobody reads.

Replaced with the invariant that does hold, per side summed across categories. Store-wide warnings went from thousands to zero. The module spec asserted the wrong form too, and is amended in the same pass per the working agreement on specs.

Verification

  • 226 tests pass (16 new), ruff check src tests scripts clean
  • Whole-store sweep above, run live against ~/code/cotdata_store
  • cotdata-vintage flow --market 088691 --source current end to end

Generated with Claude Code

mspinola and others added 4 commits July 30, 2026 22:36
… validation

Three things, all measured against the real store rather than assumed.

1. THE FROZEN-YEAR TRIPWIRE IS NOW AUTOMATIC

The rolling two-year regeneration window (measured in #78) means CFTC re-serves
the PRIOR year every week, byte-identical. That is the one place a content check
on closed data is free, and it is the only automated detector for the failure
mode this subsystem exists to guard against (July 2008: reports restated back to
2007-07-03).

The standing decision was "run --all monthly or quarterly". A tripwire that only
fires when somebody remembers to run it by hand is an intention, not a detector,
so the prior year moves into the DEFAULT fetch set: one ~7 MB transfer per week
(the other six days 304) and zero bytes on disk, since the content dedupes.

Three regimes now carry different expectations, recorded on every snapshot as
expectation / outcome / tripwire_alert:

  churn                 current year, new bytes are the data arriving
  frozen_in_window      prior year, expect exactly "unchanged bytes (deduped)"
  frozen_out_of_window  older, expect 304; content is never verified there

Two alert shapes, deliberately triggered differently. Content changed on a frozen
year always alerts: it is a discrete, diffable event. The detector going blind (a
frozen-in-window 304 or fetch failure) alerts on the TRANSITION only, because it
is a standing condition and a level-triggered alert would fire daily forever.

fetch still exits zero on purpose: the Windows wrapper aborts the run on a
non-zero fetch, which would skip the ingest that turns a restatement into
readable revision rows. ingest re-raises it and already writes the marker file.

2. FLOW DECOMPOSITION (module spec section 6.4)

vintage_flow.py plus `cotdata-vintage flow`. Weekly dLong vs dShort per
market/category into new_longs / short_covering / new_shorts / long_liquidation.
No prices, no contract master, so it is a clean smoke test of the canonical
schema. Classification is by DOMINANT leg, because the spec's "~0" leg never
happens in real data; parameter-free by default, with an optional min_frac_oi
dead zone that stays off since any value for it is a judgement.

Read-side only: writes no store domain, changes no producer/consumer contract.
ADR-0007 narrows cotdata by instrument domain, not derived-vs-raw, and ADR-0008
settles that COT-derived-from-COT stays inside that boundary. The positioning
engine (prices, contract master, configured weights) stays in crowdmon.

Swept over all 95 Legacy markets, 1986 to 2026: 149,412 of 149,412 weeks satisfy
the zero-sum identity (sum of long across categories == sum of short), 447,951
flow rows, zero exceptions. Two findings recorded in the docs:

  - NonComm_Positions_Spread_All is not captured and never was, so every side
    total falls short of OI by exactly the spreading (gold: 31,983 of 383,368,
    8%). Net positioning is unaffected; anything denominated as a SHARE of open
    interest is not. Not fixed here: it would change providers/cftc.py output and
    break the current/ byte-identical guarantee, so it is its own change.
  - COT was FORTNIGHTLY until 1992-10-13. days_elapsed is emitted as a column
    instead of assuming seven, so callers filter rather than discover it later.

3. CORRECTED validate()'s OPEN-INTEREST WARNING

It compared one category's long+short+spread against total OI, which is not a
bound: the two sides are counted separately. It fired on 811 of 5,778 gold rows
(14%), pure noise. Replaced with the invariant that holds, per side summed across
categories. Store-wide warnings went from thousands to zero. The module spec
asserted the wrong form too, and is amended in the same pass per the working
agreement on specs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Proposal only, no step-2 code. Three measurements against the real store
change what step 2 can be:

1. Step 2 does not belong in cotdata. Normalisation is by definition a
   joiner (COT x prices x contract specs), and ADR-0007's whole argument
   was that almost nothing joins the domains. It belongs in the consumer,
   which means creating crowdmon-futures as a sibling.

2. The registry declares 41 disagg + 8 TFF and ZERO legacy symbols, and
   the module spec's engines all key on Managed Money / Leveraged Funds,
   which exist only in those two reports. Vintage ingest wires Legacy
   only, so a point-in-time Managed Money series does not exist yet.
   That is the real step-2 prerequisite, and it is a cotdata change.

3. Notional off back-adjusted prices is wrong by +294% (gold 2002) and
   +257% (crude 2004), and crude's back-adjusted series reaches -27.52,
   which is not a price. The error is EXACTLY ZERO at the present date
   and grows monotonically backwards, so it passes every spot check
   anyone would run while corrupting the whole evaluation history.
   Volatility must come from back-adjusted returns, so the two factors
   of net_notional x sigma come from two different price series.

Also measured: all 47 contract specs are USD (no FX layer needed),
unadjusted prices exist for all 47, and the COT-to-specs-to-price join
covers 42 of 95 Legacy markets (44%).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The step 2 proposal identified this as the real prerequisite. Every engine
in the module spec keys on Managed Money and Leveraged Funds, which exist
only in these two reports, and ingest wired Legacy only. asof() now returns
both point-in-time.

Verified against the REAL first production capture (2026-07-31 01:15Z), not
fixtures: 39,235 Disaggregated and 12,500 TFF canonical rows for 2026, all
three reports ingesting in about 5 seconds, and re-ingest writing zero rows.

These two reports populate three fields Legacy never does: per-category
spreading, per-category trader counts, and CR4/CR8 net concentration. So
the zero-sum identity closes COMPLETELY for Disaggregated (7,847/7,847
weeks, oi_gap zero everywhere), which is the confirming counterpart to the
Legacy finding that its gap really is the uncaptured spreading column.

ROUNDING TOLERANCE. The residual imbalance is CFTC's own and fully
localised: every off-by-one-or-two row in both Legacy and TFF falls in
exactly three markets, S&P 500 / NASDAQ-100 / DJIA Consolidated, whose
codes carry CFTC's own '+' suffix for a unit-converted aggregate. Tolerance
is derived from that mechanism rather than fitted: summing n independently
rounded category figures admits at most n contracts of error. Without it,
48 off-by-one warnings fire on the 2026 files alone, the same cry-wolf rate
the OI check was corrected for in the previous commit.

Implementation points that will otherwise cost someone an afternoon:

  - CFTC's header has a typo. Swap__Positions_Short_All and _Spread_All
    carry a double underscore, _Long_All does not. Both spellings resolve,
    so the day CFTC fixes it is not the day Swap Dealer ingests as nulls.
  - An unresolvable column RAISES. Silent nulls would be written as real
    observations, and the next genuine value recorded as a revision that
    never happened, polluting the exact artifact this subsystem produces.
  - Trader counts arrive as strings because CFTC writes '.' for a
    suppressed count (3,578 of 7,847 Managed Money long counts on the 2026
    file). Every value field is coerced so a marker never reaches
    row_sha256 as a literal string.
  - combined is now READ from the file's FutOnly_or_Combined column, and a
    file mixing the two is refused. Constant-FutOnly today, so nothing
    changes, but adding the combined files becomes a fetch-list change.
  - canonicalize_legacy is deliberately NOT routed through the new shared
    helper. Its output feeds row_sha256 over rows already stored in
    production; the duplication is cheaper than risking a mass revision.

Also fixes a real cross-platform bug found in the production snapshot
index: local_path is written by the Windows producer with backslashes, so
ingest on either replica failed with "no such file" on a file plainly
present. Normalised on read, so snapshots already recorded stay readable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@mspinola mspinola changed the title feat(vintage): frozen-year tripwire, flow decomposition, corrected OI validation feat(vintage): disagg/TFF canonicalisers, frozen-year tripwire, flow decomposition Jul 31, 2026
mspinola and others added 3 commits July 30, 2026 23:19
… finding

Four questions, all answered by reading ~/code/trading_workspace/cotmetrics.

1. NO notional or multiplier code. point_value / Point Value /
   contract_specs / read_metadata all return zero hits; every `multiplier`
   hit is sigma_multiplier, a z-score threshold. The one apparent
   exception is options_data.calculate_intrinsic_curve, which is EQUITY
   ETF option chains with a hardcoded *100 shares per contract, not a
   futures multiplier. Step 2 writes the notional path fresh, and no
   existing consumer needs fixing.

2. NO market-code resolver either. cftc_code / CotSymbolCodeMap /
   contract_market_code return zero hits, and cotmetrics never imports
   cotdata.registry, confirming ADR-0007 observation #3 against the code.
   Its whole cotdata surface is get_cot / get_prices / schema_version,
   each taking the INTERNAL SYMBOL and letting cotdata resolve privately.
   cotmetrics never holds a market code. crowdmon's contract_master runs
   the other direction (canonical rows are keyed by market_code), so it
   becomes the first consumer of cotdata.registry.

3. COVERAGE WAS OVERSTATED, and the headline is withdrawn. "42 of 95,
   44%" was measured against the Legacy store and framed against the full
   CFTC list, which step 2 does not aim at. Against the deployed universe:
   params.yaml is 47 (42 Role: deploy + 5 Role: heldout), 45 of 47 are
   joinable, and ALL 42 deploy markets are. The only two failures are MME
   and MFS, both heldout, both without Norgate specs or prices.
   Also caught: the two 42s (legacy-joinable, and Role: deploy) are the
   same number AND the same set for unrelated reasons, and EMD/KE/NKD were
   excluded only for having no LEGACY table, so the canonicalisers just
   shipped move the total to 45.

4. BACKADJ VERIFIED. Three production call sites, signals.py:1028,
   CotIndexer.py:655, options_data.py:366, all explicitly backadj. The
   ADR claim is true of the code, not just of the ADR. Stronger result
   from reading what each DOES with the series: none multiplies a
   historical price by a quantity (two use returns or bar shape, the third
   only the latest close, where the two series are equal by construction).
   So the +294% error is NOT shipping anywhere and there is nothing to go
   back and fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Reviewed by an agent given the spec and diff and NO prior context, because
every earlier review was done by the author and shared its blind spots.
Six confirmed defects, four in code written the same day. All reproduced
before fixing, and each now has a regression test naming the failure.

1. A FETCH FAILURE POISONED THE DEDUPE COMPARISON. Failure records carry
   no sha, etag or Last-Modified, so the next fetch sent no
   If-Modified-Since, could not match the dedupe test, and was classified
   as changed content. One network blip on a frozen year therefore
   produced a FALSE RESTATEMENT ALERT, the single alarm this subsystem
   exists to raise, and re-retained megabytes already on disk under a
   second filename carrying the identical sha. Split into
   _latest_with_content (content questions: is this the same file?) and
   _latest_for_url (outcome questions: what happened last time?).

2. AN ABSOLUTE local_path WAS RE-ROOTED UNDER THE STORE, turning
   /Volumes/ext/... into <store>/Volumes/ext/... COTDATA_VINTAGE_ROOT is
   MANDATORY on a mirrored replica, so on exactly those hosts every ingest
   raised FileNotFoundError, was swallowed, and marked the snapshot
   failed, which --pending never re-selects. One run drained the entire
   backlog with no CLI able to reset it.

3. A BARE `ingest` REPLAYED EVERY SNAPSHOT EVER RECORDED with
   observed_at = now. Replaying an older snapshot after a revision wrote
   the superseded value back with a newer timestamp, emitting a reversed
   revision (35 -> 30) then a re-revision (30 -> 35) and inflating
   age_days on both. revisions/ is the primary artifact and age_days is
   what tells a consumer whether a restatement reached its calibration
   window. Now defaults to pending, --all-snapshots opts into a replay,
   and observed_at comes from the snapshot's own retrieved_at, which is
   what it always meant.

4. `schedule backfill` BYPASSED THE WRITE LOCK while read-modify-writing
   every observations partition, so running it beside an ingest silently
   dropped that ingest's rows and left a revision row asserting a change
   to a value no longer present in observations/. Per-file atomicity was
   never the missing piece.

5. EVERY FLAT WEEK WAS LABELLED long_liquidation, because d_long <= 0
   swallows zero and the dead zone is off by default. Measured on the real
   2026 Legacy file: 3,308 of 29,787 transitions (11.1%), making
   long_liquidation the modal state with a third of its bucket being weeks
   where nothing happened, and that value_counts() is the CLI's headline.
   Zero is now `quiet` unconditionally: it is the absence of a change, not
   a small number, so recognising it needs no parameter.

6. validate() ACCEPTED DUPLICATE NATURAL KEYS and had NO NULL-RATE BAND
   despite spec section 5 requiring one. Duplicates get identical
   (observed_at, snapshot_id), exhausting the tie-break so append order
   decides the stored value. The null band closes the hole errors="coerce"
   opened: a changed value FORMAT in a column whose NAME never moved would
   coerce to nulls, get written as observations, and be read as revisions
   when the values returned. Checked per category, since one broken column
   is only 1/n of a melted frame and a frame-wide band would wave it past.

NOT FIXED, recorded as a real gap in cot_vintage.md section 9: the
`announced` release-date tier is unreachable in production and acceptance
criterion 5 is UNMET. Nothing writes source="announced" rows, so the
Oct-Dec 2025 backlog weeks can only resolve to `derived`. The plumbing and
precedence are correct; what is missing is an extractor pulling a
(report_date, release_date) pair out of free-text announcement prose, and
writing that against no corpus of real announcements would be guessing. A
release date resolved by a guess is worse than a derived one, because it
carries a provenance flag claiming it was announced.

Also corrected two places where the code contradicted its own docs: the
comment claiming an interrupted fetch "leaves a consistent store" (it can
leave an orphan blob, which is the right direction to fail in but not what
was claimed), and cot_vintage.md still describing the planned
manifests/cot.json block rather than the self-owned snapshots.json.

Verified: 249 tests pass, ruff clean, and the full ingest re-run against
the real captured snapshots is unchanged at 82,776 observations with a
second run a no-op.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…view

The six fixes from 9a34750 were reviewed by a second cold agent, since
they were written by the same person who wrote the code being fixed and
nobody had looked at them. It downloaded EVERY real CFTC annual file (41
Legacy years, 17 Disaggregated, 17 TFF, about 1.6M canonical rows) and put
all of it through the two new raising checks.

That sweep is the most valuable result of either review: NO legitimate
duplicate natural key exists in 40 years of any report type, and the null
rate on long/short/open_interest is exactly 0.0 in every
(report_type, category) group of every year. Neither new raise can
hard-fail a historical backfill. Fixes 4 and 5 came back sound, with fix
5's 11.1% flat-week figure reproducing to the row.

TWO FIXES CLOSED THEIR REPRODUCER BUT NOT THEIR CLASS.

1. Fix 1 re-pointed restatement_suspect at the content-bearing snapshot
   but left the FIRST/CHANGED classification keyed to `prev`. So when the
   FIRST record for a URL was a fetch failure, the recovering fetch still
   fired the false restatement alert, with outcome and restatement_suspect
   disagreeing. _annotate now takes `content` and classifies FIRST off it.

2. Fix 3 corrected the observed_at stamp but not the COMPARISON:
   ingest_canonical still diffed against whatever was newest in the store,
   so --all-snapshots fabricated a reversed revision (300 -> 100) plus a
   re-revision and grew revisions/ without bound on every pass. The diff
   is now filtered to observed_at <= observed_ts, which makes it genuinely
   bitemporal. A replay is a true no-op, verified three passes deep on the
   real captured snapshots.

FOUR FURTHER FINDINGS, ALL FIXED.

- The blind-detector alert COULD NOT FIRE AT THE YEAR ROLLOVER. It needed
  prev_outcome == deduped, but a year that churned through 2025 has
  prev_outcome == changed on the January morning it becomes
  frozen-in-window. The rollover is the likeliest moment for CFTC's window
  to shift, so that was the worst possible blind spot. Now keys on whether
  the previous fetch saw bytes at all.

- A FETCH FAILURE WAS A TRIPWIRE CONDITION. Connectivity is not
  provenance: a dropped connection says nothing about whether CFTC
  restated anything, and routing it there turned one blip into a
  frozen-year restatement alarm. Failures are counted and printed by
  fetch, where an operational problem belongs.

- A RUN WHERE EVERY SNAPSHOT FAILED TO PARSE EXITED 0, reporting success
  to Task Scheduler while the store gained nothing and `failed` stayed
  terminal. Now non-zero, in ONE combined message: a first draft raised on
  the failure and silently swallowed a restatement suspect in the same
  run, which is the more serious of the two. `ingest --retry-failed` is
  the way back, and it is the only one, since nothing else ever wrote
  parse_status back to pending and two separate defects had stranded whole
  backlogs there.

- DISAGG AND TFF WERE FETCHED FROM 2006, but cftc.gov 404s for
  fut_disagg_txt_2006..2009 and fut_fin_txt_2006..2009 (verified live).
  Every `fetch --all` recorded eight permanent, unfixable failure
  snapshots. Corrected to 2010.

- THE NULL BAND WAS VACUOUS FOR LEGACY, because canonicalize_legacy never
  coerced: a value arriving as "200,000" stayed an object column, passed
  every check, and hashed differently from the numeric form. Exactly the
  fabricated-revision failure the band prevents, on the one report type
  where it could not see it. Legacy now coerces like the other two, and it
  is PROVED not to move any stored hash: across 95 markets and 448,236
  canonical rows the real values are already int64, so it is an identity.

Verified: 257 tests pass, ruff clean, and a full re-ingest of the real
captured snapshots is unchanged at 82,776 observations with both a plain
re-run and an --all-snapshots replay writing zero rows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@mspinola
mspinola merged commit 5318b41 into main Jul 31, 2026
5 checks passed
@mspinola
mspinola deleted the claude/cot-flow-decomposition branch July 31, 2026 14:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant