feat(vintage): disagg/TFF canonicalisers, frozen-year tripwire, flow decomposition - #83
Merged
Merged
Conversation
… validation Three things, all measured against the real store rather than assumed. 1. THE FROZEN-YEAR TRIPWIRE IS NOW AUTOMATIC The rolling two-year regeneration window (measured in #78) means CFTC re-serves the PRIOR year every week, byte-identical. That is the one place a content check on closed data is free, and it is the only automated detector for the failure mode this subsystem exists to guard against (July 2008: reports restated back to 2007-07-03). The standing decision was "run --all monthly or quarterly". A tripwire that only fires when somebody remembers to run it by hand is an intention, not a detector, so the prior year moves into the DEFAULT fetch set: one ~7 MB transfer per week (the other six days 304) and zero bytes on disk, since the content dedupes. Three regimes now carry different expectations, recorded on every snapshot as expectation / outcome / tripwire_alert: churn current year, new bytes are the data arriving frozen_in_window prior year, expect exactly "unchanged bytes (deduped)" frozen_out_of_window older, expect 304; content is never verified there Two alert shapes, deliberately triggered differently. Content changed on a frozen year always alerts: it is a discrete, diffable event. The detector going blind (a frozen-in-window 304 or fetch failure) alerts on the TRANSITION only, because it is a standing condition and a level-triggered alert would fire daily forever. fetch still exits zero on purpose: the Windows wrapper aborts the run on a non-zero fetch, which would skip the ingest that turns a restatement into readable revision rows. ingest re-raises it and already writes the marker file. 2. FLOW DECOMPOSITION (module spec section 6.4) vintage_flow.py plus `cotdata-vintage flow`. Weekly dLong vs dShort per market/category into new_longs / short_covering / new_shorts / long_liquidation. No prices, no contract master, so it is a clean smoke test of the canonical schema. Classification is by DOMINANT leg, because the spec's "~0" leg never happens in real data; parameter-free by default, with an optional min_frac_oi dead zone that stays off since any value for it is a judgement. Read-side only: writes no store domain, changes no producer/consumer contract. ADR-0007 narrows cotdata by instrument domain, not derived-vs-raw, and ADR-0008 settles that COT-derived-from-COT stays inside that boundary. The positioning engine (prices, contract master, configured weights) stays in crowdmon. Swept over all 95 Legacy markets, 1986 to 2026: 149,412 of 149,412 weeks satisfy the zero-sum identity (sum of long across categories == sum of short), 447,951 flow rows, zero exceptions. Two findings recorded in the docs: - NonComm_Positions_Spread_All is not captured and never was, so every side total falls short of OI by exactly the spreading (gold: 31,983 of 383,368, 8%). Net positioning is unaffected; anything denominated as a SHARE of open interest is not. Not fixed here: it would change providers/cftc.py output and break the current/ byte-identical guarantee, so it is its own change. - COT was FORTNIGHTLY until 1992-10-13. days_elapsed is emitted as a column instead of assuming seven, so callers filter rather than discover it later. 3. CORRECTED validate()'s OPEN-INTEREST WARNING It compared one category's long+short+spread against total OI, which is not a bound: the two sides are counted separately. It fired on 811 of 5,778 gold rows (14%), pure noise. Replaced with the invariant that holds, per side summed across categories. Store-wide warnings went from thousands to zero. The module spec asserted the wrong form too, and is amended in the same pass per the working agreement on specs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Proposal only, no step-2 code. Three measurements against the real store change what step 2 can be: 1. Step 2 does not belong in cotdata. Normalisation is by definition a joiner (COT x prices x contract specs), and ADR-0007's whole argument was that almost nothing joins the domains. It belongs in the consumer, which means creating crowdmon-futures as a sibling. 2. The registry declares 41 disagg + 8 TFF and ZERO legacy symbols, and the module spec's engines all key on Managed Money / Leveraged Funds, which exist only in those two reports. Vintage ingest wires Legacy only, so a point-in-time Managed Money series does not exist yet. That is the real step-2 prerequisite, and it is a cotdata change. 3. Notional off back-adjusted prices is wrong by +294% (gold 2002) and +257% (crude 2004), and crude's back-adjusted series reaches -27.52, which is not a price. The error is EXACTLY ZERO at the present date and grows monotonically backwards, so it passes every spot check anyone would run while corrupting the whole evaluation history. Volatility must come from back-adjusted returns, so the two factors of net_notional x sigma come from two different price series. Also measured: all 47 contract specs are USD (no FX layer needed), unadjusted prices exist for all 47, and the COT-to-specs-to-price join covers 42 of 95 Legacy markets (44%). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The step 2 proposal identified this as the real prerequisite. Every engine
in the module spec keys on Managed Money and Leveraged Funds, which exist
only in these two reports, and ingest wired Legacy only. asof() now returns
both point-in-time.
Verified against the REAL first production capture (2026-07-31 01:15Z), not
fixtures: 39,235 Disaggregated and 12,500 TFF canonical rows for 2026, all
three reports ingesting in about 5 seconds, and re-ingest writing zero rows.
These two reports populate three fields Legacy never does: per-category
spreading, per-category trader counts, and CR4/CR8 net concentration. So
the zero-sum identity closes COMPLETELY for Disaggregated (7,847/7,847
weeks, oi_gap zero everywhere), which is the confirming counterpart to the
Legacy finding that its gap really is the uncaptured spreading column.
ROUNDING TOLERANCE. The residual imbalance is CFTC's own and fully
localised: every off-by-one-or-two row in both Legacy and TFF falls in
exactly three markets, S&P 500 / NASDAQ-100 / DJIA Consolidated, whose
codes carry CFTC's own '+' suffix for a unit-converted aggregate. Tolerance
is derived from that mechanism rather than fitted: summing n independently
rounded category figures admits at most n contracts of error. Without it,
48 off-by-one warnings fire on the 2026 files alone, the same cry-wolf rate
the OI check was corrected for in the previous commit.
Implementation points that will otherwise cost someone an afternoon:
- CFTC's header has a typo. Swap__Positions_Short_All and _Spread_All
carry a double underscore, _Long_All does not. Both spellings resolve,
so the day CFTC fixes it is not the day Swap Dealer ingests as nulls.
- An unresolvable column RAISES. Silent nulls would be written as real
observations, and the next genuine value recorded as a revision that
never happened, polluting the exact artifact this subsystem produces.
- Trader counts arrive as strings because CFTC writes '.' for a
suppressed count (3,578 of 7,847 Managed Money long counts on the 2026
file). Every value field is coerced so a marker never reaches
row_sha256 as a literal string.
- combined is now READ from the file's FutOnly_or_Combined column, and a
file mixing the two is refused. Constant-FutOnly today, so nothing
changes, but adding the combined files becomes a fetch-list change.
- canonicalize_legacy is deliberately NOT routed through the new shared
helper. Its output feeds row_sha256 over rows already stored in
production; the duplication is cheaper than risking a mass revision.
Also fixes a real cross-platform bug found in the production snapshot
index: local_path is written by the Windows producer with backslashes, so
ingest on either replica failed with "no such file" on a file plainly
present. Normalised on read, so snapshots already recorded stay readable.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… finding Four questions, all answered by reading ~/code/trading_workspace/cotmetrics. 1. NO notional or multiplier code. point_value / Point Value / contract_specs / read_metadata all return zero hits; every `multiplier` hit is sigma_multiplier, a z-score threshold. The one apparent exception is options_data.calculate_intrinsic_curve, which is EQUITY ETF option chains with a hardcoded *100 shares per contract, not a futures multiplier. Step 2 writes the notional path fresh, and no existing consumer needs fixing. 2. NO market-code resolver either. cftc_code / CotSymbolCodeMap / contract_market_code return zero hits, and cotmetrics never imports cotdata.registry, confirming ADR-0007 observation #3 against the code. Its whole cotdata surface is get_cot / get_prices / schema_version, each taking the INTERNAL SYMBOL and letting cotdata resolve privately. cotmetrics never holds a market code. crowdmon's contract_master runs the other direction (canonical rows are keyed by market_code), so it becomes the first consumer of cotdata.registry. 3. COVERAGE WAS OVERSTATED, and the headline is withdrawn. "42 of 95, 44%" was measured against the Legacy store and framed against the full CFTC list, which step 2 does not aim at. Against the deployed universe: params.yaml is 47 (42 Role: deploy + 5 Role: heldout), 45 of 47 are joinable, and ALL 42 deploy markets are. The only two failures are MME and MFS, both heldout, both without Norgate specs or prices. Also caught: the two 42s (legacy-joinable, and Role: deploy) are the same number AND the same set for unrelated reasons, and EMD/KE/NKD were excluded only for having no LEGACY table, so the canonicalisers just shipped move the total to 45. 4. BACKADJ VERIFIED. Three production call sites, signals.py:1028, CotIndexer.py:655, options_data.py:366, all explicitly backadj. The ADR claim is true of the code, not just of the ADR. Stronger result from reading what each DOES with the series: none multiplies a historical price by a quantity (two use returns or bar shape, the third only the latest close, where the two series are equal by construction). So the +294% error is NOT shipping anywhere and there is nothing to go back and fix. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Reviewed by an agent given the spec and diff and NO prior context, because every earlier review was done by the author and shared its blind spots. Six confirmed defects, four in code written the same day. All reproduced before fixing, and each now has a regression test naming the failure. 1. A FETCH FAILURE POISONED THE DEDUPE COMPARISON. Failure records carry no sha, etag or Last-Modified, so the next fetch sent no If-Modified-Since, could not match the dedupe test, and was classified as changed content. One network blip on a frozen year therefore produced a FALSE RESTATEMENT ALERT, the single alarm this subsystem exists to raise, and re-retained megabytes already on disk under a second filename carrying the identical sha. Split into _latest_with_content (content questions: is this the same file?) and _latest_for_url (outcome questions: what happened last time?). 2. AN ABSOLUTE local_path WAS RE-ROOTED UNDER THE STORE, turning /Volumes/ext/... into <store>/Volumes/ext/... COTDATA_VINTAGE_ROOT is MANDATORY on a mirrored replica, so on exactly those hosts every ingest raised FileNotFoundError, was swallowed, and marked the snapshot failed, which --pending never re-selects. One run drained the entire backlog with no CLI able to reset it. 3. A BARE `ingest` REPLAYED EVERY SNAPSHOT EVER RECORDED with observed_at = now. Replaying an older snapshot after a revision wrote the superseded value back with a newer timestamp, emitting a reversed revision (35 -> 30) then a re-revision (30 -> 35) and inflating age_days on both. revisions/ is the primary artifact and age_days is what tells a consumer whether a restatement reached its calibration window. Now defaults to pending, --all-snapshots opts into a replay, and observed_at comes from the snapshot's own retrieved_at, which is what it always meant. 4. `schedule backfill` BYPASSED THE WRITE LOCK while read-modify-writing every observations partition, so running it beside an ingest silently dropped that ingest's rows and left a revision row asserting a change to a value no longer present in observations/. Per-file atomicity was never the missing piece. 5. EVERY FLAT WEEK WAS LABELLED long_liquidation, because d_long <= 0 swallows zero and the dead zone is off by default. Measured on the real 2026 Legacy file: 3,308 of 29,787 transitions (11.1%), making long_liquidation the modal state with a third of its bucket being weeks where nothing happened, and that value_counts() is the CLI's headline. Zero is now `quiet` unconditionally: it is the absence of a change, not a small number, so recognising it needs no parameter. 6. validate() ACCEPTED DUPLICATE NATURAL KEYS and had NO NULL-RATE BAND despite spec section 5 requiring one. Duplicates get identical (observed_at, snapshot_id), exhausting the tie-break so append order decides the stored value. The null band closes the hole errors="coerce" opened: a changed value FORMAT in a column whose NAME never moved would coerce to nulls, get written as observations, and be read as revisions when the values returned. Checked per category, since one broken column is only 1/n of a melted frame and a frame-wide band would wave it past. NOT FIXED, recorded as a real gap in cot_vintage.md section 9: the `announced` release-date tier is unreachable in production and acceptance criterion 5 is UNMET. Nothing writes source="announced" rows, so the Oct-Dec 2025 backlog weeks can only resolve to `derived`. The plumbing and precedence are correct; what is missing is an extractor pulling a (report_date, release_date) pair out of free-text announcement prose, and writing that against no corpus of real announcements would be guessing. A release date resolved by a guess is worse than a derived one, because it carries a provenance flag claiming it was announced. Also corrected two places where the code contradicted its own docs: the comment claiming an interrupted fetch "leaves a consistent store" (it can leave an orphan blob, which is the right direction to fail in but not what was claimed), and cot_vintage.md still describing the planned manifests/cot.json block rather than the self-owned snapshots.json. Verified: 249 tests pass, ruff clean, and the full ingest re-run against the real captured snapshots is unchanged at 82,776 observations with a second run a no-op. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…view The six fixes from 9a34750 were reviewed by a second cold agent, since they were written by the same person who wrote the code being fixed and nobody had looked at them. It downloaded EVERY real CFTC annual file (41 Legacy years, 17 Disaggregated, 17 TFF, about 1.6M canonical rows) and put all of it through the two new raising checks. That sweep is the most valuable result of either review: NO legitimate duplicate natural key exists in 40 years of any report type, and the null rate on long/short/open_interest is exactly 0.0 in every (report_type, category) group of every year. Neither new raise can hard-fail a historical backfill. Fixes 4 and 5 came back sound, with fix 5's 11.1% flat-week figure reproducing to the row. TWO FIXES CLOSED THEIR REPRODUCER BUT NOT THEIR CLASS. 1. Fix 1 re-pointed restatement_suspect at the content-bearing snapshot but left the FIRST/CHANGED classification keyed to `prev`. So when the FIRST record for a URL was a fetch failure, the recovering fetch still fired the false restatement alert, with outcome and restatement_suspect disagreeing. _annotate now takes `content` and classifies FIRST off it. 2. Fix 3 corrected the observed_at stamp but not the COMPARISON: ingest_canonical still diffed against whatever was newest in the store, so --all-snapshots fabricated a reversed revision (300 -> 100) plus a re-revision and grew revisions/ without bound on every pass. The diff is now filtered to observed_at <= observed_ts, which makes it genuinely bitemporal. A replay is a true no-op, verified three passes deep on the real captured snapshots. FOUR FURTHER FINDINGS, ALL FIXED. - The blind-detector alert COULD NOT FIRE AT THE YEAR ROLLOVER. It needed prev_outcome == deduped, but a year that churned through 2025 has prev_outcome == changed on the January morning it becomes frozen-in-window. The rollover is the likeliest moment for CFTC's window to shift, so that was the worst possible blind spot. Now keys on whether the previous fetch saw bytes at all. - A FETCH FAILURE WAS A TRIPWIRE CONDITION. Connectivity is not provenance: a dropped connection says nothing about whether CFTC restated anything, and routing it there turned one blip into a frozen-year restatement alarm. Failures are counted and printed by fetch, where an operational problem belongs. - A RUN WHERE EVERY SNAPSHOT FAILED TO PARSE EXITED 0, reporting success to Task Scheduler while the store gained nothing and `failed` stayed terminal. Now non-zero, in ONE combined message: a first draft raised on the failure and silently swallowed a restatement suspect in the same run, which is the more serious of the two. `ingest --retry-failed` is the way back, and it is the only one, since nothing else ever wrote parse_status back to pending and two separate defects had stranded whole backlogs there. - DISAGG AND TFF WERE FETCHED FROM 2006, but cftc.gov 404s for fut_disagg_txt_2006..2009 and fut_fin_txt_2006..2009 (verified live). Every `fetch --all` recorded eight permanent, unfixable failure snapshots. Corrected to 2010. - THE NULL BAND WAS VACUOUS FOR LEGACY, because canonicalize_legacy never coerced: a value arriving as "200,000" stayed an object column, passed every check, and hashed differently from the numeric form. Exactly the fabricated-revision failure the band prevents, on the one report type where it could not see it. Legacy now coerces like the other two, and it is PROVED not to move any stored hash: across 95 markets and 448,236 canonical rows the real values are already int64, so it is an identity. Verified: 257 tests pass, ruff clean, and a full re-ingest of the real captured snapshots is unchanged at 82,776 observations with both a plain re-run and an --all-snapshots replay writing zero rows. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Four commits, in the order they were asked for. Everything below was measured against the real store, and several measurements changed the plan.
1d8e1c938d6f90b1ade42Disaggregated and TFF canonicalisers
The step 2 proposal identified this as the real prerequisite rather than step 2 itself. Every engine in the module spec keys on Managed Money and Leveraged Funds, which exist only in these two reports, and ingest wired Legacy only.
asof()now returns both point-in-time.Verified against the real first production capture (2026-07-31 01:15Z), not fixtures:
All three ingest in about 5 seconds, and re-ingesting writes nothing, which is the strongest available confirmation that the permanent
row_sha256artifact is stable across the new code path.These two reports populate three fields Legacy never does: per-category spreading, per-category trader counts, and CR4/CR8 net concentration. So the zero-sum identity closes completely for Disaggregated:
oi_gapDisaggregated closing exactly, with a zero gap, is the confirming counterpart to the Legacy finding: that gap really is the missing spreading column and nothing else.
The rounding tolerance is derived, not fitted
Every off-by-one-or-two row in both Legacy and TFF falls in exactly three markets:
13874+20974+12460+The
+suffix is CFTC's own marker for a Consolidated contract, which aggregates several contract sizes onto a common unit and so involves a division. Summingnindependently rounded category figures admits at mostncontracts of error, which is whatrounding_tolerance()returns. Without it, 48 off-by-one warnings fire on the 2026 files alone, the same cry-wolf rate the OI check was corrected for in the first commit.Four implementation points
Swap__Positions_Short_Alland_Spread_Allcarry a double underscore;_Long_Alldoes not. Both spellings resolve, so the day CFTC fixes it is not the day Swap Dealer positions ingest as nulls..for a suppressed count (3,578 of 7,847 Managed Money long counts on the 2026 file, so "null" is routine, not an error). Every value field is coerced so a marker never reachesrow_sha256as a literal string.combinedis now read from the file'sFutOnly_or_Combinedcolumn, and a file mixing the two is refused. Constant-FutOnlytoday, so nothing changes, but adding the combined files becomes purely a fetch-list change.canonicalize_legacyis deliberately not routed through the new shared helper. Its output feedsrow_sha256over rows already stored in production, and rewriting that path to save a little duplication risks registering every stored row as revised at once.Fixed in passing: raw paths were unreadable off the producer
snapshots.jsonrecordslocal_pathas written by the capturing machine, and the producer is Windows, so the real store carriesvintage\raw\annual_zip\2026\....zip. On macOS or Linux that is one filename containing backslashes, so ingest on either replica failed with "no such file" on a file plainly sitting there. Normalised on read, so every snapshot already recorded on the producer stays readable.Three things, all measured against the real store rather than assumed.
1. The frozen-year tripwire is now automatic
The rolling two-year regeneration window (measured in #78) means CFTC re-serves the prior year every week, byte-identical. That is the one place a content check on closed data is free, and it is the only automated detector for the failure mode this subsystem exists to guard against (July 2008: reports restated back to 2007-07-03).
The standing decision was "run
--allmonthly or quarterly". A tripwire that only fires when somebody remembers to run it by hand is an intention, not a detector, so the prior year moves into the default fetch set: one ~7 MB transfer per week (the other six days 304) and zero bytes on disk, since the content dedupes.churnfrozen_in_windowunchanged bytes (deduped)frozen_out_of_window304 not-modifiedTwo alert shapes, deliberately triggered differently:
fetchstill exits zero on purpose: the Windows wrapper aborts the run on a non-zero fetch, which would skip theingestthat turns a restatement into readable revision rows.ingestre-raises it and already writes theREVISIONS_<date>.txtmarker.2. Flow decomposition (module spec section 6.4)
vintage_flow.pypluscotdata-vintage flow. Weekly dLong vs dShort per market/category intonew_longs/short_covering/new_shorts/long_liquidation. No prices, no contract master, so it is a clean smoke test of the canonical schema.Classification is by dominant leg, because the spec's "~0" leg never happens in real data. Parameter-free by default, with an optional
min_frac_oidead zone that stays off, since any value for it is a judgement in the same class as the fragility weights.Read-side only: writes no store domain, changes no producer/consumer contract. ADR-0007 narrows cotdata by instrument domain, not derived-vs-raw, and ADR-0008 settles that COT-derived-from-COT stays inside that boundary. The positioning engine (prices, contract master, configured weights) stays in crowdmon.
Worked example, gold non-commercial, week ending 2026-06-02: dNet +21,760, which reads as heavy fresh buying. Decomposed it is dShort -16,368 against dLong +5,392 with OI down 27,437. Short covering: a rally with a finite fuel supply. That distinction is the whole point of section 6.4, and it showed up on the first real week anyone looked at.
Swept over the whole store
95 Legacy markets, 1986 to 2026-07-21:
validate()warningsFinding: non-commercial SPREADING is not captured, and never was
NonComm_Positions_Spread_Allis absent fromproviders/cftc.py'sTARGET_COLS, so every side total falls short of OI by exactly the spreading. Gold on 2026-07-21: OI 383,368, both sides 351,385, gap 31,983 (8%). 64 of 95 markets have at least one week where the gap is zero, which is the confirming case.Net positioning is unaffected (spreading is long and short in equal measure). Anything denominated as a share of open interest is not, and that is section 5.2 step 2, the next build step.
Not fixed here: adding the column changes
providers/cftc.pyoutput, which breaks thecurrent/byte-identical guarantee and the golden baseline that exists to protect it. It is its own change, not a drive-by.Finding: COT was FORTNIGHTLY until 1992-10-13
415,908 of 447,951 intervals are 7 days; the 15-day (9,057) and 14-day (5,775) intervals are almost all pre-October-1992.
days_elapsedis emitted as a column rather than assumed, so callers filter on it instead of discovering it in a result. This is section 5.3's holiday hazard showing up in a second place: not only in the release date, but in the differencing interval.3. Corrected
validate()'s open-interest warningIt compared one category's
long + short + spreadagainst total OI, which is not a bound: the two sides of a market are counted separately, so a category holding 80k long and 294k short in a 383k-OI market sums to 374k with nothing wrong. Measured, it fired on 811 of 5,778 gold rows (14%), and a soft warning at that rate is one nobody reads.Replaced with the invariant that does hold, per side summed across categories. Store-wide warnings went from thousands to zero. The module spec asserted the wrong form too, and is amended in the same pass per the working agreement on specs.
Verification
ruff check src tests scriptsclean~/code/cotdata_storecotdata-vintage flow --market 088691 --source currentend to endGenerated with Claude Code