feat(vintage): COT vintage store & revision tracking (capture + change-only ingest/PIT) - #78
Merged
Merged
Conversation
Step 1 of the vintage/revision-tracking subsystem (handoff v0.2). Capture is the
only time-critical piece: CFTC serves current-state files only and git-history
recovery found nothing, so the vintage series can only accumulate forward and an
uncaptured weekly release is permanently lost. Ingest/diff/PIT are deferred and
will run retroactively over the retained raw bytes.
- cotdata-vintage fetch: conditional GET (If-None-Match/If-Modified-Since),
sha256, immutable retention under vintage/raw/{source_kind}/{year}/, provenance
in a self-owned vintage/manifest.json. Byte-identical regenerations and 304s are
recorded but deduped (a changed download is not itself a revision).
- Provenance lives in its own manifest file, not the cot-half manifest, so
store.reconcile_manifest() cannot ghost-prune snapshot ids as missing parquet.
- current/ output guarded byte-identical (tests/test_current_baseline.py) with a
golden generated from the clean tree before the subsystem lands.
- Design: docs/design/cot_vintage.md (incl. the Last-Modified negative result).
Persistence/scope decision: cotdata docs/adr stub -> crucible-stack ADR-0008.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…elease-date backfill Commit 2 of the vintage subsystem — the substantive layer, built on top of the raw snapshots captured in the previous commit. Runs retroactively over retained bytes, so nothing here was time-critical. - ingest_canonical: change-only bitemporal writes keyed on (report_date, market_code, report_type, combined, category). A row is written only when its value-hash (row_sha256 over value fields, excluding provenance and release_date) differs from the latest for its key — re-ingesting identical data, including a byte-changed regeneration, is a no-op. - field-level revisions with age_days (revision depth); asof(t) = greatest observed_at <= t per key (no valid_to column). - release-date resolution with provenance (observed > announced > scheduled > derived); backfill flags the Oct-Dec 2025 backlog weeks as `announced`, not silently `derived`. observed is honoured only for live-captured rows. - best-effort Special Announcements scrape (raw text always retained). - CLI: cotdata-vintage ingest|diff|asof, cotdata-schedule sync|backfill. - report_date stored as-reported (Monday holiday weeks not normalised to Tuesday); timestamps tz-naive UTC to match the repo's convention. Tests cover handoff §8: idempotent ingest, byte-change/data-same, single-field revision, PIT query, revision depth, holiday week, backlog-week announced resolution, release-date precedence, and validation-failure-before-write. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review follow-ups. - asof / change-detection: _latest_by_key ordered by (observed_at, snapshot_id) with a stable sort, take last per key. Previously idxmax() returned the FIRST occurrence of the max observed_at, which is file/append-order dependent and undefined when two snapshots share a timestamp. Capture snapshot_ids lead with a compact retrieved_at, so the lexicographic tie-break also tracks retrieval time. - test: two snapshots sharing observed_at resolve to the greater snapshot_id regardless of append order (append order set opposite to the winner). - test: schedule backfill is idempotent and does not downgrade an `observed` release date to `scheduled` on re-run (precedence enforced on every write, observed first). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…-date semantics
Second review pass. Three real defects, one invariant, one lock.
1. _latest_by_key returned obs.loc[idx], a duplicate-label hazard: a frame built by a
naive concat has repeated index labels and .loc returns EVERY match, silently
yielding multiple rows per natural key. Verified: 6 rows where 3 were correct.
Now returns the grouped rows directly.
2. _norm no longer depends on numpy/pandas scalar formatting: values unwrap via .item()
to Python natives and floats format explicitly (".10g", int-valued as int). The hash
is a permanent artifact; a numpy repr change would otherwise appear as a mass
revision event across every market at once, worst in the cr4/cr8 ratio fields.
3. backfill resolved `observed` from each ROW's observed_at. A row is one vintage, so a
row revised months later produced a release date that late. Now uses the FIRST
sighting per report_date. Verified: offsets were [3, 103] days for one report.
The window check is also directional (0 <= delta <= window) rather than abs():
observed_at cannot precede report_date, so a negative offset is clock skew or a tz
bug and must not be absorbed as a valid release date.
Added: ConsistencyError raised when row_sha256 says "changed" but no VALUE_FIELDS
differ — the hash and field-diff use different comparison paths, and disagreement would
write an observation with no revision detail. Turns future dtype drift into a failure.
Added: advisory _WriteLock over the vintage subtree. Writes are read-concat-rewrite, so
two concurrent ingests would last-writer-wins; the producer is single-writer by design
and this makes a second writer fail loudly instead.
Added: `published` precedence slot above `observed` for the weekly-static Last-Modified
(a true publication timestamp vs a polling-interval approximation). Not yet populated —
mapping a weekly static to its report_date needs that file parsed, deferred by §10 — but
the slot exists so the wiring lands without a taxonomy migration.
Also documents that market_name is excluded from the hash by design, and therefore is
never updated once a natural key is written.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… id collision, per-source guard
Third review pass on the capture path.
1. A corrupt manifest no longer degrades to an empty one. _read_manifest quarantined
the damaged file and returned {} — and the next write then overwrote the record with
that empty structure. The manifest is the ONLY map from snapshot_id to url/sha/
retrieval time/parse status, so losing it leaves a directory of opaque blobs while
the bytes survive. Now moves it aside to manifest.json.corrupt.<ts> and raises
CorruptManifestError, per the fail-loudly rule.
2. Raw bytes are written .part + os.replace. A crash during a plain write_bytes left a
TRUNCATED file already carrying the full sha in its name, so any sha-keyed recovery
would adopt it as valid.
3. snapshot_id carries a url discriminator. _utcnow() truncates to whole seconds, so
several sources returning 304 in one second shared "{retrieved_at}_304" — and
update_snapshot patches EVERY record with a matching id. Rate limiting hid it in
production; rate_limit_s=0 in tests walks into it.
4. fetch guards per source. One 404 propagated out and killed an entire --all run
(120+ requests) with nothing recorded. Failures are now recorded as a record with a
note and the run continues, matching the ingest path.
Also: the manifest is written after EVERY source rather than once at the end, shrinking
the crash window from a whole run to one source (small file, atomic replace, no
measurable cost) — this removes the need for a reconcile companion. _latest_for_url
takes an explicit max over retrieved_at instead of the last list element. The weekly
static partitions by CAPTURE year rather than "current/". A minimum-byte floor per
source_kind refuses truncated/empty 200 bodies (content is None does not catch b"").
The User-Agent moves to COTDATA_USER_AGENT with a repo-URL default, out of source.
Documented two measured findings: annual-zip sha churn (closed years are content-frozen
and dedupe to nothing even though Last-Modified is re-touched weekly; only the current
year genuinely churns, ~20MB/week) and the deliberate futures-only scope that leaves
`combined` constant-False.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ipwire, live smoke Corrects a wrong claim in the previous commit's doc and records the operational decisions the measurement settles. - CORRECTION: "closed years are re-touched weekly" was wrong. A header sweep across 2020-2026 shows only the current year and the immediately-prior year move (2025 and 2026 share an identical Last-Modified — a rolling two-year regeneration window); 2022-2024 got one bulk re-touch in Jan 2026, and 2020/2021 are static. - cftc.gov serves NO ETag on any of these files, so If-None-Match can never fire. If-Modified-Since DOES return 304 (verified against a static year and the current year), so the conditional GET is not decorative: on an --all sweep only ~2 of ~40 annual files transfer. Noted in _http_get so nobody assumes If-None-Match works. - Retention decision: keep everything, no pruning. ~1GB/yr is immaterial against irreplaceability, and those weekly copies ARE the vintage series. - --all is a restatement TRIPWIRE, not a backfill: a content_sha256 change on a frozen closed year is the 2008-style retroactive-restatement signature, and this is the only automated detector for it. Monthly/quarterly, not weekly. - Caveat recorded: inner entry mtime is evidence, not proof (a regeneration preserving source mtimes looks identical). content_sha256 across two weeks is definitive, and capture now does that automatically. - Pre-merge live smoke: default path run against cftc.gov into a throwaway store. All four URLs resolved, UA accepted, four 200s retained with provenance, legacy sha matched an independent manual download; a second run returned 304 on all four and retained nothing. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…T escape hatch Two things, one of which blocks scheduling capture on a replica. DEPLOYMENT HAZARD found while siting the scheduled job: the Mac replica is fed by robocopy /MIR, which deletes destination-only files and excludes only _cache/_raw/citpy. A vintage/ tree written on the Mac would therefore be DELETED by the next producer sync — the exact trap sync-store.cmd already documents for citpy, except vintage data is irreplaceable (CFTC serves current state only, so a deleted vintage never comes back). - COTDATA_VINTAGE_ROOT relocates the whole tree outside the mirrored store, so capture can run on a replica safely. Preferred placement remains the producer, where the tree syncs outward normally and matches the ADR-0007 cot-half seam. Adding `vintage` to /XD is explicitly NOT recommended: it holds only while every future invocation remembers the flag. - Restatement tripwire: a CLOSED year whose content sha changes is flagged restatement_suspect with a loud message. Closed years are frozen (measured), so this is the 2008-style retroactive-restatement signature, and it falls out of the ordinary dedupe path with no extra machinery. 2025 sits in the rolling regeneration window but is byte-frozen, so it should emit exactly one "unchanged bytes (deduped)" record per week indefinitely; the week it does not is the alert. January year-end finalization is called out in the message as the benign case. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sites the scheduled capture on the Windows producer and makes it propagate correctly. - run-vintage.cmd: fetch + ingest --pending, to chain after run-cot.cmd. Documents why it runs on the producer (capture is a producer action, and a replica-written tree is deleted by /MIR) and why the schedule is DAILY not weekly (almost everything 304s, so it is nearly free, and it catches holiday shifts and backlog publications while tightening observed_at from a 7-day to a 1-day bound). - RENAME vintage/manifest.json -> vintage/snapshots.json. Both sync scripts exclude "manifest.json" UNANCHORED (robocopy /XF and rsync --exclude both match by name at any depth), so the provenance index would have been silently stripped in transit and each replica would have received raw archives with no index. The rename means existing deployed sync scripts carry it correctly with no edit. - sync-store.cmd (Mac): carries vintage in full, deliberately. The bytes are irreplaceable and the Mac is the natural second copy (~1GB/yr). Adds *.tmp/*.part to /XF so a sync mid-capture cannot land a partial. - push-to-server.cmd (VPS): excludes vintage/ entirely. cot-analyzer reads prices and COT only, so the dash would carry ~1GB/yr of archives it never opens. Notes that excluding just vintage/raw/ is the right change if the dash ever consumes revisions. - SYNCING.md: new section on where vintage may be written, the per-replica table, and the naming gotcha, next to the existing citpy precedent. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The README was untouched by the previous commits while the branch added two CLI entry points, two env vars, a new store subtree and a new scheduled task — all categories the README already documents. - New "COT vintage tracking (as-published history)" section under Concepts & design: why revisions matter, the full CLI, how change-only storage / PIT reads / release-date provenance work, the producer-vs-replica placement rule, and why --all is a restatement tripwire rather than a backfill. - Store layout gains vintage/, flagged opt-in and additive. - Producer command list gains cotdata-vintage fetch; Contents gains the section link. - Sync section notes the irreplaceability rule and the snapshots.json naming, since the usual manifest.json exclusion matches at any depth. - WINDOWS_SCHEDULING.md: run-vintage.cmd wrapper, an optional fourth schtasks entry at 17:00, and the daily-not-weekly rationale. Corrected "Create three tasks", which my addition had made wrong. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
§10 deferred this on the assumption it needed the weekly static parsed. Measuring the file showed otherwise: it is a headerless positional CSV covering exactly ONE report date (365 rows, 129 columns, a single distinct value in field 2), so mapping a snapshot to its report date reads one field. The publication timestamp was already being captured into snapshots.json. The whole thing was ~30 lines, not a canonicaliser, so it ships. - published_from_snapshots(): retained weekly-static file -> report_date (field 2), its HTTP Last-Modified -> release_date converted to ET (CFTC publishes on ET, so the conversion decides the date for any post-midnight-UTC timestamp). - backfill folds published rows in automatically on the production path, so plain `cotdata-schedule backfill` picks up true publication times with no extra step; tests still pass an explicit schedule to isolate precedence. - _schedule_map now ranks published > announced > scheduled rather than special-casing announced-beats-scheduled. - New `cotdata-schedule published`, and run-vintage.cmd gains published + backfill after ingest (both idempotent). Verified live: report date 2026-07-21 resolves to a 2026-07-24 ET publication date, source=published, beating the poll-derived observed bound. Forward-only by nature: the weekly static holds one week and is overwritten, so this covers weeks from first capture onward. Full canonicalisation of the weekly static INTO observations is genuinely larger and remains out of scope. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…spec Committed verbatim as authored, before any amendment, so the specs-as-written are a distinct point in history. The v0.2 handoff is the document the vintage subsystem was built against; the module design is its parent (the handoff implements §5.3). Both predate implementation and are amended in the following commit with what the build actually established. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Moves findings out of commit messages and chat and into the docs, so the docs are the source of truth. The specs above the new sections are left unedited: they are the documents as written before implementation, and the amendments record where they were wrong rather than quietly correcting them. handoff §12 "Outcome" (following the ADR-0006 house pattern): - The Last-Modified spike as an explicit NEGATIVE result: historical release dates cannot be recovered from headers, so §4.6's fallback chain is the confirmed path. - Header sweep 2020-2026 and the ROLLING TWO-YEAR WINDOW: 2025 and 2026 share an identical Last-Modified, so CFTC regenerates current + prior year and everything older is static. Corrects an earlier claim that all closed years were re-touched weekly, and explains why weekly downloads have always grabbed both years. Also: no ETag is served on any file, so If-None-Match can never fire, while If-Modified-Since does 304. - Decisions: retention (keep everything, no pruning), --all as a restatement TRIPWIRE rather than a backfill, daily producer-side capture, and the futures-only carry-forward that leaves `combined` constant-False with half the reportable universe absent. - Deviations table (polars->pandas, manifest.json->snapshots.json, synthetic fixtures, no --from-git, and `published` shipping despite §10 deferring it). - Answers to §11, including that backfill coverage is legitimately empty because no production vintages exist yet. crowdmon_futures_cot_module.md: resolves the §4 open items (schema, backfill span, revisions were overwritten, no release_date, futures-only) and notes that `vintage: int` is not how it was built — the implementation is bitemporal. §5.3 gains the two findings that change what it can assume: release date is resolved with provenance rather than read, and `derived` fails on exactly the weeks that matter; and vintage history is forward-only, so PIT protection cannot be backfilled. Links to cot_vintage.md and ADR-0008 dangle on main until PRs #78 and #13 merge; called out inline rather than left to surprise a reader. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The warning was written while these docs sat on main and cot_vintage.md did not. They now travel in the same PR, so that link never dangles. Only the cross-repo ADR-0008 reference still depends on a separate merge (crucible-stack #13). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The byte-hash guard passed locally and failed on all five CI Python versions, with a DIFFERENT hash on each (3.10 ebc56905, 3.13 0478a1ca, local 53e939b8). Parquet encoding is a property of the pandas/pyarrow build -- writer metadata, compression, dictionary encoding -- not of this repo's code, so the assertion was about the library rather than about anything the change could keep true. It could never have gone green in CI. What consumers actually depend on is the data read back, so that is what is compared now, against the same pre-vintage golden file. dtypes are compared loosely because CI spans Python 3.10-3.14 and therefore pandas 2 and 3, which disagree about the string dtype; values, columns, order and index are still compared exactly. Byte-stability is kept as a separate test where it IS a real property: writing identical data twice in one environment must produce identical bytes, which is what an atomic write could plausibly break. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Second half of the same lesson. check_dtype=False does not cover the INDEX, which assert_frame_equal checks separately, so 3.10 still failed on datetime64[ns] vs datetime64[us]: pandas 3 wrote the golden at microsecond resolution and pandas 2 reads it back at nanosecond. Again a property of the library, not of this code. Index and column CONTENT are still asserted exactly, immediately after, so relaxing the dtype check does not quietly widen what this guard accepts. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A scheduled run's stdout goes nowhere, so a retroactive restatement -- the exact event this subsystem exists to detect -- would have been recorded and then silently swallowed, with exit 0 and no way to know without running 'diff' by hand. - 'ingest' now prints the changed fields (with age_days, flagging any reaching more than 30 days back into the calibration window) and exits NON-ZERO when it records a revision or sees a closed-year restatement suspect. The message says plainly that this is a notification and the data is committed, so it is not mistaken for a broken run. - The suspect scan is scoped to the snapshots THIS run processed. A store-wide scan would re-fire on every run forever once one restatement was seen, and an alert that never clears is one that gets switched off. Test pins that a prior run's suspect stays quiet. - run-vintage.cmd tees every step to a log, treats ingest's non-zero as 'revised' rather than 'failed' so later steps still run, and re-raises at the end so Task Scheduler shows the task as attention-needed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
mspinola
added a commit
that referenced
this pull request
Jul 31, 2026
Closes the notification gap left by #78. Task Scheduler's built-in Send-an-e-mail / Display-a-message actions are deprecated and non-functional on Windows 8 / Server 2012+, so ingest exiting non-zero only produced a passive Last Run Result 0x1 that someone had to go and look at. A comment in run-vintage.cmd implied Task Scheduler could alert on it; corrected. run-vintage.cmd now notifies itself: on any run that records a revision it appends that run's output to <store>\\vintage\\REVISIONS_<yyyy-MM-dd>.txt, which rides the existing robocopy /MIR push to the Mac. The marker is written immediately after ingest rather than at the end, so a later unrelated failure cannot suppress an alert for a revision that is already committed; it appends so two runs in one day both survive; the date comes from PowerShell rather than locale-formatted %DATE%; and the ingest output goes to a .tmp scratch file the existing sync exclusions skip. scripts/vintage_alert_selftest.py forces a revision in a throwaway store and drives the real CLI, so the loud path can be seen working now instead of being trusted until CFTC eventually restates something.
This was referenced Jul 31, 2026
mspinola
added a commit
that referenced
this pull request
Jul 31, 2026
…decomposition (#83) * feat(vintage): frozen-year tripwire, flow decomposition, corrected OI validation Three things, all measured against the real store rather than assumed. 1. THE FROZEN-YEAR TRIPWIRE IS NOW AUTOMATIC The rolling two-year regeneration window (measured in #78) means CFTC re-serves the PRIOR year every week, byte-identical. That is the one place a content check on closed data is free, and it is the only automated detector for the failure mode this subsystem exists to guard against (July 2008: reports restated back to 2007-07-03). The standing decision was "run --all monthly or quarterly". A tripwire that only fires when somebody remembers to run it by hand is an intention, not a detector, so the prior year moves into the DEFAULT fetch set: one ~7 MB transfer per week (the other six days 304) and zero bytes on disk, since the content dedupes. Three regimes now carry different expectations, recorded on every snapshot as expectation / outcome / tripwire_alert: churn current year, new bytes are the data arriving frozen_in_window prior year, expect exactly "unchanged bytes (deduped)" frozen_out_of_window older, expect 304; content is never verified there Two alert shapes, deliberately triggered differently. Content changed on a frozen year always alerts: it is a discrete, diffable event. The detector going blind (a frozen-in-window 304 or fetch failure) alerts on the TRANSITION only, because it is a standing condition and a level-triggered alert would fire daily forever. fetch still exits zero on purpose: the Windows wrapper aborts the run on a non-zero fetch, which would skip the ingest that turns a restatement into readable revision rows. ingest re-raises it and already writes the marker file. 2. FLOW DECOMPOSITION (module spec section 6.4) vintage_flow.py plus `cotdata-vintage flow`. Weekly dLong vs dShort per market/category into new_longs / short_covering / new_shorts / long_liquidation. No prices, no contract master, so it is a clean smoke test of the canonical schema. Classification is by DOMINANT leg, because the spec's "~0" leg never happens in real data; parameter-free by default, with an optional min_frac_oi dead zone that stays off since any value for it is a judgement. Read-side only: writes no store domain, changes no producer/consumer contract. ADR-0007 narrows cotdata by instrument domain, not derived-vs-raw, and ADR-0008 settles that COT-derived-from-COT stays inside that boundary. The positioning engine (prices, contract master, configured weights) stays in crowdmon. Swept over all 95 Legacy markets, 1986 to 2026: 149,412 of 149,412 weeks satisfy the zero-sum identity (sum of long across categories == sum of short), 447,951 flow rows, zero exceptions. Two findings recorded in the docs: - NonComm_Positions_Spread_All is not captured and never was, so every side total falls short of OI by exactly the spreading (gold: 31,983 of 383,368, 8%). Net positioning is unaffected; anything denominated as a SHARE of open interest is not. Not fixed here: it would change providers/cftc.py output and break the current/ byte-identical guarantee, so it is its own change. - COT was FORTNIGHTLY until 1992-10-13. days_elapsed is emitted as a column instead of assuming seven, so callers filter rather than discover it later. 3. CORRECTED validate()'s OPEN-INTEREST WARNING It compared one category's long+short+spread against total OI, which is not a bound: the two sides are counted separately. It fired on 811 of 5,778 gold rows (14%), pure noise. Replaced with the invariant that holds, per side summed across categories. Store-wide warnings went from thousands to zero. The module spec asserted the wrong form too, and is amended in the same pass per the working agreement on specs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: propose the step 2 approach (contract master + normalisation) Proposal only, no step-2 code. Three measurements against the real store change what step 2 can be: 1. Step 2 does not belong in cotdata. Normalisation is by definition a joiner (COT x prices x contract specs), and ADR-0007's whole argument was that almost nothing joins the domains. It belongs in the consumer, which means creating crowdmon-futures as a sibling. 2. The registry declares 41 disagg + 8 TFF and ZERO legacy symbols, and the module spec's engines all key on Managed Money / Leveraged Funds, which exist only in those two reports. Vintage ingest wires Legacy only, so a point-in-time Managed Money series does not exist yet. That is the real step-2 prerequisite, and it is a cotdata change. 3. Notional off back-adjusted prices is wrong by +294% (gold 2002) and +257% (crude 2004), and crude's back-adjusted series reaches -27.52, which is not a price. The error is EXACTLY ZERO at the present date and grows monotonically backwards, so it passes every spot check anyone would run while corrupting the whole evaluation history. Volatility must come from back-adjusted returns, so the two factors of net_notional x sigma come from two different price series. Also measured: all 47 contract specs are USD (no FX layer needed), unadjusted prices exist for all 47, and the COT-to-specs-to-price join covers 42 of 95 Legacy markets (44%). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: drop the stray em dash from the step 2 proposal title * feat(vintage): disaggregated and TFF canonicalisers The step 2 proposal identified this as the real prerequisite. Every engine in the module spec keys on Managed Money and Leveraged Funds, which exist only in these two reports, and ingest wired Legacy only. asof() now returns both point-in-time. Verified against the REAL first production capture (2026-07-31 01:15Z), not fixtures: 39,235 Disaggregated and 12,500 TFF canonical rows for 2026, all three reports ingesting in about 5 seconds, and re-ingest writing zero rows. These two reports populate three fields Legacy never does: per-category spreading, per-category trader counts, and CR4/CR8 net concentration. So the zero-sum identity closes COMPLETELY for Disaggregated (7,847/7,847 weeks, oi_gap zero everywhere), which is the confirming counterpart to the Legacy finding that its gap really is the uncaptured spreading column. ROUNDING TOLERANCE. The residual imbalance is CFTC's own and fully localised: every off-by-one-or-two row in both Legacy and TFF falls in exactly three markets, S&P 500 / NASDAQ-100 / DJIA Consolidated, whose codes carry CFTC's own '+' suffix for a unit-converted aggregate. Tolerance is derived from that mechanism rather than fitted: summing n independently rounded category figures admits at most n contracts of error. Without it, 48 off-by-one warnings fire on the 2026 files alone, the same cry-wolf rate the OI check was corrected for in the previous commit. Implementation points that will otherwise cost someone an afternoon: - CFTC's header has a typo. Swap__Positions_Short_All and _Spread_All carry a double underscore, _Long_All does not. Both spellings resolve, so the day CFTC fixes it is not the day Swap Dealer ingests as nulls. - An unresolvable column RAISES. Silent nulls would be written as real observations, and the next genuine value recorded as a revision that never happened, polluting the exact artifact this subsystem produces. - Trader counts arrive as strings because CFTC writes '.' for a suppressed count (3,578 of 7,847 Managed Money long counts on the 2026 file). Every value field is coerced so a marker never reaches row_sha256 as a literal string. - combined is now READ from the file's FutOnly_or_Combined column, and a file mixing the two is refused. Constant-FutOnly today, so nothing changes, but adding the combined files becomes a fetch-list change. - canonicalize_legacy is deliberately NOT routed through the new shared helper. Its output feeds row_sha256 over rows already stored in production; the duplication is cheaper than risking a mass revision. Also fixes a real cross-platform bug found in the production snapshot index: local_path is written by the Windows producer with backslashes, so ingest on either replica failed with "no such file" on a file plainly present. Normalised on read, so snapshots already recorded stay readable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: check cotmetrics for step 2 prior art, and correct the coverage finding Four questions, all answered by reading ~/code/trading_workspace/cotmetrics. 1. NO notional or multiplier code. point_value / Point Value / contract_specs / read_metadata all return zero hits; every `multiplier` hit is sigma_multiplier, a z-score threshold. The one apparent exception is options_data.calculate_intrinsic_curve, which is EQUITY ETF option chains with a hardcoded *100 shares per contract, not a futures multiplier. Step 2 writes the notional path fresh, and no existing consumer needs fixing. 2. NO market-code resolver either. cftc_code / CotSymbolCodeMap / contract_market_code return zero hits, and cotmetrics never imports cotdata.registry, confirming ADR-0007 observation #3 against the code. Its whole cotdata surface is get_cot / get_prices / schema_version, each taking the INTERNAL SYMBOL and letting cotdata resolve privately. cotmetrics never holds a market code. crowdmon's contract_master runs the other direction (canonical rows are keyed by market_code), so it becomes the first consumer of cotdata.registry. 3. COVERAGE WAS OVERSTATED, and the headline is withdrawn. "42 of 95, 44%" was measured against the Legacy store and framed against the full CFTC list, which step 2 does not aim at. Against the deployed universe: params.yaml is 47 (42 Role: deploy + 5 Role: heldout), 45 of 47 are joinable, and ALL 42 deploy markets are. The only two failures are MME and MFS, both heldout, both without Norgate specs or prices. Also caught: the two 42s (legacy-joinable, and Role: deploy) are the same number AND the same set for unrelated reasons, and EMD/KE/NKD were excluded only for having no LEGACY table, so the canonicalisers just shipped move the total to 45. 4. BACKADJ VERIFIED. Three production call sites, signals.py:1028, CotIndexer.py:655, options_data.py:366, all explicitly backadj. The ADR claim is true of the code, not just of the ADR. Stronger result from reading what each DOES with the series: none multiplies a historical price by a quantity (two use returns or bar shape, the third only the latest close, where the two series are equal by construction). So the +294% error is NOT shipping anywhere and there is nothing to go back and fix. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(vintage): six defects found by adversarial review Reviewed by an agent given the spec and diff and NO prior context, because every earlier review was done by the author and shared its blind spots. Six confirmed defects, four in code written the same day. All reproduced before fixing, and each now has a regression test naming the failure. 1. A FETCH FAILURE POISONED THE DEDUPE COMPARISON. Failure records carry no sha, etag or Last-Modified, so the next fetch sent no If-Modified-Since, could not match the dedupe test, and was classified as changed content. One network blip on a frozen year therefore produced a FALSE RESTATEMENT ALERT, the single alarm this subsystem exists to raise, and re-retained megabytes already on disk under a second filename carrying the identical sha. Split into _latest_with_content (content questions: is this the same file?) and _latest_for_url (outcome questions: what happened last time?). 2. AN ABSOLUTE local_path WAS RE-ROOTED UNDER THE STORE, turning /Volumes/ext/... into <store>/Volumes/ext/... COTDATA_VINTAGE_ROOT is MANDATORY on a mirrored replica, so on exactly those hosts every ingest raised FileNotFoundError, was swallowed, and marked the snapshot failed, which --pending never re-selects. One run drained the entire backlog with no CLI able to reset it. 3. A BARE `ingest` REPLAYED EVERY SNAPSHOT EVER RECORDED with observed_at = now. Replaying an older snapshot after a revision wrote the superseded value back with a newer timestamp, emitting a reversed revision (35 -> 30) then a re-revision (30 -> 35) and inflating age_days on both. revisions/ is the primary artifact and age_days is what tells a consumer whether a restatement reached its calibration window. Now defaults to pending, --all-snapshots opts into a replay, and observed_at comes from the snapshot's own retrieved_at, which is what it always meant. 4. `schedule backfill` BYPASSED THE WRITE LOCK while read-modify-writing every observations partition, so running it beside an ingest silently dropped that ingest's rows and left a revision row asserting a change to a value no longer present in observations/. Per-file atomicity was never the missing piece. 5. EVERY FLAT WEEK WAS LABELLED long_liquidation, because d_long <= 0 swallows zero and the dead zone is off by default. Measured on the real 2026 Legacy file: 3,308 of 29,787 transitions (11.1%), making long_liquidation the modal state with a third of its bucket being weeks where nothing happened, and that value_counts() is the CLI's headline. Zero is now `quiet` unconditionally: it is the absence of a change, not a small number, so recognising it needs no parameter. 6. validate() ACCEPTED DUPLICATE NATURAL KEYS and had NO NULL-RATE BAND despite spec section 5 requiring one. Duplicates get identical (observed_at, snapshot_id), exhausting the tie-break so append order decides the stored value. The null band closes the hole errors="coerce" opened: a changed value FORMAT in a column whose NAME never moved would coerce to nulls, get written as observations, and be read as revisions when the values returned. Checked per category, since one broken column is only 1/n of a melted frame and a frame-wide band would wave it past. NOT FIXED, recorded as a real gap in cot_vintage.md section 9: the `announced` release-date tier is unreachable in production and acceptance criterion 5 is UNMET. Nothing writes source="announced" rows, so the Oct-Dec 2025 backlog weeks can only resolve to `derived`. The plumbing and precedence are correct; what is missing is an extractor pulling a (report_date, release_date) pair out of free-text announcement prose, and writing that against no corpus of real announcements would be guessing. A release date resolved by a guess is worse than a derived one, because it carries a provenance flag claiming it was announced. Also corrected two places where the code contradicted its own docs: the comment claiming an interrupted fetch "leaves a consistent store" (it can leave an orphan blob, which is the right direction to fail in but not what was claimed), and cot_vintage.md still describing the planned manifests/cot.json block rather than the self-owned snapshots.json. Verified: 249 tests pass, ruff clean, and the full ingest re-run against the real captured snapshots is unchanged at 82,776 observations with a second run a no-op. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(vintage): finish two incomplete fixes, plus four from a second review The six fixes from 9a34750 were reviewed by a second cold agent, since they were written by the same person who wrote the code being fixed and nobody had looked at them. It downloaded EVERY real CFTC annual file (41 Legacy years, 17 Disaggregated, 17 TFF, about 1.6M canonical rows) and put all of it through the two new raising checks. That sweep is the most valuable result of either review: NO legitimate duplicate natural key exists in 40 years of any report type, and the null rate on long/short/open_interest is exactly 0.0 in every (report_type, category) group of every year. Neither new raise can hard-fail a historical backfill. Fixes 4 and 5 came back sound, with fix 5's 11.1% flat-week figure reproducing to the row. TWO FIXES CLOSED THEIR REPRODUCER BUT NOT THEIR CLASS. 1. Fix 1 re-pointed restatement_suspect at the content-bearing snapshot but left the FIRST/CHANGED classification keyed to `prev`. So when the FIRST record for a URL was a fetch failure, the recovering fetch still fired the false restatement alert, with outcome and restatement_suspect disagreeing. _annotate now takes `content` and classifies FIRST off it. 2. Fix 3 corrected the observed_at stamp but not the COMPARISON: ingest_canonical still diffed against whatever was newest in the store, so --all-snapshots fabricated a reversed revision (300 -> 100) plus a re-revision and grew revisions/ without bound on every pass. The diff is now filtered to observed_at <= observed_ts, which makes it genuinely bitemporal. A replay is a true no-op, verified three passes deep on the real captured snapshots. FOUR FURTHER FINDINGS, ALL FIXED. - The blind-detector alert COULD NOT FIRE AT THE YEAR ROLLOVER. It needed prev_outcome == deduped, but a year that churned through 2025 has prev_outcome == changed on the January morning it becomes frozen-in-window. The rollover is the likeliest moment for CFTC's window to shift, so that was the worst possible blind spot. Now keys on whether the previous fetch saw bytes at all. - A FETCH FAILURE WAS A TRIPWIRE CONDITION. Connectivity is not provenance: a dropped connection says nothing about whether CFTC restated anything, and routing it there turned one blip into a frozen-year restatement alarm. Failures are counted and printed by fetch, where an operational problem belongs. - A RUN WHERE EVERY SNAPSHOT FAILED TO PARSE EXITED 0, reporting success to Task Scheduler while the store gained nothing and `failed` stayed terminal. Now non-zero, in ONE combined message: a first draft raised on the failure and silently swallowed a restatement suspect in the same run, which is the more serious of the two. `ingest --retry-failed` is the way back, and it is the only one, since nothing else ever wrote parse_status back to pending and two separate defects had stranded whole backlogs there. - DISAGG AND TFF WERE FETCHED FROM 2006, but cftc.gov 404s for fut_disagg_txt_2006..2009 and fut_fin_txt_2006..2009 (verified live). Every `fetch --all` recorded eight permanent, unfixable failure snapshots. Corrected to 2010. - THE NULL BAND WAS VACUOUS FOR LEGACY, because canonicalize_legacy never coerced: a value arriving as "200,000" stayed an object column, passed every check, and hashed differently from the numeric form. Exactly the fabricated-revision failure the band prevents, on the one report type where it could not see it. Legacy now coerces like the other two, and it is PROVED not to move any stored hash: across 95 markets and 448,236 canonical rows the real values are already int64, so it is an identity. Verified: 257 tests pass, ruff clean, and a full re-ingest of the real captured snapshots is unchanged at 82,776 observations with both a plain re-run and an --all-snapshots replay writing zero rows. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Step 1 of the crowdmon-futures build (handoff v0.2). Adds an as-published (vintage) provenance layer under
cotdata, purely additive alongside the current-state Parquet store — existing consumers see byte-identical output.Two commits, capture-first:
264eca2capture —cotdata-vintage fetch: immutable hashed landing zone (vintage/raw/{source_kind}/{year}/), provenance in a self-ownedvintage/manifest.json, conditional GET with 304 / byte-identical dedupe. This is the only time-critical piece: CFTC serves current-state only and git recovery found nothing, so vintages can only accumulate forward and an uncaptured weekly release is irrecoverable.0ba2a3csubsystem — change-only bitemporalobservations/(a row written only when its value-hash differs from the latest for its natural key), field-levelrevisions/withage_daysrevision depth, point-in-timeasof(t), and release-date resolution/backfill with explicit provenance (observed > announced > scheduled > derived). Runs retroactively over retained raw bytes.Design & decisions
docs/adr/; the full ADR is committed separately in the crucible-stack worktree). Design + the Last-Modified negative-result spike:docs/design/cot_vintage.md.vintage/manifest.json, not the cot-half manifest, sostore.reconcile_manifest()can't ghost-prune snapshot IDs as missing parquet.report_datestored as-reported (Monday holiday weeks not normalized to Tuesday); timestamps tz-naive UTC to match the repo.Scope / follow-ups
ingestis wired end-to-end for Legacy only; disagg/TFF have controlled vocab defined but no canonicalizer yet (the change-only/revision machinery is report-type-agnostic once one is added).vintage stats, category-migration detection, tombstone write logic (column present), weekly-static fetching beyond the spike.Tests
announcedresolution, release-date precedence, validation-failure-before-write.current/byte-identical guard (tests/test_current_baseline.py) with a golden generated from the clean tree before the subsystem landed.Bottom line: capture is ready to start accumulating vintages the moment it runs; the full ingest/revision/PIT/backfill layer is built and tested on top of it. The genuine null stands — no historical vintages are recoverable, so the series begins from the first forward capture.
🤖 Generated with Claude Code