Skip to content

feat(vintage): COT vintage store & revision tracking (capture + change-only ingest/PIT) - #78

Merged
mspinola merged 16 commits into
mainfrom
claude/cot-revision-snapshots-9b196f
Jul 31, 2026
Merged

mspinola merged 16 commits into
mainfrom
claude/cot-revision-snapshots-9b196f

Conversation

@mspinola

Copy link
Copy Markdown
Owner

Step 1 of the crowdmon-futures build (handoff v0.2). Adds an as-published (vintage) provenance layer under cotdata, purely additive alongside the current-state Parquet store — existing consumers see byte-identical output.

Two commits, capture-first:

  • 264eca2 capture — cotdata-vintage fetch: immutable hashed landing zone (vintage/raw/{source_kind}/{year}/), provenance in a self-owned vintage/manifest.json, conditional GET with 304 / byte-identical dedupe. This is the only time-critical piece: CFTC serves current-state only and git recovery found nothing, so vintages can only accumulate forward and an uncaptured weekly release is irrecoverable.
  • 0ba2a3c subsystem — change-only bitemporal observations/ (a row written only when its value-hash differs from the latest for its natural key), field-level revisions/ with age_days revision depth, point-in-time asof(t), and release-date resolution/backfill with explicit provenance (observed > announced > scheduled > derived). Runs retroactively over retained raw bytes.

Design & decisions

  • No database — pandas/pyarrow throughout, per the repo's Parquet + manifest contract. Recorded as crucible-stack ADR-0008 (cotdata carries a stub in docs/adr/; the full ADR is committed separately in the crucible-stack worktree). Design + the Last-Modified negative-result spike: docs/design/cot_vintage.md.
  • Vintage provenance lives in its own vintage/manifest.json, not the cot-half manifest, so store.reconcile_manifest() can't ghost-prune snapshot IDs as missing parquet.
  • report_date stored as-reported (Monday holiday weeks not normalized to Tuesday); timestamps tz-naive UTC to match the repo.

Scope / follow-ups

  • CLI ingest is wired end-to-end for Legacy only; disagg/TFF have controlled vocab defined but no canonicalizer yet (the change-only/revision machinery is report-type-agnostic once one is added).
  • Deferred by the handoff (§10), not gaps: vintage stats, category-migration detection, tombstone write logic (column present), weekly-static fetching beyond the spike.
  • Backfill has no real coverage numbers yet — no production vintages exist until forward capture runs.

Tests

  • 174 pass (13 new), ruff clean. Covers handoff §8: idempotent ingest, byte-change/data-same, single-field revision, PIT query, revision depth, holiday week, backlog-week announced resolution, release-date precedence, validation-failure-before-write.
  • current/ byte-identical guard (tests/test_current_baseline.py) with a golden generated from the clean tree before the subsystem landed.

Bottom line: capture is ready to start accumulating vintages the moment it runs; the full ingest/revision/PIT/backfill layer is built and tested on top of it. The genuine null stands — no historical vintages are recoverable, so the series begins from the first forward capture.

🤖 Generated with Claude Code

mspinola and others added 10 commits July 30, 2026 13:50
Step 1 of the vintage/revision-tracking subsystem (handoff v0.2). Capture is the
only time-critical piece: CFTC serves current-state files only and git-history
recovery found nothing, so the vintage series can only accumulate forward and an
uncaptured weekly release is permanently lost. Ingest/diff/PIT are deferred and
will run retroactively over the retained raw bytes.

- cotdata-vintage fetch: conditional GET (If-None-Match/If-Modified-Since),
  sha256, immutable retention under vintage/raw/{source_kind}/{year}/, provenance
  in a self-owned vintage/manifest.json. Byte-identical regenerations and 304s are
  recorded but deduped (a changed download is not itself a revision).
- Provenance lives in its own manifest file, not the cot-half manifest, so
  store.reconcile_manifest() cannot ghost-prune snapshot ids as missing parquet.
- current/ output guarded byte-identical (tests/test_current_baseline.py) with a
  golden generated from the clean tree before the subsystem lands.
- Design: docs/design/cot_vintage.md (incl. the Last-Modified negative result).
  Persistence/scope decision: cotdata docs/adr stub -> crucible-stack ADR-0008.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…elease-date backfill

Commit 2 of the vintage subsystem — the substantive layer, built on top of the raw
snapshots captured in the previous commit. Runs retroactively over retained bytes,
so nothing here was time-critical.

- ingest_canonical: change-only bitemporal writes keyed on
  (report_date, market_code, report_type, combined, category). A row is written only
  when its value-hash (row_sha256 over value fields, excluding provenance and
  release_date) differs from the latest for its key — re-ingesting identical data,
  including a byte-changed regeneration, is a no-op.
- field-level revisions with age_days (revision depth); asof(t) = greatest
  observed_at <= t per key (no valid_to column).
- release-date resolution with provenance (observed > announced > scheduled >
  derived); backfill flags the Oct-Dec 2025 backlog weeks as `announced`, not
  silently `derived`. observed is honoured only for live-captured rows.
- best-effort Special Announcements scrape (raw text always retained).
- CLI: cotdata-vintage ingest|diff|asof, cotdata-schedule sync|backfill.
- report_date stored as-reported (Monday holiday weeks not normalised to Tuesday);
  timestamps tz-naive UTC to match the repo's convention.

Tests cover handoff §8: idempotent ingest, byte-change/data-same, single-field
revision, PIT query, revision depth, holiday week, backlog-week announced
resolution, release-date precedence, and validation-failure-before-write.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review follow-ups.

- asof / change-detection: _latest_by_key ordered by (observed_at, snapshot_id) with
  a stable sort, take last per key. Previously idxmax() returned the FIRST occurrence
  of the max observed_at, which is file/append-order dependent and undefined when two
  snapshots share a timestamp. Capture snapshot_ids lead with a compact retrieved_at,
  so the lexicographic tie-break also tracks retrieval time.
- test: two snapshots sharing observed_at resolve to the greater snapshot_id
  regardless of append order (append order set opposite to the winner).
- test: schedule backfill is idempotent and does not downgrade an `observed` release
  date to `scheduled` on re-run (precedence enforced on every write, observed first).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…-date semantics

Second review pass. Three real defects, one invariant, one lock.

1. _latest_by_key returned obs.loc[idx], a duplicate-label hazard: a frame built by a
   naive concat has repeated index labels and .loc returns EVERY match, silently
   yielding multiple rows per natural key. Verified: 6 rows where 3 were correct.
   Now returns the grouped rows directly.
2. _norm no longer depends on numpy/pandas scalar formatting: values unwrap via .item()
   to Python natives and floats format explicitly (".10g", int-valued as int). The hash
   is a permanent artifact; a numpy repr change would otherwise appear as a mass
   revision event across every market at once, worst in the cr4/cr8 ratio fields.
3. backfill resolved `observed` from each ROW's observed_at. A row is one vintage, so a
   row revised months later produced a release date that late. Now uses the FIRST
   sighting per report_date. Verified: offsets were [3, 103] days for one report.
   The window check is also directional (0 <= delta <= window) rather than abs():
   observed_at cannot precede report_date, so a negative offset is clock skew or a tz
   bug and must not be absorbed as a valid release date.

Added: ConsistencyError raised when row_sha256 says "changed" but no VALUE_FIELDS
differ — the hash and field-diff use different comparison paths, and disagreement would
write an observation with no revision detail. Turns future dtype drift into a failure.

Added: advisory _WriteLock over the vintage subtree. Writes are read-concat-rewrite, so
two concurrent ingests would last-writer-wins; the producer is single-writer by design
and this makes a second writer fail loudly instead.

Added: `published` precedence slot above `observed` for the weekly-static Last-Modified
(a true publication timestamp vs a polling-interval approximation). Not yet populated —
mapping a weekly static to its report_date needs that file parsed, deferred by §10 — but
the slot exists so the wiring lands without a taxonomy migration.

Also documents that market_name is excluded from the hash by design, and therefore is
never updated once a natural key is written.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… id collision, per-source guard

Third review pass on the capture path.

1. A corrupt manifest no longer degrades to an empty one. _read_manifest quarantined
   the damaged file and returned {} — and the next write then overwrote the record with
   that empty structure. The manifest is the ONLY map from snapshot_id to url/sha/
   retrieval time/parse status, so losing it leaves a directory of opaque blobs while
   the bytes survive. Now moves it aside to manifest.json.corrupt.<ts> and raises
   CorruptManifestError, per the fail-loudly rule.
2. Raw bytes are written .part + os.replace. A crash during a plain write_bytes left a
   TRUNCATED file already carrying the full sha in its name, so any sha-keyed recovery
   would adopt it as valid.
3. snapshot_id carries a url discriminator. _utcnow() truncates to whole seconds, so
   several sources returning 304 in one second shared "{retrieved_at}_304" — and
   update_snapshot patches EVERY record with a matching id. Rate limiting hid it in
   production; rate_limit_s=0 in tests walks into it.
4. fetch guards per source. One 404 propagated out and killed an entire --all run
   (120+ requests) with nothing recorded. Failures are now recorded as a record with a
   note and the run continues, matching the ingest path.

Also: the manifest is written after EVERY source rather than once at the end, shrinking
the crash window from a whole run to one source (small file, atomic replace, no
measurable cost) — this removes the need for a reconcile companion. _latest_for_url
takes an explicit max over retrieved_at instead of the last list element. The weekly
static partitions by CAPTURE year rather than "current/". A minimum-byte floor per
source_kind refuses truncated/empty 200 bodies (content is None does not catch b"").
The User-Agent moves to COTDATA_USER_AGENT with a repo-URL default, out of source.

Documented two measured findings: annual-zip sha churn (closed years are content-frozen
and dedupe to nothing even though Last-Modified is re-touched weekly; only the current
year genuinely churns, ~20MB/week) and the deliberate futures-only scope that leaves
`combined` constant-False.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ipwire, live smoke

Corrects a wrong claim in the previous commit's doc and records the operational
decisions the measurement settles.

- CORRECTION: "closed years are re-touched weekly" was wrong. A header sweep across
  2020-2026 shows only the current year and the immediately-prior year move (2025 and
  2026 share an identical Last-Modified — a rolling two-year regeneration window);
  2022-2024 got one bulk re-touch in Jan 2026, and 2020/2021 are static.
- cftc.gov serves NO ETag on any of these files, so If-None-Match can never fire.
  If-Modified-Since DOES return 304 (verified against a static year and the current
  year), so the conditional GET is not decorative: on an --all sweep only ~2 of ~40
  annual files transfer. Noted in _http_get so nobody assumes If-None-Match works.
- Retention decision: keep everything, no pruning. ~1GB/yr is immaterial against
  irreplaceability, and those weekly copies ARE the vintage series.
- --all is a restatement TRIPWIRE, not a backfill: a content_sha256 change on a frozen
  closed year is the 2008-style retroactive-restatement signature, and this is the only
  automated detector for it. Monthly/quarterly, not weekly.
- Caveat recorded: inner entry mtime is evidence, not proof (a regeneration preserving
  source mtimes looks identical). content_sha256 across two weeks is definitive, and
  capture now does that automatically.
- Pre-merge live smoke: default path run against cftc.gov into a throwaway store. All
  four URLs resolved, UA accepted, four 200s retained with provenance, legacy sha
  matched an independent manual download; a second run returned 304 on all four and
  retained nothing.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…T escape hatch

Two things, one of which blocks scheduling capture on a replica.

DEPLOYMENT HAZARD found while siting the scheduled job: the Mac replica is fed by
robocopy /MIR, which deletes destination-only files and excludes only _cache/_raw/citpy.
A vintage/ tree written on the Mac would therefore be DELETED by the next producer sync
— the exact trap sync-store.cmd already documents for citpy, except vintage data is
irreplaceable (CFTC serves current state only, so a deleted vintage never comes back).

- COTDATA_VINTAGE_ROOT relocates the whole tree outside the mirrored store, so capture
  can run on a replica safely. Preferred placement remains the producer, where the tree
  syncs outward normally and matches the ADR-0007 cot-half seam. Adding `vintage` to
  /XD is explicitly NOT recommended: it holds only while every future invocation
  remembers the flag.
- Restatement tripwire: a CLOSED year whose content sha changes is flagged
  restatement_suspect with a loud message. Closed years are frozen (measured), so this
  is the 2008-style retroactive-restatement signature, and it falls out of the ordinary
  dedupe path with no extra machinery. 2025 sits in the rolling regeneration window but
  is byte-frozen, so it should emit exactly one "unchanged bytes (deduped)" record per
  week indefinitely; the week it does not is the alert. January year-end finalization is
  called out in the message as the benign case.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sites the scheduled capture on the Windows producer and makes it propagate correctly.

- run-vintage.cmd: fetch + ingest --pending, to chain after run-cot.cmd. Documents why
  it runs on the producer (capture is a producer action, and a replica-written tree is
  deleted by /MIR) and why the schedule is DAILY not weekly (almost everything 304s, so
  it is nearly free, and it catches holiday shifts and backlog publications while
  tightening observed_at from a 7-day to a 1-day bound).
- RENAME vintage/manifest.json -> vintage/snapshots.json. Both sync scripts exclude
  "manifest.json" UNANCHORED (robocopy /XF and rsync --exclude both match by name at any
  depth), so the provenance index would have been silently stripped in transit and each
  replica would have received raw archives with no index. The rename means existing
  deployed sync scripts carry it correctly with no edit.
- sync-store.cmd (Mac): carries vintage in full, deliberately. The bytes are
  irreplaceable and the Mac is the natural second copy (~1GB/yr). Adds *.tmp/*.part to
  /XF so a sync mid-capture cannot land a partial.
- push-to-server.cmd (VPS): excludes vintage/ entirely. cot-analyzer reads prices and
  COT only, so the dash would carry ~1GB/yr of archives it never opens. Notes that
  excluding just vintage/raw/ is the right change if the dash ever consumes revisions.
- SYNCING.md: new section on where vintage may be written, the per-replica table, and
  the naming gotcha, next to the existing citpy precedent.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The README was untouched by the previous commits while the branch added two CLI entry
points, two env vars, a new store subtree and a new scheduled task — all categories the
README already documents.

- New "COT vintage tracking (as-published history)" section under Concepts & design:
  why revisions matter, the full CLI, how change-only storage / PIT reads / release-date
  provenance work, the producer-vs-replica placement rule, and why --all is a
  restatement tripwire rather than a backfill.
- Store layout gains vintage/, flagged opt-in and additive.
- Producer command list gains cotdata-vintage fetch; Contents gains the section link.
- Sync section notes the irreplaceability rule and the snapshots.json naming, since the
  usual manifest.json exclusion matches at any depth.
- WINDOWS_SCHEDULING.md: run-vintage.cmd wrapper, an optional fourth schtasks entry at
  17:00, and the daily-not-weekly rationale. Corrected "Create three tasks", which my
  addition had made wrong.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
§10 deferred this on the assumption it needed the weekly static parsed. Measuring the
file showed otherwise: it is a headerless positional CSV covering exactly ONE report
date (365 rows, 129 columns, a single distinct value in field 2), so mapping a snapshot
to its report date reads one field. The publication timestamp was already being captured
into snapshots.json. The whole thing was ~30 lines, not a canonicaliser, so it ships.

- published_from_snapshots(): retained weekly-static file -> report_date (field 2), its
  HTTP Last-Modified -> release_date converted to ET (CFTC publishes on ET, so the
  conversion decides the date for any post-midnight-UTC timestamp).
- backfill folds published rows in automatically on the production path, so plain
  `cotdata-schedule backfill` picks up true publication times with no extra step; tests
  still pass an explicit schedule to isolate precedence.
- _schedule_map now ranks published > announced > scheduled rather than special-casing
  announced-beats-scheduled.
- New `cotdata-schedule published`, and run-vintage.cmd gains published + backfill after
  ingest (both idempotent).

Verified live: report date 2026-07-21 resolves to a 2026-07-24 ET publication date,
source=published, beating the poll-derived observed bound.

Forward-only by nature: the weekly static holds one week and is overwritten, so this
covers weeks from first capture onward. Full canonicalisation of the weekly static INTO
observations is genuinely larger and remains out of scope.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
mspinola and others added 6 commits July 30, 2026 19:56
…spec

Committed verbatim as authored, before any amendment, so the specs-as-written are a
distinct point in history. The v0.2 handoff is the document the vintage subsystem was
built against; the module design is its parent (the handoff implements §5.3).

Both predate implementation and are amended in the following commit with what the build
actually established.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Moves findings out of commit messages and chat and into the docs, so the docs are the
source of truth. The specs above the new sections are left unedited: they are the
documents as written before implementation, and the amendments record where they were
wrong rather than quietly correcting them.

handoff §12 "Outcome" (following the ADR-0006 house pattern):
- The Last-Modified spike as an explicit NEGATIVE result: historical release dates
  cannot be recovered from headers, so §4.6's fallback chain is the confirmed path.
- Header sweep 2020-2026 and the ROLLING TWO-YEAR WINDOW: 2025 and 2026 share an
  identical Last-Modified, so CFTC regenerates current + prior year and everything older
  is static. Corrects an earlier claim that all closed years were re-touched weekly, and
  explains why weekly downloads have always grabbed both years. Also: no ETag is served
  on any file, so If-None-Match can never fire, while If-Modified-Since does 304.
- Decisions: retention (keep everything, no pruning), --all as a restatement TRIPWIRE
  rather than a backfill, daily producer-side capture, and the futures-only carry-forward
  that leaves `combined` constant-False with half the reportable universe absent.
- Deviations table (polars->pandas, manifest.json->snapshots.json, synthetic fixtures,
  no --from-git, and `published` shipping despite §10 deferring it).
- Answers to §11, including that backfill coverage is legitimately empty because no
  production vintages exist yet.

crowdmon_futures_cot_module.md: resolves the §4 open items (schema, backfill span,
revisions were overwritten, no release_date, futures-only) and notes that `vintage: int`
is not how it was built — the implementation is bitemporal. §5.3 gains the two findings
that change what it can assume: release date is resolved with provenance rather than
read, and `derived` fails on exactly the weeks that matter; and vintage history is
forward-only, so PIT protection cannot be backfilled.

Links to cot_vintage.md and ADR-0008 dangle on main until PRs #78 and #13 merge; called
out inline rather than left to surprise a reader.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The warning was written while these docs sat on main and cot_vintage.md did not. They
now travel in the same PR, so that link never dangles. Only the cross-repo ADR-0008
reference still depends on a separate merge (crucible-stack #13).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The byte-hash guard passed locally and failed on all five CI Python versions, with a
DIFFERENT hash on each (3.10 ebc56905, 3.13 0478a1ca, local 53e939b8). Parquet encoding
is a property of the pandas/pyarrow build -- writer metadata, compression, dictionary
encoding -- not of this repo's code, so the assertion was about the library rather than
about anything the change could keep true. It could never have gone green in CI.

What consumers actually depend on is the data read back, so that is what is compared
now, against the same pre-vintage golden file. dtypes are compared loosely because CI
spans Python 3.10-3.14 and therefore pandas 2 and 3, which disagree about the string
dtype; values, columns, order and index are still compared exactly.

Byte-stability is kept as a separate test where it IS a real property: writing identical
data twice in one environment must produce identical bytes, which is what an atomic
write could plausibly break.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Second half of the same lesson. check_dtype=False does not cover the INDEX, which
assert_frame_equal checks separately, so 3.10 still failed on datetime64[ns] vs
datetime64[us]: pandas 3 wrote the golden at microsecond resolution and pandas 2 reads
it back at nanosecond. Again a property of the library, not of this code.

Index and column CONTENT are still asserted exactly, immediately after, so relaxing the
dtype check does not quietly widen what this guard accepts.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A scheduled run's stdout goes nowhere, so a retroactive restatement -- the exact event
this subsystem exists to detect -- would have been recorded and then silently swallowed,
with exit 0 and no way to know without running 'diff' by hand.

- 'ingest' now prints the changed fields (with age_days, flagging any reaching more than
  30 days back into the calibration window) and exits NON-ZERO when it records a revision
  or sees a closed-year restatement suspect. The message says plainly that this is a
  notification and the data is committed, so it is not mistaken for a broken run.
- The suspect scan is scoped to the snapshots THIS run processed. A store-wide scan would
  re-fire on every run forever once one restatement was seen, and an alert that never
  clears is one that gets switched off. Test pins that a prior run's suspect stays quiet.
- run-vintage.cmd tees every step to a log, treats ingest's non-zero as 'revised' rather
  than 'failed' so later steps still run, and re-raises at the end so Task Scheduler shows
  the task as attention-needed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@mspinola
mspinola merged commit 3bee999 into main Jul 31, 2026
5 checks passed
@mspinola
mspinola deleted the claude/cot-revision-snapshots-9b196f branch July 31, 2026 00:21
mspinola added a commit that referenced this pull request Jul 31, 2026
Closes the notification gap left by #78.

Task Scheduler's built-in Send-an-e-mail / Display-a-message actions are deprecated and
non-functional on Windows 8 / Server 2012+, so ingest exiting non-zero only produced a
passive Last Run Result 0x1 that someone had to go and look at. A comment in
run-vintage.cmd implied Task Scheduler could alert on it; corrected.

run-vintage.cmd now notifies itself: on any run that records a revision it appends that
run's output to <store>\\vintage\\REVISIONS_<yyyy-MM-dd>.txt, which rides the existing
robocopy /MIR push to the Mac. The marker is written immediately after ingest rather than
at the end, so a later unrelated failure cannot suppress an alert for a revision that is
already committed; it appends so two runs in one day both survive; the date comes from
PowerShell rather than locale-formatted %DATE%; and the ingest output goes to a .tmp
scratch file the existing sync exclusions skip.

scripts/vintage_alert_selftest.py forces a revision in a throwaway store and drives the
real CLI, so the loud path can be seen working now instead of being trusted until CFTC
eventually restates something.
mspinola added a commit that referenced this pull request Jul 31, 2026
…decomposition (#83)

* feat(vintage): frozen-year tripwire, flow decomposition, corrected OI validation

Three things, all measured against the real store rather than assumed.

1. THE FROZEN-YEAR TRIPWIRE IS NOW AUTOMATIC

The rolling two-year regeneration window (measured in #78) means CFTC re-serves
the PRIOR year every week, byte-identical. That is the one place a content check
on closed data is free, and it is the only automated detector for the failure
mode this subsystem exists to guard against (July 2008: reports restated back to
2007-07-03).

The standing decision was "run --all monthly or quarterly". A tripwire that only
fires when somebody remembers to run it by hand is an intention, not a detector,
so the prior year moves into the DEFAULT fetch set: one ~7 MB transfer per week
(the other six days 304) and zero bytes on disk, since the content dedupes.

Three regimes now carry different expectations, recorded on every snapshot as
expectation / outcome / tripwire_alert:

  churn                 current year, new bytes are the data arriving
  frozen_in_window      prior year, expect exactly "unchanged bytes (deduped)"
  frozen_out_of_window  older, expect 304; content is never verified there

Two alert shapes, deliberately triggered differently. Content changed on a frozen
year always alerts: it is a discrete, diffable event. The detector going blind (a
frozen-in-window 304 or fetch failure) alerts on the TRANSITION only, because it
is a standing condition and a level-triggered alert would fire daily forever.

fetch still exits zero on purpose: the Windows wrapper aborts the run on a
non-zero fetch, which would skip the ingest that turns a restatement into
readable revision rows. ingest re-raises it and already writes the marker file.

2. FLOW DECOMPOSITION (module spec section 6.4)

vintage_flow.py plus `cotdata-vintage flow`. Weekly dLong vs dShort per
market/category into new_longs / short_covering / new_shorts / long_liquidation.
No prices, no contract master, so it is a clean smoke test of the canonical
schema. Classification is by DOMINANT leg, because the spec's "~0" leg never
happens in real data; parameter-free by default, with an optional min_frac_oi
dead zone that stays off since any value for it is a judgement.

Read-side only: writes no store domain, changes no producer/consumer contract.
ADR-0007 narrows cotdata by instrument domain, not derived-vs-raw, and ADR-0008
settles that COT-derived-from-COT stays inside that boundary. The positioning
engine (prices, contract master, configured weights) stays in crowdmon.

Swept over all 95 Legacy markets, 1986 to 2026: 149,412 of 149,412 weeks satisfy
the zero-sum identity (sum of long across categories == sum of short), 447,951
flow rows, zero exceptions. Two findings recorded in the docs:

  - NonComm_Positions_Spread_All is not captured and never was, so every side
    total falls short of OI by exactly the spreading (gold: 31,983 of 383,368,
    8%). Net positioning is unaffected; anything denominated as a SHARE of open
    interest is not. Not fixed here: it would change providers/cftc.py output and
    break the current/ byte-identical guarantee, so it is its own change.
  - COT was FORTNIGHTLY until 1992-10-13. days_elapsed is emitted as a column
    instead of assuming seven, so callers filter rather than discover it later.

3. CORRECTED validate()'s OPEN-INTEREST WARNING

It compared one category's long+short+spread against total OI, which is not a
bound: the two sides are counted separately. It fired on 811 of 5,778 gold rows
(14%), pure noise. Replaced with the invariant that holds, per side summed across
categories. Store-wide warnings went from thousands to zero. The module spec
asserted the wrong form too, and is amended in the same pass per the working
agreement on specs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: propose the step 2 approach (contract master + normalisation)

Proposal only, no step-2 code. Three measurements against the real store
change what step 2 can be:

1. Step 2 does not belong in cotdata. Normalisation is by definition a
   joiner (COT x prices x contract specs), and ADR-0007's whole argument
   was that almost nothing joins the domains. It belongs in the consumer,
   which means creating crowdmon-futures as a sibling.

2. The registry declares 41 disagg + 8 TFF and ZERO legacy symbols, and
   the module spec's engines all key on Managed Money / Leveraged Funds,
   which exist only in those two reports. Vintage ingest wires Legacy
   only, so a point-in-time Managed Money series does not exist yet.
   That is the real step-2 prerequisite, and it is a cotdata change.

3. Notional off back-adjusted prices is wrong by +294% (gold 2002) and
   +257% (crude 2004), and crude's back-adjusted series reaches -27.52,
   which is not a price. The error is EXACTLY ZERO at the present date
   and grows monotonically backwards, so it passes every spot check
   anyone would run while corrupting the whole evaluation history.
   Volatility must come from back-adjusted returns, so the two factors
   of net_notional x sigma come from two different price series.

Also measured: all 47 contract specs are USD (no FX layer needed),
unadjusted prices exist for all 47, and the COT-to-specs-to-price join
covers 42 of 95 Legacy markets (44%).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: drop the stray em dash from the step 2 proposal title

* feat(vintage): disaggregated and TFF canonicalisers

The step 2 proposal identified this as the real prerequisite. Every engine
in the module spec keys on Managed Money and Leveraged Funds, which exist
only in these two reports, and ingest wired Legacy only. asof() now returns
both point-in-time.

Verified against the REAL first production capture (2026-07-31 01:15Z), not
fixtures: 39,235 Disaggregated and 12,500 TFF canonical rows for 2026, all
three reports ingesting in about 5 seconds, and re-ingest writing zero rows.

These two reports populate three fields Legacy never does: per-category
spreading, per-category trader counts, and CR4/CR8 net concentration. So
the zero-sum identity closes COMPLETELY for Disaggregated (7,847/7,847
weeks, oi_gap zero everywhere), which is the confirming counterpart to the
Legacy finding that its gap really is the uncaptured spreading column.

ROUNDING TOLERANCE. The residual imbalance is CFTC's own and fully
localised: every off-by-one-or-two row in both Legacy and TFF falls in
exactly three markets, S&P 500 / NASDAQ-100 / DJIA Consolidated, whose
codes carry CFTC's own '+' suffix for a unit-converted aggregate. Tolerance
is derived from that mechanism rather than fitted: summing n independently
rounded category figures admits at most n contracts of error. Without it,
48 off-by-one warnings fire on the 2026 files alone, the same cry-wolf rate
the OI check was corrected for in the previous commit.

Implementation points that will otherwise cost someone an afternoon:

  - CFTC's header has a typo. Swap__Positions_Short_All and _Spread_All
    carry a double underscore, _Long_All does not. Both spellings resolve,
    so the day CFTC fixes it is not the day Swap Dealer ingests as nulls.
  - An unresolvable column RAISES. Silent nulls would be written as real
    observations, and the next genuine value recorded as a revision that
    never happened, polluting the exact artifact this subsystem produces.
  - Trader counts arrive as strings because CFTC writes '.' for a
    suppressed count (3,578 of 7,847 Managed Money long counts on the 2026
    file). Every value field is coerced so a marker never reaches
    row_sha256 as a literal string.
  - combined is now READ from the file's FutOnly_or_Combined column, and a
    file mixing the two is refused. Constant-FutOnly today, so nothing
    changes, but adding the combined files becomes a fetch-list change.
  - canonicalize_legacy is deliberately NOT routed through the new shared
    helper. Its output feeds row_sha256 over rows already stored in
    production; the duplication is cheaper than risking a mass revision.

Also fixes a real cross-platform bug found in the production snapshot
index: local_path is written by the Windows producer with backslashes, so
ingest on either replica failed with "no such file" on a file plainly
present. Normalised on read, so snapshots already recorded stay readable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: check cotmetrics for step 2 prior art, and correct the coverage finding

Four questions, all answered by reading ~/code/trading_workspace/cotmetrics.

1. NO notional or multiplier code. point_value / Point Value /
   contract_specs / read_metadata all return zero hits; every `multiplier`
   hit is sigma_multiplier, a z-score threshold. The one apparent
   exception is options_data.calculate_intrinsic_curve, which is EQUITY
   ETF option chains with a hardcoded *100 shares per contract, not a
   futures multiplier. Step 2 writes the notional path fresh, and no
   existing consumer needs fixing.

2. NO market-code resolver either. cftc_code / CotSymbolCodeMap /
   contract_market_code return zero hits, and cotmetrics never imports
   cotdata.registry, confirming ADR-0007 observation #3 against the code.
   Its whole cotdata surface is get_cot / get_prices / schema_version,
   each taking the INTERNAL SYMBOL and letting cotdata resolve privately.
   cotmetrics never holds a market code. crowdmon's contract_master runs
   the other direction (canonical rows are keyed by market_code), so it
   becomes the first consumer of cotdata.registry.

3. COVERAGE WAS OVERSTATED, and the headline is withdrawn. "42 of 95,
   44%" was measured against the Legacy store and framed against the full
   CFTC list, which step 2 does not aim at. Against the deployed universe:
   params.yaml is 47 (42 Role: deploy + 5 Role: heldout), 45 of 47 are
   joinable, and ALL 42 deploy markets are. The only two failures are MME
   and MFS, both heldout, both without Norgate specs or prices.
   Also caught: the two 42s (legacy-joinable, and Role: deploy) are the
   same number AND the same set for unrelated reasons, and EMD/KE/NKD were
   excluded only for having no LEGACY table, so the canonicalisers just
   shipped move the total to 45.

4. BACKADJ VERIFIED. Three production call sites, signals.py:1028,
   CotIndexer.py:655, options_data.py:366, all explicitly backadj. The
   ADR claim is true of the code, not just of the ADR. Stronger result
   from reading what each DOES with the series: none multiplies a
   historical price by a quantity (two use returns or bar shape, the third
   only the latest close, where the two series are equal by construction).
   So the +294% error is NOT shipping anywhere and there is nothing to go
   back and fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(vintage): six defects found by adversarial review

Reviewed by an agent given the spec and diff and NO prior context, because
every earlier review was done by the author and shared its blind spots.
Six confirmed defects, four in code written the same day. All reproduced
before fixing, and each now has a regression test naming the failure.

1. A FETCH FAILURE POISONED THE DEDUPE COMPARISON. Failure records carry
   no sha, etag or Last-Modified, so the next fetch sent no
   If-Modified-Since, could not match the dedupe test, and was classified
   as changed content. One network blip on a frozen year therefore
   produced a FALSE RESTATEMENT ALERT, the single alarm this subsystem
   exists to raise, and re-retained megabytes already on disk under a
   second filename carrying the identical sha. Split into
   _latest_with_content (content questions: is this the same file?) and
   _latest_for_url (outcome questions: what happened last time?).

2. AN ABSOLUTE local_path WAS RE-ROOTED UNDER THE STORE, turning
   /Volumes/ext/... into <store>/Volumes/ext/... COTDATA_VINTAGE_ROOT is
   MANDATORY on a mirrored replica, so on exactly those hosts every ingest
   raised FileNotFoundError, was swallowed, and marked the snapshot
   failed, which --pending never re-selects. One run drained the entire
   backlog with no CLI able to reset it.

3. A BARE `ingest` REPLAYED EVERY SNAPSHOT EVER RECORDED with
   observed_at = now. Replaying an older snapshot after a revision wrote
   the superseded value back with a newer timestamp, emitting a reversed
   revision (35 -> 30) then a re-revision (30 -> 35) and inflating
   age_days on both. revisions/ is the primary artifact and age_days is
   what tells a consumer whether a restatement reached its calibration
   window. Now defaults to pending, --all-snapshots opts into a replay,
   and observed_at comes from the snapshot's own retrieved_at, which is
   what it always meant.

4. `schedule backfill` BYPASSED THE WRITE LOCK while read-modify-writing
   every observations partition, so running it beside an ingest silently
   dropped that ingest's rows and left a revision row asserting a change
   to a value no longer present in observations/. Per-file atomicity was
   never the missing piece.

5. EVERY FLAT WEEK WAS LABELLED long_liquidation, because d_long <= 0
   swallows zero and the dead zone is off by default. Measured on the real
   2026 Legacy file: 3,308 of 29,787 transitions (11.1%), making
   long_liquidation the modal state with a third of its bucket being weeks
   where nothing happened, and that value_counts() is the CLI's headline.
   Zero is now `quiet` unconditionally: it is the absence of a change, not
   a small number, so recognising it needs no parameter.

6. validate() ACCEPTED DUPLICATE NATURAL KEYS and had NO NULL-RATE BAND
   despite spec section 5 requiring one. Duplicates get identical
   (observed_at, snapshot_id), exhausting the tie-break so append order
   decides the stored value. The null band closes the hole errors="coerce"
   opened: a changed value FORMAT in a column whose NAME never moved would
   coerce to nulls, get written as observations, and be read as revisions
   when the values returned. Checked per category, since one broken column
   is only 1/n of a melted frame and a frame-wide band would wave it past.

NOT FIXED, recorded as a real gap in cot_vintage.md section 9: the
`announced` release-date tier is unreachable in production and acceptance
criterion 5 is UNMET. Nothing writes source="announced" rows, so the
Oct-Dec 2025 backlog weeks can only resolve to `derived`. The plumbing and
precedence are correct; what is missing is an extractor pulling a
(report_date, release_date) pair out of free-text announcement prose, and
writing that against no corpus of real announcements would be guessing. A
release date resolved by a guess is worse than a derived one, because it
carries a provenance flag claiming it was announced.

Also corrected two places where the code contradicted its own docs: the
comment claiming an interrupted fetch "leaves a consistent store" (it can
leave an orphan blob, which is the right direction to fail in but not what
was claimed), and cot_vintage.md still describing the planned
manifests/cot.json block rather than the self-owned snapshots.json.

Verified: 249 tests pass, ruff clean, and the full ingest re-run against
the real captured snapshots is unchanged at 82,776 observations with a
second run a no-op.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(vintage): finish two incomplete fixes, plus four from a second review

The six fixes from 9a34750 were reviewed by a second cold agent, since
they were written by the same person who wrote the code being fixed and
nobody had looked at them. It downloaded EVERY real CFTC annual file (41
Legacy years, 17 Disaggregated, 17 TFF, about 1.6M canonical rows) and put
all of it through the two new raising checks.

That sweep is the most valuable result of either review: NO legitimate
duplicate natural key exists in 40 years of any report type, and the null
rate on long/short/open_interest is exactly 0.0 in every
(report_type, category) group of every year. Neither new raise can
hard-fail a historical backfill. Fixes 4 and 5 came back sound, with fix
5's 11.1% flat-week figure reproducing to the row.

TWO FIXES CLOSED THEIR REPRODUCER BUT NOT THEIR CLASS.

1. Fix 1 re-pointed restatement_suspect at the content-bearing snapshot
   but left the FIRST/CHANGED classification keyed to `prev`. So when the
   FIRST record for a URL was a fetch failure, the recovering fetch still
   fired the false restatement alert, with outcome and restatement_suspect
   disagreeing. _annotate now takes `content` and classifies FIRST off it.

2. Fix 3 corrected the observed_at stamp but not the COMPARISON:
   ingest_canonical still diffed against whatever was newest in the store,
   so --all-snapshots fabricated a reversed revision (300 -> 100) plus a
   re-revision and grew revisions/ without bound on every pass. The diff
   is now filtered to observed_at <= observed_ts, which makes it genuinely
   bitemporal. A replay is a true no-op, verified three passes deep on the
   real captured snapshots.

FOUR FURTHER FINDINGS, ALL FIXED.

- The blind-detector alert COULD NOT FIRE AT THE YEAR ROLLOVER. It needed
  prev_outcome == deduped, but a year that churned through 2025 has
  prev_outcome == changed on the January morning it becomes
  frozen-in-window. The rollover is the likeliest moment for CFTC's window
  to shift, so that was the worst possible blind spot. Now keys on whether
  the previous fetch saw bytes at all.

- A FETCH FAILURE WAS A TRIPWIRE CONDITION. Connectivity is not
  provenance: a dropped connection says nothing about whether CFTC
  restated anything, and routing it there turned one blip into a
  frozen-year restatement alarm. Failures are counted and printed by
  fetch, where an operational problem belongs.

- A RUN WHERE EVERY SNAPSHOT FAILED TO PARSE EXITED 0, reporting success
  to Task Scheduler while the store gained nothing and `failed` stayed
  terminal. Now non-zero, in ONE combined message: a first draft raised on
  the failure and silently swallowed a restatement suspect in the same
  run, which is the more serious of the two. `ingest --retry-failed` is
  the way back, and it is the only one, since nothing else ever wrote
  parse_status back to pending and two separate defects had stranded whole
  backlogs there.

- DISAGG AND TFF WERE FETCHED FROM 2006, but cftc.gov 404s for
  fut_disagg_txt_2006..2009 and fut_fin_txt_2006..2009 (verified live).
  Every `fetch --all` recorded eight permanent, unfixable failure
  snapshots. Corrected to 2010.

- THE NULL BAND WAS VACUOUS FOR LEGACY, because canonicalize_legacy never
  coerced: a value arriving as "200,000" stayed an object column, passed
  every check, and hashed differently from the numeric form. Exactly the
  fabricated-revision failure the band prevents, on the one report type
  where it could not see it. Legacy now coerces like the other two, and it
  is PROVED not to move any stored hash: across 95 markets and 448,236
  canonical rows the real values are already int64, so it is an identity.

Verified: 257 tests pass, ruff clean, and a full re-ingest of the real
captured snapshots is unchanged at 82,776 observations with both a plain
re-run and an --all-snapshots replay writing zero rows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant