Skip to content

feat: reconcile the ingest manifest in BOTH directions - #62

Merged
mspinola merged 2 commits into
mainfrom
feat/reconcile-ingest
Jul 27, 2026
Merged

mspinola merged 2 commits into
mainfrom
feat/reconcile-ingest

Conversation

@mspinola

Copy link
Copy Markdown
Owner

Stacked on #40 (databento-ingest-hardening), branched from ce63560.

The gap

reconcile_manifest() already fixes a manifest that has fallen behind the disk: an interrupted run leaves fetched tables unrecorded and a restart re-downloads them.

The mirror image was unhandled, and it is the worse one. ingest() derives each table's start from the manifest and skips it when already current:

rec = manifest.get(key, {})
start = rec["last_date"] + 1 day if rec else cold_start
if start > end: continue          # "already current"

It never checks the parquet exists. A manifest entry whose file is gone (partial copy, store move, deleted file) is therefore treated as current and never fetched again. No error, no rows, a permanent hole in a paid dataset.

The asymmetry matters: manifest-behind costs money re-downloading data you already have. Manifest-ahead leaves you believing you have data you do not.

Change

reconcile_manifest() now prunes entries with no parquet on disk, so the next run re-fetches them. Returns {"recorded": {...}, "pruned": [...]}. prune=False keeps the old behaviour.

--reconcile-databento reports the prune separately and states what would otherwise have happened:

databento reconcile: PRUNED 3 manifest entries with no parquet on disk, across 2
symbol(s). The next --ingest-databento will re-fetch these; without the prune it
would have skipped them as 'already current'.

Important: this is a between-runs repair

The manifest legitimately runs ahead of the disk while a batch ingest is in flight, because the batch API completes and then lands its files in bulk. Observed on a live run: 11 files at one moment, 145 ten minutes later.

So do not run this against a live ingest. It would prune entries whose files are still landing, and the running process holds the manifest in memory and would overwrite the result anyway.

Tests

Three new, covering both directions plus the opt-out. Suite 121 passed, ruff clean.

Note the fixture had to write ts_event as the index, not a column, matching what _append_raw produces. A column-based fixture passes construction and then fails in _date_bounds.

Development note

Built in a git worktree rather than by checking out this branch, because a --ingest-databento run has been live in the main checkout for 19 hours and cotdata is an editable install there, so a branch switch would have changed the code under a running paid job.

Base automatically changed from databento-ingest-hardening to main July 26, 2026 22:44
@mspinola
mspinola force-pushed the feat/reconcile-ingest branch from bbb568b to 9ddd405 Compare July 27, 2026 02:39
mspinola and others added 2 commits July 27, 2026 19:18
reconcile_manifest() backfilled the manifest from the raw parquets on disk, which
fixes a manifest that has fallen BEHIND the disk: an interrupted run leaves
fetched tables unrecorded and a restart re-downloads them.

The mirror image was unhandled and is the more dangerous one. ingest() derives
each table's start date from the manifest's last_date and skips the table when
that is already current:

    rec = manifest.get(key, {})
    start = rec["last_date"] + 1 day if rec else cold_start
    if start > end: continue          # "already current"

It never checks the parquet exists. So a manifest entry whose file is gone (a
partial copy, a store move, a deleted file) is treated as current and NEVER
fetched again. No error, no rows, a permanent hole in a paid dataset. The first
failure costs money re-downloading data you already have. This one leaves you
believing you have data you do not.

reconcile_manifest now prunes those entries so the next run re-fetches them, and
returns {"recorded": {...}, "pruned": [...]}. --reconcile-databento reports the
prune separately and says plainly what would otherwise have happened. prune=False
keeps the old behaviour.

Note the manifest legitimately runs ahead of the disk WHILE a batch ingest is in
flight, since the batch API completes and then lands its files in bulk. Observed
on a live run: 11 files at one moment, 145 ten minutes later. So this is a
between-runs repair, not something to run against a live ingest.

Suite 121 passed (3 new tests), ruff clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The prune direction deleted a manifest entry whenever its parquet was
missing. That is wrong for one entry shape, and it is the shape this very
feature makes dangerous.

A paged pull that meets an empty window writes `{last_date: ...}` with NO
parquet, deliberately (the `elif chunk_days` branch of _ingest_range):
"Advance past it so a re-run doesn't refetch the empty span forever." For a
symbol whose data begins after the requested floor, EVERY window is empty,
so that marker is the only thing the manifest ever gets and no parquet
directory is created at all. Reproduced against the real writer:

    MANIFEST: {"ES.n.0:statistics": {"last_date": "2020-01-05"},
               "ES.n.1:statistics": {"last_date": "2020-01-05"}}
    ohlcv/: (absent)   statistics/: (absent)

Pruning on "no file" alone deletes exactly those markers and restores the
refetch loop they exist to prevent, on every run, against a paid API. This
function exists to stop a silent hole in paid data; as written it would have
traded that for a silent spend.

The test is now "claims rows AND has no file". `_claims_rows` reads n_rows,
so a record of fetched data is a candidate and a bare advance marker is not.
It errs toward KEEPING on any shape it does not recognise, including the
newer `windowed` / `batch` / `reconciled` fields: a wrongly-kept entry costs
one skipped table that --ingest-databento reports, a wrongly-pruned marker
costs a paid refetch every run thereafter.

Five tests, and all five fail against the previous predicate (verified by
reverting it). The end-to-end one drives the real ingest rather than a
hand-built manifest; its first version used data starting INSIDE the range,
which produced a marker the next window immediately overwrote, so it skipped
and proved nothing.

Found via a cross-session note from the databento work. That note predicted
a merge conflict, which did not exist (GitHub reports MERGEABLE/CLEAN), but
was right that the two reconcile directions had to be checked against each
other's manifest fields.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@mspinola
mspinola force-pushed the feat/reconcile-ingest branch from 9ddd405 to ee2c67c Compare July 27, 2026 23:24
@mspinola
mspinola merged commit a952028 into main Jul 27, 2026
5 checks passed
@mspinola
mspinola deleted the feat/reconcile-ingest branch July 27, 2026 23:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant