Skip to content

fix(vintage): --retry recovers skipped snapshots, not just failed ones - #86

Merged
mspinola merged 1 commit into
mainfrom
claude/vintage-retry-skipped
Aug 1, 2026
Merged

mspinola merged 1 commit into
mainfrom
claude/vintage-retry-skipped

Conversation

@mspinola

@mspinola mspinola commented Aug 1, 2026

Copy link
Copy Markdown
Owner

skipped means no canonicaliser existed for that report type at ingest time, and it was exactly as terminal as failed: --pending selects neither, and nothing else ever wrote parse_status back. The design notes promise that adding a canonicaliser later just means re-marking these pending and re-running, and there was no way to do the re-marking.

It nearly bit today

The disagg and TFF bytes captured 2026-07-30 sat unparsed for two days because no canonicaliser existed yet. This morning's producer run recovered 51,735 observations from them, which is the immutable landing zone working as designed.

That only worked because the producer's install predated the drain that marks unparseable sources skipped, so they were still pending. On a current install they would have been stranded with no command able to reach them.

The live case is the weekly static: every retained copy of it is skipped, so all of them would be unreachable the day it gets a canonicaliser, which is listed as still-deferred work.

Change

--retry resets both failed and skipped. Re-skipping is harmless and quiet: a source that still has no canonicaliser simply drains again, and the alert fatigue the drain exists to prevent does not apply to a flag somebody typed on purpose. Output names the breakdown so the operator sees which kind was recovered.

--retry-failed kept as an alias. It shipped earlier the same day under the narrower name and the new behaviour is a superset.

259 tests pass (2 new), ruff clean.

Generated with Claude Code

`skipped` means no canonicaliser existed for that report type at ingest
time, and it was exactly as terminal as `failed`: --pending selects
neither, and nothing else ever wrote parse_status back. The design notes
promise that adding a canonicaliser later just means re-marking these
pending and re-running, and there was no way to do the re-marking.

Nearly bit today. The disagg and TFF bytes captured 2026-07-30 sat
unparsed for two days because no canonicaliser existed yet, and this
morning's run recovered 51,735 observations from them. That only worked
because the producer's install predated the drain that marks them
`skipped`, so they were still `pending`. On a current install they would
have been stranded with no command able to reach them.

The live case is the weekly static: every retained copy is `skipped`, so
all of them would be unreachable the day it gets a canonicaliser, which is
listed as still-deferred work.

--retry now resets both states. Re-skipping is harmless and quiet: a
source that still has no canonicaliser simply drains again, and the alert
fatigue the drain was added to prevent does not apply to a flag somebody
typed on purpose. The output names the breakdown so an operator can see
which kind was recovered.

--retry-failed is kept as an alias. It shipped earlier the same day under
the narrower name and the new behaviour is a superset, so nothing anyone
already wrote breaks.

259 tests pass, ruff clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@mspinola
mspinola merged commit 1e63ce5 into main Aug 1, 2026
5 checks passed
@mspinola
mspinola deleted the claude/vintage-retry-skipped branch August 1, 2026 14:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant