Skip to content

Repair the failure that actually costs correctness - #20

Merged
Wired4ncer merged 2 commits into
mainfrom
feat/block-catchup
Aug 9, 2026
Merged

Wired4ncer merged 2 commits into
mainfrom
feat/block-catchup

Conversation

@Wired4ncer

Copy link
Copy Markdown
Owner

Writing the reconciler meant first working out which failure still needs repairing, and the write rule from #17/#19 changed the answer.

Which drop costs what

  • A dropped rawtx no longer costs the record anything. All 832 of them in the measured 40-minute window arrive again in the blocks that confirm them, and the record ends up identical. What is lost is the mempool alert — latency, permanently, because no repair recovers a moment that has passed.
  • A dropped rawblock costs correctness outright. Blocks are the only writer, so coins that arrived stay invisible and coins that left stay armed, while the sequence counters stay perfectly contiguous throughout.

So this follows the chain rather than diffing the UTXO set: it notices the block in hand is not the child of the last one applied, and fetches the difference by height before letting it through. No descriptors, no scan, no plaintext.

What is deliberately not here

⚠️ The UTXO-set diff this issue asks for cannot be built against the current schema. scantxoutset takes descriptors, a descriptor contains the scriptPubKey, and the schema stores spk_hmac — one-way. There is nothing to reconstruct a descriptor from.

The second commit decides that rather than leaving it open: an spk_ct column, AEAD(k_store, spk). The published claim is untouched — the database alone stays unmatchable noise, because ciphertext without the key is noise — and it barely moves the database-plus-master case, which §3b already classes as total, since an attacker holding k_match can hash any candidate script and test it against the public UTXO set. That makes every watched script holding coins already recoverable with effort; this turns recoverable-with-effort into readable, under a compromise already conceded.

Rejected on cost, not on merit: a dumptxoutset walk needs no plaintext at rest at all, and costs a multi-gigabyte dump plus HMACing ~166M outputs every pass, on a host chosen for being lean.

No code for it in this PR — the column lands with the enrolment API that writes it.

Reorgs are detected and refused, not repaired. The record keeps no per-block provenance — the outpoint set records nothing about which block added which coin — so there is nothing to roll back, and a reorg can un-confirm a spend and put back a coin already dropped. Saying so loudly is the honest behaviour; the alternative is a record that is wrong and a service that believes it is fine. Closing it properly needs the diff above.

Three details that are load-bearing rather than defensive

  • Blocks fetched during catch-up are chained to each other, not trusted by height. A reorg while walking would otherwise splice two branches into a record that matches neither.
  • A pruned gap fails rather than skipping. The prune target is the catch-up window. Losing precision about when something happened is survivable; inventing a record is not.
  • A fresh follower does not invent a starting point. Guessing the node's current tip would silently skip everything that happened before the service started — the hole the enrolment baseline scan exists to fill.

The repair layer is wired in as a hook rather than an import: nothing under match/ talks to the node, which is what lets the whole package run against fixtures with no network. expand_block=None is the fixture path and applies exactly what arrives.

What the sweep caught

The chaining check was not actually tested. The test I wrote for it replaced the last fetched block, which the outer check catches anyway — so the inner check could have been deleted with the suite still green. Replacing an earlier block is what isolates it.

It also turned up an equivalent mutant worth recording rather than deleting: clearing the reconciliation flag before the work rather than after leaves every exit in the same state, because nothing inside the loop reads it. It stops being equivalent the moment something does, which is why the source keeps the safe ordering.

test_a_dropped_funding_transaction_disarms_the_alarm_on_the_spend is rewritten rather than deleted, as its own docstring asked when the reconciler landed: it now asserts that the alert is lost and the record is not.

Checks

165 tests, ruff clean, mutation sweep 60/60 with 3 documented equivalents.

Refs #4

🤖 Generated with Claude Code

Wired4ncer and others added 2 commits August 9, 2026 13:14
Writing the reconciler first meant working out which failure still needs
repairing, and the write rule changed the answer.

A dropped `rawtx` no longer costs the record anything. All 832 of them in the
measured window arrive again in the blocks that confirm them, and the record
ends up identical. What is lost is the mempool alert -- latency, permanently,
because no repair recovers a moment that has passed. A dropped `rawblock` is the
one that costs correctness outright: blocks are the only writer, so coins that
arrived stay invisible and coins that left stay armed, while the sequence
counters stay perfectly contiguous throughout.

So this follows the chain rather than diffing the UTXO set. It notices the block
in hand is not the child of the last one applied and fetches the difference by
height before letting it through. No descriptors, no scan, no plaintext.

⛔ The UTXO-set diff issue #4 asks for is NOT here, and cannot be built against
the current schema. `scantxoutset` takes descriptors, a descriptor contains the
scriptPubKey, and we store `spk_hmac` -- one-way. There is nothing to
reconstruct a descriptor from. The two ways out are both decisions rather than
code: an `spk_ct` column under `k_store`, or a whole-UTXO-set walk that needs no
plaintext but costs a multi-gigabyte dump on a host chosen for being lean. Both
are recorded in docs/architecture.md 4 rather than quietly worked around, and
the second is what the reorg path is waiting on.

Reorgs are detected and refused, not repaired. The record keeps no per-block
provenance -- the outpoint set records nothing about which block added which
coin -- so there is nothing to roll back, and a reorg can un-confirm a spend and
put back a coin already dropped. Saying so loudly is the honest behaviour; the
alternative is a record that is wrong and a service that believes it is fine.

Three details that are load-bearing rather than defensive. Blocks fetched during
catch-up are chained to each other rather than trusted by height, because a
reorg *while walking* would otherwise splice two branches into a record that
matches neither. A pruned gap fails rather than skipping: the prune target is
the catch-up window, and losing precision about when something happened is
survivable where inventing a record is not. And a fresh follower does not invent
a starting point -- guessing the node's current tip would silently skip
everything before the service started.

The repair layer is a hook rather than an import. Nothing under `match/` talks
to the node, which is what lets the whole package run against fixtures with no
network; `expand_block=None` is the fixture path and applies exactly what
arrives.

The sweep found the chaining check was not actually tested -- the test I wrote
for it replaced the *last* fetched block, which the outer check catches anyway,
so the inner one could be deleted with the suite still green. Replacing an
earlier block is what isolates it. It also turned up an equivalent mutant worth
recording: clearing the reconciliation flag before the work rather than after
leaves every exit in the same state, since nothing inside the loop reads it --
and it stops being equivalent the moment something does.

`test_a_dropped_funding_transaction_disarms_the_alarm_on_the_spend` is rewritten
rather than deleted, as its own docstring asked: it now asserts the alert is
lost and the record is not.

Refs #4

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The diff cannot be built while the watched script exists only as a one-way
HMAC: scantxoutset takes descriptors, and there is nothing to reconstruct one
from. Deciding it now rather than at the moment of writing that code, because
it is a threat-model call and not a coding one.

An spk_ct column, AEAD(k_store, spk). The published claim is untouched -- the
database alone stays unmatchable noise, since ciphertext without the key is
noise. It barely moves the database-plus-master case either, which 3b already
classes as total: an attacker holding k_match can hash any candidate script and
test it against the public UTXO set, so every watched script holding coins is
already recoverable with effort. This turns that into readable, under a
compromise already conceded.

The rejected option is genuinely stronger and was rejected on cost: a
dumptxoutset walk needs no plaintext at rest at all, and costs a multi-gigabyte
dump plus HMACing ~166M outputs per pass on a host chosen for being lean.

Schema and enrolment updated to match. No code yet -- the column lands with the
enrolment API that writes it.

Refs #4

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant