You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
In cluster mode, deleting a record that has a unique constraint is two non-atomic steps: the record delete and the release of its committed unique claim. If the coordinator process dies between them, the record is gone but the committed unique claim survives, and nothing ever reclaims it. The unique value becomes permanently unclaimable — every future create with that value fails with 409 unique constraint violation, with no operator remedy short of manual store surgery.
This is the residue of the delete-leak fix from #101 (e3ca96e), not a regression: before that commit the cluster delete never released the claim at all; #101 added the release but it is non-atomic with the record delete, and no crash-recovery mechanism repairs the gap.
Where
crates/mqdb-cluster/src/cluster/node_controller/db_ops.rs, execute_json_delete — the record is committed deleted, then the claim is released, in that order:
self.db_commit(write, outbox.clone()).await; // the RECORD is deleted
self.release_unique_for_deleted_record(entity, id, &data).await; // the CLAIM is released
Three properties make the residue permanent:
The release is driven by data (the record's contents), which the delete destroys — so it cannot be retried after a restart.
The record-driven reconciler (collect_unique_reconcile_keys) enumerates existing records only, so a claim whose record is gone is never revisited.
Committed claims are TTL-exempt (cleanup_expired reclaims only uncommitted reservations).
Agent mode is not affected by this path — there the record delete and guard release are one atomic batch (crates/mqdb-agent/src/database/crud.rs). (A second, independent variant of the same root cause — a record removed without releasing its guard — reached agent mode via the TTL sweep and was closed separately in #132/#133; only this crash-mid-delete cluster path remains.)
Impact
Cluster mode, any entity with a unique constraint, on delete (including cascade delete).
The window is two adjacent awaits with no I/O between them, so per-delete probability is low, but it is unbounded under a node stall, OOM kill, or a rolling restart landing mid-delete, and it accumulates permanently — every occurrence burns one unique value for the life of the cluster.
Single-node clusters self-heal (the value-partition primary does not persist its own claims, so a restart finds nothing); multi-node clusters do not (peers persist the claim and a promoting primary re-learns it through the seal).
Reproduction
Repro tests exist on the local branch unique-delete-seam-repro (871e858, not pushed), in crates/mqdb-cluster/src/cluster/node_controller/tests.rs:
cargo test -p mqdb-cluster --lib -- crash_between orphaned_claim relearns
crash_between_record_delete_and_claim_release_wedges_the_unique_value performs exactly db_delete_prepare + db_commit and stops (simulating the crash), then runs both repair paths (reconcile_unique_claims() and unique_cleanup_expired far past the reserve TTL) and asserts a subsequent create still gets 409.
References
Design doc and proposed fix: docs/design/unique-delete-seam.md (untracked in the working tree)
Model: specs/UniqueDeleteSeam.tla (untracked). The model shows an orphan sweep alone cannot fix it — it cannot distinguish a crashed release from an in-flight one; the recommended direction is a durable release tombstone plus claim-first ordering.
Summary
In cluster mode, deleting a record that has a unique constraint is two non-atomic steps: the record delete and the release of its committed unique claim. If the coordinator process dies between them, the record is gone but the committed unique claim survives, and nothing ever reclaims it. The unique value becomes permanently unclaimable — every future create with that value fails with
409 unique constraint violation, with no operator remedy short of manual store surgery.This is the residue of the delete-leak fix from #101 (
e3ca96e), not a regression: before that commit the cluster delete never released the claim at all; #101 added the release but it is non-atomic with the record delete, and no crash-recovery mechanism repairs the gap.Where
crates/mqdb-cluster/src/cluster/node_controller/db_ops.rs,execute_json_delete— the record is committed deleted, then the claim is released, in that order:Three properties make the residue permanent:
data(the record's contents), which the delete destroys — so it cannot be retried after a restart.collect_unique_reconcile_keys) enumerates existing records only, so a claim whose record is gone is never revisited.cleanup_expiredreclaims only uncommitted reservations).Agent mode is not affected by this path — there the record delete and guard release are one atomic batch (
crates/mqdb-agent/src/database/crud.rs). (A second, independent variant of the same root cause — a record removed without releasing its guard — reached agent mode via the TTL sweep and was closed separately in #132/#133; only this crash-mid-delete cluster path remains.)Impact
Reproduction
Repro tests exist on the local branch
unique-delete-seam-repro(871e858, not pushed), incrates/mqdb-cluster/src/cluster/node_controller/tests.rs:crash_between_record_delete_and_claim_release_wedges_the_unique_valueperforms exactlydb_delete_prepare+db_commitand stops (simulating the crash), then runs both repair paths (reconcile_unique_claims()andunique_cleanup_expiredfar past the reserve TTL) and asserts a subsequent create still gets409.References
docs/design/unique-delete-seam.md(untracked in the working tree)specs/UniqueDeleteSeam.tla(untracked). The model shows an orphan sweep alone cannot fix it — it cannot distinguish a crashed release from an in-flight one; the recommended direction is a durable release tombstone plus claim-first ordering.e3ca96e; reconciler contract indocs/design/cluster-unique-hardening.md§5.Proposed fix (from the design doc)
release_epochon the value key) before touching either store.Flipping
crash_between_...anda_restarted_primary_relearns_...to expect success would validate the fix.