Skip to content

cluster delete: crash between record delete and unique-claim release strands the unique value permanently #136

Description

@fabracht

Summary

In cluster mode, deleting a record that has a unique constraint is two non-atomic steps: the record delete and the release of its committed unique claim. If the coordinator process dies between them, the record is gone but the committed unique claim survives, and nothing ever reclaims it. The unique value becomes permanently unclaimable — every future create with that value fails with 409 unique constraint violation, with no operator remedy short of manual store surgery.

This is the residue of the delete-leak fix from #101 (e3ca96e), not a regression: before that commit the cluster delete never released the claim at all; #101 added the release but it is non-atomic with the record delete, and no crash-recovery mechanism repairs the gap.

Where

crates/mqdb-cluster/src/cluster/node_controller/db_ops.rs, execute_json_delete — the record is committed deleted, then the claim is released, in that order:

self.db_commit(write, outbox.clone()).await;                    // the RECORD is deleted
self.release_unique_for_deleted_record(entity, id, &data).await; // the CLAIM is released

Three properties make the residue permanent:

  1. The release is driven by data (the record's contents), which the delete destroys — so it cannot be retried after a restart.
  2. The record-driven reconciler (collect_unique_reconcile_keys) enumerates existing records only, so a claim whose record is gone is never revisited.
  3. Committed claims are TTL-exempt (cleanup_expired reclaims only uncommitted reservations).

Agent mode is not affected by this path — there the record delete and guard release are one atomic batch (crates/mqdb-agent/src/database/crud.rs). (A second, independent variant of the same root cause — a record removed without releasing its guard — reached agent mode via the TTL sweep and was closed separately in #132/#133; only this crash-mid-delete cluster path remains.)

Impact

  • Cluster mode, any entity with a unique constraint, on delete (including cascade delete).
  • The window is two adjacent awaits with no I/O between them, so per-delete probability is low, but it is unbounded under a node stall, OOM kill, or a rolling restart landing mid-delete, and it accumulates permanently — every occurrence burns one unique value for the life of the cluster.
  • Single-node clusters self-heal (the value-partition primary does not persist its own claims, so a restart finds nothing); multi-node clusters do not (peers persist the claim and a promoting primary re-learns it through the seal).

Reproduction

Repro tests exist on the local branch unique-delete-seam-repro (871e858, not pushed), in crates/mqdb-cluster/src/cluster/node_controller/tests.rs:

cargo test -p mqdb-cluster --lib -- crash_between orphaned_claim relearns

crash_between_record_delete_and_claim_release_wedges_the_unique_value performs exactly db_delete_prepare + db_commit and stops (simulating the crash), then runs both repair paths (reconcile_unique_claims() and unique_cleanup_expired far past the reserve TTL) and asserts a subsequent create still gets 409.

References

  • Design doc and proposed fix: docs/design/unique-delete-seam.md (untracked in the working tree)
  • Model: specs/UniqueDeleteSeam.tla (untracked). The model shows an orphan sweep alone cannot fix it — it cannot distinguish a crashed release from an in-flight one; the recommended direction is a durable release tombstone plus claim-first ordering.
  • Prior art: harden cluster unique constraints: delete-leak fix, non-blocking reconciler, failover epoch fence #101, commit e3ca96e; reconciler contract in docs/design/cluster-unique-hardening.md §5.

Proposed fix (from the design doc)

  1. Write a durable release tombstone (a monotone release_epoch on the value key) before touching either store.
  2. Release the claim first, then delete the record.
  3. Do not add an unfenced orphan sweep.

Flipping crash_between_... and a_restarted_primary_relearns_... to expect success would validate the fix.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions