Skip to content

P0: recover abandoned running report jobs with durable leases #32

Description

@seonghobae

Buyer-visible gap

The standalone queue persists queued and running state, but a worker crash after claim_next() can leave a report in running indefinitely. The browser correctly shows the state, yet neither the user nor an operator has a safe way to prove the worker is gone, reclaim the job, resume it idempotently, or classify it terminal. For a commercial asynchronous report service this is a durability/recovery gap, not merely an observability issue.

Required behavior

Design and implement an explicit durable execution-lease/recovery contract without weakening the current single-node SQLite guarantees or the modular MSA ports.

  1. Add a versioned lease/attempt model for a claimed job: worker/attempt identity, claim/lease time, expiry/heartbeat policy, and bounded attempt count. New durable database objects/fields/indexes must use descriptive two-or-more-word snake_case names.
  2. A healthy worker must be able to renew only its own current lease. A stale worker must not complete/fail a job after ownership was transferred.
  3. Recovery must distinguish transient abandonment from deterministic quality_failed/failed; never turn an unknown in-flight model/artifact side effect into duplicate paid work without an explicit idempotency/effect boundary.
  4. Reclaim must be atomic. SQLite standalone behavior and any future multi-node repository must implement equivalent compare-and-swap/lease semantics through a versioned repository capability, not direct cross-service DB access.
  5. Preserve request idempotency independently from execution-attempt identity. Retrying a leased execution must not create a second public report job.
  6. Stage artifacts by attempt and publish only the current winning attempt atomically. A stale attempt cannot overwrite the completed artifact set.
  7. Expose privacy-safe operator/user recovery state without leaking worker IDs, internal paths, birth data, prompts, provider bodies, credentials, or raw traces in the public history schema.
  8. Add explicit cancellation/recovery semantics only if they can be made race-safe with the same lease design; otherwise keep cancellation out of the first PR rather than inventing a separate state machine.
  9. Update PRD/TRD, docs/architecture/DATA_MODEL.md, UML, threat model, test strategy, operability/runbook, ADR, traceability and CHANGELOG when implementation begins.
  10. Preserve standalone operation and replaceable MSA repository/artifact/interpreter ports.

Realistic acceptance tests

  • worker claims a queued job, process crashes, lease expires, a later worker atomically reclaims the same public job;
  • original stale worker attempts finish/fail after reclaim and is rejected without publishing artifacts;
  • heartbeat before expiry prevents another worker from reclaiming;
  • two recovery workers racing after expiry produce exactly one new current attempt;
  • API restart does not lose lease metadata;
  • report request Idempotency-Key semantics remain unchanged across execution recovery;
  • staged artifact from abandoned attempt remains non-public and cannot overwrite the winning attempt;
  • cleanup/recovery failures remain observable and retryable;
  • public history/status remains purpose-bound and redacted;
  • crash/recovery suite is deterministic under Python 3.11/3.12 and preserves 100% owned production statement/branch coverage and public docstrings.

Standards / architecture evidence

Ground the implementation in SQLite transaction/locking semantics and current primary secure/reliable software guidance. Record APA 7 references and claim boundaries in doctoring. This issue does not claim SOC 2/CSAP certification; it targets evidence useful for availability, processing integrity, change governance, recovery and incident controls.

Dependency order

Do not implement on a branch that races PR #31's current API/deletion/data-model work. Prefer implementation after the canonical architecture baseline merges so lease/recovery semantics can update the new ERD/operability graph coherently. PR #29's PR-steward branch is source-disjoint but should also be allowed to settle before adding unnecessary PR queue pressure.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions