Buyer-visible gap
The standalone queue persists queued and running state, but a worker crash after claim_next() can leave a report in running indefinitely. The browser correctly shows the state, yet neither the user nor an operator has a safe way to prove the worker is gone, reclaim the job, resume it idempotently, or classify it terminal. For a commercial asynchronous report service this is a durability/recovery gap, not merely an observability issue.
Required behavior
Design and implement an explicit durable execution-lease/recovery contract without weakening the current single-node SQLite guarantees or the modular MSA ports.
- Add a versioned lease/attempt model for a claimed job: worker/attempt identity, claim/lease time, expiry/heartbeat policy, and bounded attempt count. New durable database objects/fields/indexes must use descriptive two-or-more-word snake_case names.
- A healthy worker must be able to renew only its own current lease. A stale worker must not complete/fail a job after ownership was transferred.
- Recovery must distinguish transient abandonment from deterministic
quality_failed/failed; never turn an unknown in-flight model/artifact side effect into duplicate paid work without an explicit idempotency/effect boundary.
- Reclaim must be atomic. SQLite standalone behavior and any future multi-node repository must implement equivalent compare-and-swap/lease semantics through a versioned repository capability, not direct cross-service DB access.
- Preserve request idempotency independently from execution-attempt identity. Retrying a leased execution must not create a second public report job.
- Stage artifacts by attempt and publish only the current winning attempt atomically. A stale attempt cannot overwrite the completed artifact set.
- Expose privacy-safe operator/user recovery state without leaking worker IDs, internal paths, birth data, prompts, provider bodies, credentials, or raw traces in the public history schema.
- Add explicit cancellation/recovery semantics only if they can be made race-safe with the same lease design; otherwise keep cancellation out of the first PR rather than inventing a separate state machine.
- Update PRD/TRD,
docs/architecture/DATA_MODEL.md, UML, threat model, test strategy, operability/runbook, ADR, traceability and CHANGELOG when implementation begins.
- Preserve standalone operation and replaceable MSA repository/artifact/interpreter ports.
Realistic acceptance tests
- worker claims a queued job, process crashes, lease expires, a later worker atomically reclaims the same public job;
- original stale worker attempts
finish/fail after reclaim and is rejected without publishing artifacts;
- heartbeat before expiry prevents another worker from reclaiming;
- two recovery workers racing after expiry produce exactly one new current attempt;
- API restart does not lose lease metadata;
- report request
Idempotency-Key semantics remain unchanged across execution recovery;
- staged artifact from abandoned attempt remains non-public and cannot overwrite the winning attempt;
- cleanup/recovery failures remain observable and retryable;
- public history/status remains purpose-bound and redacted;
- crash/recovery suite is deterministic under Python 3.11/3.12 and preserves 100% owned production statement/branch coverage and public docstrings.
Standards / architecture evidence
Ground the implementation in SQLite transaction/locking semantics and current primary secure/reliable software guidance. Record APA 7 references and claim boundaries in doctoring. This issue does not claim SOC 2/CSAP certification; it targets evidence useful for availability, processing integrity, change governance, recovery and incident controls.
Dependency order
Do not implement on a branch that races PR #31's current API/deletion/data-model work. Prefer implementation after the canonical architecture baseline merges so lease/recovery semantics can update the new ERD/operability graph coherently. PR #29's PR-steward branch is source-disjoint but should also be allowed to settle before adding unnecessary PR queue pressure.
Buyer-visible gap
The standalone queue persists
queuedandrunningstate, but a worker crash afterclaim_next()can leave a report inrunningindefinitely. The browser correctly shows the state, yet neither the user nor an operator has a safe way to prove the worker is gone, reclaim the job, resume it idempotently, or classify it terminal. For a commercial asynchronous report service this is a durability/recovery gap, not merely an observability issue.Required behavior
Design and implement an explicit durable execution-lease/recovery contract without weakening the current single-node SQLite guarantees or the modular MSA ports.
quality_failed/failed; never turn an unknown in-flight model/artifact side effect into duplicate paid work without an explicit idempotency/effect boundary.docs/architecture/DATA_MODEL.md, UML, threat model, test strategy, operability/runbook, ADR, traceability and CHANGELOG when implementation begins.Realistic acceptance tests
finish/failafter reclaim and is rejected without publishing artifacts;Idempotency-Keysemantics remain unchanged across execution recovery;Standards / architecture evidence
Ground the implementation in SQLite transaction/locking semantics and current primary secure/reliable software guidance. Record APA 7 references and claim boundaries in doctoring. This issue does not claim SOC 2/CSAP certification; it targets evidence useful for availability, processing integrity, change governance, recovery and incident controls.
Dependency order
Do not implement on a branch that races PR #31's current API/deletion/data-model work. Prefer implementation after the canonical architecture baseline merges so lease/recovery semantics can update the new ERD/operability graph coherently. PR #29's PR-steward branch is source-disjoint but should also be allowed to settle before adding unnecessary PR queue pressure.