Skip to content

Unify DB write serialization across sync/server/CLI - #16

Merged
thrr87 merged 5 commits into
mainfrom
claude/analyze-db-agent-locking-M0MZv
Feb 7, 2026
Merged

thrr87 merged 5 commits into
mainfrom
claude/analyze-db-agent-locking-M0MZv

Conversation

@thrr87

@thrr87 thrr87 commented Feb 7, 2026

Copy link
Copy Markdown
Owner

Summary:

  • add a shared internal write execution contract (WriteSubmit) with direct and temporary coordinator-backed submitters
  • harden WriteCoordinator with configurable lock timeout, retry budget, and retry backoff for lock/busy contention
  • route sync connector writes through submit closures and remove the direct sync write-path bypass
  • run background sync through the server writer queue and add deterministic server/background/worker shutdown with explicit singleton lock release
  • normalize transaction ownership by removing internal commits from memory store writes and centralizing commit/rollback in execution wrappers
  • add config defaults for lock_timeout_ms, retry_budget_ms, and retry_backoff_ms; update docs and analysis

Testing:

  • .venv/bin/python -m ruff check hoard/cli/main.py hoard/core/config.py hoard/core/db/connection.py hoard/core/db/lock.py hoard/core/db/writer.py hoard/core/db/write_exec.py hoard/core/ingest/sync.py hoard/core/mcp/server.py hoard/core/memory/store.py hoard/core/sync/background.py hoard/core/sync/service.py tests/conftest.py tests/test_cli.py tests/test_write_lock.py
  • .venv/bin/python -m pytest tests/test_write_lock.py -q
  • .venv/bin/python -m pytest -q

claude and others added 5 commits February 7, 2026 09:32
Detailed analysis of how Hoard prevents write conflicts when multiple
agents share a single SQLite database, covering WAL mode, the
WriteCoordinator single-writer thread, busy timeout, optimistic
concurrency checks, and async conflict detection. Identifies potential
gaps including lack of concurrent write tests.

https://claude.ai/code/session_012Zb8mCrSHJtkuCu9mqZCZ2
Introduces flock-based advisory locks to guarantee that database writes
are serialized across all processes, not just within a single server.

Changes:
- New hoard/core/db/lock.py with DatabaseWriteLock (per-write) and
  ServerSingletonLock (prevents two hoard serve on same DB)
- WriteCoordinator now acquires DatabaseWriteLock around each write task
- CLI write commands (memory put, memory prune, db migrate) use
  write_locked() context manager instead of bare connect()
- run_sync_with_lock() holds the write lock during sync operations
- MCP server acquires ServerSingletonLock on startup, fails fast if
  another server is already running

Includes 9 new tests covering lock semantics, mutual exclusion,
WriteCoordinator integration, and concurrent HTTP writes from multiple
agents.

All 142 existing tests continue to pass.

https://claude.ai/code/session_012Zb8mCrSHJtkuCu9mqZCZ2
…inator

BackgroundSync runs inside the hoard serve process. When it called
write_locked(), it acquired flock(fd_A) on hoard.db.lock. Any MCP
write then hit WriteCoordinator._run which tried flock(fd_B) on the
same file. On Linux flock is per-open-file-description not per-process,
so fd_B blocked. After the 10s timeout the writer thread died from an
unhandled DatabaseLockError, and every submit() caller hung forever on
event.wait().

Fixes:
- Revert run_sync_with_lock to plain connect() since it runs inside
  the server process where WriteCoordinator already serializes writes.
  Cross-process safety relies on SQLite WAL + busy_timeout (original
  behavior).
- Wrap the lock acquisition in WriteCoordinator._run with a try/except
  so that if the lock cannot be acquired, the error is reported to the
  caller via task.error and the writer thread keeps running.
- Move task.event.set() to an outer finally block so callers are always
  unblocked, even if lock acquisition fails.

Adds a regression test that holds the flock externally, verifies the
WriteCoordinator reports the error (not hangs), and then verifies the
writer still works after the contention clears.

All 143 tests pass.

https://claude.ai/code/session_012Zb8mCrSHJtkuCu9mqZCZ2
Updates the analysis to document all seven layers of write protection
now in place, including the new flock-based DatabaseWriteLock,
ServerSingletonLock, and the flock caveat for same-process usage.
Adds flow diagrams for CLI-vs-server write serialization and server
singleton rejection. Moves resolved gaps to a "Previously Fixed" table.

https://claude.ai/code/session_012Zb8mCrSHJtkuCu9mqZCZ2
@thrr87
thrr87 merged commit f03c354 into main Feb 7, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants