Problem
Recover root access without deleting data. ChangePassword requires the current password, and trusted issuers cannot map to user id 0.
Who can reset
Trust operators who control a node's server process, environment and data directory. No client reset endpoint is added.
Design
- Trigger
--reset-root-password on one node with a username and new password. Reject a username differing from stored root. Resolve the reset-input issue below before choosing environment variables.
- Recover, catch up metadata and join a normal view. Replica joins use
cluster.auth, independently of root credentials. The reset needs a functioning quorum; other nodes need no reset flag.
- Verify the password against root's stored hash; if it matches, submit nothing. Otherwise hash it once with Argon2 and build
ChangePasswordRequest for user id 0 with an empty current password.
- Use an internal proposal path that bypasses client password verification; apply in
core/metadata/src/stm/user.rs remains unchanged. The primary proposes directly; a backup forwards one hop over the replica channel, following Register.
- Success requires quorum commitment and successful application. Warn on flagged boots, log the committed operation number, never the password. Remove the flag afterward.
All replicas apply the same hash; snapshots, WAL replay and state transfer preserve it. Metadata records, client protocol and SDKs stay unchanged. Forwarding adds a header in core/binary_protocol.
Alternatives rejected
- Offline metadata edits bypass consensus and can be overwritten by state transfer.
- Bootstrap seeding cannot replace a committed reset once metadata is checkpointed.
- Network recovery secrets avoid restarting but add a command and authentication surface.
- Per-node credential files add another login source without avoiding current-password bypass.
- Requiring every node to reset adds coordination without helping replica authentication.
Risks and follow-ups
- Reusing
IGGY_ROOT_PASSWORD can seed the new password before WAL replay when no snapshot exists, causing the later check to skip replication. Separate reset input from bootstrap credentials.
- A retained flag restores the configured password after later changes.
- Backup forwarding requires an upgraded primary; define mixed-version failures.
- Root's tokens and sessions remain valid; this covers loss, not compromise.
- Without
cluster.auth, replica peers are unauthenticated.
- Define crash/view-change retries and deduplication; local hash comparison cannot establish exactly-once commitment.
Tests
- Single node: preserve data, reject old password, accept new; unchanged restart adds no operation.
- Cluster: reset through a backup, authenticate everywhere after catch-up, including state transfer.
- Exercise snapshotless recovery, crashes before/after commit, and primary changes during forwarding; verify deduplication.
- Reject missing credentials and username mismatch; client password verification remains enforced.
Problem
Recover root access without deleting data.
ChangePasswordrequires the current password, and trusted issuers cannot map to user id 0.Who can reset
Trust operators who control a node's server process, environment and data directory. No client reset endpoint is added.
Design
--reset-root-passwordon one node with a username and new password. Reject a username differing from stored root. Resolve the reset-input issue below before choosing environment variables.cluster.auth, independently of root credentials. The reset needs a functioning quorum; other nodes need no reset flag.ChangePasswordRequestfor user id 0 with an empty current password.core/metadata/src/stm/user.rsremains unchanged. The primary proposes directly; a backup forwards one hop over the replica channel, followingRegister.All replicas apply the same hash; snapshots, WAL replay and state transfer preserve it. Metadata records, client protocol and SDKs stay unchanged. Forwarding adds a header in
core/binary_protocol.Alternatives rejected
Risks and follow-ups
IGGY_ROOT_PASSWORDcan seed the new password before WAL replay when no snapshot exists, causing the later check to skip replication. Separate reset input from bootstrap credentials.cluster.auth, replica peers are unauthenticated.Tests