Two nodes can be Raft leader in the same term (election safety violation). On a fresh 5-node start the second leader issues its own partition assignments, and the nodes then report different partition maps for minutes.
Evidence
Fresh 5-node cluster (cluster start, nodes started 0.5 s apart, partial-mesh --peers), polling $SYS/mqdb/cluster/status on every node every 2 s. One captured run on main (abcaf6c) with mqdb_cluster::cluster::raft::coordinator=debug:
n1 19:52:40.929 became Raft leader node=1 (term 1)
n2 19:52:41.902 sending RequestVote to=1 term=1 (48 ms after n2 started)
n2 19:52:43.462 received AppendEntries from=1 term=1 entries=40 prev_log_index=257
n2 19:52:43.486 became Raft follower node=2 leader=Some(1)
n5 19:52:43.543 sending RequestVote to=1..4 term=1 (85 ms after n5 started)
n2 19:52:43.559 responding to RequestVote granted=true
n3 19:52:43.559 responding to RequestVote granted=true
n4 19:52:43.559 responding to RequestVote granted=true
n5 19:52:43.570 became Raft leader node=5 (term 1)
n1 19:52:43.995 sending AppendEntries to=2 term=1 entries=128 (n1 still leading term 1)
n1 19:52:44.512 sending AppendEntries to=2 term=1 entries=208
Both then found unique-constraint voters with term=1.
Cause
RaftState::become_follower clears voted_for even when the term does not change (cluster/raft/state.rs, self.voted_for = None). Node 2 voted for itself in term 1, then became a follower of node 1 in term 1 and forgot that vote, so it could vote again in term 1. Nodes 3 and 4 did the same. Raft only resets the vote when the term increases.
- A joining node starts an election on its first tick.
last_heartbeat_time starts at 0, so now >= last_heartbeat_time + election_timeout is already true, and can_start_election returns true at once when peers are configured (they are added from --peers at startup, cluster_agent/init.rs). Every joiner campaigned in term 1 within 26–85 ms of starting, before hearing from the existing leader.
- Node 2's log was empty (it had rejected node 1's first AppendEntries on
prev_log_index mismatch), so node 5's empty log passed the up-to-date check.
Contributing: node 1 sent its first AppendEntries to the new peers 2.5 s after becoming leader and applied about one partition command per 10 ms, which leaves a long window for 1–3.
Frequency (16 fresh starts each, 180 s each, settled = all nodes report the same complete map with every node holding primaries for 3 consecutive polls)
|
second election |
never settled within 180 s |
| #139 (3c94ea5) |
1/16 |
1/16 |
main (abcaf6c) |
10/16 |
8/16 |
Fisher exact p = 0.002 and 0.016. Every never-settled run had a second election. The faulty Raft code predates #145–#151; which of those changes raises the frequency is not yet established.
Possibly related: #153 (a record missing on the new primary after its partition moved). Not verified.
Plan
TLA+ model of election with same-term step-down and first-tick elections (expect a two-leaders-in-one-term counterexample), then fix both defects, then verify with the same 16-run settle measurement on 5 nodes (expect 0 never-settled), then bisect #145–#151 with the settle method.
Two nodes can be Raft leader in the same term (election safety violation). On a fresh 5-node start the second leader issues its own partition assignments, and the nodes then report different partition maps for minutes.
Evidence
Fresh 5-node cluster (
cluster start, nodes started 0.5 s apart, partial-mesh--peers), polling$SYS/mqdb/cluster/statuson every node every 2 s. One captured run onmain(abcaf6c) withmqdb_cluster::cluster::raft::coordinator=debug:Both then found unique-constraint voters with
term=1.Cause
RaftState::become_followerclearsvoted_foreven when the term does not change (cluster/raft/state.rs,self.voted_for = None). Node 2 voted for itself in term 1, then became a follower of node 1 in term 1 and forgot that vote, so it could vote again in term 1. Nodes 3 and 4 did the same. Raft only resets the vote when the term increases.last_heartbeat_timestarts at 0, sonow >= last_heartbeat_time + election_timeoutis already true, andcan_start_electionreturns true at once when peers are configured (they are added from--peersat startup,cluster_agent/init.rs). Every joiner campaigned in term 1 within 26–85 ms of starting, before hearing from the existing leader.prev_log_indexmismatch), so node 5's empty log passed the up-to-date check.Contributing: node 1 sent its first AppendEntries to the new peers 2.5 s after becoming leader and applied about one partition command per 10 ms, which leaves a long window for 1–3.
Frequency (16 fresh starts each, 180 s each, settled = all nodes report the same complete map with every node holding primaries for 3 consecutive polls)
main(abcaf6c)Fisher exact p = 0.002 and 0.016. Every never-settled run had a second election. The faulty Raft code predates #145–#151; which of those changes raises the frequency is not yet established.
Possibly related: #153 (a record missing on the new primary after its partition moved). Not verified.
Plan
TLA+ model of election with same-term step-down and first-tick elections (expect a two-leaders-in-one-term counterexample), then fix both defects, then verify with the same 16-run settle measurement on 5 nodes (expect 0 never-settled), then bisect #145–#151 with the settle method.