Skip to content

cluster raft: two leaders can be elected in the same term #155

Description

@fabracht

Two nodes can be Raft leader in the same term (election safety violation). On a fresh 5-node start the second leader issues its own partition assignments, and the nodes then report different partition maps for minutes.

Evidence

Fresh 5-node cluster (cluster start, nodes started 0.5 s apart, partial-mesh --peers), polling $SYS/mqdb/cluster/status on every node every 2 s. One captured run on main (abcaf6c) with mqdb_cluster::cluster::raft::coordinator=debug:

n1 19:52:40.929 became Raft leader node=1                        (term 1)
n2 19:52:41.902 sending RequestVote to=1 term=1                  (48 ms after n2 started)
n2 19:52:43.462 received AppendEntries from=1 term=1 entries=40 prev_log_index=257
n2 19:52:43.486 became Raft follower node=2 leader=Some(1)
n5 19:52:43.543 sending RequestVote to=1..4 term=1               (85 ms after n5 started)
n2 19:52:43.559 responding to RequestVote granted=true
n3 19:52:43.559 responding to RequestVote granted=true
n4 19:52:43.559 responding to RequestVote granted=true
n5 19:52:43.570 became Raft leader node=5                        (term 1)
n1 19:52:43.995 sending AppendEntries to=2 term=1 entries=128    (n1 still leading term 1)
n1 19:52:44.512 sending AppendEntries to=2 term=1 entries=208

Both then found unique-constraint voters with term=1.

Cause

  1. RaftState::become_follower clears voted_for even when the term does not change (cluster/raft/state.rs, self.voted_for = None). Node 2 voted for itself in term 1, then became a follower of node 1 in term 1 and forgot that vote, so it could vote again in term 1. Nodes 3 and 4 did the same. Raft only resets the vote when the term increases.
  2. A joining node starts an election on its first tick. last_heartbeat_time starts at 0, so now >= last_heartbeat_time + election_timeout is already true, and can_start_election returns true at once when peers are configured (they are added from --peers at startup, cluster_agent/init.rs). Every joiner campaigned in term 1 within 26–85 ms of starting, before hearing from the existing leader.
  3. Node 2's log was empty (it had rejected node 1's first AppendEntries on prev_log_index mismatch), so node 5's empty log passed the up-to-date check.

Contributing: node 1 sent its first AppendEntries to the new peers 2.5 s after becoming leader and applied about one partition command per 10 ms, which leaves a long window for 1–3.

Frequency (16 fresh starts each, 180 s each, settled = all nodes report the same complete map with every node holding primaries for 3 consecutive polls)

second election never settled within 180 s
#139 (3c94ea5) 1/16 1/16
main (abcaf6c) 10/16 8/16

Fisher exact p = 0.002 and 0.016. Every never-settled run had a second election. The faulty Raft code predates #145–#151; which of those changes raises the frequency is not yet established.

Possibly related: #153 (a record missing on the new primary after its partition moved). Not verified.

Plan

TLA+ model of election with same-term step-down and first-tick elections (expect a two-leaders-in-one-term counterexample), then fix both defects, then verify with the same 16-run settle measurement on 5 nodes (expect 0 never-settled), then bisect #145–#151 with the settle method.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions