Skip to content

cluster raft: node joining after log compaction never catches up #159

Description

@fabracht

A node that joins after the Raft leader has compacted its log never receives the partition map through Raft. It applies no Raft commands at all and keeps only what heartbeats tell it, so some partitions show no primary on that node indefinitely.

The leader keeps the last 1000 log entries (compact_log(1000) in raft/node.rs) and there is no snapshot transfer. For a new follower (next index 1) after compaction:

  1. The leader sends prev_log_index 0, prev_log_term 0 and entries_from(1). Index 1 is below the log base, so entries_from returns the whole remaining log, starting at the first kept index (for example 272).
  2. The follower's append_entries skips the prev check for index 0. append_entry then drops every entry because its index is not last_log_index() + 1, and still returns success.
  3. The follower replies success with last_log_index 0 and sets commit_index to 0. The leader learns nothing and the same exchange repeats on every heartbeat.

Seen live: 5 nodes on a partial mesh started 12 s apart. Rebalances produce more than 1000 entries before node 5 joins, and node 5 then reports raft_log_len 0, raft_commit_index 0 while the leader is at index 1272 with 1001 entries, and 58 partitions with no primary. Node 5 applied no Raft command in 6/6 runs on main and 12/12 on #158. The same can happen to any node added to, or replacing a node in, a cluster that has been running long enough.

Repro test through RaftNode (fails on main: follower commit index 0, leader 1201; passes with 500 entries, below the compaction threshold):

#[test]
fn follower_joining_after_compaction_catches_up() {
    let peer1 = NodeId::validated(1).unwrap();
    let peer2 = NodeId::validated(2).unwrap();
    let mut leader = make_node(1);
    leader.tick(0);
    leader.tick(1000);
    for _ in 0..1200 {
        leader.propose(RaftCommand::Noop);
    }
    leader.tick(1001);
    let mut follower = make_node(2);
    follower.add_peer(peer1);
    leader.add_peer(peer2);
    for round in 0..20 {
        let now = 2000 + round * 100;
        for output in leader.tick(now) {
            if let RaftOutput::SendAppendEntries { request, .. } = output {
                let (response, _) = follower.handle_append_entries(peer1, request, now);
                leader.handle_append_entries_response(peer2, response);
            }
        }
    }
    assert_eq!(follower.commit_index(), leader.commit_index());
}

Possible direction: when a follower needs an entry below the log base, send a snapshot of the applied state (the partition map plus the last included index and term) instead of entries, and have the follower install it and continue from there. The follower should also reject entries that do not follow its log instead of reporting success.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions