A node that joins after the Raft leader has compacted its log never receives the partition map through Raft. It applies no Raft commands at all and keeps only what heartbeats tell it, so some partitions show no primary on that node indefinitely.
The leader keeps the last 1000 log entries (compact_log(1000) in raft/node.rs) and there is no snapshot transfer. For a new follower (next index 1) after compaction:
- The leader sends
prev_log_index 0, prev_log_term 0 and entries_from(1). Index 1 is below the log base, so entries_from returns the whole remaining log, starting at the first kept index (for example 272).
- The follower's
append_entries skips the prev check for index 0. append_entry then drops every entry because its index is not last_log_index() + 1, and still returns success.
- The follower replies success with
last_log_index 0 and sets commit_index to 0. The leader learns nothing and the same exchange repeats on every heartbeat.
Seen live: 5 nodes on a partial mesh started 12 s apart. Rebalances produce more than 1000 entries before node 5 joins, and node 5 then reports raft_log_len 0, raft_commit_index 0 while the leader is at index 1272 with 1001 entries, and 58 partitions with no primary. Node 5 applied no Raft command in 6/6 runs on main and 12/12 on #158. The same can happen to any node added to, or replacing a node in, a cluster that has been running long enough.
Repro test through RaftNode (fails on main: follower commit index 0, leader 1201; passes with 500 entries, below the compaction threshold):
#[test]
fn follower_joining_after_compaction_catches_up() {
let peer1 = NodeId::validated(1).unwrap();
let peer2 = NodeId::validated(2).unwrap();
let mut leader = make_node(1);
leader.tick(0);
leader.tick(1000);
for _ in 0..1200 {
leader.propose(RaftCommand::Noop);
}
leader.tick(1001);
let mut follower = make_node(2);
follower.add_peer(peer1);
leader.add_peer(peer2);
for round in 0..20 {
let now = 2000 + round * 100;
for output in leader.tick(now) {
if let RaftOutput::SendAppendEntries { request, .. } = output {
let (response, _) = follower.handle_append_entries(peer1, request, now);
leader.handle_append_entries_response(peer2, response);
}
}
}
assert_eq!(follower.commit_index(), leader.commit_index());
}
Possible direction: when a follower needs an entry below the log base, send a snapshot of the applied state (the partition map plus the last included index and term) instead of entries, and have the follower install it and continue from there. The follower should also reject entries that do not follow its log instead of reporting success.
A node that joins after the Raft leader has compacted its log never receives the partition map through Raft. It applies no Raft commands at all and keeps only what heartbeats tell it, so some partitions show no primary on that node indefinitely.
The leader keeps the last 1000 log entries (
compact_log(1000)inraft/node.rs) and there is no snapshot transfer. For a new follower (next index 1) after compaction:prev_log_index 0, prev_log_term 0andentries_from(1). Index 1 is below the log base, soentries_fromreturns the whole remaining log, starting at the first kept index (for example 272).append_entriesskips the prev check for index 0.append_entrythen drops every entry because its index is notlast_log_index() + 1, and still returns success.last_log_index0 and setscommit_indexto 0. The leader learns nothing and the same exchange repeats on every heartbeat.Seen live: 5 nodes on a partial mesh started 12 s apart. Rebalances produce more than 1000 entries before node 5 joins, and node 5 then reports
raft_log_len 0,raft_commit_index 0while the leader is at index 1272 with 1001 entries, and 58 partitions with no primary. Node 5 applied no Raft command in 6/6 runs on main and 12/12 on #158. The same can happen to any node added to, or replacing a node in, a cluster that has been running long enough.Repro test through
RaftNode(fails on main: follower commit index 0, leader 1201; passes with 500 entries, below the compaction threshold):Possible direction: when a follower needs an entry below the log base, send a snapshot of the applied state (the partition map plus the last included index and term) instead of entries, and have the follower install it and continue from there. The follower should also reject entries that do not follow its log instead of reporting success.