Unverified lead, noted while investigating the peer-mesh assumption behind #139. Filing so it is not lost — this is not a confirmed bug.
SubscriptionCache::reconcile (cluster/subscription_cache.rs) appears to treat TopicIndex / WildcardStore as authoritative and the replicated SUBSCRIPTIONS cache as derived, removing cached subscriptions that are absent from the local index. If that reading is right, then on a node that missed a TopicSubscriptionBroadcast but did receive the partition-replicated write, reconciliation would delete the correct replicated record rather than restore the missing index entry — turning a missed broadcast into permanent loss. It runs on a 5-minute timer and again after partition takeover.
Worth checking at the same time: handle_wildcard_reconciliation iterates WildcardPendingStore, whose add_pending has no production callers — the only calls are inside its own #[cfg(test)] module (cluster/wildcard_pending.rs:193+). That is verified, and it means the wildcard anti-entropy timer is a permanent no-op.
To confirm or dismiss: construct a node that missed a subscription broadcast but received the replicated write, run reconcile, and observe whether the replicated record survives.
Unverified lead, noted while investigating the peer-mesh assumption behind #139. Filing so it is not lost — this is not a confirmed bug.
SubscriptionCache::reconcile(cluster/subscription_cache.rs) appears to treatTopicIndex/WildcardStoreas authoritative and the replicatedSUBSCRIPTIONScache as derived, removing cached subscriptions that are absent from the local index. If that reading is right, then on a node that missed aTopicSubscriptionBroadcastbut did receive the partition-replicated write, reconciliation would delete the correct replicated record rather than restore the missing index entry — turning a missed broadcast into permanent loss. It runs on a 5-minute timer and again after partition takeover.Worth checking at the same time:
handle_wildcard_reconciliationiteratesWildcardPendingStore, whoseadd_pendinghas no production callers — the only calls are inside its own#[cfg(test)]module (cluster/wildcard_pending.rs:193+). That is verified, and it means the wildcard anti-entropy timer is a permanent no-op.To confirm or dismiss: construct a node that missed a subscription broadcast but received the replicated write, run reconcile, and observe whether the replicated record survives.