Skip to content

cluster: investigate whether SubscriptionCache::reconcile deletes correct replicated state #141

Description

@fabracht

Unverified lead, noted while investigating the peer-mesh assumption behind #139. Filing so it is not lost — this is not a confirmed bug.

SubscriptionCache::reconcile (cluster/subscription_cache.rs) appears to treat TopicIndex / WildcardStore as authoritative and the replicated SUBSCRIPTIONS cache as derived, removing cached subscriptions that are absent from the local index. If that reading is right, then on a node that missed a TopicSubscriptionBroadcast but did receive the partition-replicated write, reconciliation would delete the correct replicated record rather than restore the missing index entry — turning a missed broadcast into permanent loss. It runs on a 5-minute timer and again after partition takeover.

Worth checking at the same time: handle_wildcard_reconciliation iterates WildcardPendingStore, whose add_pending has no production callers — the only calls are inside its own #[cfg(test)] module (cluster/wildcard_pending.rs:193+). That is verified, and it means the wildcard anti-entropy timer is a permanent no-op.

To confirm or dismiss: construct a node that missed a subscription broadcast but received the replicated write, run reconcile, and observe whether the replicated record survives.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions