Skip to content

cluster: subscription record writes are built from the handling node's local copy #152

Description

@fabracht

Found during the quorum review of #151. After #151 the replicated subscription record (_mqtt_subs) is the source of truth for repairing the routing index, so these existing problems matter more than they did before.

Every record write is built on the node that handles the MQTT event, from that node's local SubscriptionCache, and sent as a whole-snapshot write (mqtt_state.rs add_subscription_replicated / remove_subscription_replicated). A node's local copy is only guaranteed current if it is the partition's primary or replica, or the node the client has always been connected to.

  1. Subscribe on a node without the record overwrites it. add_subscription_with_data builds a snapshot from scratch when the client is absent locally (subscription_cache.rs), so the primary's record is replaced by a single-topic snapshot, losing the client's other subscriptions.
  2. Unsubscribe on a node without the record never updates it. remove_subscription_with_data returns NotFound and nothing is written (broker_events.rs unsubscribe path), so the record keeps the topic while the index broadcast removes it everywhere.
  3. Session expiry on a stale node pushes a stale snapshot. clear_expired_session_subscriptions (event_loop.rs) sends that node's cached snapshot and broadcasts unsubscribes for the client's live entries.
  4. Copies are never garbage-collected. SubscriptionCache::clear_partition has no production caller, so a node that loses a partition keeps its copy. make replicated subscriptions authoritative in subscription reconcile #151 stops reconcile from using copies for partitions the node no longer holds, but they still feed restart recovery (rebuild_topic_index rebuilds from every locally persisted record) and snapshot export.

Before #151, reconcile overwrote the record from the index, which masked 1–3 while causing #141.

Possible direction, not yet modelled: build record changes at the partition primary (send a per-topic add/remove op instead of a whole snapshot), and clear the cache for a partition when the node stops holding it.

Related: the mqtt5 broker's ClientConnectEvent.clean_start is always true for a plain CONNECT (it is read from pending_connect, which is only set during enhanced auth; mqtt5 0.39.2 broker/client_handler/mod.rs fire_connect_event). MQDB therefore treats every session as clean and clears its subscriptions on every disconnect, which currently hides the multi-node persistent-session variants of 1–3.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions