shuffle: refresh the disk-backlog gauge on the LogActor tick - #3361
Open
williamhbaker wants to merge 1 commit into
Open
shuffle: refresh the disk-backlog gauge on the LogActor tick#3361williamhbaker wants to merge 1 commit into
williamhbaker wants to merge 1 commit into
Conversation
`shuffle_log_disk_backlog_bytes` was written only on segment seal and on reclaim. service-kit configures the Prometheus exporter with a 600s idle timeout across all metric kinds, which deletes a metric from the registry once its generation stops advancing -- the series ends, rather than holding a stale value. A Log wedged by disk back-pressure seals nothing (`may_buffer` gates on it) and reclaims nothing (sealed segments compress within seconds, then only poll for an unlink that never comes). It stops writing the gauge entirely, so the series disappears about ten minutes in. Re-set the gauge on the actor's existing 60s tick, tying the metric's lifetime to the actor's rather than to its rate of change. Eviction still bounds `shard_id` cardinality once the actor exits, which keeps series presence a usable liveness signal.
williamhbaker
force-pushed
the
wb/shuffle-log-backlog-gauge-refresh
branch
from
August 14, 2026 22:32
cfe2637 to
b5ca088
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description:
shuffle_log_disk_backlog_byteswas written only on segment seal and on reclaim. service-kit configures the Prometheus exporter with a 600s idle timeout across all metric kinds, which deletes a metric from the registry once its generation stops advancing -- the series ends, rather than holding a stale value.A Log wedged by disk back-pressure seals nothing (
may_buffergates on it) and reclaims nothing (sealed segments compress within seconds, then only poll for an unlink that never comes). It stops writing the gauge entirely, so the series disappears about ten minutes in.Re-set the gauge on the actor's existing 60s tick, tying the metric's lifetime to the actor's rather than to its rate of change. Eviction still bounds
shard_idcardinality once the actor exits, which keeps series presence a usable liveness signal.Workflow steps:
(How does one use this feature, and how has it changed)
Documentation links affected:
(list any documentation links that you created, or existing ones that you've identified as needing updates, along with a brief description)
Notes for reviewers:
(anything that might help someone review this PR)