Skip to content

[fix][video-gateway] harden discovery recovery, task admission and worker retries - #53

Closed
DJSaa wants to merge 12 commits into
githubgxll:DingoRouter-basefrom
DJSaa:DingoRouter-base
Closed

DJSaa wants to merge 12 commits into
githubgxll:DingoRouter-basefrom
DJSaa:DingoRouter-base

Conversation

@DJSaa

@DJSaa DJSaa commented Sep 10, 2026 •

Copy link
Copy Markdown
Collaborator

Overview:

Improve the resilience of etcd-backed Video Gateway deployments and detached vLLM-Omni workers under discovery stalls, Runtime termination, admission contention, and worker loss.

Details:

Changes

  • Detect stale discovery watches and compare the Runtime worker view with etcd.
    Coordinate recovery across Gateway replicas using an etcd lease lock while
    preserving graceful drain behavior.
  • Exit detached video workers after unexpected permanent Runtime termination,
    allowing the supervisor to restart them instead of leaving stale processes.
  • Harden concurrent task admission and cancellation against etcd CAS races.
  • Add bounded, persistent retry accounting and a retry waiting queue for one
    retry after confirmed worker loss or known execution-engine unavailability.
  • Independently check worker liveness while waiting for results, avoiding
    indefinite waits when the result stream does not report worker loss.
  • Keep parameter, media, model, and unknown execution errors non-retryable.
  • Add targeted tests and recovery/retry telemetry.

Compatibility

  • Discovery watchdog and worker retry are opt-in.
  • The Runtime termination guard applies only to detached video workers.
  • No changes to ordinary non-video worker startup paths.

Validation

Historical validation of the corresponding image/source tree:

  • 227 unit and real-etcd contract tests passed.
  • 14 retry strategy scenarios and 6 cancellation-log scenarios passed.
  • 14 regression groups passed through per-case acceptance, including retests:
    API behavior, concurrent creation/cancellation, Gateway failover, etcd
    outages, and worker-loss retry.

These results combine the initial run and corrective retests; they do not
represent a single clean full run or production capacity/SLO certification.

Included commits

  • 3ca3302 — recover stale discovery watches safely
  • 6954663 — restart detached workers after Runtime termination
  • ce0a802 — harden admission and retry lost workers

Related Issues

🚫 This PR is NOT linked to an issue:

  • Confirmed — no related issue

Detect silent etcd watch stalls and compare the Runtime worker view with etcd discovery truth. Coordinate HA replica restarts with an etcd lease lock, preserve drain behavior, and expose bounded consistency metrics.

The watchdog is opt-in and validated for etcd-backed Video Gateway deployments, so ordinary Router, DeepSeek, GLM, and other LLM paths keep their existing behavior.
Add a process-level Runtime termination guard for detached vLLM-Omni video workers. Unexpected permanent Runtime shutdown cancels the component cleanly and exits for supervisor restart, while normal signal-driven shutdown retains graceful drain behavior.

The guard is enabled only when detached_video_task_root is configured; ordinary Omni and non-video worker startup paths are unchanged.
@DJSaa
DJSaa deployed to external_collaborator September 10, 2026 09:38 — with GitHub Actions Active
@DJSaa
DJSaa deployed to external_collaborator September 10, 2026 09:38 — with GitHub Actions Active
@github-actions

Copy link
Copy Markdown

👋 Hi DJSaa! Thank you for contributing to githubgxll/dynamo.

Just a reminder: The NVIDIA Test Github Validation CI runs an essential subset of the testing framework to quickly catch errors.Your PR reviewers may elect to test the changes comprehensively before approving your changes.

🚀

Pin vLLM 0.29.0 and Omni 0.29.0rc1, retain compiled framework constraints, include and validate ffprobe, and honor the selected native build parallelism.
Add fenced slot capacity, durable result handoff, independent recoverable finalization, bounded Worker prefetch, binary artifacts and native reusable frame conversion. Keep blocking filesystem work off the event loop and harden shared-counter CAS transitions. Document capacity metrics and upgrade boundaries. Validated with 451 Gateway tests, detached manager tests, 40/80/120 virtual Worker stages and real Omni continuous-execution fault cases.
@DJSaa
DJSaa deployed to external_collaborator September 20, 2026 16:40 — with GitHub Actions Active
@github-actions github-actions Bot added documentation Improvements or additions to documentation backend::vllm container labels Sep 20, 2026
@DJSaa
DJSaa deployed to external_collaborator September 23, 2026 09:38 — with GitHub Actions Active
@DJSaa DJSaa closed this Sep 24, 2026

This branch was successfully deployed

1 active deployment
external_collaborator — 4e8c3242 Deployed Sep 23, 2026 by DJSaa via ok-to-test #129
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant