You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
fix: WebUI Activity gate collides with host GPU work and stale incremental cache #2111
The WebUI installed artifact job (.github/workflows/ci.yml:632) on the GB10 runner lablup-dgxspark21 fails for two environmental reasons.
Lock collision. The Activity gate (:829-857) takes its own flock under $HOME/.cache/mlxcel/ci-locks, then fails closed at :833-837 if any GPU compute process exists. Development sessions on this host use gpu-lock (/usr/local/bin/gpu-lock, a flock on /tmp/gpu-lock-$(id -u)/lock), and the gate ignores that lock. Run 36982300183 attempt 1 (PR fix(tts): cast bf16 tables explicitly where the CUDA overlay demotes #2093) failed this way and passed on a rerun made while holding gpu-lock. The runner runs as uid 1000 with PrivateTmp=no, so it sees the same lock file.
Stale incremental state. The target dir $HOME/.cargo-target/mlxcel-webui-installed-ci (:659, profile test-fast with incremental = true, Cargo.toml:346-351, about 60 GB) hit a reproducible link failure: undefined serde_json generics in CGU c0jzxrmgkpaphvspsqkdqpw09. Run 36988516810 (PR fix(server): forward echoed reasoning under reasoning_content too #2094) failed at the build step on attempts 1-5 and passed on attempt 6 only after the incremental state was cleared.
Proposed fix
When command -v gpu-lock succeeds, run the gate body under gpu-lock run --tag webui-activity --wait 600 -- ... inside the existing flock. On timeout, print gpu-lock status. With no gpu-lock on the host, keep the current behavior. The oom_score_adj=1000 that gpu-lock sets is acceptable.
Set CARGO_INCREMENTAL=0 for this job only. Rejected: clean-and-retry on link failure, which hides the cause and costs a full rebuild. Other GB10 target dirs are out of scope.
Document both behaviors in docs/webui-integration-matrix.md.
This edits the same block as open #1949 (quiet-host wait). Land it with #1949 or stack it on that PR, with the gpu-lock wait placed first. Coordinate with #1925.
Acceptance Criteria
A run started while gpu-lock is held waits, which the job log shows, then passes after release.
The job builds with CARGO_INCREMENTAL=0; the PR records the warm build time.
actionlint passes and WebUI installed artifact is green on the PR.
Do it in one PR and keep pushes to a minimum, because each push costs a full GB10 gate cycle.
Problem
The
WebUI installed artifactjob (.github/workflows/ci.yml:632) on the GB10 runnerlablup-dgxspark21fails for two environmental reasons.:829-857) takes its own flock under$HOME/.cache/mlxcel/ci-locks, then fails closed at:833-837if any GPU compute process exists. Development sessions on this host usegpu-lock(/usr/local/bin/gpu-lock, a flock on/tmp/gpu-lock-$(id -u)/lock), and the gate ignores that lock. Run 36982300183 attempt 1 (PR fix(tts): cast bf16 tables explicitly where the CUDA overlay demotes #2093) failed this way and passed on a rerun made while holdinggpu-lock. The runner runs as uid 1000 withPrivateTmp=no, so it sees the same lock file.$HOME/.cargo-target/mlxcel-webui-installed-ci(:659, profiletest-fastwithincremental = true,Cargo.toml:346-351, about 60 GB) hit a reproducible link failure: undefinedserde_jsongenerics in CGUc0jzxrmgkpaphvspsqkdqpw09. Run 36988516810 (PR fix(server): forward echoed reasoning under reasoning_content too #2094) failed at the build step on attempts 1-5 and passed on attempt 6 only after the incremental state was cleared.Proposed fix
command -v gpu-locksucceeds, run the gate body undergpu-lock run --tag webui-activity --wait 600 -- ...inside the existing flock. On timeout, printgpu-lock status. With nogpu-lockon the host, keep the current behavior. Theoom_score_adj=1000thatgpu-locksets is acceptable.CARGO_INCREMENTAL=0for this job only. Rejected: clean-and-retry on link failure, which hides the cause and costs a full rebuild. Other GB10 target dirs are out of scope.docs/webui-integration-matrix.md.This edits the same block as open #1949 (quiet-host wait). Land it with #1949 or stack it on that PR, with the
gpu-lockwait placed first. Coordinate with #1925.Acceptance Criteria
gpu-lockis held waits, which the job log shows, then passes after release.CARGO_INCREMENTAL=0; the PR records the warm build time.actionlintpasses andWebUI installed artifactis green on the PR.Do it in one PR and keep pushes to a minimum, because each push costs a full GB10 gate cycle.