From 42586bc4fa6ff12fbd45daf65da6341aff6c5904 Mon Sep 17 00:00:00 2001 From: Claude Date: Fri, 14 Aug 2026 17:47:05 +0000 Subject: [PATCH] Bound the VAX image test's wait for sshd on the clock The wait counted attempts -- 180 of them, spaced by sleep 20 -- and the comment claimed that came to about an hour. It only does when every attempt fails instantly, which is what happens while the guest's ssh port is still closed: slirp answers with an RST. An attempt against a guest that accepts the connection and then stalls costs the whole ConnectTimeout instead, which is 60s here, making the same loop nearly four hours. That is not hypothetical: a guest that stalled mid-boot sat in this step for over two hours before anyone noticed, and would have kept the runner for four. Take the deadline off the clock so the bound holds regardless of what an individual attempt costs. The overshoot is now just the attempt in flight when the hour passes, rather than a multiple of the whole budget. Also drop the "~15 minutes" figure from the comment while rewriting it. A healthy guest reaches sshd in about two minutes; the host keys stopped being generated on first boot a while ago, which is what that number described. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01M9LHkNMNNfDf2Fa1ieGT8g --- .github/workflows/build.yml | 22 +++++++++++++++------- changelog.md | 5 +++++ 2 files changed, 20 insertions(+), 7 deletions(-) diff --git a/.github/workflows/build.yml b/.github/workflows/build.yml index 7bd835e..1bff727 100644 --- a/.github/workflows/build.yml +++ b/.github/workflows/build.yml @@ -305,14 +305,22 @@ jobs: -o ConnectTimeout=60 runner@127.0.0.1 "$@" } - # The host keys are pre-generated into the image, but booting - # an emulated VAX to sshd still takes ~15 minutes — and a loaded - # CI runner (or a boot that falls back to regenerating a host key) - # can push well past that, so poll for up to ~60 minutes. Attempts - # while the port is still closed fail instantly via slirp's RST, so - # a generous cap doesn't slow the common case. + # The host keys are pre-generated into the image, so a healthy guest + # reaches sshd in a couple of minutes. A loaded CI runner, or a boot + # that falls back to regenerating a host key, can take considerably + # longer, so poll for an hour, plus whatever the attempt in flight + # when the deadline passes still costs. + # + # Bounded on the clock rather than by counting attempts, because an + # attempt does not cost a fixed amount of time. One against a closed + # port fails instantly via slirp's RST, but one against a guest that + # accepts the connection and then stalls costs the whole + # ConnectTimeout, so `180` attempts spaced by `sleep 20` is an hour + # only in the first case and nearly four in the second. A guest that + # stalled mid-boot held a runner for over two hours that way. + deadline=$(($(date +%s) + 3600)) ok= - for _ in $(seq 1 180); do + while [ "$(date +%s)" -lt "$deadline" ]; do if ssh_vax true 2>/dev/null; then ok=1; break; fi sleep 20 done diff --git a/changelog.md b/changelog.md index 3609ba2..165ee12 100644 --- a/changelog.md +++ b/changelog.md @@ -12,6 +12,11 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 lookup that went unanswered left the guest sitting at `Setting date via ntp.` It is now given numeric addresses, keeping DNS off the boot path, and a per-query timeout +- Bound the VAX image test's wait for `sshd` on the clock. It counted attempts + instead, and an attempt costs nothing against a closed port but a whole + `ConnectTimeout` against a guest that stalls after accepting, so the + intended hour was really anywhere up to four. One stalled guest held a + runner for over two hours ## [0.7.0] - 2026-08-09 ### Added