You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A release tag's docs deployment can silently lose to the main-push deployment for the same commit, leaving the released version rendered on gh-pages but never served. A release PR always edits docs/guide/** (the {{reactorVersion}} substitutions), so the push that merges it matches the docs workflow's path filter, and the release tag then points at that same merge commit. Both runs hand actions/deploy-pages the same pages_build_version (it is github.sha, with no input to override it), so the two deployments collide under one identity and one is stranded. Cutting 0.1.0-preview.16 that is exactly what happened: every job green, the deployment reporting success, gh-pages byte-correct, and the live site still serving preview.15 for 2h25m.
Nothing in the pipeline looked at the live site, so no amount of green CI could have caught it.
A verify job now polls the deployed site and fails unless it is serving the artifact that run produced. Its oracle is deploy-stamp.json, written into the Pages artifact by publish and carrying the run id and attempt. Comparing versions.json alone would be vacuous for a main push, since mike deploy main republishes an existing version and leaves the version set unchanged. Beyond the stamp it asserts that every version that run published serves its own index.html (not just the latest holder, which a backported tag deliberately does not move and a main push never touches), that the latest/ alias tree is served, and that the site root parses to a redirect at the expected default. Every probe uses a unique ?nc= key so a pass cannot come from cache, rejects a followed redirect so a bounce cannot masquerade as a 200, and is bounded by an abort timeout so one stalled socket cannot eat the polling window. An already-published version fetched with the identical request shape is the positive control that separates a stranded deploy (exit 1) from a broken probe (exit 2).
The duplicate deployment is tolerated rather than avoided. Standing one of the two runs down was implemented and then removed during review. Every form of it requires predicting that the other run will deploy, and each way that prediction fails is silent in the direction that matters: a pending tag run can be evicted from the concurrency queue, a tag deleted or re-cut on origin leaves a stale local ref, an unreachable origin makes the check answer "no tag here", and the two runs are not guaranteed to enter the concurrency group in event order. Each of those skips the deployment and its verification together, leaving the site stale with every job green — strictly worse than a duplicate that verify catches. Issue #1268 itself noted this mitigation reduces the window rather than closing it, with verification carrying correctness.
The concurrency group now sets queue: max. The Actions default keeps at most one pending run per group and cancels the previous one when a new run queues, so a docs push landing while a release tag's run waited behind main would evict the tag run before it started, and the release would never reach gh-pages at all.
45 node cases in .github/scripts/verify-docs-deployment.test.mjs; run them with node --test --test-reporter=tap .github/scripts/verify-docs-deployment.test.mjs. The expected count is pinned by MinimumCases in DocsDeploymentVerifierTests, which is the source of truth if these two ever disagree — it exists so the suite cannot be silently gutted while the xUnit gate stays green.
Every comparison mutation-checked. Disabling each check in evaluate(), probe(), selectProbeTargets() and parseRootRedirect() in turn reddens exactly the intended cases and no others. Three mutations initially survived — nothing asserted the driver actually requests the version, root or alias URLs — and assertions were added until they reddened.
Workflow seams mutation-checked against DocsDeployWiringTests. Two assertions initially survived: one matched a cat line rather than the write, and one checked only that a publishing step mentions the record file, so moving the write back above mike deploy still passed. Both tightened until the regression reddens.
Ran the real Assemble the Pages artifact step with jq against origin/gh-pages: valid stamp, single-line versions= output, and the published/versions.json consistency guard exercised across consistent, divergent and empty inputs (exit 0, 1, 1). The first version of that guard's jq was wrong — inside select() the dot rebinds, so the filter errored and the guard fired on the passing case — caught only by running it.
Ran the verifier end to end over HTTP against the assembled artifact: pass (0), stranded (1, naming both run ids), probe-broken (2), and both usage-error paths (3). With 0.1.0-preview.14/index.html removed, the old latest-only target selection exits 0 while the new one exits 1 naming the 404.
Root redirect verified against real bytes: healthy passes; rewriting the stub to forward to main/ exits 1 with root-mistargeted.
dotnet test tests/Reactor.DocPipeline.Tests -> 472/472 pass.
dotnet test tests/Reactor.Tests -> 14144 total, 0 failed, 64 skipped.
docs.yml parses as YAML and every run block passes bash -n.
dotnet restore Reactor.slnx then dotnet build Reactor.slnx --no-restore -c Release. See the note below, since this needed a short-path checkout to be meaningful.
Three instrument failures are worth recording, because in each case a check reported a pass it had not earned — or could not have earned:
Probes originally requested the bare directory, and python -m http.server answers a directory with no index by generating a 200 listing while Pages 404s it — so the local end-to-end check masked the very break it was meant to demonstrate. Probes now request <version>/index.html.
The first request-timeout implementation used AbortSignal.timeout(), whose timer is unref'd and therefore never fires when nothing else holds the event loop open. It passed locally and hung on CI, cancelling four cases. Measured directly: AbortSignal.timeout(300) lets the process exit at +1ms without firing, while an owned setTimeout(300) fires at +302ms.
The process.exit() → process.exitCode change rests on documented POSIX behaviour, not a local measurement: Windows pipes are synchronous in node, so 4000 stderr lines survived process.exit() here. That is silence, not a negative result, and the commit says so.
On the Release gate: run from the session worktree it reported 71 errors. 66 are PRI/APPX0002 MakePri failures whose output paths measure 260-265 characters against a MAX_PATH of 260, and they all vanish when the same commit is built at a 16-character root (the documented long-path gotcha in AGENTS.md; the same path at CI's checkout root is 179 characters). The remaining 5 are NU1301 restore failures in src/vs-reactor caused by api.nuget.org being unreachable in that environment, and they reproduce identically at origin/main without this commit. No error references any file or project in this diff, and CI's own Build solution job passes.
Risk / breaking changes
No public API change and no product code touched; this is release/CI infrastructure plus contributor docs.
Worth a reviewer's attention:
deploy-stamp.json is now published at the site root. It carries the commit SHA, run id, attempt and a timestamp, all already public for a public repo. It sits beside versions.json and 404.html, and like 404.html it is refreshed from the workflow on every deploy rather than committed to gh-pages. Named without a leading . or _ on purpose, since Pages has historically excluded those.
verify can turn a release run red. That is the intent, but Pages propagation latency is a flakiness surface. The poll window is 10 minutes at 15-second intervals with a 20-second per-request ceiling, the job times out at 15 minutes, and every failure prints what was observed live so a transient is distinguishable from a real strand on sight. A red verify has two distinct meanings — stranded deploy versus unreadable site — and the runbook directs the operator to read the annotation rather than assume the first.
A release commit still produces two deployments. If Pages strands one, that run's verify goes red naming the run id and attempt actually being served, and the runbook's dispatch-on-main remedy resolves it in one command. The deploy job carries a comment enumerating the four failure modes that ruled out suppressing the duplicate, so the decision is not re-litigated from scratch.
The first Publish docs run after this merges is also its own proof: this PR touches .github/workflows/docs.yml, which is in the workflow's path filter, so merging it publishes the first stamp and exercises the gate.
A release PR always edits docs/guide/** (the {{reactorVersion}} substitutions),
so the push that merges it matches the docs workflow's path filter, and the
release tag then points at that same merge commit. Both runs hand
actions/deploy-pages the same pages_build_version -- it is github.sha, with no
input to override it -- so the two deployments collide under one identity and
one is silently stranded. Cutting 0.1.0-preview.16 that is what happened: every
job green, the deployment reporting success, gh-pages byte-correct, and the
live site still serving preview.15 for 2h25m.
Nothing in the pipeline looked at the live site, so no amount of green CI could
have caught it. Two changes:
A verify job now polls the deployed site and fails unless it is serving the
artifact this run produced. Its oracle is deploy-stamp.json, written into the
Pages artifact by publish and carrying the run id. Comparing versions.json
alone would be vacuous for a main push, since mike deploy main republishes an
existing version and leaves the version set unchanged. Every probe uses a
unique ?nc= key so a pass cannot come from cache, and an already-published
version fetched with the same request shape is the positive control that
separates a stranded deploy (exit 1) from a broken probe (exit 2).
A main push whose commit already carries a v* tag now stands down from
deploying and leaves it to the tag run. publish still runs, so gh-pages is
unaffected. A tag pushed after that check is invisible to it, so this narrows
the race rather than closing it, which is why verify carries correctness.
The decision logic is a pure function with 19 node cases; every comparison was
mutation-checked, as were the five workflow seams asserted by
DocsDeployWiringTests.
Fixes: #1268
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: ff518ae2-0ea3-466c-a9c0-4ea2265937c5
Artifact sizes for d67b29e vs the base branch (459f723).
Packages (compressed .nupkg)
Artifact
base
PR
Δ
Microsoft.UI.Reactor.nupkg
1.71 MB
1.71 MB
+3 B (0.00%)
≈
Microsoft.UI.Reactor.Advanced.nupkg
487.6 KB
487.6 KB
+9 B (0.00%)
≈
Microsoft.UI.Reactor.Devtools.nupkg
284.0 KB
284.0 KB
-3 B (0.00%)
≈
Assemblies in Microsoft.UI.Reactor
Artifact
base
PR
Δ
Reactor.Analyzers.dll
373.0 KB
373.0 KB
+0 B (0.00%)
≈
Reactor.dll
2.44 MB
2.44 MB
+0 B (0.00%)
≈
Reactor.Localization.Generator.dll
16.0 KB
16.0 KB
+0 B (0.00%)
≈
Reactor.Wrappers.Abstractions.dll
10.5 KB
10.5 KB
+0 B (0.00%)
≈
Reactor.Wrappers.Generator.dll
99.5 KB
99.5 KB
+0 B (0.00%)
≈
Assemblies in Microsoft.UI.Reactor.Advanced
Artifact
base
PR
Δ
Reactor.Advanced.dll
1021.5 KB
1021.5 KB
+0 B (0.00%)
≈
Assemblies in Microsoft.UI.Reactor.Devtools
Artifact
base
PR
Δ
Microsoft.UI.Reactor.Devtools.dll
784.5 KB
784.5 KB
+0 B (0.00%)
≈
No size change beyond the noise floor. ✅
✅ smaller / ⚠️ larger / ≈ within noise. Sizes come from a Release dotnet pack on the CI runner: packages are the compressed .nupkg download size, assemblies the uncompressed DLL inside it. workflow run.
Coverage for d67b29e vs the base branch (459f723) — unit + selftest merged.
Metric
base
PR
Δ
Line
85.78%
85.79%
+0.01 pp
≈
Branch
77.66% (963/1240)
77.58% (962/1240)
-0.08 pp
≈
No coverage change beyond the noise floor. ✅
✅ higher / ⚠️ lower / ≈ within noise. Δ is in percentage points; coverage is unit + selftest merged (Debug x64) on the CI runner. Cobertura reports attached to the workflow run as artifacts.
Three findings from github-code-quality on DocsDeployWiringTests.cs:
- LoadWorkflow leaked a StringReader; it is now a using declaration.
- Map did a ContainsKey lookup followed by an indexer lookup; it now
uses a single TryGetValue. The missing-key path still fails with the
same assertion message rather than returning null, confirmed by
mutating the publish job's outputs key out of docs.yml.
- The_verifier_and_its_regression_suite_are_present mapped its loop
variable on the first line of the body; the projection moved to a
Select.
No assertion changed. 469/469 doc-pipeline tests pass and the five
workflow-seam mutations still redden their intended cases.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: ff518ae2-0ea3-466c-a9c0-4ea2265937c5
Copilot review finding on #1269: the deploy stand-down could lose the
only release deployment.
The publish job stands main down when the pushed commit already carries a
v* tag, on the assumption that the tag run will deploy. Under the Actions
default concurrency behaviour that assumption does not hold. \queue: single\
keeps at most one pending run per group and cancels the previous one when a
new run queues, so a docs push landing while the tag run waited behind main
would evict the tag run before it ever started. The release version would
then never reach gh-pages, and main had already deferred its own deployment
to a run that no longer existed, so neither the deployment nor its
verification happened.
Setting queue: max makes runs wait in FIFO order (up to 100 pending) instead
of evicting each other, which makes the stand-down sound rather than
speculative. cancel-in-progress is dropped: false is the default, and
combining queue: max with cancel-in-progress: true is a workflow validation
error.
Because the two are now coupled, DocsDeployWiringTests asserts the queue
setting; removing it would silently re-arm the failure. Both mutations
(reverting to the default queue, and adding the illegal cancel-in-progress)
redden that test.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: ff518ae2-0ea3-466c-a9c0-4ea2265937c5
Copilot review finding on #1269: the verifier could miss a broken newly
published version.
selectProbeTargets picked the version holding the latest alias, so it never
touched the directory a run had just written whenever those differ. Two
reachable cases: publishing a backported or re-cut tag deliberately does not
move latest (see the Publish the release version step), and a main push
republishes main while latest sits on a release. In both, a 404 or otherwise
broken new directory passed as long as versions.json listed it and the stamp
matched.
Every publishing step now appends the version it wrote to
\/\, the artifact step turns that into a
published job output, and the verify job passes it as DOCS_PUBLISHED_VERSIONS.
The verifier probes each of those plus the latest alias holder, keeping a
distinct control. DOCS_PUBLISHED_VERSIONS is required rather than defaulted,
and the artifact step hard-fails on an empty list, so a wiring mistake is loud
instead of quietly narrowing the gate back to the alias holder.
Probes now request <version>/index.html rather than the bare directory. That
was found by the local end-to-end check reporting a pass it should not have:
python http.server answers a directory with no index by generating a 200
listing, while Pages 404s it, so the harness masked the very break it was
meant to show. Against real gh-pages bytes with 0.1.0-preview.14/index.html
removed, the old selection exits 0 and the new one exits 1 naming the 404.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: ff518ae2-0ea3-466c-a9c0-4ea2265937c5
Avoid asserting an unverified deployment collision cause
.github/scripts/verify-docs-deployment.mjs:152
The verifier is described as cause-agnostic, but this message asserts that every stale stamp was caused by two deployments colliding under one pages_build_version. A stale run ID can also result from a failed or overwritten deployment for another reason, so the annotation would report an unverified root cause. Keep the message tied to the observed mismatch and leave the collision explanation to the surrounding runbook.
Three Copilot review findings on #1269.
1. The published list and mike's versions.json are produced independently, so
selectProbeTargets intersecting them could silently shrink the probe set on
exactly the run that needed it. The artifact step now fails when a recorded
version is absent from site/versions.json, the verifier no longer filters,
and evaluate reports published-version-unlisted. The release step also
recorded its version before mike deploy ran; it now records after.
2. Nothing probed the site root, which mike writes only via set-default and no
version directory covers, so a stale or 404 root passed. The verifier now
fetches the root index.html, asserts 200, and asserts it still forwards to
the expected default (latest, else main, else unjudged rather than guessed).
3. The stamp-stale message asserted a cause the gate cannot observe. It now
states the mismatch and leaves the collision explanation to the runbook.
The jq for (1) was wrong on first write: inside select() the dot rebinds to the
element, so map(.version) ran against a string and the filter errored, which
made the guard fire on the passing case too. Caught by running the real step
against origin/gh-pages rather than assuming; replaced with array difference and
re-tested consistent, divergent, and empty inputs (exit 0, 1, 1).
Root probing verified against real bytes as well: healthy passes, and rewriting
the root stub to forward to main/ instead of latest/ exits 1 with
root-mistargeted. All five new comparisons are mutation-checked, including one
that initially survived because no test proved the driver actually requests the
root URL.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: ff518ae2-0ea3-466c-a9c0-4ea2265937c5
The CI wrapper only checks that the TAP output contains a positive pass count. If this regression file were accidentally reduced from its current 29 cases to a single passing case, the xUnit gate would still pass and the live-site verifier could be left largely untested. Assert the expected test count (or another stable completeness marker) so this safety net cannot be silently gutted.
Two Copilot review findings on #1269.
1. mike deploy --alias-type copy publishes latest/ as its own path in the
gh-pages tree, and the site root forwards there, so a missing or stale
latest/index.html 404s every reader arriving at the bare URL. Probing the
version that owns the alias does not cover it: confirmed on origin/gh-pages
that latest/index.html and 0.1.0-preview.16/index.html are distinct tree
entries (they happen to share a blob today only because git deduplicates
identical content). The verifier now probes the alias path explicitly and
reports alias-unreachable, without treating latest as a published version.
2. DocsDeploymentVerifierTests only asserted a positive TAP pass count, so a
suite gutted to one passing case would still have kept the xUnit gate green.
It now asserts a floor, with a message naming the constant to lower if cases
are deliberately removed. Verified by replacing the suite with a single test:
the gate fails with 'Only 1 cases ran, below the 30 this suite is expected
to carry.'
observe() also stopped using positional destructuring. The probe set is
conditional, so an off-by-one there would have silently swapped two results
rather than failing; it now builds a labelled job list.
Both new comparisons are mutation-checked, including that the driver actually
requests the alias URL rather than merely deciding it should.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: ff518ae2-0ea3-466c-a9c0-4ea2265937c5
…l-open
Three Copilot review findings on #1269.
1. The site-root check searched for the expected target anywhere in the HTML.
That accepted a root redirecting to not-latest/, since it contains latest/,
and accepted one whose only mention of the target was the human-visible
<a href> while the real script redirect pointed elsewhere. The redirect
target is now parsed from the location.replace call, falling back to the
noscript meta refresh, and compared exactly; an unrecognisable root is
reported as root-unparseable rather than passing.
2. The stand-down gate failed open on a tag-refresh failure. git fetch --tags
was allowed to fail with only a log line, after which git tag --points-at
would find nothing on a stale checkout and write deploy=true, restoring the
very main/tag pages_build_version collision the gate prevents, silently. It
now retries three times and fails the job, which costs a re-run rather than
a publish: gh-pages is already written by that point.
3. Probes were unbounded. Neither fetch nor response.text() times out on its
own and observe() awaits them together, so one stalled connection held the
polling loop past its deadline and swallowed the exit-2 diagnostic, leaving
only the outer job timeout. Each probe now carries an AbortSignal with a
20s ceiling, well under the 15s-interval polling window.
Suite is 37 cases; the DocsDeploymentVerifierTests floor moved with it. All
four new comparisons are mutation-checked and redden only their intended
cases.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: ff518ae2-0ea3-466c-a9c0-4ea2265937c5
Two low-severity review follow-ups on #1269.
Pass --test alongside --test-reporter=tap. The review claimed the gate 'cannot
pass and never validates the 41 cases' without it; that premise is false and
the CI log disproves it -- the previous failing Docs build printed the full TAP
summary (# tests 37 / # pass 33 / # cancelled 4), which it could only do in
test-runner mode. Both forms were re-checked locally and report 41/41. Adding
--test changes no behaviour here, it just stops the invocation depending on the
implicit 'registers tests on import' path.
The concurrency assertion's comment still explained itself in terms of the
deploy stand-down, which no longer exists. Rewritten to describe the failure it
actually guards: an evicted pending tag run never publishes its version to
gh-pages, so the release is lost outright with nothing failing.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: ff518ae2-0ea3-466c-a9c0-4ea2265937c5
Copilot review finding on #1269: verification ignored publish attempt
identity on full workflow reruns.
GITHUB_RUN_ID is stable across re-run attempts, so 'Re-run all jobs'
republishes under the same id. If that deployment were stranded while the
previous attempt's artifact stayed live, an id-only stamp comparison passed --
a false green on exactly the recovery path an operator reaches for after a
failure.
The stamp already carried run_attempt; it is now compared. Crucially the
expectation comes from needs.publish.outputs.run_attempt, the attempt that
built the artifact, not the verify job's own github.run_attempt: re-running
only the failed deploy job reuses the original publish, so the live stamp
legitimately carries the older attempt and a self-referential check would
redden every partial re-run. That asymmetry is why I had originally compared
run_id alone, and it is what the review's suggestion resolves.
Three cases: a full re-run is not satisfied by the previous attempt's
artifact (naming both attempts), a deploy-only re-run still matches the
artifact publish built, and a stamp missing run_attempt is stranded rather
than assumed. Both halves of the comparison are independently
mutation-checked -- dropping either the id or the attempt term reddens two
distinct cases.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: ff518ae2-0ea3-466c-a9c0-4ea2265937c5
This seam test only checks that each step mentions the output file; it does not enforce the ordering that the workflow relies on. Moving the echo before mike deploy (the previously fixed failure mode) would still pass this assertion, so a failed publish could again be reported as a directory to verify. Assert that the recording write occurs after the step's mike deploy command.
Copilot review finding on #1269 (previously-missed block): the seam test only
checked that each publishing step mentions the output file, so moving the echo
back above mike deploy -- the exact regression fixed two rounds ago -- would
still pass. A failed publish would then name a version as published and verify
would blame the deployment for a directory that was never written.
The test now locates the last mike deploy and the last recording write and
asserts the write comes after. Comments are stripped first: these steps
explain themselves in prose that names the very commands being located, and a
trailing comment mentioning mike deploy inverted the comparison on the backfill
step -- caught by running it, not by reading it.
Mutation-checked properly on the second attempt. Adding a redundant early echo
does not redden it and should not: the invariant is that a write happens after
a successful deploy, not that no write happens before. Actually moving the
write reddens the release-step case with the intended message.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: ff518ae2-0ea3-466c-a9c0-4ea2265937c5
The verifier's CI wrapper awaited node with only the ambient xUnit
cancellation token, so a hung suite would sit there until the job timeout,
reporting nothing useful and leaving the process behind. Not hypothetical:
this suite produced exactly that failure two commits ago, when an unref'd
abort timer left a probe pending and node cancelled the remaining cases
instead of failing them.
The wait is now bounded at three minutes -- the suite runs in well under a
second, so it only ever fires on a hang -- and on expiry the process tree is
killed and the test fails naming the likely cause, since a hang is invisible
in the TAP summary.
Verified by dropping the ceiling to 1ms: the timeout path fires, the failure
message appears, and no node process survives.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: ff518ae2-0ea3-466c-a9c0-4ea2265937c5
The CLI ended with process.exit(report(...)). On POSIX -- which is where the
verify job runs -- writes to a piped stdout or stderr are asynchronous, and
process.exit() tears the process down without draining them. The entire value
of a failing run is the ::error:: annotations naming what the live site was
actually serving, so truncating them loses exactly the output this gate exists
to produce.
Now sets process.exitCode and lets node exit once the loop drains, which
flushes. The usage-error paths throw a UsageError caught at the top of the CLI
block rather than calling process.exit(3) inline, so the same applies to them.
Honest about the evidence: I could not reproduce the truncation locally,
because Windows pipes are synchronous in node -- 4000 stderr lines survived
process.exit() here. That makes the local result silence, not a negative, so
the change rests on documented POSIX behaviour rather than a measurement. What
I did verify: all five exit paths still propagate (0/1/2 end to end against a
live server and a dead port, 3 for both usage errors), and a source assertion
now fails if process.exit( returns to the file.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: ff518ae2-0ea3-466c-a9c0-4ea2265937c5
The reason will be displayed to describe this comment to others. Learn more.
Copilot review overview
🔵 Needs a closer look
Moderate test metadata and run_attempt wiring-guard issues remain unresolved.
Review effort: Balanced Findings: None
Previously missed (1)
In code that hasn't changed since last review
Distinguish probe-broken from stranded deployment in runbook
docs/contributing/release-runbook.md:184
The release runbook's top-level interpretation is too strong: verify also exits red with the probe-broken status when no known-good response can be read, which does not establish that the live site is serving another artifact. Please direct the operator to distinguish the two annotations, consistent with the table below, rather than prescribing the stranded-deployment explanation for every red verification.
Copilot review finding on #1269 (previously-missed block): the monitor step
told an operator that a red verify means the live site is serving something
else. That is only one of its two outcomes. probe-broken also exits red and
establishes the opposite -- that no probe reached known-good content, so the
run says nothing about what is being served.
The step now names both annotations and points at the table that gives each
its own response, instead of prescribing the stranded-deployment explanation
for a run that may simply have been unable to read the site.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: ff518ae2-0ea3-466c-a9c0-4ea2265937c5
The PR test plan still says the Node suite has 44 cases, but this file now requires 45 and the test script currently registers 45 test(...) calls. Update the PR description so reviewers run and expect the current suite size.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
A release tag's docs deployment can silently lose to the
main-push deployment for the same commit, leaving the released version rendered ongh-pagesbut never served. A release PR always editsdocs/guide/**(the{{reactorVersion}}substitutions), so the push that merges it matches the docs workflow's path filter, and the release tag then points at that same merge commit. Both runs handactions/deploy-pagesthe samepages_build_version(it isgithub.sha, with no input to override it), so the two deployments collide under one identity and one is stranded. Cutting 0.1.0-preview.16 that is exactly what happened: every job green, the deployment reportingsuccess,gh-pagesbyte-correct, and the live site still serving preview.15 for 2h25m.Nothing in the pipeline looked at the live site, so no amount of green CI could have caught it.
A
verifyjob now polls the deployed site and fails unless it is serving the artifact that run produced. Its oracle isdeploy-stamp.json, written into the Pages artifact bypublishand carrying the run id and attempt. Comparingversions.jsonalone would be vacuous for amainpush, sincemike deploy mainrepublishes an existing version and leaves the version set unchanged. Beyond the stamp it asserts that every version that run published serves its ownindex.html(not just thelatestholder, which a backported tag deliberately does not move and amainpush never touches), that thelatest/alias tree is served, and that the site root parses to a redirect at the expected default. Every probe uses a unique?nc=key so a pass cannot come from cache, rejects a followed redirect so a bounce cannot masquerade as a 200, and is bounded by an abort timeout so one stalled socket cannot eat the polling window. An already-published version fetched with the identical request shape is the positive control that separates a stranded deploy (exit 1) from a broken probe (exit 2).The duplicate deployment is tolerated rather than avoided. Standing one of the two runs down was implemented and then removed during review. Every form of it requires predicting that the other run will deploy, and each way that prediction fails is silent in the direction that matters: a pending tag run can be evicted from the concurrency queue, a tag deleted or re-cut on
originleaves a stale local ref, an unreachableoriginmakes the check answer "no tag here", and the two runs are not guaranteed to enter the concurrency group in event order. Each of those skips the deployment and its verification together, leaving the site stale with every job green — strictly worse than a duplicate thatverifycatches. Issue #1268 itself noted this mitigation reduces the window rather than closing it, with verification carrying correctness.The concurrency group now sets
queue: max. The Actions default keeps at most one pending run per group and cancels the previous one when a new run queues, so a docs push landing while a release tag's run waited behindmainwould evict the tag run before it started, and the release would never reachgh-pagesat all.Linked issue / spec
Fixes: #1268
Test plan
.github/scripts/verify-docs-deployment.test.mjs; run them withnode --test --test-reporter=tap .github/scripts/verify-docs-deployment.test.mjs. The expected count is pinned byMinimumCasesinDocsDeploymentVerifierTests, which is the source of truth if these two ever disagree — it exists so the suite cannot be silently gutted while the xUnit gate stays green.evaluate(),probe(),selectProbeTargets()andparseRootRedirect()in turn reddens exactly the intended cases and no others. Three mutations initially survived — nothing asserted the driver actually requests the version, root or alias URLs — and assertions were added until they reddened.DocsDeployWiringTests. Two assertions initially survived: one matched acatline rather than the write, and one checked only that a publishing step mentions the record file, so moving the write back abovemike deploystill passed. Both tightened until the regression reddens.Assemble the Pages artifactstep withjqagainstorigin/gh-pages: valid stamp, single-lineversions=output, and the published/versions.jsonconsistency guard exercised across consistent, divergent and empty inputs (exit 0, 1, 1). The first version of that guard'sjqwas wrong — insideselect()the dot rebinds, so the filter errored and the guard fired on the passing case — caught only by running it.0.1.0-preview.14/index.htmlremoved, the oldlatest-only target selection exits 0 while the new one exits 1 naming the 404.main/exits 1 withroot-mistargeted.dotnet test tests/Reactor.DocPipeline.Tests-> 472/472 pass.dotnet test tests/Reactor.Tests-> 14144 total, 0 failed, 64 skipped.docs.ymlparses as YAML and everyrunblock passesbash -n.dotnet restore Reactor.slnxthendotnet build Reactor.slnx --no-restore -c Release. See the note below, since this needed a short-path checkout to be meaningful.Three instrument failures are worth recording, because in each case a check reported a pass it had not earned — or could not have earned:
python -m http.serveranswers a directory with no index by generating a 200 listing while Pages 404s it — so the local end-to-end check masked the very break it was meant to demonstrate. Probes now request<version>/index.html.AbortSignal.timeout(), whose timer is unref'd and therefore never fires when nothing else holds the event loop open. It passed locally and hung on CI, cancelling four cases. Measured directly:AbortSignal.timeout(300)lets the process exit at +1ms without firing, while an ownedsetTimeout(300)fires at +302ms.process.exit()→process.exitCodechange rests on documented POSIX behaviour, not a local measurement: Windows pipes are synchronous in node, so 4000 stderr lines survivedprocess.exit()here. That is silence, not a negative result, and the commit says so.On the Release gate: run from the session worktree it reported 71 errors. 66 are
PRI/APPX0002MakePri failures whose output paths measure 260-265 characters against a MAX_PATH of 260, and they all vanish when the same commit is built at a 16-character root (the documented long-path gotcha in AGENTS.md; the same path at CI's checkout root is 179 characters). The remaining 5 areNU1301restore failures insrc/vs-reactorcaused byapi.nuget.orgbeing unreachable in that environment, and they reproduce identically atorigin/mainwithout this commit. No error references any file or project in this diff, and CI's ownBuild solutionjob passes.Risk / breaking changes
No public API change and no product code touched; this is release/CI infrastructure plus contributor docs.
Worth a reviewer's attention:
deploy-stamp.jsonis now published at the site root. It carries the commit SHA, run id, attempt and a timestamp, all already public for a public repo. It sits besideversions.jsonand404.html, and like404.htmlit is refreshed from the workflow on every deploy rather than committed togh-pages. Named without a leading.or_on purpose, since Pages has historically excluded those.verifycan turn a release run red. That is the intent, but Pages propagation latency is a flakiness surface. The poll window is 10 minutes at 15-second intervals with a 20-second per-request ceiling, the job times out at 15 minutes, and every failure prints what was observed live so a transient is distinguishable from a real strand on sight. A redverifyhas two distinct meanings — stranded deploy versus unreadable site — and the runbook directs the operator to read the annotation rather than assume the first.verifygoes red naming the run id and attempt actually being served, and the runbook's dispatch-on-mainremedy resolves it in one command. Thedeployjob carries a comment enumerating the four failure modes that ruled out suppressing the duplicate, so the decision is not re-litigated from scratch.The first
Publish docsrun after this merges is also its own proof: this PR touches.github/workflows/docs.yml, which is in the workflow's path filter, so merging it publishes the first stamp and exercises the gate.