Repository navigation
matrix: test_gap_repsel_p4a3_ptr_numarray red on main across ALL requires=move arms, untriaged — blocks clean gating of every barrier/GC PR (#7016 neighbourhood) #7194
Description
Activity
This issue understates its own surface by 7×.
Measured on an idle mini during #7196's gate run (in-place A/B: same worktree, same target dir, only the PR's runtime files reverted — identical failing sets on both arms, so main-side):
Seven tests — not just
test_gap_repsel_p4a3_ptr_numarray— fail across all tenrequires=movearms withevacuated=0. Full run either side:PASS=296 UNVER=242 XFAIL=1 FAIL=70(7 tests × 10 arms).Consequences unchanged but larger: every GC/barrier PR must hand-exonerate seventy cells, and the arms that gate evacuation correctness are the ones broken. Worth owning before #7187's Phase A, and before #7161's revert makes these arms load-bearing again.
Note also the contention hazard:
FAIL=70is also the signature a load-contended matrix run produces (#7198 saw exactly that number on a busy host,FAIL=10idle). Two different causes, same number — always re-run idle AND A/B before attributing.- added a commit that references this issue
on Aug 2, 2026 Re-measured on an idle host at
origin/main=c6f5e0192(+ PR #7233),--arms all --pressure 8, oracle node 26.5.1: still red, unchanged, and still the same single test.test \ arm evac_mi force_e force_v loop_po rep_i32 rep_str rep_str rep_ptr rep_ptr rep_spe rep_int repsel_p4a3_ptr_numarray FAIL FAIL FAIL FAIL FAIL FAIL FAIL FAIL FAIL FAIL FAILEvidence string is stable:
output-mismatch cycles=2 evacuated=0 scavenged=6514onevac_minor/force_evac/force_verify,…scavenged=6515on the rest.evacuated=0whilescavenged≈6515— the arm collected and the scavenger touched ~6.5k objects, but the evacuation counter is flat, which is the shape this issue names.The arms are not inert: per-arm liveness reports
collected 46/47, moved-objects 46/47, copy-minor 46/47for every one of them, so this is a genuine copying-minor failure and not a green-hole artifact.Also confirmed again by in-place A/B on PR #7233 (same target dir, same package set, only that PR's five production files reverted): identical 10 FAIL, identical evidence strings. That is the third PR now paying the manual-exoneration tax this issue predicted — #7193, and now #7233 twice (once for the matrix, once for #7234). The corpus-wide cost is visible in the headline:
PASS=631 UNVER=325 XFAIL=21 FAIL=10, and every one of those 10 is this.No new information on the root cause from my side; flagging only that it reproduces exactly at the current
mainand is still untriaged inknown_failures.json.Correction to my comment above — I published a table row that is wrong, and the error hides a real datum.
I wrote the FAIL row with eleven arms including
loop_polls.loop_pollsPASSES. The matrix has 21 arms, not 23 (21 × 47 = the 987 cells the summary reports); I never verified that denominator and carried it through several comments.Ground truth, straight from the run's own results table (
repsel_p4a3_ptr_numarray, all 21 columns):default evac_mi force_e verify_ force_v gen_gc_ wb_off gen_off wb_off_ all_fou cons_sc cons_sc loop_po rep_i32 rep_str rep_str rep_ptr rep_ptr rep_spe rep_int shipped UNVER FAIL FAIL UNVER FAIL UNVER UNVER UNVER UNVER UNVER UNVER UNVER PASS FAIL FAIL FAIL FAIL FAIL FAIL FAIL PASSThe exact FAIL set is 10 arms:
evac_minor force_evac force_verify rep_i32_off rep_str_off rep_str_static_off rep_ptr_shape_off rep_ptr_numarray_off rep_spec_abi_off rep_int_valued_off★ The correction is worth more than the number:
loop_pollsis arequires=movearm that PASSESFrom the same run's per-arm liveness:
loop_polls requires=move collected 46/47 moved-objects 46/47 copy-minor 46/47So
loop_pollscollects, moves objects, and runs a copying minor on 46 of 47 files — and this test passes there, while ten other arms that report the identical liveness triple fail. That rules out "any copying minor breaks it" and narrows the search: the differentiator isPERRY_GC_MOVING_LOOP_POLLS=1, i.e. where the moving collection is triggered from (the loop-back-edge safepoint versus the alloc-point trigger deferring intogc_safepoint_moving_minor).That is a bisect axis this issue did not have, and my wrong table erased it by painting
loop_pollsred.Everything else in my previous comment stands and is unaffected: the evidence strings (
cycles=2 evacuated=0 scavenged=6514on the three evac arms,6515on the sevenrep_*_offarms), the non-inert liveness, and the in-place A/B on #7233 — which ran exactly these ten arms and reproduced all ten identically with that PR's production files reverted.Recording the correction rather than editing the original, for the same reason the surrounding work does: a published number nobody re-derived is how this campaign keeps finding phantom results, and this one was mine.
Triaged: outcome (1) — a real defect under the evacuating minor, already fixed on
mainby #7249. Bisected, and re-verified above the movement level that produced the failure. PR #7253 closes the reason nobody noticed.It was real
Both arms built in one worktree, one target dir, identical package set, artifacts hashed so the A/B cannot be vacuous:
perrymd5libperry_runtime.amd5c9cd73ba5(the SHA this issue last measured)43b6fdeb…4b054e4f…64c1f56fb(origin/main)e1c9a1fd…0252c0c6…--arms all --filter p4a3_ptr_numarray --pressure 8, node 26.5.1:c9cd73ba5:PASS=2 UNVER=9 FAIL=10—output-mismatch cycles=2 evacuated=0 scavenged=6514on the three evac arms and…6515on the sevenrep_*_offarms. Byte-identical to the evidence in this issue, including the 6514/6515 split and theloop_polls/shipped_defaultgreens from the correction comment.64c1f56fb:PASS=12 UNVER=9 XFAIL=0 FAIL=0, ten consecutive runs, withcopy-minor 1/1on every one of the tenrequires=movearms — so the arms bit, they just no longer fail.
The symptom, which nobody had recorded: exactly one lost increment out of 5000.
< 153,…,164,159,147,160,168,156 ← node < 5000 > 153,…,164,158,147,160,168,156 ← perry, bkt 27 > 4999One
counts[v] = (counts[v] || 0) + 1store landed in from-space. Deterministic — 10/10 byte-identical wrong output — which is the cache/raw-reference shape rather than the stale-register shape.Bisected, one build per hop, 10 runs each
commit evac_minorc9cd73ba5#7233 FAIL 10/10 8b024958f#7242 FAIL 10/10 c7893c4ac#7250 FAIL 10/10 061b24163#7252 FAIL 10/10 64c1f56fb#7249 — globalThis bootstrap in a no-move window PASS 10/10 Minor #0 was landing inside realm construction, whose installers thread raw
*mut ObjectHeaderlocals across their own allocations; the corrupted store was downstream of that window.★ Fixed, not hidden — checked, because head collects less
#7249 changes where minor #0 lands, so head runs this file at
cycles=1 scavenged=3585where the failing base rancycles=2 scavenged=6515. A green cell under strictly less movement proves nothing, so green was re-established above the base's movement, not below it. Ten runs per arm at64c1f56fb:arm result movement evac_minorat pressure 8 / 4 / 2 / 1PASS 10/10 cycles=1 scavenged=3585— the pressure knob does not move this file at all+FORCE_EVACUATE,+VERIFY_EVACUATION,+PROTECT_FROMSPACE,+FROMSPACE_SCAN_ABORTPASS 10/10 — loop polls + PERRY_GC_ZEAL=1PASS 10/10 cycles=14373 scavenged=948514 373 evacuating minors, byte-exact throughout. The array is relocated continuously and every live read still resolves.
Why it sat here untriaged for a week: the gate had no main-line run
gc-stresscarriesscripts/gc_repsel_matrix.sh— the only CI execution of therequires=moveallocation-point arms over the representation corpus — and never ran onmain.test.yml'spush:trigger is tags-only ("Direct pushes to main do NOT trigger tests"), sopushin the job'sif:only ever meant a release tag; and the nightly cron, which the same file calls "the only backstop for integration-suite regressions a scoped PR run can't see", fires asschedule, which theif:did not list. Twelve consecutive nightlymainruns:gc-stressskippedin every one.The pre-merge half is worse, and is branch-protection state rather than a file.
gc-stressis not inmain's required contexts, and every one of the five GC/codegen PRs this campaign just landed merged with it still queued — #7242, #7243, #7250, #7252 and #7249, the commit that actually fixed this issue. #7239 likewise; #7237 merged with itFAILURE. So this issue's core claim was exactly right, and stronger than stated: it is not that the landings were gated against a red matrix — they were not gated by the matrix at all.PR #7253 adds
scheduleto theif:and addsscripts/gc_gate_wiring_check.pytolint(a required context) so the four hazards are asserted mechanically instead of re-derived during the next incident. Promotinggc-stressitself to a required context is deliberately left as a follow-up — one green scheduled run first, then promotion, or it blocks every open PR the day it lands.Not closed by this, filed separately
#7254 — under loop polls +
PERRY_GC_ZEAL=1+PERRY_GC_VERIFY_EVACUATION=1, this same file panicsstale forwarded pointer in shadow stack roots10/10. It reproduces onc9cd73ba5too, so it predates #7249 and is not what this issue was about; output is byte-exact under the same zeal without the verifier, and the from-space scan reportsmissing_rewrites=0 dangling=0on every cycle, so the slot is very likely dead-but-unrewritten. No CI arm sets that knob pair.Recommend closing this issue as fixed by #7249.
Corpus-wide confirmation, and a second instrument gap this run exposes.
Full matrix at
origin/main=64c1f56fb,--arms all --pressure 8, node 26.5.1, 49 files × 21 arms = 1029 cells:byte-exact vs node 26.5.1: 1008/1029 cells summary: PASS=676 UNVER=332 XFAIL=21 FAIL=0FAIL=0. Against thePASS=631 UNVER=325 XFAIL=21 FAIL=10recorded in this issue atc9cd73ba5, the ten cells are gone and nothing else has gone red. The 21 XFAILs are the existing triaged entries (assign_string_source_rooting×10 → #7248,regexp_receiver_rooting×10 → #7247,ptr_shape_locals × rep_ptr_shape_off→ #6976). Sogc_repsel_matrix.sh --arms allis now clean onmain, and the "every barrier/GC PR must hand-exonerate seventy cells" tax this issue was opened over is paid off.
★ But the arms that gate a PR are inert corpus-wide
From the same run's own liveness summary:
default requires=scavenge collected 24/49 copy-minor 0/49 verify_evac requires=scavenge collected 24/49 copy-minor 0/49 cons_scan_off requires=scavenge collected 24/49 copy-minor 0/49 cons_scan_off_force requires=scavenge collected 24/49 copy-minor 0/49 evac_minor requires=move collected 48/49 copy-minor 48/49 force_verify requires=move collected 48/49 copy-minor 48/49PR_ARMSisdefault,evac_minor,verify_evac,force_verify,cons_scan_off,shipped_default. Four of those six run zero copying minors on all 49 files, so on a pull request they can produce nothing butUNVER. Onlyevac_minorandforce_verifyrelocate anything.That contradicts the script's own header, which is now stale and reads as a stronger guarantee than the matrix delivers:
default… Since #7024 this is a RELOCATING arm with no env override at all beyond the pressure knob
arm copy-minor before #7024 after/default 0/22 → 12/22Measured today:
0/49. The pressure knob is inert too —evac_minorontest_gap_repsel_p4a3_ptr_numarrayreportscycles=1 scavenged=3585identically at--pressure 8,4,2and1.This is the same root as the note already in this issue about the gc-ratchet baseline (
minor_cycles=0everywhere because #7161 flipped the evacuating default) — but it lands somewhere sharper: the arm named after the shipped configuration, the one #6993/#7024 added specifically so the relocating-minor defect class would be reproducible per PR, is exercising the non-moving path. Andgc_repsel_matrix.sh's exit status counts onlyFAIL, so an all-UNVERtable exits 0.gc-moving-witnessesrejectsUNVERas hard asFAIL;gc-stressdoes not.I have not changed any of this — it is tied to the #7161 revert timeline that #7187 already owns, and turning it red today would create exactly the untriaged permanent red this gate exists to prevent. Filed as #7255 so it is not carried only in this thread.
Not verified by me: whether the
defaultarm's 0/49 is entirely #7161 or also involves the--pressureinteraction of #7024/#7025, and the other 45 corpus files under the zeal arm of #7254.- added a commit that references this issue
on Aug 2, 2026
test_gap_repsel_p4a3_ptr_numarrayis red onorigin/mainacross all tenrequires=movematrix arms (evacuated=0 scavenged=…— likely #7016's neighbourhood), untriaged: not inknown_failures.json, no issue.Found during PR #7193's matrix run: an in-place A/B (same target dir, same package set, only #7193's three production files reverted) produced the identical 10-arm failure set at
origin/main— proving it pre-existing, not barrier-change-caused.Why it matters more than a normal red
These are exactly the arms that exercise the write barrier and the evacuating minor — the arms every barrier/GC PR must gate on, and the arms that go live again when #7161 reverts. Any future barrier change lands on a matrix already red in precisely the cells that matter, forcing every author into the same manual A/B exoneration #7193 had to run. Own it before Phase A of #7187 (lazy arming) lands.
Also note (same session): the pinned gc-ratchet baseline currently has
minor_cycles=0everywhere because #7161 flipped the evacuating default — the ratchet cannot check evacuation counters until the revert (or a deliberate re-pin). Two instrument gaps, same root, both tied to the #7161 revert timeline.