Skip to content

matrix: test_gap_repsel_p4a3_ptr_numarray red on main across ALL requires=move arms, untriaged — blocks clean gating of every barrier/GC PR (#7016 neighbourhood) #7194

Description

@proggeramlug

test_gap_repsel_p4a3_ptr_numarray is red on origin/main across all ten requires=move matrix arms (evacuated=0 scavenged=… — likely #7016's neighbourhood), untriaged: not in known_failures.json, no issue.

Found during PR #7193's matrix run: an in-place A/B (same target dir, same package set, only #7193's three production files reverted) produced the identical 10-arm failure set at origin/main — proving it pre-existing, not barrier-change-caused.

Why it matters more than a normal red

These are exactly the arms that exercise the write barrier and the evacuating minor — the arms every barrier/GC PR must gate on, and the arms that go live again when #7161 reverts. Any future barrier change lands on a matrix already red in precisely the cells that matter, forcing every author into the same manual A/B exoneration #7193 had to run. Own it before Phase A of #7187 (lazy arming) lands.

Also note (same session): the pinned gc-ratchet baseline currently has minor_cycles=0 everywhere because #7161 flipped the evacuating default — the ratchet cannot check evacuation counters until the revert (or a deliberate re-pin). Two instrument gaps, same root, both tied to the #7161 revert timeline.

Activity

  1. proggeramlug commented on Aug 1, 2026

    @proggeramlug
    ContributorAuthor

    This issue understates its own surface by 7×.

    Measured on an idle mini during #7196's gate run (in-place A/B: same worktree, same target dir, only the PR's runtime files reverted — identical failing sets on both arms, so main-side):

    Seven tests — not just test_gap_repsel_p4a3_ptr_numarray — fail across all ten requires=move arms with evacuated=0. Full run either side: PASS=296 UNVER=242 XFAIL=1 FAIL=70 (7 tests × 10 arms).

    Consequences unchanged but larger: every GC/barrier PR must hand-exonerate seventy cells, and the arms that gate evacuation correctness are the ones broken. Worth owning before #7187's Phase A, and before #7161's revert makes these arms load-bearing again.

    Note also the contention hazard: FAIL=70 is also the signature a load-contended matrix run produces (#7198 saw exactly that number on a busy host, FAIL=10 idle). Two different causes, same number — always re-run idle AND A/B before attributing.

  2. proggeramlug commented on Aug 2, 2026

    @proggeramlug
    ContributorAuthor

    Re-measured on an idle host at origin/main = c6f5e0192 (+ PR #7233), --arms all --pressure 8, oracle node 26.5.1: still red, unchanged, and still the same single test.

    test \ arm                 evac_mi force_e force_v loop_po rep_i32 rep_str rep_str rep_ptr rep_ptr rep_spe rep_int
    repsel_p4a3_ptr_numarray   FAIL    FAIL    FAIL    FAIL    FAIL    FAIL    FAIL    FAIL    FAIL    FAIL    FAIL
    

    Evidence string is stable: output-mismatch cycles=2 evacuated=0 scavenged=6514 on evac_minor/force_evac/force_verify, …scavenged=6515 on the rest. evacuated=0 while scavenged≈6515 — the arm collected and the scavenger touched ~6.5k objects, but the evacuation counter is flat, which is the shape this issue names.

    The arms are not inert: per-arm liveness reports collected 46/47, moved-objects 46/47, copy-minor 46/47 for every one of them, so this is a genuine copying-minor failure and not a green-hole artifact.

    Also confirmed again by in-place A/B on PR #7233 (same target dir, same package set, only that PR's five production files reverted): identical 10 FAIL, identical evidence strings. That is the third PR now paying the manual-exoneration tax this issue predicted — #7193, and now #7233 twice (once for the matrix, once for #7234). The corpus-wide cost is visible in the headline: PASS=631 UNVER=325 XFAIL=21 FAIL=10, and every one of those 10 is this.

    No new information on the root cause from my side; flagging only that it reproduces exactly at the current main and is still untriaged in known_failures.json.

  3. proggeramlug commented on Aug 2, 2026

    @proggeramlug
    ContributorAuthor

    Correction to my comment above — I published a table row that is wrong, and the error hides a real datum.

    I wrote the FAIL row with eleven arms including loop_polls. loop_polls PASSES. The matrix has 21 arms, not 23 (21 × 47 = the 987 cells the summary reports); I never verified that denominator and carried it through several comments.

    Ground truth, straight from the run's own results table (repsel_p4a3_ptr_numarray, all 21 columns):

    default evac_mi force_e verify_ force_v gen_gc_ wb_off  gen_off wb_off_ all_fou cons_sc cons_sc loop_po rep_i32 rep_str rep_str rep_ptr rep_ptr rep_spe rep_int shipped
    UNVER   FAIL    FAIL    UNVER   FAIL    UNVER   UNVER   UNVER   UNVER   UNVER   UNVER   UNVER   PASS    FAIL    FAIL    FAIL    FAIL    FAIL    FAIL    FAIL    PASS
    

    The exact FAIL set is 10 arms:

    evac_minor  force_evac  force_verify
    rep_i32_off  rep_str_off  rep_str_static_off  rep_ptr_shape_off
    rep_ptr_numarray_off  rep_spec_abi_off  rep_int_valued_off
    

    ★ The correction is worth more than the number: loop_polls is a requires=move arm that PASSES

    From the same run's per-arm liveness:

    loop_polls   requires=move   collected 46/47   moved-objects 46/47   copy-minor 46/47
    

    So loop_polls collects, moves objects, and runs a copying minor on 46 of 47 files — and this test passes there, while ten other arms that report the identical liveness triple fail. That rules out "any copying minor breaks it" and narrows the search: the differentiator is PERRY_GC_MOVING_LOOP_POLLS=1, i.e. where the moving collection is triggered from (the loop-back-edge safepoint versus the alloc-point trigger deferring into gc_safepoint_moving_minor).

    That is a bisect axis this issue did not have, and my wrong table erased it by painting loop_polls red.

    Everything else in my previous comment stands and is unaffected: the evidence strings (cycles=2 evacuated=0 scavenged=6514 on the three evac arms, 6515 on the seven rep_*_off arms), the non-inert liveness, and the in-place A/B on #7233 — which ran exactly these ten arms and reproduced all ten identically with that PR's production files reverted.

    Recording the correction rather than editing the original, for the same reason the surrounding work does: a published number nobody re-derived is how this campaign keeps finding phantom results, and this one was mine.

  4. proggeramlug commented on Aug 2, 2026

    @proggeramlug
    ContributorAuthor

    Triaged: outcome (1) — a real defect under the evacuating minor, already fixed on main by #7249. Bisected, and re-verified above the movement level that produced the failure. PR #7253 closes the reason nobody noticed.

    It was real

    Both arms built in one worktree, one target dir, identical package set, artifacts hashed so the A/B cannot be vacuous:

    perry md5 libperry_runtime.a md5
    c9cd73ba5 (the SHA this issue last measured) 43b6fdeb… 4b054e4f…
    64c1f56fb (origin/main) e1c9a1fd… 0252c0c6…

    --arms all --filter p4a3_ptr_numarray --pressure 8, node 26.5.1:

    • c9cd73ba5: PASS=2 UNVER=9 FAIL=10 — output-mismatch cycles=2 evacuated=0 scavenged=6514 on the three evac arms and …6515 on the seven rep_*_off arms. Byte-identical to the evidence in this issue, including the 6514/6515 split and the loop_polls/shipped_default greens from the correction comment.
    • 64c1f56fb: PASS=12 UNVER=9 XFAIL=0 FAIL=0, ten consecutive runs, with copy-minor 1/1 on every one of the ten requires=move arms — so the arms bit, they just no longer fail.

    The symptom, which nobody had recorded: exactly one lost increment out of 5000.

    < 153,…,164,159,147,160,168,156   ← node          < 5000
    > 153,…,164,158,147,160,168,156   ← perry, bkt 27 > 4999
    

    One counts[v] = (counts[v] || 0) + 1 store landed in from-space. Deterministic — 10/10 byte-identical wrong output — which is the cache/raw-reference shape rather than the stale-register shape.

    Bisected, one build per hop, 10 runs each

    commit evac_minor
    c9cd73ba5 #7233 FAIL 10/10
    8b024958f #7242 FAIL 10/10
    c7893c4ac #7250 FAIL 10/10
    061b24163 #7252 FAIL 10/10
    64c1f56fb #7249 — globalThis bootstrap in a no-move window PASS 10/10

    Minor #0 was landing inside realm construction, whose installers thread raw *mut ObjectHeader locals across their own allocations; the corrupted store was downstream of that window.

    ★ Fixed, not hidden — checked, because head collects less

    #7249 changes where minor #0 lands, so head runs this file at cycles=1 scavenged=3585 where the failing base ran cycles=2 scavenged=6515. A green cell under strictly less movement proves nothing, so green was re-established above the base's movement, not below it. Ten runs per arm at 64c1f56fb:

    arm result movement
    evac_minor at pressure 8 / 4 / 2 / 1 PASS 10/10 cycles=1 scavenged=3585 — the pressure knob does not move this file at all
    +FORCE_EVACUATE, +VERIFY_EVACUATION, +PROTECT_FROMSPACE, +FROMSPACE_SCAN_ABORT PASS 10/10 —
    loop polls + PERRY_GC_ZEAL=1 PASS 10/10 cycles=14373 scavenged=9485

    14 373 evacuating minors, byte-exact throughout. The array is relocated continuously and every live read still resolves.

    Why it sat here untriaged for a week: the gate had no main-line run

    gc-stress carries scripts/gc_repsel_matrix.sh — the only CI execution of the requires=move allocation-point arms over the representation corpus — and never ran on main. test.yml's push: trigger is tags-only ("Direct pushes to main do NOT trigger tests"), so push in the job's if: only ever meant a release tag; and the nightly cron, which the same file calls "the only backstop for integration-suite regressions a scoped PR run can't see", fires as schedule, which the if: did not list. Twelve consecutive nightly main runs: gc-stress skipped in every one.

    The pre-merge half is worse, and is branch-protection state rather than a file. gc-stress is not in main's required contexts, and every one of the five GC/codegen PRs this campaign just landed merged with it still queued — #7242, #7243, #7250, #7252 and #7249, the commit that actually fixed this issue. #7239 likewise; #7237 merged with it FAILURE. So this issue's core claim was exactly right, and stronger than stated: it is not that the landings were gated against a red matrix — they were not gated by the matrix at all.

    PR #7253 adds schedule to the if: and adds scripts/gc_gate_wiring_check.py to lint (a required context) so the four hazards are asserted mechanically instead of re-derived during the next incident. Promoting gc-stress itself to a required context is deliberately left as a follow-up — one green scheduled run first, then promotion, or it blocks every open PR the day it lands.

    Not closed by this, filed separately

    #7254 — under loop polls + PERRY_GC_ZEAL=1 + PERRY_GC_VERIFY_EVACUATION=1, this same file panics stale forwarded pointer in shadow stack roots 10/10. It reproduces on c9cd73ba5 too, so it predates #7249 and is not what this issue was about; output is byte-exact under the same zeal without the verifier, and the from-space scan reports missing_rewrites=0 dangling=0 on every cycle, so the slot is very likely dead-but-unrewritten. No CI arm sets that knob pair.

    Recommend closing this issue as fixed by #7249.

  5. proggeramlug commented on Aug 2, 2026

    @proggeramlug
    ContributorAuthor

    Corpus-wide confirmation, and a second instrument gap this run exposes.

    Full matrix at origin/main = 64c1f56fb, --arms all --pressure 8, node 26.5.1, 49 files × 21 arms = 1029 cells:

    byte-exact vs node 26.5.1: 1008/1029 cells
    summary: PASS=676 UNVER=332 XFAIL=21 FAIL=0
    

    FAIL=0. Against the PASS=631 UNVER=325 XFAIL=21 FAIL=10 recorded in this issue at c9cd73ba5, the ten cells are gone and nothing else has gone red. The 21 XFAILs are the existing triaged entries (assign_string_source_rooting ×10 → #7248, regexp_receiver_rooting ×10 → #7247, ptr_shape_locals × rep_ptr_shape_off → #6976). So gc_repsel_matrix.sh --arms all is now clean on main, and the "every barrier/GC PR must hand-exonerate seventy cells" tax this issue was opened over is paid off.


    ★ But the arms that gate a PR are inert corpus-wide

    From the same run's own liveness summary:

    default                  requires=scavenge  collected 24/49   copy-minor  0/49
    verify_evac              requires=scavenge  collected 24/49   copy-minor  0/49
    cons_scan_off            requires=scavenge  collected 24/49   copy-minor  0/49
    cons_scan_off_force      requires=scavenge  collected 24/49   copy-minor  0/49
    evac_minor               requires=move      collected 48/49   copy-minor 48/49
    force_verify             requires=move      collected 48/49   copy-minor 48/49
    

    PR_ARMS is default,evac_minor,verify_evac,force_verify,cons_scan_off,shipped_default. Four of those six run zero copying minors on all 49 files, so on a pull request they can produce nothing but UNVER. Only evac_minor and force_verify relocate anything.

    That contradicts the script's own header, which is now stale and reads as a stronger guarantee than the matrix delivers:

    default … Since #7024 this is a RELOCATING arm with no env override at all beyond the pressure knob
    arm copy-minor before #7024 after / default 0/22 → 12/22

    Measured today: 0/49. The pressure knob is inert too — evac_minor on test_gap_repsel_p4a3_ptr_numarray reports cycles=1 scavenged=3585 identically at --pressure 8, 4, 2 and 1.

    This is the same root as the note already in this issue about the gc-ratchet baseline (minor_cycles=0 everywhere because #7161 flipped the evacuating default) — but it lands somewhere sharper: the arm named after the shipped configuration, the one #6993/#7024 added specifically so the relocating-minor defect class would be reproducible per PR, is exercising the non-moving path. And gc_repsel_matrix.sh's exit status counts only FAIL, so an all-UNVER table exits 0. gc-moving-witnesses rejects UNVER as hard as FAIL; gc-stress does not.

    I have not changed any of this — it is tied to the #7161 revert timeline that #7187 already owns, and turning it red today would create exactly the untriaged permanent red this gate exists to prevent. Filed as #7255 so it is not carried only in this thread.

    Not verified by me: whether the default arm's 0/49 is entirely #7161 or also involves the --pressure interaction of #7024/#7025, and the other 45 corpus files under the zeal arm of #7254.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions