Repository navigation
perf(codegen): the i32 param rep defeats LLVM shrink-wrapping on fib40's leaf path — 2.3-2.5x available (was: 'IPC collapsed 6.30 -> 1.92') #8175
Description
Activity
Re-scoping. All three ranked candidates in this issue are refuted, and so is the statepoint hypothesis I went in with. The mechanism is LLVM shrink-wrapping, defeated by the i32 parameter representation — and #8171 is the same defect, not a separate one.
The statepoint hypothesis — refuted
I expected this:
classify_direct_callee(gc_call_effects.rs:30-207) is a closed allowlist ofjs_*helper names with_ => Unknownat:207, so a generated user function — including fib's own specialized clone — can never match;Unknown => leaf=false(precise_roots.rs:145-152) then statepoint-wraps it. The premise checks out:perry_fn_fib40_ts__fib$spec_i32 backend=rs4gc reserved_logical_slots: 1 bound_native_slots: 0 textual_calls: 1 calls_with_live_roots: 0The clone does reserve a logical slot it never binds —
collect_pointer_typed_localssizes the frame from the untrusted declared param type, and theSpecParamRep::I32arm thencontinues pastjs_shadow_slot_bind— and that reservation is what putsgc "statepoint-example"on the define line.It costs nothing. Deleting
gc "statepoint-example"from the clone's define line only, then running both variants through perry's exact pipeline (LLVM 22.1.4,function(mem2reg,sccp),rewrite-statepoints-for-gcthendefault<O3>,llc -O3 -mcpu=apple-m1), produces byte-identical instruction streams — 267 lines each, emptydiff. RS4GC emits two stackmap labels around the recursivebls and zero relocations, zero spills, zero extra instructions. All 331 M recursive calls carry no statepoint machinery. (Note the env spellingPERRY_STATEPOINT_REPORTno longer exists — deleted under the GC-knob kill policy;--statepoint-report=jsonis the only entry point.)What it actually is
Half of all 331 M invocations are leaves, and each executes a 48-byte frame save/restore (3
stp+ 3ldp) it never uses. LLVM cannot shrink-wrap it:subs w0,w0,#1fuses then<2test withn-1and clobbersw0in the entry block, so the leaf must readnback fromx19— a callee-saved def in the entry block, which pins the save point.This is a regression introduced by the i32 parameter representation, not by #8167's guard. With an f64 parameter the leaf value is already in
d0, nothing defines a CSR in the entry block, and LLVM shrink-wraps normally.Best-of-5 on the quiet M1 mini,
/usr/bin/time -l, identical driver, every row producing 102334155:shape instructions cycles IPC wall peak RSS perry main today ( 3be2016c1)4.660 G 2.429 G 1.918 0.75 s 4.65 MB kernel replica of that asm 4.646 G 2.425 G 1.916 0.75 s — clang -O3on equivalent C4.646 G 2.425 G 1.916 0.75 s — f64-param shape (pre-#8033, shrink-wrapped) 3.818 G 1.252 G 3.05 0.39 s — i32 spec + shrink-wrapped frame 3.487 G 1.086 G 3.21 0.33 s — i32 spec + base case inlined at call sites 3.487 G 1.039 G 3.36 0.32 s — i32 spec + leaf-guard trampoline split 2.966 G 0.974 G 3.04 0.30 s — node v26.5.1 17.16 G 3.196 G 5.37 1.00 s 57 MB The f64 row reproduces both filed numbers exactly and independently: 4.646/3.818 = 1.217 (#8171's 22%) and 0.75/0.39 = 1.92 (this issue's +93.2%). Two issues, one defect. #8171 is closed on that basis.
Perry matching
clang -O3on the same shape, to three digits, is the load-bearing control here — it rules out anything perry-specific in the codegen and localises the whole gap to the frame.IPC 1.92 is not itself the defect
Perry is at 7.33 cycles per invocation against node's 9.65. The low ratio is the arithmetic consequence of retiring 3.7x fewer instructions over an unchanged call/branch structure. The quantity that actually moved is cycles, and it moved because of stack traffic — not stalls in the guard. Framing this as an "IPC collapse" pointed three separate investigations at the wrong layer, mine included.
Prize, and why no patch is attached
0.75 s → 0.30–0.33 s (2.3–2.5x), IPC 1.92 → 3.2–3.4, instructions −25% to −36%, RSS unchanged. That moves perry from 1.33x to 3.0–3.3x faster than node on this row.
No IR reshaping reaches it. Tried and rejected, all producing identical non-shrink-wrapped MIR: branch weights in both directions, integer vs
fcmpcompare form, hoisting thesitofpinto the entry block, a singlephireturn,-enable-partial-inlining,-disable-peephole,-enable-shrink-wrap=true. The only frontend route is splitting the function into a frameless leaf guard plus anoinlinebody — verified to produce the 0.30 s row, with LLVM then inlining the guard into the recursive call sites.It is narrower than it looks, which is why I stopped rather than shipping a recognizer. A second binary-recursive shape compiled through the real compiler —
build(depth) { if (depth<=0) return 1; return 1+build(depth-1)+build(depth-1) }— does shrink-wrap (cmp w0,#0; b.leahead of the frame, 2-instruction leaf). The pathology needs a base case that uses the parameter and ≥2 recursive calls. A recognizer narrow enough to be provably safe would match essentiallyfibalone; the general form — "split any function with a call-free early exit that LLVM failed to shrink-wrap" — is a new codegen capability that needs corpus evidence first. That corpus check is exactly what #8171 asked for and what nobody has run.Consequence for #8079
fib40's entry there is a genuine regression, not a residual. The pre-#8033 build was the shrink-wrapping f64 shape, and every candidate above beats it.
One free tidy-up, measured as worth zero
The clone reserves a root slot it never binds, so
gc "statepoint-example"lands on a function with no roots. Cost is stackmap metadata only — not worth touching root-slot reservation for.- changed the title
[-]perf: fib40's IPC collapsed 6.30 → 1.92 — #8167's 45.5x instruction win delivers only 13.9x wall time[/-][+]perf(codegen): the i32 param rep defeats LLVM shrink-wrapping on fib40's leaf path — 2.3-2.5x available (was: 'IPC collapsed 6.30 -> 1.92')[/+]on Aug 16, 2026 Decision: attack the root with a calling convention, not a splitter — and this issue's stated trigger is wrong in both directions
A corpus sweep at
bfb0707be: 1,347 sources (all 157gc-handoff/includingapps/andmdapp, plus 1,190 of 1,280test-files/; 90 compile-fail, 0 unaccounted) compiled through the real compiler, every emitted binary disassembled, all 6,878 user functions classified by CFG walk. Detector self-tested both ways (fib ⇒ SW_FAIL,build⇒ SW_OK) plus a planted-positive probe matrix.The population
- 104 functions (1.5%) pay a frame LLVM failed to shrink-wrap; 51 more have the shape and LLVM handled it.
- Recursive instances: 12 sites, 8 unique functions — fib (×4 files),
tree*.ts count(×3), and 4 cold test helpers. - A recogniser scoped as this issue describes (recursive + i32-spec) reaches fib and a 16-byte
sumTo. It is a benchmark special. - The 92 non-recursive hits are overwhelmingly cold empty-loop guards ahead of hot loop bodies — frame runs once per call, body loops, amortized to noise. The "call" pinning them is often the loop's own GC poll.
- The realistic-app corpus has zero hot instances.
evalNode/parseExprininterp/iso_*— the actual recursion — are all NO_EARLY, because realistic early exits call helpers and therefore need the frame anyway.
Runtime-weighted, the population is fib (2.2x) plus tree-count (−1%, i.e. nothing).
The trigger characterization here is wrong, and a recogniser built on it would both overfire and underfire
Probe matrix through the real compiler, each verified in disassembly:
probe shape verdict fib 2 rec calls, base uses param SW_FAIL, 48 B ( subs→x19pin)sum1 1 rec call, base uses param SW_FAIL, 32 B (hoisted scvtf→d8pin)nonrec2 non-recursive, guard + 2 calls SW_FAIL, 48 B nonrec1 non-rec, 1 call, param dead after SW_OK build base ignores param SW_OK memo-fib, ack realistic recursive SW_FAIL isEven/isOdd (mutual) base is a constant SW_FAIL, 16 B — different mode Two real modes, neither of which is "≥2 recursive calls":
- Any param-derived value or hoisted invariant live across a call gets materialised into a callee-saved register in the entry block, pinning the save point — a fused
subs, a hoistedscvtf, or (intree.ts count, a pointer-param function) nanbox constants and registry pointers. So it is not i32-specific either. - The cross-function spec-call guard diamond creates two call+ret paths, and LLVM shrink-wrap supports only a single restore point.
Recommendation:
preserve_noneon recursion-participating specialized clonesMeasured on the quiet M1 mini, best-of-5, binaries verified different, every row producing 102334155. CC patched on the internal clone and its module-local call sites in
--trace llvmIR, pushed through perry's exact pipeline (LLVM 22.1.4) and linked with perry's own captured link line:arm instr cycles IPC wall peak RSS fib40 control (replicates main to 3 digits) 4.660 G 2.428 G 1.92 0.78 s 2.10 MB fib40 preserve_none3.666 G 1.091 G 3.36 0.35 s 2.10 MB trampoline split, for reference 2.966 G 0.974 G 3.04 0.30 s — tree.ts control 16.119 G 3.427 G — 1.10 s 32.9 MB tree.ts preserve_nonecount15.962 G 3.405 G — 1.10 s 32.9 MB With no callee-saved registers there is nothing to pin: the leaf becomes 5 instructions with zero frame (3 on x86-64, also verified). 2.2x on fib40 — 93% of the split's cycle win — with no new pass, no cost model and no pattern matching. It deletes the cause rather than recognising the symptom. RSS unchanged.
The safety argument is scoping, and it is already an invariant: spec clones are
internaland direct-call-only by construction (spec_abi.rs, asserted at:289), so the convention cannot escape the module. Verified compatible withinvoke/EH (perry doesinvokespec clones inside try regions) and with RS4GC statepoints. The boundary cost is a normal-CC caller saving all 20 CSRs — a 160-byte prologue once per entry, amortized under a recursive tree but a genuine pessimisation risk for a cheap non-recursive callee in a hot loop, which is why the gate is "participates in recursion".It does not address mode 2 —
isEvenkeeps its 16-byte fp/lr frame. That lever is cross-call guard elimination, i.e. #8171's flow-sensitive narrowing applied cross-function, and belongs in its own issue.tree-count's −1% says extending the clone criteria to boxed-rep recursion is not worth doing today.Implementation surface, moderate and contained: a CC token in
function.rs::define_headerand declare rendering;dialect/mod.rs(LLVMSetFunctionCallConvplus per-call-site CC); both spec dispatch tiers' call emission;force_externalunit-promotion must carry the CC or refuse to split (a mismatch is UB); target-gate off watchOSarm64_32and Windows ARM64 (same predicate family as the RS4GC target-awareness); and a subject-liveness gate asserting the clone's first instruction is not a frame store.Not implemented, deliberately
The evidence says (c) is clearly right; "clearly safe" needs EH-unwind stress, GC-schedule runs over pointer-param clones, cross-codegen-unit CC carrying, and the target gating — a real PR's validation, not a decision session's.
What would change this
- A realistic-app hit (Effect, a Next route, marked) with a hot call-free early exit → upgrade to the general split, which also covers the non-recursive shapes and the last 7% (0.30 s vs 0.35 s).
- A measured regression from the caller-CSR boundary cost under recursive scoping → retreat to a benchmark-special and say so plainly.
- No implementation budget → doing nothing is fully defensible, on the honest ground that this serves one hot function.
Fresh data from the 2026-08-17 four-engine sweep, which changes the shape of this issue:
fib40is not a steady-state gap, it is a partly-repaired catastrophe, and it is now the one benchmark where Perry lost a clear win.The three-point sequence
commit date wall instructions retired 601a02d2308-14 0.3934 s — 38cf1533608-15 10.5587 s 212.142 B 14468dcbc08-17 0.7619 s 4.660 B #8094's guarded ordinary-parameter specialization madefib4027× slower. The param-guard work since (#8201 / #8238 / #8242) recovered −92.8% wall and −97.8% instructions — 212.1 B down to 4.66 B.It is still +93.7% above the 08-14 baseline, and that is enough to lose the row:
engine wall Porffor 0.5133 Perry 0.7619 scriptc 0.7900 Node 1.0361 Perry used to win
fib40outright against all three. It now loses to Porffor and is barely ahead of scriptc. Measurement is clean: five interleaved runs, samples0.7627 0.7619 0.7633 0.7642 0.7620, spread 0.3% — this is not noise.What this means for the fix
The instruction count is already near the floor (4.66 B vs Porffor's 5.474 B — Perry executes fewer instructions than Porffor and is still slower on the clock). That points away from "emit less work" and squarely at this issue's thesis: the i32 parameter representation defeating LLVM's shrink-wrapping on the leaf, so the cost is prologue/epilogue and spill traffic rather than instruction count.
PR #8203 (
preserve_nonecalling convention for recursion participants, measured 2.2×) is still in draft. Givenfib40is now a lost row rather than a slow-but-winning one, that draft is the highest-value unlanded perf work I know of. Worth finishing, measuring against14468dcbc, and landing.Sweep evidence:
gc-handoff/current-sweep-2026-08-17/.
Summary
#8167 cut
fib40's instructions 45.5x (212.14 G → 4.66 G) but wall time only 13.9x, because IPC collapsed from 6.30 to 1.92.Perry now retires 3.5x fewer instructions than node on
fib(40)and is only 1.36x faster (0.759 s vs 1.032 s). The remaining gap is a stall problem, not an instruction-count problem, and no amount of further instruction removal will close it.The measurement
Quiet M1 mini, best-of-5, two independent interleaved runs, both
VERDICT: CLEAN, perry499e29627, node v26.5.1.38cf15336(pre-fix)499e29627(post-fix)1.92 is by far the lowest IPC in the corpus — every other row sits at 4.5–6.5. So this is specific to what #8167's fast path emits, not a general property of perry's codegen.
Why this matters beyond one row
It exposes a blind spot in how this campaign has been measuring. Instructions retired is load-independent, which makes it the right metric on a contended box, and essentially every A/B today used it — including my own remeasure, which reported
fib40 −0.01%while the row was 55x off its own baseline. But instructions are only a proxy for time, and here the proxy broke: a 45.5x instruction win delivered 13.9x.Any future claim of the form "N% fewer instructions" on a hot numeric loop should carry cycles or IPC alongside it, or it can overstate the delivered result by 3x.
Candidate causes, unranked and unverified
spec_i32_derived_windowcannot prove containment forn - 1(its window is one value wider than the slot), so each recursive call carries a range check plus a cold boxed arm. Two per invocation across ~331 M invocations, and a hard-to-predict branch would hurt IPC exactly this way. perf(codegen): #8167 leaves 22% on fib40 — the range test is provably unnecessary on a branch that already narrows the parameter #8171 proposes removing it for this shape via flow-sensitive narrowing — that issue was filed as a 22% instruction residual, but if the branch is also the stall source its real value is larger than 22%.alwaysinlinebut self-recursive, so it cannot actually inline into itself).A
perf-style stall breakdown, or simply reading the emitted assembly for the recursive edge, should separate these quickly. That has not been done.Related
fib40is still +93.2% vs 2026-08-12 even after the 45.5x recovery, and the other eight regressed rows are entirely unmoved by perf(codegen): let a specialized entry re-enter itself #8167.