Skip to content

perf(codegen): the i32 param rep defeats LLVM shrink-wrapping on fib40's leaf path — 2.3-2.5x available (was: 'IPC collapsed 6.30 -> 1.92') #8175

Description

@proggeramlug

Summary

#8167 cut fib40's instructions 45.5x (212.14 G → 4.66 G) but wall time only 13.9x, because IPC collapsed from 6.30 to 1.92.

Perry now retires 3.5x fewer instructions than node on fib(40) and is only 1.36x faster (0.759 s vs 1.032 s). The remaining gap is a stall problem, not an instruction-count problem, and no amount of further instruction removal will close it.

The measurement

Quiet M1 mini, best-of-5, two independent interleaved runs, both VERDICT: CLEAN, perry 499e29627, node v26.5.1.

instructions cycles IPC wall
perry 38cf15336 (pre-fix) 212.14 G — 6.30 10.56 s
perry 499e29627 (post-fix) 4.66 G — 1.92 0.759 s
node 26.5.1 ~16 G — 4.98 1.032 s

1.92 is by far the lowest IPC in the corpus — every other row sits at 4.5–6.5. So this is specific to what #8167's fast path emits, not a general property of perry's codegen.

Why this matters beyond one row

It exposes a blind spot in how this campaign has been measuring. Instructions retired is load-independent, which makes it the right metric on a contended box, and essentially every A/B today used it — including my own remeasure, which reported fib40 −0.01% while the row was 55x off its own baseline. But instructions are only a proxy for time, and here the proxy broke: a 45.5x instruction win delivered 13.9x.

Any future claim of the form "N% fewer instructions" on a hot numeric loop should carry cycles or IPC alongside it, or it can overstate the delivered result by 3x.

Candidate causes, unranked and unverified

A perf-style stall breakdown, or simply reading the emitted assembly for the recursive edge, should separate these quickly. That has not been done.

Related

Activity

  1. proggeramlug commented on Aug 16, 2026

    @proggeramlug
    ContributorAuthor

    Re-scoping. All three ranked candidates in this issue are refuted, and so is the statepoint hypothesis I went in with. The mechanism is LLVM shrink-wrapping, defeated by the i32 parameter representation — and #8171 is the same defect, not a separate one.

    The statepoint hypothesis — refuted

    I expected this: classify_direct_callee (gc_call_effects.rs:30-207) is a closed allowlist of js_* helper names with _ => Unknown at :207, so a generated user function — including fib's own specialized clone — can never match; Unknown => leaf=false (precise_roots.rs:145-152) then statepoint-wraps it. The premise checks out:

    perry_fn_fib40_ts__fib$spec_i32   backend=rs4gc
      reserved_logical_slots: 1      bound_native_slots: 0
      textual_calls: 1               calls_with_live_roots: 0
    

    The clone does reserve a logical slot it never binds — collect_pointer_typed_locals sizes the frame from the untrusted declared param type, and the SpecParamRep::I32 arm then continues past js_shadow_slot_bind — and that reservation is what puts gc "statepoint-example" on the define line.

    It costs nothing. Deleting gc "statepoint-example" from the clone's define line only, then running both variants through perry's exact pipeline (LLVM 22.1.4, function(mem2reg,sccp),rewrite-statepoints-for-gc then default<O3>, llc -O3 -mcpu=apple-m1), produces byte-identical instruction streams — 267 lines each, empty diff. RS4GC emits two stackmap labels around the recursive bls and zero relocations, zero spills, zero extra instructions. All 331 M recursive calls carry no statepoint machinery. (Note the env spelling PERRY_STATEPOINT_REPORT no longer exists — deleted under the GC-knob kill policy; --statepoint-report=json is the only entry point.)

    What it actually is

    Half of all 331 M invocations are leaves, and each executes a 48-byte frame save/restore (3 stp + 3 ldp) it never uses. LLVM cannot shrink-wrap it: subs w0,w0,#1 fuses the n<2 test with n-1 and clobbers w0 in the entry block, so the leaf must read n back from x19 — a callee-saved def in the entry block, which pins the save point.

    This is a regression introduced by the i32 parameter representation, not by #8167's guard. With an f64 parameter the leaf value is already in d0, nothing defines a CSR in the entry block, and LLVM shrink-wraps normally.

    Best-of-5 on the quiet M1 mini, /usr/bin/time -l, identical driver, every row producing 102334155:

    shape instructions cycles IPC wall peak RSS
    perry main today (3be2016c1) 4.660 G 2.429 G 1.918 0.75 s 4.65 MB
    kernel replica of that asm 4.646 G 2.425 G 1.916 0.75 s —
    clang -O3 on equivalent C 4.646 G 2.425 G 1.916 0.75 s —
    f64-param shape (pre-#8033, shrink-wrapped) 3.818 G 1.252 G 3.05 0.39 s —
    i32 spec + shrink-wrapped frame 3.487 G 1.086 G 3.21 0.33 s —
    i32 spec + base case inlined at call sites 3.487 G 1.039 G 3.36 0.32 s —
    i32 spec + leaf-guard trampoline split 2.966 G 0.974 G 3.04 0.30 s —
    node v26.5.1 17.16 G 3.196 G 5.37 1.00 s 57 MB

    The f64 row reproduces both filed numbers exactly and independently: 4.646/3.818 = 1.217 (#8171's 22%) and 0.75/0.39 = 1.92 (this issue's +93.2%). Two issues, one defect. #8171 is closed on that basis.

    Perry matching clang -O3 on the same shape, to three digits, is the load-bearing control here — it rules out anything perry-specific in the codegen and localises the whole gap to the frame.

    IPC 1.92 is not itself the defect

    Perry is at 7.33 cycles per invocation against node's 9.65. The low ratio is the arithmetic consequence of retiring 3.7x fewer instructions over an unchanged call/branch structure. The quantity that actually moved is cycles, and it moved because of stack traffic — not stalls in the guard. Framing this as an "IPC collapse" pointed three separate investigations at the wrong layer, mine included.

    Prize, and why no patch is attached

    0.75 s → 0.30–0.33 s (2.3–2.5x), IPC 1.92 → 3.2–3.4, instructions −25% to −36%, RSS unchanged. That moves perry from 1.33x to 3.0–3.3x faster than node on this row.

    No IR reshaping reaches it. Tried and rejected, all producing identical non-shrink-wrapped MIR: branch weights in both directions, integer vs fcmp compare form, hoisting the sitofp into the entry block, a single phi return, -enable-partial-inlining, -disable-peephole, -enable-shrink-wrap=true. The only frontend route is splitting the function into a frameless leaf guard plus a noinline body — verified to produce the 0.30 s row, with LLVM then inlining the guard into the recursive call sites.

    It is narrower than it looks, which is why I stopped rather than shipping a recognizer. A second binary-recursive shape compiled through the real compiler — build(depth) { if (depth<=0) return 1; return 1+build(depth-1)+build(depth-1) } — does shrink-wrap (cmp w0,#0; b.le ahead of the frame, 2-instruction leaf). The pathology needs a base case that uses the parameter and ≥2 recursive calls. A recognizer narrow enough to be provably safe would match essentially fib alone; the general form — "split any function with a call-free early exit that LLVM failed to shrink-wrap" — is a new codegen capability that needs corpus evidence first. That corpus check is exactly what #8171 asked for and what nobody has run.

    Consequence for #8079

    fib40's entry there is a genuine regression, not a residual. The pre-#8033 build was the shrink-wrapping f64 shape, and every candidate above beats it.

    One free tidy-up, measured as worth zero

    The clone reserves a root slot it never binds, so gc "statepoint-example" lands on a function with no roots. Cost is stackmap metadata only — not worth touching root-slot reservation for.

  2. changed the title [-]perf: fib40's IPC collapsed 6.30 → 1.92 — #8167's 45.5x instruction win delivers only 13.9x wall time[/-] [+]perf(codegen): the i32 param rep defeats LLVM shrink-wrapping on fib40's leaf path — 2.3-2.5x available (was: 'IPC collapsed 6.30 -> 1.92')[/+] on Aug 16, 2026
  3. proggeramlug commented on Aug 16, 2026

    @proggeramlug
    ContributorAuthor

    Decision: attack the root with a calling convention, not a splitter — and this issue's stated trigger is wrong in both directions

    A corpus sweep at bfb0707be: 1,347 sources (all 157 gc-handoff/ including apps/ and mdapp, plus 1,190 of 1,280 test-files/; 90 compile-fail, 0 unaccounted) compiled through the real compiler, every emitted binary disassembled, all 6,878 user functions classified by CFG walk. Detector self-tested both ways (fib ⇒ SW_FAIL, build ⇒ SW_OK) plus a planted-positive probe matrix.

    The population

    • 104 functions (1.5%) pay a frame LLVM failed to shrink-wrap; 51 more have the shape and LLVM handled it.
    • Recursive instances: 12 sites, 8 unique functions — fib (×4 files), tree*.ts count (×3), and 4 cold test helpers.
    • A recogniser scoped as this issue describes (recursive + i32-spec) reaches fib and a 16-byte sumTo. It is a benchmark special.
    • The 92 non-recursive hits are overwhelmingly cold empty-loop guards ahead of hot loop bodies — frame runs once per call, body loops, amortized to noise. The "call" pinning them is often the loop's own GC poll.
    • The realistic-app corpus has zero hot instances. evalNode/parseExpr in interp/iso_* — the actual recursion — are all NO_EARLY, because realistic early exits call helpers and therefore need the frame anyway.

    Runtime-weighted, the population is fib (2.2x) plus tree-count (−1%, i.e. nothing).

    The trigger characterization here is wrong, and a recogniser built on it would both overfire and underfire

    Probe matrix through the real compiler, each verified in disassembly:

    probe shape verdict
    fib 2 rec calls, base uses param SW_FAIL, 48 B (subs → x19 pin)
    sum1 1 rec call, base uses param SW_FAIL, 32 B (hoisted scvtf → d8 pin)
    nonrec2 non-recursive, guard + 2 calls SW_FAIL, 48 B
    nonrec1 non-rec, 1 call, param dead after SW_OK
    build base ignores param SW_OK
    memo-fib, ack realistic recursive SW_FAIL
    isEven/isOdd (mutual) base is a constant SW_FAIL, 16 B — different mode

    Two real modes, neither of which is "≥2 recursive calls":

    1. Any param-derived value or hoisted invariant live across a call gets materialised into a callee-saved register in the entry block, pinning the save point — a fused subs, a hoisted scvtf, or (in tree.ts count, a pointer-param function) nanbox constants and registry pointers. So it is not i32-specific either.
    2. The cross-function spec-call guard diamond creates two call+ret paths, and LLVM shrink-wrap supports only a single restore point.

    Recommendation: preserve_none on recursion-participating specialized clones

    Measured on the quiet M1 mini, best-of-5, binaries verified different, every row producing 102334155. CC patched on the internal clone and its module-local call sites in --trace llvm IR, pushed through perry's exact pipeline (LLVM 22.1.4) and linked with perry's own captured link line:

    arm instr cycles IPC wall peak RSS
    fib40 control (replicates main to 3 digits) 4.660 G 2.428 G 1.92 0.78 s 2.10 MB
    fib40 preserve_none 3.666 G 1.091 G 3.36 0.35 s 2.10 MB
    trampoline split, for reference 2.966 G 0.974 G 3.04 0.30 s —
    tree.ts control 16.119 G 3.427 G — 1.10 s 32.9 MB
    tree.ts preserve_none count 15.962 G 3.405 G — 1.10 s 32.9 MB

    With no callee-saved registers there is nothing to pin: the leaf becomes 5 instructions with zero frame (3 on x86-64, also verified). 2.2x on fib40 — 93% of the split's cycle win — with no new pass, no cost model and no pattern matching. It deletes the cause rather than recognising the symptom. RSS unchanged.

    The safety argument is scoping, and it is already an invariant: spec clones are internal and direct-call-only by construction (spec_abi.rs, asserted at :289), so the convention cannot escape the module. Verified compatible with invoke/EH (perry does invoke spec clones inside try regions) and with RS4GC statepoints. The boundary cost is a normal-CC caller saving all 20 CSRs — a 160-byte prologue once per entry, amortized under a recursive tree but a genuine pessimisation risk for a cheap non-recursive callee in a hot loop, which is why the gate is "participates in recursion".

    It does not address mode 2 — isEven keeps its 16-byte fp/lr frame. That lever is cross-call guard elimination, i.e. #8171's flow-sensitive narrowing applied cross-function, and belongs in its own issue. tree-count's −1% says extending the clone criteria to boxed-rep recursion is not worth doing today.

    Implementation surface, moderate and contained: a CC token in function.rs::define_header and declare rendering; dialect/mod.rs (LLVMSetFunctionCallConv plus per-call-site CC); both spec dispatch tiers' call emission; force_external unit-promotion must carry the CC or refuse to split (a mismatch is UB); target-gate off watchOS arm64_32 and Windows ARM64 (same predicate family as the RS4GC target-awareness); and a subject-liveness gate asserting the clone's first instruction is not a frame store.

    Not implemented, deliberately

    The evidence says (c) is clearly right; "clearly safe" needs EH-unwind stress, GC-schedule runs over pointer-param clones, cross-codegen-unit CC carrying, and the target gating — a real PR's validation, not a decision session's.

    What would change this

    • A realistic-app hit (Effect, a Next route, marked) with a hot call-free early exit → upgrade to the general split, which also covers the non-recursive shapes and the last 7% (0.30 s vs 0.35 s).
    • A measured regression from the caller-CSR boundary cost under recursive scoping → retreat to a benchmark-special and say so plainly.
    • No implementation budget → doing nothing is fully defensible, on the honest ground that this serves one hot function.
  4. proggeramlug commented on Aug 17, 2026

    @proggeramlug
    ContributorAuthor

    Fresh data from the 2026-08-17 four-engine sweep, which changes the shape of this issue: fib40 is not a steady-state gap, it is a partly-repaired catastrophe, and it is now the one benchmark where Perry lost a clear win.

    The three-point sequence

    commit date wall instructions retired
    601a02d23 08-14 0.3934 s —
    38cf15336 08-15 10.5587 s 212.142 B
    14468dcbc 08-17 0.7619 s 4.660 B

    #8094's guarded ordinary-parameter specialization made fib40 27× slower. The param-guard work since (#8201 / #8238 / #8242) recovered −92.8% wall and −97.8% instructions — 212.1 B down to 4.66 B.

    It is still +93.7% above the 08-14 baseline, and that is enough to lose the row:

    engine wall
    Porffor 0.5133
    Perry 0.7619
    scriptc 0.7900
    Node 1.0361

    Perry used to win fib40 outright against all three. It now loses to Porffor and is barely ahead of scriptc. Measurement is clean: five interleaved runs, samples 0.7627 0.7619 0.7633 0.7642 0.7620, spread 0.3% — this is not noise.

    What this means for the fix

    The instruction count is already near the floor (4.66 B vs Porffor's 5.474 B — Perry executes fewer instructions than Porffor and is still slower on the clock). That points away from "emit less work" and squarely at this issue's thesis: the i32 parameter representation defeating LLVM's shrink-wrapping on the leaf, so the cost is prologue/epilogue and spill traffic rather than instruction count.

    PR #8203 (preserve_none calling convention for recursion participants, measured 2.2×) is still in draft. Given fib40 is now a lost row rather than a slow-but-winning one, that draft is the highest-value unlanded perf work I know of. Worth finishing, measuring against 14468dcbc, and landing.

    Sweep evidence: gc-handoff/current-sweep-2026-08-17/.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions