Skip to content

Faster front end, [In] parser fix, layout pins, diagnostic helpers - #156

Merged
Lewin671 merged 10 commits into
mainfrom
agent/perf-21/claude
Sep 26, 2026
Merged

Lewin671 merged 10 commits into
mainfrom
agent/perf-21/claude

Conversation

@Lewin671

Copy link
Copy Markdown
Owner

Summary

  • Front end (measured 2-3x QuickJS-NG, 1-8% of many cases' totals): binary operators by precedence climbing instead of ten one-function-per-level parsers; the compiler's scope tables on the engine's name hasher; canonical UTF-8 copied straight through string_from_utf8_scalars (function source text for every function); borrowed name sets in bytecode finalization.
  • Parser fix: in inside brackets of a for initializer (for (var a = ('k' in o); ...), a function body, arguments, templates) was a SyntaxError; it is withheld only at the initializer's own bracket depth now.
  • Layout: math_binary kept out of the typed executor (a codegen-unit shift had inlined it: access-nsieve +4% at equal instructions), boxed_equality and get_named pinned, executor offset re-scanned to 0x200.
  • Workflow: tools.benchmark.front_end (front-end cost vs NG or another build), tools.benchmark.bundles (corpus bundles on disk), scripts/layout-scan.sh (executor placement scan); documented in docs/performance-workflow.md and AGENTS.md. None touches the hashed measurement protocol.

Evidence

Formal stack run 3babb93 vs main f2b21ab (30 blocks, target/comparison/perf23-3babb930): external geomean 0.994 vs main, 0.772 vs QuickJS-NG; gaussian-blur 0.952, imaging-darkroom 0.963, tinderbox 0.965. Worst: ai-astar 1.023 (two byte-identical binaries at different paths differ by 1-2% on it), regexp-dna 1.015. Sentinels 1.003.

Verification

  • cargo test -p qjs-parser, cargo test -p qjs-runtime (2246), new tests for precedence/associativity and the for-initializer in cases (expected values from V8).
  • Upstream Test262 language/statements/for, for-in, expressions/in, exponentiation, coalesce, logical-and: 0 gaps.
  • check.sh via pre-push; tools unit tests.

🤖 Generated with Claude Code

Lewin671 and others added 10 commits September 25, 2026 22:52
A for head's initializer is parsed with the grammar's [~In] so that `in`
can start a for-in, but every bracketed production inside it --
parentheses, array and object literals, arguments, member brackets,
template substitutions, function bodies -- takes [+In] again. The parser
cleared `allow_in` for the whole initializer, so `for (var a = ('k' in o);
...)` and `for (var f = function () { return k in o; }; ...)` were syntax
errors. The parser now records each token's bracket depth and withholds
`in` only at the depth the initializer started at; an unbracketed arrow
body still inherits the restriction.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Each operand of a binary expression descended through ten
one-function-per-level parsers (`||` down to `*`), each returning the
expression by value and scanning its own operator list, before the one
level that owned the operator saw it. One loop now takes an operand once
and binds each operator by its precedence, building the same
left-associative trees. Parsing imaging-darkroom (a 1.8 MB literal)
takes 21% fewer cycles; crypto-md5's parse 12% fewer.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every identifier the compiler resolves looks its name up in the lexical
scope maps and the local-slot map, which used the standard SipHash.
They now use the engine's `NameMap`/`NameSet` (FxHash), like the
runtime's name tables. Parsing and compiling cdjs takes about 10% fewer
cycles.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Converting host UTF-8 to the engine's representation pushed every scalar
through the surrogate-escape check, and function source text for
`Function.prototype.toString` is converted this way for every function a
script declares, nested ones included. Only a scalar at or above U+10000
-- a UTF-8 sequence led by 0xF0 or above -- can fall in the sentinel range,
so text without one is copied as it is. cdjs's parse and compile run 9%
fewer instructions.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Finishing each compiled body gathers the names it and its nested
closures read and write, cloning every name (and every nested closure's
whole cached list) into a BTreeSet<String>, once per store instruction.
The sets now borrow the names and the lists are cloned once, at the end.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Whether LLVM inlined vm_numeric_leaf::math_binary into the typed-loop
executor flipped with unrelated edits (here, front-end changes in other
modules): inlined, it grew the executor by 176 bytes and re-rolled its
code, and access-nsieve ran 4% and heterogeneous_property_read 5% more
cycles at identical instructions. It is now out of line, as it was, and
the order file is re-pinned.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Both are hot typed-executor callees (ai-astar's equality and named reads)
that floated in the unordered tail; they now own slots after the other
callees. With every callee fixed, a scan of the executor's offset put
0x200 best across the canaries (string_key_map_churn 0.98,
heterogeneous_property_read 1.00, imaging-desaturate 0.97 against the
0x0 placement's 1.00, 1.06, 0.97). Byte-identical binaries at different
paths still measure ai-astar 1-2% apart, which bounds what a scan can
resolve there.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Three diagnostics this session's work kept rebuilding by hand in /tmp:
`tools.benchmark.front_end` measures lexing, parsing and compiling of each
external bundle (wrapped in a function never called, process start
subtracted) against QuickJS-NG or another build, with each case's share
of its whole run; `tools.benchmark.bundles` writes the corpus bundles to
target/bundles for profiling; `scripts/layout-scan.sh` pins the
typed-loop executor at each offset, relinks to a fixed point and screens
the layout canaries. None touches the hashed measurement protocol.
docs/performance-workflow.md lists them with the existing counter
screen, which is the whole-corpus A/B to use instead of ad-hoc loops.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A binary's first process after a build or a pause runs slower, which
inflated the empty-script baseline past small front ends (crypto-md5 read
negative). One run is now discarded, and the least of the pairs estimates
each cost. Also records the perf23 units and stack run in T033.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@Lewin671
Lewin671 merged commit 8310883 into main Sep 26, 2026
7 of 10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant