Skip to content

perf(json): tape materialization allocates field_count=0, so every 5+-field parsed record spills to overflow storage #7267

Description

@proggeramlug

Found while fixing #7264 (PR #7265). Perf, not correctness.

json_tape.rs::materialize_object allocates every object with js_object_alloc(0, 0) and then adds each property with js_object_set_field_by_name:

let obj = crate::object::js_object_alloc(0, 0);
...
crate::object::js_object_set_field_by_name(obj, key_ptr, value);

js_object_alloc_with_parent reserves max(field_count, INLINE_SLOT_FLOOR) slots — with field_count = 0 that is the floor, currently 4. Slots 4..N therefore go to overflow storage (spill.rs), and field_count stays pinned at 4. So every JSON record with 5 or more properties — the single most common API shape there is — keeps its 5th and later values out of line, paid on every read and every write, plus a per-object spill entry that every GC cycle visits, rekeys and finalizes.

The direct parser does not have this problem: both parse_object and parse_object_shaped pass the real key count to js_object_alloc_class_inline_keys, so everything lands inline.

Measurement

benchmarks/json_polyglot/bench_field_access.ts (10k records × 5 fields, 50 iterations; nested is index 4, i.e. the first overflow slot, and the workload reads it every iteration), same binary, --profile perry-dev:

default (tape)        2950–2983 ms
PERRY_JSON_TAPE=0      870–877 ms

3.4×. Some of that gap is other tape work, but the overflow lookup is on the hot read path by construction here, and #7265 puts it on the stringify path as well (correctly — the values genuinely live there).

Suggested fix

Count the object's top-level keys before allocating. The tape already carries what is needed: each KIND_OBJ_START/KIND_ARR_START entry stores its matching end index in link, so a top-level walk from idx to end_idx that counts KIND_KEY and skips nested values via link is O(top-level entries). Then js_object_alloc(0, key_count) and every field lands inline.

This is a behaviour change to the parser (allocation sizing, field_count in the header, spill-map population) rather than a correctness fix, which is why #7265 left it alone. Worth benchmarking on its own.

Note the interaction with #6712: the floor moved 8 → 4, which doubled the population of records that spill. Anything that measured tape-path JSON before 2026-07-20 was measuring a materially different allocation profile.

Activity

proggeramlug commented on Aug 10, 2026

@proggeramlug
ContributorAuthor

Implemented the suggested fix and measured it as a regression — posting the numbers rather than the patch, because the premise does not survive measurement.

What I built

Exactly what the issue proposes: a top-level walk from idx to end_idx counting KIND_KEY and skipping each value, stepping over a nested container in O(1) through its link, then js_object_alloc(0, key_count) instead of js_object_alloc(0, 0).

What it measured

benchmarks/json_polyglot/bench_field_access.ts — the benchmark this issue names — perry-dev, PERRY_NO_AUTO_OPTIMIZE=1, archives rebuilt and mtime-verified, 7 interleaved runs per arm (A/B/A/B, not batched, because this host is noisy enough that batched runs invert):

arm min median
clean main 1.557 s 1.625 s
with the pre-sizing fix 1.704 s 1.759 s

+9.5% on min, and the distributions do not overlap (baseline 1.56–2.05, fixed 1.70–1.79).

Which half costs

To separate "the counting walk is expensive" from "the bigger objects are expensive", I replaced the walk with a constant — let key_count = 8 — so the allocation grows with no walk at all. Fresh interleaved batch:

arm min
clean main 1.764 s
constant 8 slots, no walk 2.003 s (+13.6%)

So it is not the walk. Reserving more inline slots is itself the cost, and the walk is roughly free next to it. Whatever the overflow lookup costs on the read path, it is smaller than the extra per-object footprint across 10k records × 50 iterations — more bytes per object is more allocation and more for every GC cycle to move.

(Absolute numbers drift between batches — the baseline min moved 1.557 → 1.764 as the host warmed — which is why only within-batch interleaved comparisons are quoted.)

Note on the original 3.4×

The 2950–2983 ms vs 870–877 ms figures in the issue do not reproduce: on current main the same benchmark is ~1.56 s with the tape and ~1.55 s with PERRY_JSON_TAPE=0, i.e. the tape/no-tape gap is now roughly 1.0×, not 3.4×. Something between the filing commit and db44b31b7 closed that gap, so the motivating measurement is stale independently of the fix.

This is the same shape as #7714's re-profile: per-object footprint arithmetic predicts a win that allocation behaviour then reverses. Worth keeping the two together — "spill is expensive, so reserve more slots" is the intuition both tickets are built on, and it has now measured negative twice.

Not closing, since the underlying observation (5+-field records spill) is still true and a different remedy — one that does not grow the object, e.g. sizing only when the count exceeds the floor by enough to pay for itself, or reducing the per-spill-entry GC cost instead — could still win. But the fix as specified should not be implemented; it makes the named benchmark slower.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions