Skip to content

feat: Benchmarking infrastructure, perf optimisations, and GC/latency investigation - #7

Merged
benmandrew merged 13 commits into
mainfrom
feat/benchmarking
Jun 29, 2026
Merged

benmandrew merged 13 commits into
mainfrom
feat/benchmarking

Conversation

@benmandrew

Copy link
Copy Markdown
Owner

Overview

This branch does two things: it builds a Landmark-based benchmarking framework that makes the heap vs zero-alloc stack allocation paths measurable, and uses that framework to drive a sequence of performance optimisations. A final investigation pass then goes deeper on GC behaviour and latency distribution across all 40 QR versions.

Cumulative speedup from the original baseline: 15.06G → 3.74G cycles (4.03×) across a 120 000-call benchmark (3 inputs × 4 ECLs × 10 000 iterations).


Benchmarking infrastructure

Landmark profiling (bench/bench.ml)

A bench/ executable comparing generate_qr (heap path, Qr.make each call) against generate_qr_stack (zero-alloc path, pre-allocated Qr.t in an Arena). Each pipeline phase is wrapped with its own Landmark node so the profiler emits a full call-graph breakdown across encode → split_into_blocks → interleave_blocks → place_pattern_modules → place_format_info → place_data_and_apply_mask.

Because generate_qr_stack is [@zero_alloc], its sub-functions cannot share Landmark IDs with the heap path. A generate_qr_stack_bench variant carries its own stack/*-prefixed Landmark set for instrumented profiling, leaving the original generate_qr_stack as the clean zero-alloc reference.

A third benchmark path calls generate_qr_stack directly with only one outer Landmark enter/exit, isolating Landmark overhead from allocation overhead:

Path Cycles Minor bytes/call Major bytes/call
generate_qr (heap) 3.72G 2,275 6.1
generate_qr_stack (landmark-wrapped) 3.63G 197 0.03
generate_qr_stack (direct) 3.45G 35 0

The 197 bytes/call on the Landmark-wrapped stack path is entirely the six inner enter/exit pairs. The direct path is genuinely zero-alloc.


Performance optimisations

1. Precompute reserved-cell bitmap

Added a reserved : bytes field to Qr.t, populated once by mark_reserved. place_data and apply_mask_pattern previously called is_reserved on every cell, which walked alignment coordinate lists. They now do an O(1) bitmap lookup. Also fixes a latent bug where apply_mask_pattern passed hardcoded version=1 to is_reserved, silently skipping alignment-pattern exclusion for QR codes version 2+.

  • place_data: 7.15G → 2.45G cycles (2.9×)
  • apply_mask_pattern: 3.65G → 1.04G cycles (3.5×)

2. Rewrite mark_reserved to stamp regions

The old implementation iterated all width² cells testing each via is_reserved → is_in_alignment_pattern (O(alignment_coords²) per cell). The rewrite stamps known regions directly — three 9×9 corner rectangles, two timing strips, each alignment 5×5 square — all O(reserved_cells). is_in_alignment_pattern and is_reserved are removed entirely. Also replaces bit_pos/8 and bit_pos%8 in place_data with running refs (removing integer division from the inner loop) and (x+y) % 2 with (x+y) land 1 in apply_mask_pattern.

  • place_pattern_modules: 6.68G → 761M cycles (8.8×)

3. Combine place_data + apply_mask into one pass; optimise Reed-Solomon

Both functions performed an identical full-matrix zigzag scan. The combined place_data_and_apply_mask does a single pass applying mask pattern 0 at write time, and bypasses the set_module bounds checks since the scan guarantees valid coordinates.

Reed-Solomon: polynomial_mult(generator[j], coef) was recomputing log_table[coef] on every inner-loop iteration. Caching it once eliminates that redundancy. log_table[generator[j]] is now precomputed into a generator_log_polynomials table at startup, removing one lookup per inner step and the polynomial_mult call overhead entirely.

  • place_data_and_apply_mask: 3.40G → 1.54G cycles (2.2×)
  • split_into_blocks (RS): 1.51G → 716M cycles (2.1×)

4. Replace assoc-list lookups with flat array indexing in Config

capacity_table and ec_table were association lists built by prepending, so version 1 sat at position 159. Both are now flat int arrays indexed by (version - 1) * 4 + ecl_idx — O(1) per lookup. get_ec_info reconstructs its record inside exclave_ (stack-allocated, zero heap cost). find_version no longer allocates in its loop.

  • encode: 913M → 223M cycles (4.1×)
  • Cumulative from baseline: 15.06G → 3.74G (4.03×)

GC and latency investigation

Major-heap spill threshold (bench/bench.ml — GC stat section)

Runs 1 000 iterations per version for inputs spanning v1–v21, capturing Gc.stat deltas for minor_gc, major_gc, minor_words, promoted_words, and major_direct.

OCaml's Max_young_wosize is 256 words (2 048 bytes). Qr.t carries two Bytes.make buffers of width² bytes:

Version Per-buffer size GC path
v6 1 681 B (210 words) minor heap
v10 3 249 B (406 words) direct to major heap

At v21, the heap path triggers 5× more major GC cycles per 1 000 calls than the stack path. The stack path shows constant minor_gc=6, major_gc=3 at every version.

Mean runtime vs QR version (bench/bench_versions.ml)

Scans increasing alphanumeric string lengths to find the minimum input forcing each version 1–40 under ECL L. Uses an adaptive iteration count targeting 300 ms of wall time (min 300, max 100 000). Output: CSV with version, input_chars, heap_ns, stack_ns, heap_iters, stack_iters.

Finding: mean latency is essentially identical for both paths at every version. The heap overhead from Qr.make is 1–5% at small versions and disappears into noise above v15 as Reed-Solomon and O(width²) matrix operations dominate.

Per-call latency distribution (bench/bench_dist.ml + bench/time_ns.c)

time_ns.c wraps clock_gettime(CLOCK_MONOTONIC) as a [@@noalloc] OCaml external returning an unboxed int (nanoseconds). [@@noalloc] prevents GC safe-point insertion around the call, so no collection can slip between the timer read and the function under test.

bench_dist.ml collects 5 000 individual call latencies per path per version and outputs p50/p90/p95/p99/p99.9/max as CSV.

Tail latency diverges sharply at v8–v25 — the range where Qr.t spills to the major heap and per-call time is still short enough for GC pauses to dominate the tail:

Version heap p50 heap p99.9 stack p50 stack p99.9 ratio
v8 32 µs 211 µs 32 µs 83 µs 2.5×
v9 40 µs 287 µs 40 µs 104 µs 2.8×
v10 38 µs 265 µs 38 µs 98 µs 2.7×
v15 76 µs 321 µs 75 µs 181 µs 1.8×
v40 461 µs 956 µs 457 µs 914 µs ~1.0×

Above v30, per-call time (~270–460 µs) is long enough that OS scheduler jitter dominates both paths equally. Remaining spikes on the stack path at p99.9/max are pure OS preemption — the stack path does zero allocation inside the timed call.


Running the benchmarks

# Detailed Landmark profile (heap vs stack, sub-function breakdown, GC stats)
dune exec bench/bench.exe

# Mean runtime vs version — outputs CSV to stdout
taskset -c 7 dune exec bench/bench_versions.exe > versions.csv

# Latency distribution (p50–max) vs version — outputs CSV to stdout
taskset -c 7 dune exec bench/bench_dist.exe > dist.csv

taskset -c N pins the process to a single core to eliminate CPU-migration jitter. For cleaner tail-latency numbers, combine with sudo chrt -f 99 or configure isolcpus=N nohz_full=N in the kernel boot parameters.

Adds a bench/ executable comparing generate_qr (heap) vs
generate_qr_stack (stack) across three input sizes and all four ECLs.
Registers a landmark per pipeline phase (encode, split_into_blocks,
interleave_blocks, place_pattern_modules, place_format_info, place_data,
apply_mask_pattern) so the profiler shows a full call-graph breakdown.
Wraps call sites inside generate_qr to avoid touching [@zero_alloc]
sub-functions.
Adds a non-zero_alloc variant of generate_qr_stack with its own set of
landmark instances (stack/* prefix) so both functions expand fully in
the call graph without triggering recursive-call warnings from shared
landmark IDs.
Adds a `reserved : bytes` field to `Qr.t`, populated once by
`mark_reserved` at the end of `place_pattern_modules`. `place_data`
and `apply_mask_pattern` now do an O(1) bitmap lookup per cell instead
of calling `is_reserved` (which walked alignment coordinate lists) on
every cell.

place_data:        7.15G → 2.45G cycles (2.9x)
apply_mask_pattern: 3.65G → 1.04G cycles (3.5x)

Also fixes a latent bug where apply_mask_pattern passed hardcoded
version=1 to is_reserved, incorrectly skipping the alignment-pattern
exclusion for QR codes version 2+.
Three changes, driven by perf profiling after the reserved-cell bitmap
precomputation commit:

  1. Rewrite mark_reserved to stamp known regions directly instead of
     calling is_reserved (which called is_in_alignment_pattern) for
     every cell.  The old path iterated width² cells, each paying an
     O(coords²) alignment search.  The new path marks three fixed 9×9
     corner rectangles, the timing strips, and each alignment 5×5 square
     — all O(reserved_cells).  is_in_alignment_pattern and is_reserved
     become dead code and are removed.

  2. place_data: replace bit_pos/8 and bit_pos%8 with running byte-index
     and shift-counter refs, removing integer division from the inner
     loop.

  3. apply_mask_pattern: replace (x+y) % 2 with (x+y) land 1.

Benchmark results (120 000 calls across 3 inputs × 4 ECLs × 10 000 iters):

  ┌───────────────────────┬───────────────┬───────────────┬───────────────────────┐
  │         Phase         │    Before     │     After     │        Speedup        │
  ├───────────────────────┼───────────────┼───────────────┼───────────────────────┤
  │ place_pattern_modules │ 6.68G cycles  │  761M cycles  │ 8.8×                  │
  ├───────────────────────┼───────────────┼───────────────┼───────────────────────┤
  │ apply_mask_pattern    │ 1.04G cycles  │  881M cycles  │ 1.18×                 │
  ├───────────────────────┼───────────────┼───────────────┼───────────────────────┤
  │ generate_qr total     │ 13.80G cycles │ 7.84G cycles  │ 1.76×                 │
  └───────────────────────┴───────────────┴───────────────┴───────────────────────┘

The dominant win is mark_reserved: is_in_alignment_pattern was 22% of
total samples and is_reserved a further 12%.  Both are now gone.
oxqr.opam is generated by dune from dune-project. landmarks was added
to the (depends ...) stanza in dune-project alongside the benchmarking
commit (e480e7e) but the regenerated opam file was never staged.
…ner loop

Two independent optimisations, each targeting a separate perf hotspot.

place_data + apply_mask_pattern → place_data_and_apply_mask
  Both functions performed the same full-matrix zigzag scan, checking the
  same reserved bitmap.  The new combined function does a single pass,
  applying mask pattern 0 (XOR with (x+y+1) land 1) at write time.
  set_module's four bounds checks are also bypassed since the scan
  guarantees valid coordinates.  apply_mask_pattern is retained as a
  public API but is no longer called by the pipeline.

Reed-Solomon: hoist log_coef + precompute log-generator
  polynomial_mult(generator[j], coef) recomputed log_table[coef] on
  every iteration of the inner j loop (ec_count+1 times per data byte).
  Caching it once before the inner loop eliminates that redundancy.
  Additionally, log_table[generator[j]] is now precomputed into a
  generator_log_polynomials table at startup, removing one table lookup
  per inner step and the polynomial_mult function-call overhead entirely.

Benchmark results (120 000 calls across 3 inputs × 4 ECLs × 10 000 iters):

  ┌───────────────────────────────┬───────────────┬───────────────┬───────────┐
  │             Phase             │    Before     │     After     │  Speedup  │
  ├───────────────────────────────┼───────────────┼───────────────┼───────────┤
  │ place_data + apply_mask       │ 3.40G cycles  │               │           │
  │ → place_data_and_apply_mask   │               │ 1.54G cycles  │ 2.2×      │
  ├───────────────────────────────┼───────────────┼───────────────┼───────────┤
  │ split_into_blocks (RS)        │ 1.51G cycles  │  716M cycles  │ 2.1×      │
  ├───────────────────────────────┼───────────────┼───────────────┼───────────┤
  │ generate_qr total             │ 7.84G cycles  │ 5.12G cycles  │ 1.53×     │
  └───────────────────────────────┴───────────────┴───────────────┴───────────┘

Cumulative speedup from the original baseline: 15.06G → 5.12G (2.94×).
capacity_table and ec_table were association lists built by prepending,
so version 1 sat at position 159 in each.  Every get_capacity call
scanned all 160 entries; get_ec_info did the same.  find_version also
allocated a throwaway config record on every iteration just to look up
the capacity for that version.

Replace both tables with flat int arrays indexed by
  (version - 1) * 4 + ecl_idx
O(1) per lookup.  get_ec_info reconstructs the ec_info record from five
consecutive ints inside exclave_ (stack-allocated in the caller's frame,
zero heap cost).  find_version now avoids the throwaway make_local call
on every iteration, and get_config computes ecl_idx once before the loop.

Benchmark results (120 000 calls across 3 inputs × 4 ECLs × 10 000 iters):

  ┌──────────────────┬──────────────┬──────────────┬───────────┐
  │      Phase       │    Before    │    After     │  Speedup  │
  ├──────────────────┼──────────────┼──────────────┼───────────┤
  │ encode           │  913M cycles │  223M cycles │ 4.1×      │
  ├──────────────────┼──────────────┼──────────────┼───────────┤
  │ generate_qr      │ 5.12G cycles │ 3.74G cycles │ 1.37×     │
  └──────────────────┴──────────────┴──────────────┴───────────┘

Cumulative speedup from the original baseline: 15.06G → 3.74G (4.03×).
Three additions to the benchmarking suite, each investigating a
different aspect of heap vs zero-alloc stack allocation:

bench.ml — Landmark overhead isolation + GC mode investigation
- Add a third benchmark path, generate_qr_stack (direct), which calls
  the [@zero_alloc] function directly with one outer Landmark
  enter/exit per call rather than the six inner pairs in the
  _bench variant. Measured result: the six inner Landmark calls cost
  ~0.18G cycles (5%) and ~162 bytes/call; the direct path shows
  exactly 0 major-heap bytes across 120 000 calls.
- Add a GC-stat section that runs 1 000 iterations per version for
  inputs spanning v1–v21 (ECL L) and prints minor_gc, major_gc,
  minor_words, promoted_words, and major_direct (direct-to-major
  allocations that bypass the minor heap). Key finding: Qr.t buf and
  reserved cross the Max_young_wosize threshold (~256 words / 2 KB)
  somewhere between v6 (1 681 B per buffer, stays in minor heap) and
  v10 (3 249 B, goes directly to major heap). At v21, the heap path
  triggers 5× more major GC cycles than the stack path per 1 000 calls.

bench_versions.ml — Mean runtime vs QR version (v1–v40)
- Scans increasing alphanumeric string lengths to find the minimum
  input that forces each version 1–40 under ECL L.
- Uses an adaptive iteration count (target 300 ms of wall time,
  min 300, max 100 000 iters) so small versions get ~30 000–60 000
  iterations and large versions get ~600–1 000, giving consistent
  sub-1% noise across the full range.
- Outputs CSV (version, input_chars, heap_ns, stack_ns, heap_iters,
  stack_iters) to stdout for plotting.
- Finding: p50 heap ≈ p50 stack at every version; the mean overhead of
  Qr.make is only 1–5% and vanishes into noise above v15 as
  Reed-Solomon computation and matrix operations dominate.

bench_dist.ml + time_ns.c — Per-call latency distribution (v1–v40)
- time_ns.c wraps clock_gettime(CLOCK_MONOTONIC) as a [@@noalloc]
  OCaml external returning an unboxed int (nanoseconds). The noalloc
  attribute prevents GC safe-point insertion around the call, so no
  collection can slip between the timer read and the function under
  test.
- bench_dist.ml collects 5 000 individual call latencies per path per
  version, sorts them, and outputs p50/p90/p95/p99/p99.9/max as CSV.
- Finding: median latency is equal for both paths. The tail diverges
  sharply at v8–v25 (the range where Qr.t spills to the major heap):
  at v10, heap p99.9 = 223 µs vs stack p99.9 = 85 µs (2.6× worse).
  At v8 the heap max reaches 254 µs vs 95 µs for the stack path.
  The gap narrows at large versions (v30+) because the ~270–460 µs
  per-call time makes GC pauses a smaller fraction of total latency,
  and OS scheduler jitter dominates both paths equally.
bench_versions.ml was deleted in f21244c but its dune stanza was not,
breaking the build.
@benmandrew
benmandrew merged commit f8dbf02 into main Jun 29, 2026
1 check passed
@benmandrew
benmandrew deleted the feat/benchmarking branch June 29, 2026 22:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant