Skip to content

perf(tokenizer): reuse PCRE2 scratch for each scan - #73

Merged
AlonKejzman merged 3 commits into
crusoecloud:mainfrom
jthomson04:perf/pcre2-scratch-reuse
Sep 22, 2026
Merged

AlonKejzman merged 3 commits into
crusoecloud:mainfrom
jthomson04:perf/pcre2-scratch-reuse

Conversation

@jthomson04

@jthomson04 jthomson04 commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Summary

find_matches_pcre2 previously acquired and returned pooled PCRE2 scratch for every match. With the default JIT stack, reuse one CaptureLocations within each nonempty scan to reduce that pool overhead.

When max_jit_stack_size is configured, keep the original pooled find_at path. Creating fresh capture storage in that configuration also creates a custom JIT stack, so allocating it per scan can cause a large regression. The guard retains stock custom-stack allocation behavior without adding a cache.

Keep the existing scan loop, offsets, ordering, empty-match handling, errors, and termination. In particular, preserve the end-of-input case where replacing the loop with find_iter changes success into an error. Boundary repair, prefix reuse, BPE, Rayon scheduling, and vector preallocation remain unchanged.

Default scans still allocate match storage, and allocation failure in the dependency can panic. This change does not introduce a recoverable allocation-error API.

Validation

  • Formatting and library Clippy with warnings denied passed.
  • All 51 affected Split tests passed locally and on NVIDIA Grace.
  • The differential offset/error matrix uses the production compilation helper and covers default and 1 MiB stack settings.
  • The existing chunk-boundary integration test now forces two PCRE2 matchers and covers both settings.
  • Exact stock/candidate token IDs matched for all 14 fixed cases under default, 1 MiB, and 64 MiB settings, including concurrent calls and the 775,168-token prompt.

Grace complete-tokenizer microbenchmark

Stock Fastokens 0.3.2 c1e193ea7754306c394863e09d0def3aea211d4b versus guarded scratch reuse 5541d6eff9acc9e1b02611998b5141257ca54fbc, with PCRE2 0.2.11. Both release binaries were built on the same Grace node with identical settings and separate target directories.

Results below use the default JIT stack.

Input Callers Stock encodes/s Guarded encodes/s Speedup CPU/encode change
short 1 210,740.5 227,585.5 1.08× -7.4%
short 8 603,136.9 1,319,097.2 2.19× -53.4%
typical 1 306.7 392.2 1.28× -14.4%
typical 8 1,205.5 1,507.3 1.25× -45.2%
long 1 60.1 83.0 1.38× -14.3%
long 8 228.4 410.1 1.80× -32.4%

We also tested 1 MiB and 64 MiB custom JIT stacks on the same six cases. These retain stock pooled behavior, with no meaningful performance regression observed.

  • Qwen3-0.6B tokenizer revision c1899de289a04d12100db370d81485cdf75e47ca.
  • Three fixed AgentX-derived inputs per group: short 256 bytes each; typical 0.34–0.43 MB; long 0.88–1.61 MB. These are decoded prompt fixtures and short excerpts, not captured frontend cache misses.
  • One shared tokenizer; one or eight callers; eight BPE threads; 72 Rayon threads; NUMA0 CPUs 0–71; jemalloc. Normal internal caches enabled; optional whole-input cache disabled.
  • Fresh process per variant. Input preparation and tokenizer construction precede timing. Warm up on the actual caller threads. Measure wall time and process CPU around complete encode calls, excluding thread creation and teardown; consume output vectors with black_box.
  • Rotate through each group in a fixed order. Both variants process identical encode counts and complete input cycles, selected by pilots for approximately five seconds on the faster variant.
  • One initial measured pair per condition; 0 reverse-order repeats. Repeats are reserved for suspected regressions, inconclusive claimed gains, or material CPU/wall disagreement. No profiling or allocation instrumentation.

Acceptance passed: exact parity, retained default-stack short-input CPU savings with eight callers, and no confirmed CPU or elapsed-time regression above 5% against stock. Custom-stack cases retain the stock algorithm; differences in their one-pair measurements are not presented as a separate optimization.

These results measure warmed complete tokenizer calls. They have no statistical confidence intervals and do not predict end-to-end frontend speedup. Benchmark source and raw results remain outside this PR.

Signed-off-by: jthomson04 <jwillthomson19@gmail.com>

@jthomson04 jthomson04 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: perf(tokenizer): reuse PCRE2 scratch for each scan

I read the pcre2 0.2.11 crate internals and ran microbenchmarks against this branch to check the headline claim. Summary: the optimization is a real ~22% win in exactly one configuration — pcre2_max_jit_stack_size = None — and a 14x-113x regression in every other configuration. That option is public and user-settable, and it happens to be the only case the PR's benchmark campaign doesn't cover.

1. capture_locations() does not take from the scratch pool (split.rs:1001)

Regex::capture_locations() calls MatchData::new(...), which allocates:

  • a match context (pcre2_match_context_create_8),
  • a match data block (pcre2_match_data_create_from_pattern_8),
  • and, when max_jit_stack_size is Some, a fresh PCRE2 JIT stack (pcre2_jit_stack_create_8), freed again when the CaptureLocations drops.

So the change replaces a pool get/put with an mmap/munmap pair plus two mallocs, once per find_matches_pcre2 call.

pcre2_max_jit_stack_size is a public, user-settable option — python/src/lib.rs:389, python/fastokens/__init__.py:36.

Measured on this branch, release build, cl100k-style pattern, 31-byte input:

max_jit_stack_size find_at (before) captures_read_at (after)
None 706 ns 548 ns −22%
1 MiB 213 ns 2.96 µs 14x slower
64 MiB 142 ns 16.0 µs 113x slower

The regression isn't confined to tiny inputs: at 6.4 KB with a 1 MiB stack it is still a net loss (28.8 µs → 30.6 µs). And under the 512-concurrency profile in the description, the per-call mmap/munmap serializes on the kernel's mmap lock across every frontend thread, so the cost compounds rather than amortizes.

2. Suggested fix

The pattern this file already uses ~980 lines above is the one that actually delivers the win: a thread_local! cache (see SPLIT_CACHE, line 16) holding a CaptureLocations, reset when the regex identity changes.

That removes the per-call allocation entirely — the JIT stack stays alive across calls, the short-input win applies in every configuration, the bytes.is_empty() guard becomes unnecessary, and the same treatment extends to the find_at loop at line 602.

As written, nothing is reused across calls, despite the PR title and the // Reuse one workspace comment: capture_locations() builds a new CaptureLocations (an Arc::clone(&code) plus 2-3 allocations) on every call, and reuse happens only within a single scan.

3. Allocation failure is an assert!, now on the per-request path

pcre2 0.2.11 asserts on allocation failure inside MatchData::new — failed to allocate match context / failed to allocate match data block / failed to allocate JIT stack. That allocation used to happen roughly once per concurrent-search slot; it now happens per scan. With a large max_jit_stack_size and N parallel chunks in find_matches_pcre2_parallel, a single mmap failure (VA pressure, vm.max_map_count exhaustion) panics from inside a rayon par_iter closure instead of returning Err(Error::Unsupported), unwinding the whole encode.

4. Outside the diff: the repair path still pays the pool cost (split.rs:602)

The find_from closure in find_matches_pcre2_parallel still loops on regex.regex.find_at(bytes, p) — a pool get + PoolGuard::put per iteration, exactly the churn this PR sets out to eliminate. Any input >= 16 KB with a cross-boundary match takes that path, and empty matches advance it one byte at a time, so a repair over a long region pays the full per-byte cost. The optimization currently covers one of the file's two scan loops.

Test coverage

The new differential test is a good shape — 15 patterns x ~1557 inputs is thorough — but it misses the two cases that matter here:

  • max_jit_stack_size: None is hardcoded (line 1092), so the branch where capture_locations() behaves differently is never exercised.
  • The 139 KB input (line 1066) is built large enough to chunk, but the test calls find_matches_pcre2 directly and bypasses find_segments, so the parallel and repair paths get no coverage.

Remaining inline comments are smaller: builder-config duplication, the preallocation ahead of the empty guard, a dead match arm, and banner placement.

Comment thread src/pre_tokenizers/split.rs Outdated
Comment thread src/pre_tokenizers/split.rs Outdated
Comment thread src/pre_tokenizers/split.rs Outdated
Comment thread src/pre_tokenizers/split.rs
Comment thread src/pre_tokenizers/split.rs Outdated
Comment thread src/pre_tokenizers/split.rs Outdated
Comment thread src/pre_tokenizers/split.rs
Comment thread src/pre_tokenizers/split.rs Outdated
Comment thread src/pre_tokenizers/split.rs
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
Comment thread src/pre_tokenizers/split.rs Outdated
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
@AlonKejzman
AlonKejzman merged commit 66e634f into crusoecloud:main Sep 22, 2026
20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants