perf(tokenizer): reuse PCRE2 scratch for each scan - #73
Conversation
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
jthomson04
left a comment
There was a problem hiding this comment.
Review: perf(tokenizer): reuse PCRE2 scratch for each scan
I read the pcre2 0.2.11 crate internals and ran microbenchmarks against this branch to check the headline claim. Summary: the optimization is a real ~22% win in exactly one configuration — pcre2_max_jit_stack_size = None — and a 14x-113x regression in every other configuration. That option is public and user-settable, and it happens to be the only case the PR's benchmark campaign doesn't cover.
1. capture_locations() does not take from the scratch pool (split.rs:1001)
Regex::capture_locations() calls MatchData::new(...), which allocates:
- a match context (
pcre2_match_context_create_8), - a match data block (
pcre2_match_data_create_from_pattern_8), - and, when
max_jit_stack_sizeisSome, a fresh PCRE2 JIT stack (pcre2_jit_stack_create_8), freed again when theCaptureLocationsdrops.
So the change replaces a pool get/put with an mmap/munmap pair plus two mallocs, once per find_matches_pcre2 call.
pcre2_max_jit_stack_size is a public, user-settable option — python/src/lib.rs:389, python/fastokens/__init__.py:36.
Measured on this branch, release build, cl100k-style pattern, 31-byte input:
max_jit_stack_size |
find_at (before) |
captures_read_at (after) |
|
|---|---|---|---|
None |
706 ns | 548 ns | −22% |
1 MiB |
213 ns | 2.96 µs | 14x slower |
64 MiB |
142 ns | 16.0 µs | 113x slower |
The regression isn't confined to tiny inputs: at 6.4 KB with a 1 MiB stack it is still a net loss (28.8 µs → 30.6 µs). And under the 512-concurrency profile in the description, the per-call mmap/munmap serializes on the kernel's mmap lock across every frontend thread, so the cost compounds rather than amortizes.
2. Suggested fix
The pattern this file already uses ~980 lines above is the one that actually delivers the win: a thread_local! cache (see SPLIT_CACHE, line 16) holding a CaptureLocations, reset when the regex identity changes.
That removes the per-call allocation entirely — the JIT stack stays alive across calls, the short-input win applies in every configuration, the bytes.is_empty() guard becomes unnecessary, and the same treatment extends to the find_at loop at line 602.
As written, nothing is reused across calls, despite the PR title and the // Reuse one workspace comment: capture_locations() builds a new CaptureLocations (an Arc::clone(&code) plus 2-3 allocations) on every call, and reuse happens only within a single scan.
3. Allocation failure is an assert!, now on the per-request path
pcre2 0.2.11 asserts on allocation failure inside MatchData::new — failed to allocate match context / failed to allocate match data block / failed to allocate JIT stack. That allocation used to happen roughly once per concurrent-search slot; it now happens per scan. With a large max_jit_stack_size and N parallel chunks in find_matches_pcre2_parallel, a single mmap failure (VA pressure, vm.max_map_count exhaustion) panics from inside a rayon par_iter closure instead of returning Err(Error::Unsupported), unwinding the whole encode.
4. Outside the diff: the repair path still pays the pool cost (split.rs:602)
The find_from closure in find_matches_pcre2_parallel still loops on regex.regex.find_at(bytes, p) — a pool get + PoolGuard::put per iteration, exactly the churn this PR sets out to eliminate. Any input >= 16 KB with a cross-boundary match takes that path, and empty matches advance it one byte at a time, so a repair over a long region pays the full per-byte cost. The optimization currently covers one of the file's two scan loops.
Test coverage
The new differential test is a good shape — 15 patterns x ~1557 inputs is thorough — but it misses the two cases that matter here:
max_jit_stack_size: Noneis hardcoded (line 1092), so the branch wherecapture_locations()behaves differently is never exercised.- The 139 KB input (line 1066) is built large enough to chunk, but the test calls
find_matches_pcre2directly and bypassesfind_segments, so the parallel and repair paths get no coverage.
Remaining inline comments are smaller: builder-config duplication, the preallocation ahead of the empty guard, a dead match arm, and banner placement.
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
Summary
find_matches_pcre2previously acquired and returned pooled PCRE2 scratch for every match. With the default JIT stack, reuse oneCaptureLocationswithin each nonempty scan to reduce that pool overhead.When
max_jit_stack_sizeis configured, keep the original pooledfind_atpath. Creating fresh capture storage in that configuration also creates a custom JIT stack, so allocating it per scan can cause a large regression. The guard retains stock custom-stack allocation behavior without adding a cache.Keep the existing scan loop, offsets, ordering, empty-match handling, errors, and termination. In particular, preserve the end-of-input case where replacing the loop with
find_iterchanges success into an error. Boundary repair, prefix reuse, BPE, Rayon scheduling, and vector preallocation remain unchanged.Default scans still allocate match storage, and allocation failure in the dependency can panic. This change does not introduce a recoverable allocation-error API.
Validation
Grace complete-tokenizer microbenchmark
Stock Fastokens 0.3.2
c1e193ea7754306c394863e09d0def3aea211d4bversus guarded scratch reuse5541d6eff9acc9e1b02611998b5141257ca54fbc, with PCRE2 0.2.11. Both release binaries were built on the same Grace node with identical settings and separate target directories.Results below use the default JIT stack.
We also tested 1 MiB and 64 MiB custom JIT stacks on the same six cases. These retain stock pooled behavior, with no meaningful performance regression observed.
c1899de289a04d12100db370d81485cdf75e47ca.black_box.Acceptance passed: exact parity, retained default-stack short-input CPU savings with eight callers, and no confirmed CPU or elapsed-time regression above 5% against stock. Custom-stack cases retain the stock algorithm; differences in their one-pair measurements are not presented as a separate optimization.
These results measure warmed complete tokenizer calls. They have no statistical confidence intervals and do not predict end-to-end frontend speedup. Benchmark source and raw results remain outside this PR.