Skip to content

eager: fail a request that would pass RUNNER_RSS_LIMIT_MB - #67

Merged
olydis merged 2 commits into
mainfrom
claude/tc-performance-optimizations-rqtyzl-runner-memory-limit
Oct 7, 2026
Merged

olydis merged 2 commits into
mainfrom
claude/tc-performance-optimizations-rqtyzl-runner-memory-limit

Conversation

@olydis

@olydis olydis commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Bug

Nothing makes a runner stay within the memory a build grants it. forest #620 and arboretum #78 give each core one thread with 3 GB, but no limit enforces it. forest#535's src/tree_type/derivation.lamb:72 (render △) recurses without end under eager evaluation, and the continuation stack (std::vector<Frame>) grows until the machine runs out.

  • The test on its own. The test is extracted from the bundle and run in one runner (rss-after.py runner-eager.exe deriv44.lib.dag deriv44.test.dag). VmHWM passes 4,000 MB in 5.6 s, when the watch kills it. Earlier, with no kill, it reached std::bad_alloc at 8.2 GB.
  • The forest build. forest main a62d3225 plus #535's src/tree_type, ./build.sh, 4 threads: the session reaches 7,042 MB at 106 s and is killed. No test is named.

Fix

RUNNER_RSS_LIMIT_MB=<MiB> works for the eager runner only. Unset or 0 means no limit, so the CLI, www and lazy callers are unchanged.

A reduction grows two things. Each gets a ceiling, fixed once in set_limit(). Because the ceilings never change, whether a request fits does not depend on what the runner ran before it.

  • Budget: the largest power of two whose footprint fits in ¾ of the limit. The footprint is the arena plus the hash-cons table and memo that budget sizes.
    • It reuses the existing 2^31 ceiling branch in collect_if_over_budget; that constant is now the member _max_budget.
    • set_budget also clamps to it. So a limit lowers RUNNER_RSS_THRESHOLD_MB, and makes 0 collect.
  • Stack: the frames that fit in the rest of the limit, leaving room for the copy a growing vector makes. The check sits in the stack's allocator, so a step that does not grow the stack pays nothing.
  • Error: out of memory: the stack would pass RUNNER_RSS_LIMIT_MB=2048 (or the live set …). It goes through apply()'s existing catch, and the runner then serves the next request.
  • Releasing the stack: apply()'s catch at base 0 now calls shrink_to_fit on the stack. After the #535 failure the runner holds 4.6 MB VmRSS, against 529 MB without the release.
  • Effect at 2048 with the default 512 MB threshold: the budget is unchanged (64M nodes, a 1,120 MiB footprint), and stacks stop at 2^25 frames. The deepest stack forest or arboretum reaches is 2^18.

Validation

All runs compare base (5679507/b6b016b) against this branch at RUNNER_RSS_LIMIT_MB=2048. Everything ran under the bench lock and an RSS watch.

check base this branch
render △ alone >4,000 MB in 5.6 s (killed) err out of memory: the stack would pass …=2048 in 1.06 s, VmHWM 529 MB; the next request answers data 11, and a repeat fails the same way
forest + #535, ./build.sh 7,042 MB at 106 s, killed, no test named lambada: src/tree_type/derivation.lamb:72: runner: out of memory: the stack would pass RUNNER_RSS_LIMIT_MB=2048, rc=1 at 92 s, peak 5,664 MB
forest ./build.sh, empty cache rc 0, 101 s, tracked files clean rc 0, 93 s, clean apart from the submodule pin
arboretum ./build.sh, empty cache rc 0, src clean rc 0, src clean; its lambada export is byte-identical to base's
test phase, 1 thread: steps (forest / arboretum) 2,975,223,708 / 2,044,726,073 identical; all per-command counters identical (2,101 / 1,262 commands)
test phase, 1 thread: wall (forest / arboretum) 199.8 / 131.3 s 197.2 / 125.2 s
test phase, 4 threads: wall (forest / arboretum) 77.3, 72.8 / 45.0, 44.5 s 71.3, 71.3 / 46.8, 44.0 s
runner VmHWM, test phases 1,153–1,166 MB the same
cachegrind Ir (Nat.Bench.68, Audio.Wav.28, Certify.Size.Test.48) 1 0.994, 0.994, 0.995
CPU time, 4 heaviest arboretum tests, 3 rounds 1 1.010 (0.993–1.023 per test); steps equal
  • Test suite: implementation/cpp/dag-machine/test.sh passes under LC_ALL=C and LC_ALL=C.UTF-8.
  • Compilers: builds with g++ and clang++ (libstdc++, -std=c++17/20, -D_GLIBCXX_DEBUG) and against libc++ 18, macOS's standard library (-std=c++17/20/23, hardened debug mode), where it passes the new check too.
  • New check in test-runner.sh: at a 1 MB limit, a 20,000-deep recursion (a list of stems applied to △) fails as "the stack", an 80,000-cell bound string fails as "the live set", and the next request still answers. Both requests end, so a runner that ignores the limit fails the check in 0.3 s at 3 MB (base answers data 3 △ and the string) rather than running away.
  • Re-run in review (b6b016b vs this branch, own builds): render △ failed at 2048 in 0.91 s, VmHWM 529 MB, VmRSS 4.6 MB after, next request answered; forest + #535 test phase on 4 threads named derivation.lamb:72 at 74 s; arboretum test phase on 1 thread had identical steps and per-command counters; cachegrind Ir 0.994 / 0.995 (Nat.Bench.68, Certify.Size.Test.48).

Caveats

  • What the limit counts: what the budget and stack can reach, not RSS. Not counted: the collector's mark stack, the module maps, rendered output, and a bound string's run, which can sit above the budget until the next collection.
  • Tail loops count as "the stack": a memoized tail loop pushes MEMOIZE frames, so a loop of about 1.3G steps also fails at 2048. No forest or arboretum test comes close, and runner.md says so.
  • Lazy runner: it ignores the variable.

Merge order

🤖 Generated with Claude Code

https://claude.ai/code/session_018ffv5AybPTpV56D5vuayQr

Nothing held a runner to the memory a build grants it (forest #620,
arboretum #78: a thread per core, 3 GB each). forest#535's
derivation.lamb:72, `render △`, recurses without end under eager
evaluation: its stack passed 4 GB in 5.6 s, and a forest build with it
reached 7 GB in 106 s (killed) without naming a test.

RUNNER_RSS_LIMIT_MB (eager only; unset or 0, no limit) gives the two
things a reduction grows a ceiling each, fixed once in set_limit so
that whether a request fits does not depend on what ran before it: the
budget, at the largest power of two whose arena, table and memo fit in
3/4 of the limit (the 2^31 ceiling of collect_if_over_budget, now a
member, which set_budget clamps to as well); the stack, at what is
left, checked in its allocator so that a step pays nothing. A failed
request gives its stack back.

At 2048, `render △` answers "out of memory: the stack would pass
RUNNER_RSS_LIMIT_MB=2048" in 1.06 s at VmHWM 529 MB, with VmRSS 4.6 MB
after (529 MB without the release), and the runner answers the next
request; the forest build fails naming derivation.lamb:72. Forest and
arboretum test phases at 2048: per-command counters identical at one
thread (2,975,223,708 and 2,044,726,073 steps), no source changed, wall
time within noise; cachegrind Ir 0.994-0.995x.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ffv5AybPTpV56D5vuayQr
The check's stack case was W W, a recursion that never returns, so on a
runner that ignores the limit it fails by running away: base b6b016b
passed 2 GB in 4.6 s before the watch killed it, and unwatched it runs
to std::bad_alloc at about 8 GB. Applied to △, a list of 20,000 stems
recurses as deep, a frame a cell, past the 2^14 frames the stack gets at
1 MB, and ends: base answers both requests and the check fails in 0.3 s
at 3 MB, while this branch passes it as before (both locales, and built
against libc++).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ffv5AybPTpV56D5vuayQr
@olydis
olydis merged commit 2632b6d into main Oct 7, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants