Repository navigation
eager: fail a request that would pass RUNNER_RSS_LIMIT_MB - #67
Merged
olydis merged 2 commits intoOct 7, 2026
Merged
Conversation
Nothing held a runner to the memory a build grants it (forest #620, arboretum #78: a thread per core, 3 GB each). forest#535's derivation.lamb:72, `render △`, recurses without end under eager evaluation: its stack passed 4 GB in 5.6 s, and a forest build with it reached 7 GB in 106 s (killed) without naming a test. RUNNER_RSS_LIMIT_MB (eager only; unset or 0, no limit) gives the two things a reduction grows a ceiling each, fixed once in set_limit so that whether a request fits does not depend on what ran before it: the budget, at the largest power of two whose arena, table and memo fit in 3/4 of the limit (the 2^31 ceiling of collect_if_over_budget, now a member, which set_budget clamps to as well); the stack, at what is left, checked in its allocator so that a step pays nothing. A failed request gives its stack back. At 2048, `render △` answers "out of memory: the stack would pass RUNNER_RSS_LIMIT_MB=2048" in 1.06 s at VmHWM 529 MB, with VmRSS 4.6 MB after (529 MB without the release), and the runner answers the next request; the forest build fails naming derivation.lamb:72. Forest and arboretum test phases at 2048: per-command counters identical at one thread (2,975,223,708 and 2,044,726,073 steps), no source changed, wall time within noise; cachegrind Ir 0.994-0.995x. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018ffv5AybPTpV56D5vuayQr
The check's stack case was W W, a recursion that never returns, so on a runner that ignores the limit it fails by running away: base b6b016b passed 2 GB in 4.6 s before the watch killed it, and unwatched it runs to std::bad_alloc at about 8 GB. Applied to △, a list of 20,000 stems recurses as deep, a frame a cell, past the 2^14 frames the stack gets at 1 MB, and ends: base answers both requests and the check fails in 0.3 s at 3 MB, while this branch passes it as before (both locales, and built against libc++). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018ffv5AybPTpV56D5vuayQr
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Bug
Nothing makes a runner stay within the memory a build grants it. forest #620 and arboretum #78 give each core one thread with 3 GB, but no limit enforces it. forest#535's
src/tree_type/derivation.lamb:72(render △) recurses without end under eager evaluation, and the continuation stack (std::vector<Frame>) grows until the machine runs out.rss-after.py runner-eager.exe deriv44.lib.dag deriv44.test.dag). VmHWM passes 4,000 MB in 5.6 s, when the watch kills it. Earlier, with no kill, it reachedstd::bad_allocat 8.2 GB.src/tree_type,./build.sh, 4 threads: the session reaches 7,042 MB at 106 s and is killed. No test is named.Fix
RUNNER_RSS_LIMIT_MB=<MiB>works for the eager runner only. Unset or0means no limit, so the CLI, www and lazy callers are unchanged.A reduction grows two things. Each gets a ceiling, fixed once in
set_limit(). Because the ceilings never change, whether a request fits does not depend on what the runner ran before it.collect_if_over_budget; that constant is now the member_max_budget.set_budgetalso clamps to it. So a limit lowersRUNNER_RSS_THRESHOLD_MB, and makes0collect.out of memory: the stack would pass RUNNER_RSS_LIMIT_MB=2048(orthe live set …). It goes throughapply()'s existing catch, and the runner then serves the next request.apply()'s catch at base 0 now callsshrink_to_fiton the stack. After the #535 failure the runner holds 4.6 MB VmRSS, against 529 MB without the release.Validation
All runs compare base (5679507/b6b016b) against this branch at
RUNNER_RSS_LIMIT_MB=2048. Everything ran under the bench lock and an RSS watch.render △aloneerr out of memory: the stack would pass …=2048in 1.06 s, VmHWM 529 MB; the next request answersdata 11, and a repeat fails the same way./build.shlambada: src/tree_type/derivation.lamb:72: runner: out of memory: the stack would pass RUNNER_RSS_LIMIT_MB=2048, rc=1 at 92 s, peak 5,664 MB./build.sh, empty cache./build.sh, empty cacheimplementation/cpp/dag-machine/test.shpasses underLC_ALL=CandLC_ALL=C.UTF-8.-std=c++17/20,-D_GLIBCXX_DEBUG) and against libc++ 18, macOS's standard library (-std=c++17/20/23, hardened debug mode), where it passes the new check too.test-runner.sh: at a 1 MB limit, a 20,000-deep recursion (a list of stems applied to △) fails as "the stack", an 80,000-cell bound string fails as "the live set", and the next request still answers. Both requests end, so a runner that ignores the limit fails the check in 0.3 s at 3 MB (base answersdata 3 △and the string) rather than running away.render △failed at 2048 in 0.91 s, VmHWM 529 MB, VmRSS 4.6 MB after, next request answered; forest + #535 test phase on 4 threads namedderivation.lamb:72at 74 s; arboretum test phase on 1 thread had identical steps and per-command counters; cachegrind Ir 0.994 / 0.995 (Nat.Bench.68, Certify.Size.Test.48).Caveats
Merge order
RUNNER_RSS_LIMIT_MB=2048next toTREE_CALCULUS_RUNNER=eager.🤖 Generated with Claude Code
https://claude.ai/code/session_018ffv5AybPTpV56D5vuayQr