Add per-tool output caps and run_r worker options - #211
Merged
Merged
Conversation
config$tool_output_caps raises the universal 50-line / 5000-char tool
result cap for named tools whose results the model must see whole, keyed
by tool name with max_chars and/or max_lines. Other tools keep the cap and
a result past the raised cap still stashes to a handle.
config$run_r_worker_options passes named arguments to
callr::r_session_options() for the supervised run_r worker (env, libpath,
cmdargs, arch), so a strict host can run only the model's R under a
sandbox wrapper at R.home("bin")/<arch>/R while the host stays put.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two small host hooks needed by the corteza MazeBench harness (and useful to any embedded host).
Per-tool output caps.
config$tool_output_capsraises the universal 50-line / 5000-char tool-result cap for named tools whose results the model must see whole, keyed by tool name withmax_charsand/ormax_lines. Other tools keep the cap, a field left out keeps the tool's default budget (the read budget forread_fileand friends), and a result past the raised cap still stashes to a handle. Validated once per handler so a malformed entry fails before any tool runs.Why: a MazeBench board observation is about 80 lines. Under the universal cap the model saw only the top of every board in the first smoke runs (
[tool output truncated] ... showing: first 40 lines). The ARC harness never hit this because its frame lived in an R binding with metadata-only tool text.Worker process options.
config$run_r_worker_optionspasses named arguments tocallr::r_session_options()for the supervisedrun_rworker (env,libpath,cmdargs,arch). Witharch, callr resolves the binary toR.home("bin")/<arch>/R, so a strict host can install a bubblewrap wrapper there and run only the model's R confined while the host process keeps its engine access.envreplaces callr's default (TERM=dumb).Tests: two new files (
test_tool_output_caps_config.R,test_run_r_worker_options.R); full suite 4773 asserts green. Vignetteconfiguration.mddocuments both keys.