Repository navigation
Return what run_r code prints along with its value - #212
Merged
Merged
Conversation
config$tool_output_caps raises the universal 50-line / 5000-char tool
result cap for named tools whose results the model must see whole, keyed
by tool name with max_chars and/or max_lines. Other tools keep the cap and
a result past the raised cap still stashes to a handle.
config$run_r_worker_options passes named arguments to
callr::r_session_options() for the supervised run_r worker (env, libpath,
cmdargs, arch), so a strict host can run only the model's R under a
sandbox wrapper at R.home("bin")/<arch>/R while the host stays put.
tool_run_r() captured only the print of a visible final value, so text a model wrote with cat(), print(), message(), or a warning never reached it, in either run_r_mode. The evaluation now runs inside capture.output() with handlers for messages and warnings, and the result reads like a console: the streamed output, then the value's print. Output written before an error is kept ahead of the Error: line. The supervised worker calls the same function, so both modes agree.
Messages and warnings were appended after all stdout, so message() before cat() printed after it; and warnings were muffled unconditionally, so options(warn=2) no longer turned a warning into an error. Both streams now go to one connection in emission order, and a warning is muffled only when warn<2, leaving warn=2 to error and halt. tool_run_r() also returns an r_error flag so a caller can tell a failed evaluation from a successful one; the model-facing text and isError are unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
# Conflicts: # DESCRIPTION # NEWS.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #211 (
mazebench-hooks); retarget to main once that merges.tool_run_r()captured only the print of a visible final value, so text a model wrote withcat(),print(),message(), or a warning never reached it, in eitherrun_r_mode. Models narrate withcat()constantly, and in the MazeBench smoke runs that text was silently lost.The evaluation now runs inside
capture.output()with handlers for messages and warnings, and the result reads like a console: the streamed output, then the value's print. Output written before an error is kept ahead of theError:line, and the whole thing still goes through the tool-output cap. The supervised worker calls the same function, so both modes agree. Therun_rtool description and CLAUDE.md say so.Tests:
test_run_r_streams.R(15 asserts covering cat, print, messages, warnings, partial output before an error, parse errors, balanced sinks, the cap, and the worker path). Full suite 4788 asserts green. Verified live inside the MazeBench strict-mode container: the model'scat()summary of the board came back in the tool result.