Skip to content

Return what run_r code prints along with its value - #212

Merged
TroyHernandez merged 7 commits into
mainfrom
run-r-streams
Sep 16, 2026
Merged

TroyHernandez merged 7 commits into
mainfrom
run-r-streams

Conversation

@TroyHernandez

Copy link
Copy Markdown
Contributor

Stacked on #211 (mazebench-hooks); retarget to main once that merges.

tool_run_r() captured only the print of a visible final value, so text a model wrote with cat(), print(), message(), or a warning never reached it, in either run_r_mode. Models narrate with cat() constantly, and in the MazeBench smoke runs that text was silently lost.

The evaluation now runs inside capture.output() with handlers for messages and warnings, and the result reads like a console: the streamed output, then the value's print. Output written before an error is kept ahead of the Error: line, and the whole thing still goes through the tool-output cap. The supervised worker calls the same function, so both modes agree. The run_r tool description and CLAUDE.md say so.

Tests: test_run_r_streams.R (15 asserts covering cat, print, messages, warnings, partial output before an error, parse errors, balanced sinks, the cap, and the worker path). Full suite 4788 asserts green. Verified live inside the MazeBench strict-mode container: the model's cat() summary of the board came back in the tool result.

TroyHernandez and others added 6 commits September 15, 2026 17:56
config$tool_output_caps raises the universal 50-line / 5000-char tool
result cap for named tools whose results the model must see whole, keyed
by tool name with max_chars and/or max_lines. Other tools keep the cap and
a result past the raised cap still stashes to a handle.

config$run_r_worker_options passes named arguments to
callr::r_session_options() for the supervised run_r worker (env, libpath,
cmdargs, arch), so a strict host can run only the model's R under a
sandbox wrapper at R.home("bin")/<arch>/R while the host stays put.
tool_run_r() captured only the print of a visible final value, so text a
model wrote with cat(), print(), message(), or a warning never reached
it, in either run_r_mode. The evaluation now runs inside capture.output()
with handlers for messages and warnings, and the result reads like a
console: the streamed output, then the value's print. Output written
before an error is kept ahead of the Error: line. The supervised worker
calls the same function, so both modes agree.
Messages and warnings were appended after all stdout, so message() before
cat() printed after it; and warnings were muffled unconditionally, so
options(warn=2) no longer turned a warning into an error. Both streams now
go to one connection in emission order, and a warning is muffled only when
warn<2, leaving warn=2 to error and halt. tool_run_r() also returns an
r_error flag so a caller can tell a failed evaluation from a successful
one; the model-facing text and isError are unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@TroyHernandez
TroyHernandez changed the base branch from mazebench-hooks to main September 16, 2026 16:08
@TroyHernandez
TroyHernandez merged commit 37a1069 into main Sep 16, 2026
2 checks passed
@TroyHernandez
TroyHernandez deleted the run-r-streams branch September 16, 2026 16:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant