Repository navigation
ci: Run Terminal Bench through ug on Linux and Windows - #1074
Open
max-rozen-oss-db wants to merge 16 commits into
Open
max-rozen-oss-db wants to merge 16 commits into
max-rozen-oss-db wants to merge 16 commits into
Conversation
Adds a nightly and manual Terminal Bench workflow. Linux runs a TB2 subset and ug's own tasks in Harbor containers through a custom ug agent. Linux and Windows also run ug's tasks natively on the host. Auth mints a per-task token from the CUJ service principal when its secrets are set, else uses DATABRICKS_BEARER. Co-authored-by: Isaac <no-reply@databricks.com>
Co-authored-by: Isaac <no-reply@databricks.com>
Co-authored-by: Isaac <no-reply@databricks.com>
Co-authored-by: Isaac <no-reply@databricks.com>
Seven tasks check that agent features keep working through ug: project hooks, a project MCP server, CLAUDE.md/AGENTS.md conventions, a haiku subagent, image input, web fetch, and resuming a session across steps. Both runners filter tasks per agent and support multi-step tasks. Co-authored-by: Isaac <no-reply@databricks.com>
The CUJ workspace's managed config disallows haiku, so Claude Code runs the subagent on the fallback it names. Require that notice and the fallback model instead of failing. Co-authored-by: Isaac <no-reply@databricks.com>
project-skill checks a project skill's procedure runs, background-shell needs a job past the default command timeout, and pdf-extract reads a compressed PDF. The TB2 default now favors realistic tasks on the same paths (images, documents, long edits, services) over model-skill tasks. Co-authored-by: Isaac <no-reply@databricks.com>
Co-authored-by: Isaac <no-reply@databricks.com>
max-rozen-oss-db
marked this pull request as ready for review
October 9, 2026 21:25
max-rozen-oss-db
requested review from
AarushiShah-db,
lilly-luo and
rohita5l
as code owners
October 9, 2026 21:25
README covers running the workflow, reading results, running locally, credentials, and adding a task. TASKS.md documents the task format, the variables verifiers read, and what each task checks. Co-authored-by: Isaac <no-reply@databricks.com>
Co-authored-by: Isaac <no-reply@databricks.com>
…EADME Co-authored-by: Isaac <no-reply@databricks.com>
One agent miss out of 14 tasks no longer fails a native job. The README gets its own section on why Harbor and native lanes both exist. Co-authored-by: Isaac <no-reply@databricks.com>
Co-authored-by: Isaac <no-reply@databricks.com>
Co-authored-by: Isaac <no-reply@databricks.com>
Re-adds the push trigger to verify smart routing before merge. Both runners copy ~/.ucode/*.log into task artifacts as evidence that routing ran. Co-authored-by: Isaac <no-reply@databricks.com>
Co-authored-by: Isaac <no-reply@databricks.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds a
Terminal Benchworkflow that runs agents throughug claudeandug codex, nightly and onworkflow_dispatch. PR CI is unchanged. A temporarypushtrigger on this branch runs it from the PR while smart routing is verified. Remove that commit before merge.Harbor lane (Linux): runs a TB2 subset plus ug's own tasks in Harbor's Docker containers.
bench/terminal_bench/ug_agent.pyis a custom Harbor agent that installs the ug wheel, the Databricks CLI (andunzipfor its installer), and the agent CLI in each container. It then runsug configureand a headless launch. It supports Harbor's--resume-trajectorywith--continue/exec resume --last. Codex launches with--enable unified_execto match Harbor's own Codex agent.Native lane (Linux and Windows): Windows runners can't run Linux containers, so
run_native.pyruns ug's own tasks on the host, including multi-step tasks. A native job passes when at least 75% of its tasks pass (--min-pass-rate), because agents miss a different task or two from run to run.ug's own tasks check that agent features keep working through ug. Each task leaves evidence only that feature can produce, such as a hook log, an MCP call log, or a subagent's model in the transcript:
user-hooks(Claude): projectUserPromptSubmitandPostToolUsehooks runmcp-server(Claude): a project stdio MCP server loads and is calledsubagent-model(Claude): amodel: haikusubagent runs on the gateway's haiku modelproject-skill(Claude): a project skill's procedure runs, including its scriptbackground-shell: a job longer than the default command timeout runs in the background and its output is capturedproject-instructions,image-code,pdf-extract,web-fetch,resume-sessionfix-failing-tests,log-triage,scrub-git-secret,background-serviceTasks declare their agents in
[metadata] agents, andbench_tasks.pylists them per agent.Docs:
bench/terminal_bench/README.mdcovers running the workflow, reading results, running locally, and adding a task, with Mermaid diagrams of the two lanes and of one task run. It also documents that each run gets its instruction once, as in Terminal Bench, and lists what the bench doesn't test.bench/terminal_bench/TASKS.mddocuments the task format and what each task checks.Smart routing: every launch sets
ENABLE_SMART_ROUTING_V2=1. On Linux, ug routes Claude's first prompt and its subagents. On Windows, ug routes only Claude's subagents. Codex sends ug's router header. Routed headless-plaunches have no other test coverage yet. Both runners save ug's~/.ucode/*.logrouting logs in each task's artifacts. The first routed run found that ug's PTY router swallowed headless Claude prompts on Linux, which failed every Claude task there. fix: Route headless Claude prompts before launch under smart routing #1096, stacked on this PR, fixes it. The temporary trigger also covers that branch.Auth:
bench_auth.pymints a fresh M2M token per task on the host fromUG_CUJ_SP_CLIENT_ID/SECRETagainstUG_CUJ1_WORKSPACE. Without those secrets it falls back toDATABRICKS_BEARERandUCODE_TEST_WORKSPACE. The SP secret never reaches the agent.scrub_secrets.pyredacts tokens before artifacts upload.Verification
These results predate smart routing, which no run has exercised yet. Latest run 37986169782, with no harness errors:
Passing jobs from the latest run:
The Harbor jobs also passed in 37965359339 and 37967423275. Harbor jobs pass when there are no harness errors. Native jobs fail if any task fails, so every run so far shows as failed overall.
image-code. That includes hooks, the project MCP server, project instructions, the project skill, the haiku subagent (with the CUJ1 policy fallback), PDF and web input, a background long job, and session resume.image-codein both of the last two runs, and never ran a command to inspect the file. A Codex run without ug on the same task would show whether that's model vision or image handling in the gateway.subagent-model: the CUJ1 workspace's policy disallows haiku, so Claude Code runs the subagent on the fallback it names. The check accepts that when Claude prints the policy notice.chess-best-move(Claude) andkv-store-grpc(Codex). Codex's servers die when it exits, from its own process cleanup, and ug hands off withexecvp. Earlier runs with the old TB2 subset had no harness errors either.Taskloader accepts them all.This pull request and its description were written by the Isaac LLM tool (Claude Code, session
35c1ab71-d945-4de6-97b9-b35d5758edb4).