Skip to content

ci: Run Terminal Bench through ug on Linux and Windows - #1074

Open
max-rozen-oss-db wants to merge 16 commits into
mainfrom
max-rozen-oss-db/stack/terminal-bench
Open

max-rozen-oss-db wants to merge 16 commits into
mainfrom
max-rozen-oss-db/stack/terminal-bench

Conversation

@max-rozen-oss-db

@max-rozen-oss-db max-rozen-oss-db commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Adds a Terminal Bench workflow that runs agents through ug claude and ug codex, nightly and on workflow_dispatch. PR CI is unchanged. A temporary push trigger on this branch runs it from the PR while smart routing is verified. Remove that commit before merge.

  • Harbor lane (Linux): runs a TB2 subset plus ug's own tasks in Harbor's Docker containers. bench/terminal_bench/ug_agent.py is a custom Harbor agent that installs the ug wheel, the Databricks CLI (and unzip for its installer), and the agent CLI in each container. It then runs ug configure and a headless launch. It supports Harbor's --resume-trajectory with --continue / exec resume --last. Codex launches with --enable unified_exec to match Harbor's own Codex agent.

  • Native lane (Linux and Windows): Windows runners can't run Linux containers, so run_native.py runs ug's own tasks on the host, including multi-step tasks. A native job passes when at least 75% of its tasks pass (--min-pass-rate), because agents miss a different task or two from run to run.

  • ug's own tasks check that agent features keep working through ug. Each task leaves evidence only that feature can produce, such as a hook log, an MCP call log, or a subagent's model in the transcript:

    • user-hooks (Claude): project UserPromptSubmit and PostToolUse hooks run
    • mcp-server (Claude): a project stdio MCP server loads and is called
    • subagent-model (Claude): a model: haiku subagent runs on the gateway's haiku model
    • project-skill (Claude): a project skill's procedure runs, including its script
    • background-shell: a job longer than the default command timeout runs in the background and its output is captured
    • project-instructions, image-code, pdf-extract, web-fetch, resume-session
    • plus fix-failing-tests, log-triage, scrub-git-secret, background-service

    Tasks declare their agents in [metadata] agents, and bench_tasks.py lists them per agent.

  • Docs: bench/terminal_bench/README.md covers running the workflow, reading results, running locally, and adding a task, with Mermaid diagrams of the two lanes and of one task run. It also documents that each run gets its instruction once, as in Terminal Bench, and lists what the bench doesn't test. bench/terminal_bench/TASKS.md documents the task format and what each task checks.

  • Smart routing: every launch sets ENABLE_SMART_ROUTING_V2=1. On Linux, ug routes Claude's first prompt and its subagents. On Windows, ug routes only Claude's subagents. Codex sends ug's router header. Routed headless -p launches have no other test coverage yet. Both runners save ug's ~/.ucode/*.log routing logs in each task's artifacts. The first routed run found that ug's PTY router swallowed headless Claude prompts on Linux, which failed every Claude task there. fix: Route headless Claude prompts before launch under smart routing #1096, stacked on this PR, fixes it. The temporary trigger also covers that branch.

  • Auth: bench_auth.py mints a fresh M2M token per task on the host from UG_CUJ_SP_CLIENT_ID/SECRET against UG_CUJ1_WORKSPACE. Without those secrets it falls back to DATABRICKS_BEARER and UCODE_TEST_WORKSPACE. The SP secret never reaches the agent. scrub_secrets.py redacts tokens before artifacts upload.

Verification

These results predate smart routing, which no run has exercised yet. Latest run 37986169782, with no harness errors:

Lane TB2 ug tasks
Harbor · Claude 6/7 14/14
Harbor · Codex 6/7 9/10

Passing jobs from the latest run:

The Harbor jobs also passed in 37965359339 and 37967423275. Harbor jobs pass when there are no harness errors. Native jobs fail if any task fails, so every run so far shows as failed overall.

  • Every harness-feature task passes for every agent that runs it, except Codex image-code. That includes hooks, the project MCP server, project instructions, the project skill, the haiku subagent (with the CUJ1 policy fallback), PDF and web input, a background long job, and session resume.
  • Codex misread the same digit in image-code in both of the last two runs, and never ran a command to inspect the file. A Codex run without ug on the same task would show whether that's model vision or image handling in the gateway.
  • subagent-model: the CUJ1 workspace's policy disallows haiku, so Claude Code runs the subagent on the fallback it names. The check accepts that when Claude prints the policy notice.
  • TB2 misses are chess-best-move (Claude) and kv-store-grpc (Codex). Codex's servers die when it exits, from its own process cleanup, and ug hands off with execvp. Earlier runs with the old TB2 subset had no harness errors either.
  • Locally, the reference solutions pass all 14 ug tasks, every verifier fails on the starting state, and Harbor's Task loader accepts them all.

This pull request and its description were written by the Isaac LLM tool (Claude Code, session 35c1ab71-d945-4de6-97b9-b35d5758edb4).

max-rozen-oss-db and others added 8 commits October 9, 2026 10:50
Adds a nightly and manual Terminal Bench workflow. Linux runs a TB2 subset and
ug's own tasks in Harbor containers through a custom ug agent. Linux and
Windows also run ug's tasks natively on the host. Auth mints a per-task token
from the CUJ service principal when its secrets are set, else uses
DATABRICKS_BEARER.

Co-authored-by: Isaac <no-reply@databricks.com>
Co-authored-by: Isaac <no-reply@databricks.com>
Co-authored-by: Isaac <no-reply@databricks.com>
Co-authored-by: Isaac <no-reply@databricks.com>
Seven tasks check that agent features keep working through ug: project
hooks, a project MCP server, CLAUDE.md/AGENTS.md conventions, a haiku
subagent, image input, web fetch, and resuming a session across steps.
Both runners filter tasks per agent and support multi-step tasks.

Co-authored-by: Isaac <no-reply@databricks.com>
The CUJ workspace's managed config disallows haiku, so Claude Code runs the
subagent on the fallback it names. Require that notice and the fallback model
instead of failing.

Co-authored-by: Isaac <no-reply@databricks.com>
project-skill checks a project skill's procedure runs, background-shell needs
a job past the default command timeout, and pdf-extract reads a compressed
PDF. The TB2 default now favors realistic tasks on the same paths (images,
documents, long edits, services) over model-skill tasks.

Co-authored-by: Isaac <no-reply@databricks.com>
Co-authored-by: Isaac <no-reply@databricks.com>
@max-rozen-oss-db
max-rozen-oss-db marked this pull request as ready for review October 9, 2026 21:25
max-rozen-oss-db and others added 8 commits October 9, 2026 18:05
README covers running the workflow, reading results, running locally,
credentials, and adding a task. TASKS.md documents the task format, the
variables verifiers read, and what each task checks.

Co-authored-by: Isaac <no-reply@databricks.com>
Co-authored-by: Isaac <no-reply@databricks.com>
…EADME

Co-authored-by: Isaac <no-reply@databricks.com>
One agent miss out of 14 tasks no longer fails a native job. The README
gets its own section on why Harbor and native lanes both exist.

Co-authored-by: Isaac <no-reply@databricks.com>
Co-authored-by: Isaac <no-reply@databricks.com>
Co-authored-by: Isaac <no-reply@databricks.com>
Re-adds the push trigger to verify smart routing before merge. Both runners
copy ~/.ucode/*.log into task artifacts as evidence that routing ran.

Co-authored-by: Isaac <no-reply@databricks.com>
Co-authored-by: Isaac <no-reply@databricks.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant