Skip to content

pi-lab: a harness whose mechanisms are switches, starting with SoL-Pi's four (#205, #280) - #318

Open
sakurahello1 wants to merge 2 commits into
HarnessRouter:mainfrom
sakurahello1:feat/pi-lab-rebased
Open

sakurahello1 wants to merge 2 commits into
HarnessRouter:mainfrom
sakurahello1:feat/pi-lab-rebased

Conversation

@sakurahello1

@sakurahello1 sakurahello1 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

#231 put half of what I had in mind for #205 into the repository: public task suites, run against harness × model, graded the same way for every configuration. The other half is the harness itself. If the question is what makes a harness cheaper or more accurate, and not only which of the existing ones does better on a given pack, there has to be a harness whose mechanisms can be turned on and off one at a time and run through those same packs. This PR is the first piece of that: Pi Lab, a separate pi-lab backend whose first switches are the four mechanisms from NVIDIA's SoL-Pi.

I wrote up the design and the trade-offs in #280 first. There's no reply there yet, so I'm opening the PR to have the code to talk about; if the direction is wrong, I'll change it.

Pi Lab's normalizer follows the rule from #285 (pi's agent_end with willRetry is not the end of the turn), which is merged now, so this sits on top of it.

Why SoL-Pi first

I don't think SoL-Pi is the best possible harness, and the paper doesn't claim that either. It is a good place to start. Its four mechanisms were kept by an automated search that tested candidate changes against Pi, each sits behind its own boolean, and each goes after a cost the harness controls rather than the model:

  • Action Fusion adds a then_run to write and edit, so writing a file and running it is one call instead of two turns.
  • Online Context Compact gives the agent an update_plan tool and considers compaction at completed steps, weighing it against the cache write/read price.
  • ObservationPack archives large tool outputs and leaves a handle the agent can obs_recall.
  • Evidence-Preserving Reducer hands diagnostic logs to a (possibly cheaper) model for extraction and checks that what comes back is quoted exactly.

The paper's comparison of all four against plain Pi (Tables 1–3, at its API-equivalent prices):

Pi → SoL-Pi, score tokens model cost
EdgeBench, 51 tasks, GPT-5.6 Sol 44.83 → 42.00 2.154B → 1.099B (−49%) $1,339 → $894 (−33%)
EdgeBench, 51 tasks, Opus 5 44.76 → 42.22 2.370B → 1.310B (−45%) $1,741 → $1,158 (−34%)
Terminal-Bench 4, 63 CPU-only tasks 18 → 15 solved not reported $286.45 → $211.12 (−26%)

About a third off the bill for a few points of score, and three fewer Terminal-Bench tasks solved. The mechanisms were developed on GPT-5.6 Sol; Opus 5 was held out.

Why a backend of its own, not -e on pi

On #231 I described this as one -e path on the pi base plus four switches. Building it, I went the other way, for three reasons:

  1. Pi stays untouched as the control arm. A pi harness never loads the extension, never reads its config and never shares its sessions, so a pi column before and after this PR is the same pi.
  2. The switches have to come from the harness definition and nowhere else. With the extension loaded into ordinary pi, a .pi/sol-pi.json in the workspace could override them and the column would measure something other than what it says. Here the extension is wrapped with a fixed config and its state lives under .harness/home/.pi-lab.
  3. This is where the next mechanisms go (subagents and how they talk, handoff, long-term memory, CodeAct versus function calls, the context window), each as its own switch, without piling them onto pi. That is also why it is called Pi Lab and not SoL-Pi.

What's in it

  • Runtime. Pi 0.85.1, NVlabs/SoL-Pi at 1559b5c and the full npm lockfile are pinned, the upstream source is hash-checked, and the result is installed to its own directory on the data volume, read-only to isolated session users. The runner starts that runtime's own pi, never the one on PATH.
  • Runner (runner/pi_lab.py). Its own agent, session and skills directories, credential exclusions from checkpoints, and event normalizer. Native compaction can keep going after agent_end, so the result waits for the process to exit; an agent_end with willRetry drops the failed attempt's error and text, so a turn that pi retried successfully isn't recorded as failed. MCP tools are registered directly, as for pi since Runner: pi registers MCP tools directly, as every other CLI does #277.
  • Gateway. The base, model/provider routing, a pi_lab block on the harness config API (checked once, when the harness is saved), and a per-turn record of which mechanisms were on and what the reducer model used. The pi_lab block is implementation config, not a UHP change.
  • Console. The four switches, the cache write/read price ratio and an optional reducer model. The ratio defaults to 12.5, SoL-Pi's own default rather than any provider's price; DeepSeek, where a cache hit costs 1/50 of a miss, wants 50. Built-in harnesses show these read-only; forks can edit them.
  • Early errors for settings that contradict each other: Action Fusion with bash/edit/write disabled, ObservationPack without obs_recall, compaction without update_plan.
  • docs/pi-lab.md, tests, and a pi-lab-runtime CI job (lockfile check, install, isolated-uid permissions, CLI smoke against a mock provider, no paid calls).

Of the diff's 7.6k lines, 5.9k are the pinned npm lockfile and 0.7k are live verification records; the product code is about three hundred lines across the gateway, runner, console and install script, and the rest is tests, docs and CI.

The branch is rebased onto main at b525279. CheetahClaws and the gpt-6 model rows landed earlier and both sides keep their entries; Claude Sonnet 5.5 landed since, and Pi Lab's console model list now carries it beside Sonnet 5 like every other base (the catalog/placeholder parity test caught the missing entry). The latest rebase had one conflict, the set of backends that take a custom Anthropic-format endpoint: upstream added goose, hermes and openhands, and Pi Lab stays in the set beside them. Pi moved to pi-mcp-adapter 3.x (mcp-adapter.json) while Pi Lab keeps its locked 2.37.0 and mcp.json, so the MCP writer now takes the agent dir and the file name. Pi Lab's console model list is Pi's, and it shows Pi's mark. The new console placeholder test only matched backend ids made of letters, so it read pi-lab as missing; I let it accept hyphens.

Testing

Offline, on main at b525279 (the CI run on this head is green): gateway 731 passed, 17 skipped; runner 610 passed, 1 skipped; conformance 104 and the benchmark scripts' 18 pass; console type-check, Jest (6 tests) and build pass; the pi-lab-runtime job (lockfile, install, permissions, CLI against a mock provider) passes. A fresh install verifies all 23 SoL-Pi source hashes, and the installed CLI against a mock provider passes with all four on, all four off and the reducer on its own: fused write-and-run, produced files, a reducer failure falling back safely with its usage recorded once, resume and a model switch and back, checkpoint restore. The settings panel has no unit tests of its own; it is covered by the build and a browser smoke (save/reload, built-ins read-only).

Live, DeepSeek V4.1 Flash on one custom connection. After the rebase I built the lean image from this branch and started a fresh container (HR_BACKENDS=pi,pi-lab, new data volume). The four applicable support-matrix scenarios pass (first turn, follow-up, file card against the stored file, recycle and recall), and every turn is confirmed on that connection with deepseek-v4.1-flash served. The cross-model switch is n/a because Flash is the only model on it. The custom-harness checks pass as well: bundled skill and script, a produced file, the disabled edit not called, a real DeepWiki MCP call, which is also the check that Pi Lab still finds its MCP server after Pi's adapter move. The records are in docs/verification/pi-lab-2026-09-28/; the 2026-09-25 ones under the old sol-pi name stay, and their README points to the new ones. These live runs were made on 51e420f; what landed since is model lists, plugs, console pages, session titles and changes to the Codex and Hermes paths, and I haven't repeated them on b525279.

The lean image is WITH_DOC_PREVIEW=0, WITH_MEDIA=0, WITH_STARTER_KITS=0, WITH_BUILTIN_SKILLS=0. The first run in it, on 2026-09-25, found one real bug: mktemp left the runtime directory at 0700, so the isolated uid couldn't read it. That's fixed, and reusing an existing install repairs the mode. A responsive pass over 15 widths from 390 to 1440 px found only the 4 px sidebar overflow that the plain pi page has too.

Not covered: any model other than Flash, the full optional-feature image. The reducer model's usage is recorded per turn, but the existing cost widget still shows the main model only, so for Pi Lab it undercounts.

Pi and SoL-Pi are both MIT.

What the first runs show

Outside this PR, I ran a few comparisons with images from this branch, all on DeepSeek V4.1 Flash through the same relay, pi and Pi Lab in the same time window on the same tasks, one run per task:

  • SpreadsheetBench, 50 tasks. Too short (median 48 s) for the mechanisms to fire, so no difference shows; it only says that having them on doesn't break simple tasks.
  • GDPval, 20 office tasks (Excel / Word / PDF deliverables). On the official rubrics, graded twice by GPT (97.8% agreement, κ 0.93): pi 83.1%, Pi Lab 85.9%, and Pi Lab 10 better, 8 worse, 2 tied over the 20. In a blind human ranking of looks among four harnesses (1 is best), mean rank pi 2.05, Pi Lab 1.95. Pi Lab cost about 7% less. All of this is within the noise of a single run.
  • How often the mechanisms actually fired. On GDPval: ObservationPack in 7 of 20 sessions, Action Fusion in 4, a real compaction in 1, the Reducer in none. On office work the four mostly sat idle, so the small gap above can't be credited to them. The Reducer only takes test/build output over 4 KB, and the agent tends to | tail it first, which is probably why it never fired.

SWE-bench Pro, 22 tasks (8 Python, 8 Go, 6 JS), through Harbor. The first attempts leaked the answer. When the task container has the internet, or the image keeps the upstream git history, the agent goes and gets the official fix (git log --all -S, git show <fix> | git apply, pip download of a newer release), and all five harnesses I ran did it. The run below has the image's remotes, refs and reflog removed and the network closed except for the model endpoint, which here is DeepSeek's official API instead of the relay. V4.1 Flash, effort high for all five, one run per task:

resolved (of 22) cost per run
dsh 15 ¥0.45
codex 14 ¥0.32
pi 13 ¥0.36
claude-code 13 ¥0.24
pi-lab 12 ¥0.37

One run per task cannot separate the solved counts. The cost per run does differ, by 1.9× between the cheapest and the dearest harness, and it is mostly output tokens (cache hits were 98–99%). Pi Lab's two losses against pi were both one-hour timeouts: the agent ran grep -r … / and pi's bash tool has no default timeout, so one blocking command used up the task. A default tool timeout looks like the next switch to add. The four mechanisms mostly go after input cost, and on DeepSeek a cache read is 1/50 of a miss, so there is little for them to save here; the same comparison on a model where that ratio is nearer 10 is what I want to see next.

Next: run the mechanisms one at a time on long tasks (ObservationPack first, since it's the only one that fires reliably so far), with several runs per task to get the noise down, add the default tool timeout, and then add new mechanisms to this backend.

Copilot AI lite review requested due to automatic review settings September 28, 2026 12:48
@vercel

vercel Bot commented Sep 28, 2026

Copy link
Copy Markdown

@sakurahello1 is attempting to deploy a commit to the Future HR Team on Vercel.

A member of the Team first needs to authorize it.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@sakurahello1
sakurahello1 force-pushed the feat/pi-lab-rebased branch 2 times, most recently from 8de3909 to 3509083 Compare September 30, 2026 01:10
@vercel

vercel Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
unified-harness-protocol Ready Ready Preview Oct 1, 2026 7:43am UTC

Request Review

…oL-Pi's four (HarnessRouter#205, HarnessRouter#280)

A separate `pi-lab` backend: a pinned Pi 0.85.1 runtime with NVlabs/SoL-Pi at 1559b5c and a full npm
lockfile, hash-checked and installed read-only for the isolated session user. Action Fusion,
Online Context Compact, ObservationPack and the Evidence-Preserving Reducer are switches on the harness
definition only (`pi_lab`, checked once when the harness is saved); ordinary pi never loads them and
stays the control column. Runner: its own agent, session and skills dirs, checkpoints without
credentials, its own event normalization (a retried attempt does not fail the turn), MCP tools as pi
after HarnessRouter#277. Gateway: base, model and provider routing, per-turn record of the switches and the reducer
model's tokens. Console: the four switches, the cache write/read price ratio and the reducer model;
Pi's model list and mark. Contradictory settings fail early. docs/pi-lab.md, tests, and a CI job that
checks the lockfile, installs to /tmp (RUNNER_TEMP sits under /home/runner, which the isolated user
cannot traverse), checks the isolated user's permissions and runs the CLI against a
scripted provider.

Rebased onto main at 7d0fa14 as one commit, after HarnessRouter#285 (which Pi Lab relies on) was merged.
CheetahClaws and the gpt-6 rows arrived meanwhile and both
sides keep their entries; pi moved to pi-mcp-adapter 3.x (mcp-adapter.json) while Pi Lab keeps its
locked 2.37.0 and mcp.json, so the MCP writer takes the agent dir and file name. The console placeholder
test now reads hyphenated backend ids.
… the offline suite was re-run on 7d0fa14)

A fresh lean container from this branch, one Banban integration serving deepseek-v4.1-flash: the four
applicable support-matrix scenarios and the custom-harness suite (skill, script, disabled edit, DeepWiki
MCP) pass; the MCP call covers the move of plain Pi to pi-mcp-adapter 3.x.

This branch was successfully deployed

1 active deployment
Preview — bb5ffef6 Deployed Oct 1, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants