pi-lab: a harness whose mechanisms are switches, starting with SoL-Pi's four (#205, #280) - #318
Open
sakurahello1 wants to merge 2 commits into
Open
sakurahello1 wants to merge 2 commits into
sakurahello1 wants to merge 2 commits into
Conversation
|
@sakurahello1 is attempting to deploy a commit to the Future HR Team on Vercel. A member of the Team first needs to authorize it. |
sakurahello1
force-pushed
the
feat/pi-lab-rebased
branch
2 times, most recently
from
September 30, 2026 01:10
8de3909 to
3509083
Compare
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
sakurahello1
force-pushed
the
feat/pi-lab-rebased
branch
from
October 1, 2026 01:07
3509083 to
bbbe3e3
Compare
sakurahello1
force-pushed
the
feat/pi-lab-rebased
branch
from
October 1, 2026 01:11
bbbe3e3 to
7a66216
Compare
…oL-Pi's four (HarnessRouter#205, HarnessRouter#280) A separate `pi-lab` backend: a pinned Pi 0.85.1 runtime with NVlabs/SoL-Pi at 1559b5c and a full npm lockfile, hash-checked and installed read-only for the isolated session user. Action Fusion, Online Context Compact, ObservationPack and the Evidence-Preserving Reducer are switches on the harness definition only (`pi_lab`, checked once when the harness is saved); ordinary pi never loads them and stays the control column. Runner: its own agent, session and skills dirs, checkpoints without credentials, its own event normalization (a retried attempt does not fail the turn), MCP tools as pi after HarnessRouter#277. Gateway: base, model and provider routing, per-turn record of the switches and the reducer model's tokens. Console: the four switches, the cache write/read price ratio and the reducer model; Pi's model list and mark. Contradictory settings fail early. docs/pi-lab.md, tests, and a CI job that checks the lockfile, installs to /tmp (RUNNER_TEMP sits under /home/runner, which the isolated user cannot traverse), checks the isolated user's permissions and runs the CLI against a scripted provider. Rebased onto main at 7d0fa14 as one commit, after HarnessRouter#285 (which Pi Lab relies on) was merged. CheetahClaws and the gpt-6 rows arrived meanwhile and both sides keep their entries; pi moved to pi-mcp-adapter 3.x (mcp-adapter.json) while Pi Lab keeps its locked 2.37.0 and mcp.json, so the MCP writer takes the agent dir and file name. The console placeholder test now reads hyphenated backend ids.
… the offline suite was re-run on 7d0fa14) A fresh lean container from this branch, one Banban integration serving deepseek-v4.1-flash: the four applicable support-matrix scenarios and the custom-harness suite (skill, script, disabled edit, DeepWiki MCP) pass; the MCP call covers the move of plain Pi to pi-mcp-adapter 3.x.
sakurahello1
force-pushed
the
feat/pi-lab-rebased
branch
from
October 1, 2026 07:43
7a66216 to
bb5ffef
Compare
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
#231 put half of what I had in mind for #205 into the repository: public task suites, run against harness × model, graded the same way for every configuration. The other half is the harness itself. If the question is what makes a harness cheaper or more accurate, and not only which of the existing ones does better on a given pack, there has to be a harness whose mechanisms can be turned on and off one at a time and run through those same packs. This PR is the first piece of that: Pi Lab, a separate
pi-labbackend whose first switches are the four mechanisms from NVIDIA's SoL-Pi.I wrote up the design and the trade-offs in #280 first. There's no reply there yet, so I'm opening the PR to have the code to talk about; if the direction is wrong, I'll change it.
Pi Lab's normalizer follows the rule from #285 (pi's
agent_endwithwillRetryis not the end of the turn), which is merged now, so this sits on top of it.Why SoL-Pi first
I don't think SoL-Pi is the best possible harness, and the paper doesn't claim that either. It is a good place to start. Its four mechanisms were kept by an automated search that tested candidate changes against Pi, each sits behind its own boolean, and each goes after a cost the harness controls rather than the model:
then_runtowriteandedit, so writing a file and running it is one call instead of two turns.update_plantool and considers compaction at completed steps, weighing it against the cache write/read price.obs_recall.The paper's comparison of all four against plain Pi (Tables 1–3, at its API-equivalent prices):
About a third off the bill for a few points of score, and three fewer Terminal-Bench tasks solved. The mechanisms were developed on GPT-5.6 Sol; Opus 5 was held out.
Why a backend of its own, not
-eon piOn #231 I described this as one
-epath on the pi base plus four switches. Building it, I went the other way, for three reasons:.pi/sol-pi.jsonin the workspace could override them and the column would measure something other than what it says. Here the extension is wrapped with a fixed config and its state lives under.harness/home/.pi-lab.What's in it
1559b5cand the full npm lockfile are pinned, the upstream source is hash-checked, and the result is installed to its own directory on the data volume, read-only to isolated session users. The runner starts that runtime's ownpi, never the one onPATH.runner/pi_lab.py). Its own agent, session and skills directories, credential exclusions from checkpoints, and event normalizer. Native compaction can keep going afteragent_end, so the result waits for the process to exit; anagent_endwithwillRetrydrops the failed attempt's error and text, so a turn that pi retried successfully isn't recorded as failed. MCP tools are registered directly, as for pi since Runner: pi registers MCP tools directly, as every other CLI does #277.pi_labblock on the harness config API (checked once, when the harness is saved), and a per-turn record of which mechanisms were on and what the reducer model used. Thepi_labblock is implementation config, not a UHP change.bash/edit/writedisabled, ObservationPack withoutobs_recall, compaction withoutupdate_plan.docs/pi-lab.md, tests, and api-lab-runtimeCI job (lockfile check, install, isolated-uid permissions, CLI smoke against a mock provider, no paid calls).Of the diff's 7.6k lines, 5.9k are the pinned npm lockfile and 0.7k are live verification records; the product code is about three hundred lines across the gateway, runner, console and install script, and the rest is tests, docs and CI.
The branch is rebased onto
mainatb525279. CheetahClaws and the gpt-6 model rows landed earlier and both sides keep their entries; Claude Sonnet 5.5 landed since, and Pi Lab's console model list now carries it beside Sonnet 5 like every other base (the catalog/placeholder parity test caught the missing entry). The latest rebase had one conflict, the set of backends that take a custom Anthropic-format endpoint: upstream added goose, hermes and openhands, and Pi Lab stays in the set beside them. Pi moved to pi-mcp-adapter 3.x (mcp-adapter.json) while Pi Lab keeps its locked 2.37.0 andmcp.json, so the MCP writer now takes the agent dir and the file name. Pi Lab's console model list is Pi's, and it shows Pi's mark. The new console placeholder test only matched backend ids made of letters, so it readpi-labas missing; I let it accept hyphens.Testing
Offline, on
mainatb525279(the CI run on this head is green): gateway 731 passed, 17 skipped; runner 610 passed, 1 skipped; conformance 104 and the benchmark scripts' 18 pass; console type-check, Jest (6 tests) and build pass; thepi-lab-runtimejob (lockfile, install, permissions, CLI against a mock provider) passes. A fresh install verifies all 23 SoL-Pi source hashes, and the installed CLI against a mock provider passes with all four on, all four off and the reducer on its own: fused write-and-run, produced files, a reducer failure falling back safely with its usage recorded once, resume and a model switch and back, checkpoint restore. The settings panel has no unit tests of its own; it is covered by the build and a browser smoke (save/reload, built-ins read-only).Live, DeepSeek V4.1 Flash on one custom connection. After the rebase I built the lean image from this branch and started a fresh container (
HR_BACKENDS=pi,pi-lab, new data volume). The four applicable support-matrix scenarios pass (first turn, follow-up, file card against the stored file, recycle and recall), and every turn is confirmed on that connection withdeepseek-v4.1-flashserved. The cross-model switch is n/a because Flash is the only model on it. The custom-harness checks pass as well: bundled skill and script, a produced file, the disablededitnot called, a real DeepWiki MCP call, which is also the check that Pi Lab still finds its MCP server after Pi's adapter move. The records are indocs/verification/pi-lab-2026-09-28/; the 2026-09-25 ones under the oldsol-piname stay, and their README points to the new ones. These live runs were made on51e420f; what landed since is model lists, plugs, console pages, session titles and changes to the Codex and Hermes paths, and I haven't repeated them onb525279.The lean image is
WITH_DOC_PREVIEW=0,WITH_MEDIA=0,WITH_STARTER_KITS=0,WITH_BUILTIN_SKILLS=0. The first run in it, on 2026-09-25, found one real bug:mktempleft the runtime directory at0700, so the isolated uid couldn't read it. That's fixed, and reusing an existing install repairs the mode. A responsive pass over 15 widths from 390 to 1440 px found only the 4 px sidebar overflow that the plain pi page has too.Not covered: any model other than Flash, the full optional-feature image. The reducer model's usage is recorded per turn, but the existing cost widget still shows the main model only, so for Pi Lab it undercounts.
Pi and SoL-Pi are both MIT.
What the first runs show
Outside this PR, I ran a few comparisons with images from this branch, all on DeepSeek V4.1 Flash through the same relay, pi and Pi Lab in the same time window on the same tasks, one run per task:
| tailit first, which is probably why it never fired.SWE-bench Pro, 22 tasks (8 Python, 8 Go, 6 JS), through Harbor. The first attempts leaked the answer. When the task container has the internet, or the image keeps the upstream git history, the agent goes and gets the official fix (
git log --all -S,git show <fix> | git apply,pip downloadof a newer release), and all five harnesses I ran did it. The run below has the image's remotes, refs and reflog removed and the network closed except for the model endpoint, which here is DeepSeek's official API instead of the relay. V4.1 Flash, effort high for all five, one run per task:One run per task cannot separate the solved counts. The cost per run does differ, by 1.9× between the cheapest and the dearest harness, and it is mostly output tokens (cache hits were 98–99%). Pi Lab's two losses against pi were both one-hour timeouts: the agent ran
grep -r … /and pi's bash tool has no default timeout, so one blocking command used up the task. A default tool timeout looks like the next switch to add. The four mechanisms mostly go after input cost, and on DeepSeek a cache read is 1/50 of a miss, so there is little for them to save here; the same comparison on a model where that ratio is nearer 10 is what I want to see next.Next: run the mechanisms one at a time on long tasks (ObservationPack first, since it's the only one that fires reliably so far), with several runs per task to get the noise down, add the default tool timeout, and then add new mechanisms to this backend.