Following #205 and #231. The benchmark dimension measures harness × model on public packs; what it can't do yet is tell you why one harness is cheaper or more accurate than another, because every harness it runs differs from the next in a dozen ways at once. I'd like to add the other half: a harness whose mechanisms can be switched one at a time, so a column can vary exactly one thing and the result says which mechanism was worth its cost.
The first set of switches I'd use is NVIDIA's SoL-Pi (arXiv 2609.20519, NVlabs/SoL-Pi, MIT): Action Fusion, Online Context Compact, ObservationPack and the Evidence-Preserving Reducer, each behind its own boolean on pi 0.85.1. The paper's automated search kept these four; with all of them on, EdgeBench cost dropped by about a third for 2–3 points of score, and Terminal-Bench 4 went from 18 to 15 solved at 26% lower cost. It isn't necessarily the best harness, but it's a set of mechanisms with published numbers to check ours against, and other mechanisms (subagents, handoff, long-term memory, CodeAct versus function calls, the context window) can be added beside them later.
One change from what I wrote on #231: rather than one -e path on the pi base, I've built it as a separate pi-lab backend. Pi stays byte-for-byte the control arm; the switches come only from the harness definition, which a workspace .pi/sol-pi.json could otherwise override; and later mechanisms have somewhere to go that isn't pi. It's less code than the diff suggests: 5.9k of its 7.5k lines are the pinned npm lockfile, and the runner, gateway and console code is under 300 lines; the rest is docs, tests and verification records.
It's built on sakurahello1/harnessrouter@feat/pi-lab, rebased on current main. Under its first name it passed the support matrix on DeepSeek V4.1 Flash locally and in a fresh container; I'll re-run that under pi-lab before opening the PR. Two questions first: is a separate backend the shape you'd want, or would you rather keep it to -e on pi and live with the override risk? And would you rather see the pi/Pi Lab column in the same PR, or as its own afterwards?
Following #205 and #231. The benchmark dimension measures harness × model on public packs; what it can't do yet is tell you why one harness is cheaper or more accurate than another, because every harness it runs differs from the next in a dozen ways at once. I'd like to add the other half: a harness whose mechanisms can be switched one at a time, so a column can vary exactly one thing and the result says which mechanism was worth its cost.
The first set of switches I'd use is NVIDIA's SoL-Pi (arXiv 2609.20519, NVlabs/SoL-Pi, MIT): Action Fusion, Online Context Compact, ObservationPack and the Evidence-Preserving Reducer, each behind its own boolean on pi 0.85.1. The paper's automated search kept these four; with all of them on, EdgeBench cost dropped by about a third for 2–3 points of score, and Terminal-Bench 4 went from 18 to 15 solved at 26% lower cost. It isn't necessarily the best harness, but it's a set of mechanisms with published numbers to check ours against, and other mechanisms (subagents, handoff, long-term memory, CodeAct versus function calls, the context window) can be added beside them later.
One change from what I wrote on #231: rather than one
-epath on the pi base, I've built it as a separatepi-labbackend. Pi stays byte-for-byte the control arm; the switches come only from the harness definition, which a workspace.pi/sol-pi.jsoncould otherwise override; and later mechanisms have somewhere to go that isn't pi. It's less code than the diff suggests: 5.9k of its 7.5k lines are the pinned npm lockfile, and the runner, gateway and console code is under 300 lines; the rest is docs, tests and verification records.It's built on
sakurahello1/harnessrouter@feat/pi-lab, rebased on current main. Under its first name it passed the support matrix on DeepSeek V4.1 Flash locally and in a fresh container; I'll re-run that underpi-labbefore opening the PR. Two questions first: is a separate backend the shape you'd want, or would you rather keep it to-eon pi and live with the override risk? And would you rather see the pi/Pi Lab column in the same PR, or as its own afterwards?