Repository navigation
round7-8: benchmark round 7 driver — 17-case /83 pin, pooled claims, all three slots, cold protocol, chained behind the cold baseline, write-up and promotion records #3087
Description
Activity
- addedenhancementNew feature or requestNew feature or requestmlML worker and GPU scoringML worker and GPU scoringllmLLM analysis workerLLM analysis workerghidraGhidra static-analysis integrationGhidra static-analysis integrationanalysisPayload analysis pipelinePayload analysis pipeline
on Sep 6, 2026 Working conventions for this issue: grit (claim the symbols you touch with
grit claim -a issue-<N>-coder …where N is this issue number, edit only in.grit/worktrees/issue-<N>-coder/,grit doneto merge,grit initafter the merge if you added files), rtk on every shell command (rtk proxyonly for raw output;rtk gain --historyin the closing report), and the gh claim/comment/PR/merge conventions. Full text: epic #3079 → "Working conventions", and the plan's §15.Claimed for round 7 (orchestrator, 2026-09-06). Pre-flight done: plan of record merged via #3089; grit initialised at repo root; no GPU work starts until the cold re-run completes.
- changed the title
[-]round7-8: benchmark round 7 driver — 17-case /79 pin, pooled claims, all three slots, cold protocol, chained behind the cold re-run, write-up and promotion records[/-][+]round7-8: benchmark round 7 driver — 17-case /83 pin, pooled claims, all three slots, cold protocol, chained behind the cold baseline, write-up and promotion records[/+]on Sep 6, 2026 The ghidra-slot cold baseline of this round is running (since 2026-09-06 14:08Z)
Operator folded the #1947 a99e765 cold re-run into this round. Prepared and launched today, everything committed as
analysis/ghidra/benchmarks/corpus/round7_*.sh(PR #3089) with operational copies in/mnt-1/benchmarks/:pin APIARY-round7clone detached at32dbdeb1(17 rubric cases, 850 builds) — guarded byround7_coldrun.shcorpus corpus-round7/rebuilt indebian:trixie-slimviaci_verify.sh, manifest byte-identical, 280 semantic executions 0 failedTier B cache tierb-cache-round7/: 17 extracted, 0 errors, Ghidra 11.3.2, scripts63a935159abbsmoke qwen2.5:7b-instruct-q4_K_MTier A 70/83, Tier B 69/83, 17 cases, 0 empty answersdriver sweep_extra.shgainedGHIDRA_CACHEandOPERATORenv overrides (defaults unchanged);round7_coldrun.sh= coldrun for the new pin;round7_launch.shrefuses without a Tier B smoke on this pinrun 96 tags (models_all + models_extra_all + models_requant), cold, STOP_WORKERS=1, N=2 → 3 → 5, weights kept, operator tagbg-round7, resultsround7/, loground7.log,keep_and_sample.shsampling beside itThe max is 83, not 79 —
required_groups + 1per case; title and body corrected.What this issue still owns
slots_sweep.sh— the sessions + Rev·Deck legs (evaluate-models.py --slots sessions,revdeck) on the same pin, cold, afterround7_coldrun.shcompletes- pooled-claims rescoring (
claims.py) from the round-7 transcripts (untracked inAPIARY-round7/docs/benchmarks/runs/— mirror them, P8) chain_round7.shfor the training legs: positive condition = every tag inmodels_round7.txthas both tier files or anUNMEASUREDmarker inround7/, and noround7_coldrun/sweep_extra/record_baselineprocess alive (the body's1947cold/wording is superseded)- the write-up and promotion records
Runbook:
STATE-2026-09-06-round7-fold.mdon the homeserver and plan §14. Verify the leg, not the pull:tail -20 round7.logmust show bothdone Aanddone Blines.- addedin-progressActively being worked onActively being worked onand removedin-progressActively being worked onActively being worked on
on Sep 7, 2026 Claimed. Coder worktree
.grit/worktrees/issue-3086-coder, CPU-only legs (merge/convert, calibration sets, driver scripts). Scoring/exec gated behind the running cold baseline. Plan: docs/benchmarks/plans/2026-09-06-round7-unsloth-train-requant-ollama.md §7/§8/§9.
Part of #3079 (round 7). Plan §8 and §9. This is the measurement; everything else in the epic produces rows for it. The driver can be written now; it runs only when the cold re-run has finished.
Harness — one new pin for the whole round
mainat the commit the round starts on and record it in every result file and in the write-up. Do not move it mid-round (One benchmark, every model, every aspect: rebuild the model matrix on current hardware with error bars #1947 rule 1). Thea99e765numbers (14 cases / 69) are historical context — their top is saturated (16 models within one point) and cannot show a training gain.record_baseline.py --tier Aand--tier Bon the 17-case / 83 rubric,tierb-cacheregenerated for the 17 cases on the currentghidra-ghidra-1(GHIDRA_VERSIONexported — ghidra: the headless service does not publish its Ghidra version, so ghidra_cache.py's cache key depends on an env var an operator has to remember #2983), injection gate v3 paired verdicts, and pooled-claims scoring (claims.py) againsttier-a-v1.jsonextended by this round's adjudication (state thepool_version).evaluate-models.py --slots sessions,revdeckon the fixtures — this is One benchmark, every model, every aspect: rebuild the model matrix on current hardware with error bars #1947 phase 5, still unrun becauseslots_sweep.shwas lost (ops: #1947 phases 2.5/3/5 are unrunnable — gptoss_rerun.sh, requant_sweep.sh and slots_sweep.sh were never committed and did not survive the rebuild #2985). Writeslots_sweep.shas this round's driver leg, committed underanalysis/ghidra/benchmarks/corpus/, and close that part of ops: #1947 phases 2.5/3/5 are unrunnable — gptoss_rerun.sh, requant_sweep.sh and slots_sweep.sh were never committed and did not survive the rebuild #2985.STOP_WORKERS=1(hp-llm-worker,ghidra-revdeck-1stopped and restored by trap), N=2 → 3 → 5 escalation withUNRESOLVED,UNMEASURED/UNMEASURABLEmarkers,keep_and_sample.shresidency sampling,uptimeper run,KEEP_WEIGHTS_ABOVE_GBfloor, transcripts underdocs/benchmarks/runs/for synthetic runs, outside the repo for anything captured.Rows
qwen3:14b(incumbent),qwen2.5-coder:7b-instruct-q4_K_M(#159 baseline), the top 8 of the cold matrix re-scored on the 17-case pinEvery trained row is compared twice: against its own base at the same quant (isolates training) and against the incumbent (isolates the promotion decision). Every requant row: against the as-published quant and the phase-3 plain level.
Decision metric (from #2245 step 4, unchanged)
A candidate wins a slot only if it is within noise on quality at better decode speed, or better on quality at equal speed, with
schema_ok≥ base, injection gateresistedon every injection case, no polarity regressions, and theevaluate-models.pygates for its slot green. Winners promote only throughmodel-governance.py promotewith a fresh approval record (never a hand edit ofapproved-models.json), one PR per slot, and a production verification on documents indexed after the deploy.Sequencing (
chain_round7.sh)Same shape as
chain_cold.sh: wait on a positive condition — every tag in the cold roster has both tier files or anUNMEASUREDmarker in1947cold/, and nocoldrun.sh/sweep_extra.sh/record_baseline.pyis alive — then run the round-7 roster into/mnt-1/benchmarks/round7/. Never "the GPU looks idle". The training legs (round7-4/5/6/7) share this arbitration: they run in the same window, sequentially, and the chain records which leg holds the card.Deliverables
chain_round7.sh,round7_sweep.sh,slots_sweep.shcommitted with the operational-copy header;models_round7.txtroster committedtierb-cacheregenerated for 17 cases and its cache key recorded/mnt-1/benchmarks/round7/, mirrored to the workstation snapshot dir as the cold run is; transcripts committeddocs/benchmarks/2026-xx-xx-round7-results.md: one table per slot, both comparisons per row, residency + speed columns, the injection axis per the v3 protocol, the decontamination pointer, the pin, the pool version, the Ghidra cache key, uptime rangeDepends on the cold re-run finishing and on rows from round7-2/4/5/6/7. Blocks promotion.