Skip to content

round7-8: benchmark round 7 driver — 17-case /83 pin, pooled claims, all three slots, cold protocol, chained behind the cold baseline, write-up and promotion records #3087

Description

@Xore

Part of #3079 (round 7). Plan §8 and §9. This is the measurement; everything else in the epic produces rows for it. The driver can be written now; it runs only when the cold re-run has finished.

Harness — one new pin for the whole round

Rows

group rows control
controls qwen3:14b (incumbent), qwen2.5-coder:7b-instruct-q4_K_M (#159 baseline), the top 8 of the cold matrix re-scored on the 17-case pin —
T0 REx86 merged (round7-2) at Q8_0 / Q4_K_M its base at the same quant
T1 CPT-only checkpoints (round7-4) untouched base, same quant
T2 / T2′ / T3 per-slot, CPT-then-SFT, multi-slot adapters (round7-5) at Q8_0 / Q6_K / Q4_K_M / IQ4_XS untouched base, same quant
T4 / T5 DPO / GRPO checkpoints (round7-6) their SFT parent
R2 every dynamic-requant level (round7-7) the plain K-quant at matched size from phase 3

Every trained row is compared twice: against its own base at the same quant (isolates training) and against the incumbent (isolates the promotion decision). Every requant row: against the as-published quant and the phase-3 plain level.

Decision metric (from #2245 step 4, unchanged)

A candidate wins a slot only if it is within noise on quality at better decode speed, or better on quality at equal speed, with schema_ok ≥ base, injection gate resisted on every injection case, no polarity regressions, and the evaluate-models.py gates for its slot green. Winners promote only through model-governance.py promote with a fresh approval record (never a hand edit of approved-models.json), one PR per slot, and a production verification on documents indexed after the deploy.

Sequencing (chain_round7.sh)

Same shape as chain_cold.sh: wait on a positive condition — every tag in the cold roster has both tier files or an UNMEASURED marker in 1947cold/, and no coldrun.sh / sweep_extra.sh / record_baseline.py is alive — then run the round-7 roster into /mnt-1/benchmarks/round7/. Never "the GPU looks idle". The training legs (round7-4/5/6/7) share this arbitration: they run in the same window, sequentially, and the chain records which leg holds the card.

Deliverables

Depends on the cold re-run finishing and on rows from round7-2/4/5/6/7. Blocks promotion.

Activity

  1. added
    enhancementNew feature or request
    mlML worker and GPU scoring
    llmLLM analysis worker
    ghidraGhidra static-analysis integration
    analysisPayload analysis pipeline
    on Sep 6, 2026
  2. Xore commented on Sep 6, 2026

    @Xore
    OwnerAuthor

    Working conventions for this issue: grit (claim the symbols you touch with grit claim -a issue-<N>-coder … where N is this issue number, edit only in .grit/worktrees/issue-<N>-coder/, grit done to merge, grit init after the merge if you added files), rtk on every shell command (rtk proxy only for raw output; rtk gain --history in the closing report), and the gh claim/comment/PR/merge conventions. Full text: epic #3079 → "Working conventions", and the plan's §15.

  3. self-assigned this
    on Sep 6, 2026
  4. Xore commented on Sep 6, 2026

    @Xore
    OwnerAuthor

    Claimed for round 7 (orchestrator, 2026-09-06). Pre-flight done: plan of record merged via #3089; grit initialised at repo root; no GPU work starts until the cold re-run completes.

  5. changed the title [-]round7-8: benchmark round 7 driver — 17-case /79 pin, pooled claims, all three slots, cold protocol, chained behind the cold re-run, write-up and promotion records[/-] [+]round7-8: benchmark round 7 driver — 17-case /83 pin, pooled claims, all three slots, cold protocol, chained behind the cold baseline, write-up and promotion records[/+] on Sep 6, 2026
  6. Xore commented on Sep 6, 2026

    @Xore
    OwnerAuthor

    The ghidra-slot cold baseline of this round is running (since 2026-09-06 14:08Z)

    Operator folded the #1947 a99e765 cold re-run into this round. Prepared and launched today, everything committed as analysis/ghidra/benchmarks/corpus/round7_*.sh (PR #3089) with operational copies in /mnt-1/benchmarks/:

    pin APIARY-round7 clone detached at 32dbdeb1 (17 rubric cases, 850 builds) — guarded by round7_coldrun.sh
    corpus corpus-round7/ rebuilt in debian:trixie-slim via ci_verify.sh, manifest byte-identical, 280 semantic executions 0 failed
    Tier B cache tierb-cache-round7/: 17 extracted, 0 errors, Ghidra 11.3.2, scripts 63a935159abb
    smoke qwen2.5:7b-instruct-q4_K_M Tier A 70/83, Tier B 69/83, 17 cases, 0 empty answers
    driver sweep_extra.sh gained GHIDRA_CACHE and OPERATOR env overrides (defaults unchanged); round7_coldrun.sh = coldrun for the new pin; round7_launch.sh refuses without a Tier B smoke on this pin
    run 96 tags (models_all + models_extra_all + models_requant), cold, STOP_WORKERS=1, N=2 → 3 → 5, weights kept, operator tag bg-round7, results round7/, log round7.log, keep_and_sample.sh sampling beside it

    The max is 83, not 79 — required_groups + 1 per case; title and body corrected.

    What this issue still owns

    • slots_sweep.sh — the sessions + Rev·Deck legs (evaluate-models.py --slots sessions,revdeck) on the same pin, cold, after round7_coldrun.sh completes
    • pooled-claims rescoring (claims.py) from the round-7 transcripts (untracked in APIARY-round7/docs/benchmarks/runs/ — mirror them, P8)
    • chain_round7.sh for the training legs: positive condition = every tag in models_round7.txt has both tier files or an UNMEASURED marker in round7/, and no round7_coldrun / sweep_extra / record_baseline process alive (the body's 1947cold/ wording is superseded)
    • the write-up and promotion records

    Runbook: STATE-2026-09-06-round7-fold.md on the homeserver and plan §14. Verify the leg, not the pull: tail -20 round7.log must show both done A and done B lines.

  7. added
    in-progressActively being worked on
    and removed
    in-progressActively being worked on
    on Sep 7, 2026
  8. Xore commented on Sep 7, 2026

    @Xore
    OwnerAuthor

    Claimed. Coder worktree .grit/worktrees/issue-3086-coder, CPU-only legs (merge/convert, calibration sets, driver scripts). Scoring/exec gated behind the running cold baseline. Plan: docs/benchmarks/plans/2026-09-06-round7-unsloth-train-requant-ollama.md §7/§8/§9.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

analysisPayload analysis pipelineenhancementNew feature or requestghidraGhidra static-analysis integrationllmLLM analysis workermlML worker and GPU scoring

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions