feat(benchmarks): round 7 drivers for the new pin, the folded cold baseline, and the orchestrator handoff - #3091
Merged
Merged
Conversation
…; fold the #1947 cold re-run into the round-7 pin Operator decision 2026-09-06: the a99e765 cold re-run (2–4 GPU-days on a 14-case rubric whose top is saturated) is folded into round 7, which needed the same roster re-scored on the 17-case pin anyway. Stopped after 13 models (kept in 1947cold/, marked ABORTED), then prepared and launched the round-7 cold baseline on pin 32dbdeb the same afternoon. New, all with operational copies in /mnt-1/benchmarks/: - round7_prep_pin.sh second clone APIARY-round7, detached at the pin; the a99e765 clone is never touched - round7_build_corpus.sh 17-case corpus rebuilt in the provenance container through ci_verify.sh (manifest byte-identical, 280 semantic executions, 0 failed) -> corpus-round7/ - round7_cache.sh 17-case Tier B cache in its own directory - round7_smoke.sh one Tier A + one Tier B run before any launch (measured: qwen2.5:7b 70/83 A, 69/83 B) - round7_coldrun.sh coldrun.sh for the new pin: PIN guard, 17-entry cache guard, results round7/, roster models_round7.txt; self-matching process guard fixed - round7_launch.sh refuses without a Tier B smoke on this pin (17 cases, max 83, score > 0); starts the sampler - round7_prepare.sh the three prep steps in order, resume-safe sweep_extra.sh gains GHIDRA_CACHE and OPERATOR env overrides; defaults are the a99e765 sweep's, unchanged. The rubric maximum on this pin is 83, not the 79 the resume plan guessed (required_groups + 1 per case) — corrected in the plan, which also gains §0 (what changed after it was written) and the runbook for the running sweep, plus an init.md-style handoff for the orchestrator.
Dependency Review✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.Scanned FilesNone |
9 tasks
…the round-7 plan citations round7_build_corpus.sh hardcoded both LAN resolvers on its `docker run` line; one of them is the home-server address scripts/check-public-leaks.py bans from this repo, so the public-leak gate and its #2285 regression test both failed. The resolver pair is read from the host's own /etc/resolv.conf instead (the benchmark host's list is exactly those two nodes, and install-homeserver.sh asserts the first entry is one of them), loopback stubs filtered out, with a RESOLVERS="ip ip" override and a loud abort when nothing usable is found. Doc-path lint (#2458): - the round-7 hermes-init doc cited PR #3089's branch as if it were a path; the plan is merged, so it now cites the merged path alone; - the plan's grit note used `../x` / `../requirements.txt` as generic examples of relative-path dependencies, not repo citations -- rewritten as `../<crate>` / `../<file>` placeholders, which the lint skips by design; - analysis/ghidra/training and its TOOLCHAIN.md are the training tree #3080 creates, untracked by design until then, so they take allowlist entries.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follows the plan PR #3089 (merged). Operator decision 2026-09-06: the #1947 a99e765 cold re-run is folded into round 7 — stopped after 13 models (kept in
1947cold/, marked ABORTED), and the whole 96-tag roster is being re-measured cold on pin32dbdeb1, running since 14:08Z.Drivers (
analysis/ghidra/benchmarks/corpus/round7_*.sh, operational copies in/mnt-1/benchmarks/)round7_prep_pin.sh— second cloneAPIARY-round7, detached at the pin; the a99e765 clone is never touchedround7_build_corpus.sh— 17-case corpus rebuilt in the provenance container throughci_verify.sh(manifest byte-identical, 280 semantic executions, 0 failed) →corpus-round7/round7_cache.sh— 17-case Tier B cache in its own directory (17 entries, 0 errors, Ghidra 11.3.2)round7_smoke.sh— one Tier A + one Tier B run before any launch; measuredqwen2.5:7b70/83 A, 69/83 Bround7_coldrun.sh— coldrun for the new pin: PIN guard, 17-entry cache guard, resultsround7/, rostermodels_round7.txt; self-matching process guard fixedround7_launch.sh— refuses without a Tier B smoke on this pin (17 cases, max 83, score > 0); starts the VRAM samplerround7_prepare.sh— the three prep steps in order, resume-safesweep_extra.sh—GHIDRA_CACHEandOPERATORenv overrides, defaults unchangedDocs
required_groups + 1per case)docs/benchmarks/plans/2026-09-06-round7-hermes-init.md— init.md-style handoff for the orchestrator: prep mode, GPU gate, kickoff order, grit + ponytail + rtk + gh conventions, dispatch pattern, hard rulesRelated: #3079, #3087, #2245, #3090 (harness bug the old-pin ladder exposed).