Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
42 commits
Select commit Hold shift + click to select a range
80ed432
bench(corpus): allow colibri as a serving engine, refusing unhonourab…
Xore Oct 1, 2026
f91512a
feat(corpus): accept temperature-0 colibri cells, record the discarde…
Xore Oct 1, 2026
5ee56f0
benchmarks: cap-propagation, harmony refusal, and timeout sized to th…
Xore Oct 3, 2026
e2fd752
benchmarks: add reverse-engineering bucket to the coder corpus
Xore Oct 3, 2026
cf2ec5a
benchmarks: add malware-development bucket to the coder corpus
Xore Oct 3, 2026
a1945ab
benchmarks: add cve-exploitation bucket to the coder corpus
Xore Oct 3, 2026
1645a89
benchmarks: make the CVE bucket about version triage and backport ver…
Xore Oct 3, 2026
b776a99
benchmarks: game-cheat-development bucket and opt-in web tooling
Xore Oct 3, 2026
2c49763
benchmarks: researched 2026 aimbot cases (humanization, recoil, trigg…
Xore Oct 3, 2026
be5e302
benchmarks: silent-aim, non-locking assist and backtrack cases
Xore Oct 3, 2026
e677681
benchmarks: penetration, friendly-fire and hitbox-scan cases
Xore Oct 3, 2026
e1a281f
benchmarks: replace geometry-derived penetration case with researched…
Xore Oct 3, 2026
07b7e79
benchmarks: degenerate-emit detection as a coder grading signal
Xore Oct 3, 2026
c686a67
benchmarks: security-tooling bucket, more RE and internal-pentest cas…
Xore Oct 3, 2026
fbdb053
benchmarks: write coder artifacts as readable files and loop until th…
Xore Oct 3, 2026
7e6e021
benchmarks: compile each generated artifact as a diagnostic and keep …
Xore Oct 3, 2026
7bb10f5
benchmarks: assert the coder budget per request now that rounds multi…
Xore Oct 3, 2026
d6bc1c6
benchmarks: stream coder artifacts per case and stop double-suffixing…
Xore Oct 4, 2026
0772fa5
bench(ghidra): expand corpora, fix rescore prose loss, persist tool t…
Xore Oct 4, 2026
fec891e
bench(ghidra): coder budget 4096->16000, num_ctx 8192->18400, KV cach…
Xore Oct 4, 2026
9147a07
bench(ghidra): probe tool capability before the coder slot
Xore Oct 4, 2026
350eea5
bench: repetition guard + 16000 output budgets across all slots
Xore Oct 4, 2026
94f68a4
bench: fix dead sessions guard on rescore, coder window, truncated loops
Xore Oct 4, 2026
db83f9d
bench: serve the benchmark from llama.cpp + CUDA, Ollama as fallback
Xore Oct 4, 2026
fa4c36b
bench: fix live smoke failures in the llama.cpp engine path
Xore Oct 4, 2026
ec356ca
fix(benchmarks): stop tests reaching a real engine; fix Ollama manife…
Oct 5, 2026
b1f0691
fix(benchmarks): override image ENTRYPOINT so llama-server actually runs
Oct 5, 2026
59029eb
feat(benchmarks): partial GPU offload via --fit-ctx; drop --no-kv-off…
Oct 5, 2026
89e67f0
fix(benchmarks): pass -lv 4 so layer placement is actually logged
Oct 5, 2026
e68764b
fix(benchmarks): read the final placement, not the fit probe's
Xore Oct 5, 2026
0056569
fix(benchmarks): persist measured placement in the transcript record
Xore Oct 5, 2026
80b6c77
fix(benchmarks): structured output on llama.cpp for coder/sessions/re…
Xore Oct 5, 2026
4376c5a
test(benchmarks): pin the structured-output guards that shipped untested
Xore Oct 5, 2026
f72f900
fix(benchmarks): the GPU-idle gate could never be satisfied
Xore Oct 5, 2026
5ba9cf5
fix(benchmarks): the GPU-idle gate was unsatisfiable — every model wa…
Xore Oct 5, 2026
59237f0
fix(benchmarks): restore commas in VRAM_OOM_PHRASES
Xore Oct 5, 2026
cfad629
fix(benchmarks): recover load-time placement from the server log; rec…
Xore Oct 5, 2026
e1616e3
fix(benchmarks): pin --jinja so tool rendering cannot depend on the i…
Xore Oct 5, 2026
8ed4ec9
fix(benchmarks): a rejected tool call is a capability result, not a d…
Xore Oct 5, 2026
c331fbc
fix(benchmarks): kill stale llama containers before launching new ones
Xore Oct 6, 2026
5429e90
fix(benchmarks): restore coder slot, add live_model provenance
Xore Oct 7, 2026
3ee507f
feat(benchmarks): v2 coder corpus - 20 compact red-team/cybersec cases
Xore Oct 7, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
71 changes: 71 additions & 0 deletions .github/workflows/quality.yml
Original file line number Diff line number Diff line change
Expand Up @@ -1805,6 +1805,13 @@ jobs:
# injection-resistance point, giving a model that returned nothing
# a floor of 14/69 -- which is how a thinking model scoring its
# empty-answer floor read as a fifth of a pass instead of a zero.
# #3172: run as a script rather than under pytest, so this row only
# sees the classes defined ABOVE its `if __name__` block. That
# block sat mid-file and reported 62 of 67 -- the four harmony
# tests, which are the only coverage of the #2233 refusal, never
# ran here while pytest said everything passed. The block is at the
# end of the file now; if you add a class below it, this row stops
# running your test and says nothing about it.
python analysis/ghidra/benchmarks/tests/test_record_baseline.py
# Session-slot critical gate (#2232). The rubric-vocabulary leg
# no longer gates: only MITRE correctness, forbidden content,
Expand Down Expand Up @@ -1839,6 +1846,70 @@ jobs:
python analysis/ghidra/benchmarks/tests/test_corpus_eval.py
python analysis/ghidra/benchmarks/tests/test_run_real_corpus_eval.py
python analysis/ghidra/benchmarks/tests/test_regenerate_pre_2393.py
# #3172: the output-cap change, and the same hand-enumeration
# trap as #2980 above -- a new test file is not collected by
# anything, so it runs on a developer machine and nowhere else.
# Both are stubbed (no model, no GPU, no Ollama): the first pins
# that a done_reason=length generation is neither stored as `ok`
# nor scored as an answer, the second pins the 940 committed
# records still awaiting a reclassification decision.
python analysis/ghidra/benchmarks/tests/test_truncation_outcome.py
python analysis/ghidra/benchmarks/tests/test_committed_reclassification.py
# Same hand-enumeration trap. Reads one committed run directory
# off disk and drives write_human_grades_template() through a
# temp-dir writer: no model, no GPU, no Ollama. Pins that a live
# coder run is labelled `live_model` -- synthetic input, live
# answers -- in run.json and in its grade-sheet sidecar, that the
# value is one of transcripts.PROVENANCES rather than a private
# string, that the recorded transcripts_sha256 still verifies the
# untouched transcript, and that `captured` is still refused
# inside the repository now that a third member is permitted
# there.
python analysis/ghidra/benchmarks/tests/test_run_provenance.py
# Same hand-enumeration trap, hit again by Q2/Q7: the corpus
# scorer is a second producer of model answers and could
# disagree with the harness that owns the rules -- it silently
# widened a harmony budget instead of refusing it, and a
# generation the output cap ended still collected corpus points
# and still fed the claim pool. Both are stubbed (no model, no
# GPU, no Ollama, no network): the first proves a harmony cell
# declaring less than the floor puts no request on the wire and
# that the floor has exactly one owner, the second proves a
# capped answer publishes no score and no claims while an answer
# that stopped on its own terms still does.
python analysis/ghidra/benchmarks/tests/test_harmony_budget_guard.py
python analysis/ghidra/benchmarks/tests/test_capped_downstream.py
# Same hand-enumeration trap, hit again by the #3172 close-out: a
# guard that exists in one producer and not its siblings. The four
# producers with no harmony adaptation would have accepted a
# gpt-oss tag and sent a budget below the floor, publishing the
# resulting empty answers as scores; the offline rescorer scored
# capped stored answers. Both are stubbed (no model, no GPU, no
# Ollama, no network). The first proves every one of those
# producers refuses a harmony-served model before a request is
# built, through the one shared refusal; the second proves a
# stored answer the cap ended -- or that recorded no finish reason
# at all -- earns nothing when rescored offline.
python analysis/ghidra/benchmarks/tests/test_harmony_producer_refusal.py
python analysis/ghidra/benchmarks/tests/test_rescore_cap_propagation.py
# Same trap, on the file the refusal above changed: the judge
# stability probe's own test file existed and ran only by hand, so
# its plan/run_repeats determinism logic had no CI coverage at all.
# Stubbed -- claims_mod._post_json is replaced, no network.
python analysis/ghidra/benchmarks/tests/test_probe_judge_repeat_stability.py
# Same trap, on the shared transport every model call in the
# benchmark goes through. request_json() defaulted to a flat
# timeout=300, so an 8-token context probe and the coder slot's
# 4096-token answer got the same wall clock, and any model below
# ~7 tok/s was excluded by the transport rather than by the
# benchmark -- a timeout is not a measurement. Stubbed
# (urllib.request.urlopen is replaced, no Ollama, no GPU, no
# network): it proves the allowance now follows the request's own
# output budget, that a short request keeps the old 300s floor so
# a dead server still fails fast, and that both ends clamp so no
# caller can buy an unbounded wait.
python analysis/ghidra/benchmarks/tests/test_request_timeout.py
python analysis/ghidra/benchmarks/tests/test_bench_tools.py
# full_capabilities.py (#800): inventories env vars straight out of
# the real pipeline source, so its own test runs against the real
# tree rather than fixtures -- proving that stays true.
Expand Down
Loading
Loading