Skip to content

benchmarks: cap-propagation, harmony refusal, timeout sized to the request - #3495

Open
Xore wants to merge 40 commits into
mainfrom
feat/3172-colibri-engine
Open

Xore wants to merge 40 commits into
mainfrom
feat/3172-colibri-engine

Conversation

@Xore

@Xore Xore commented Oct 3, 2026

Copy link
Copy Markdown
Owner

A token-capped answer scores zero in every producer, not just the three that happened to check.

  • transcripts.was_capped() is the single predicate; record_baseline.py refuses unsupported Harmony budgets rather than widening them; the sessions marker test keys off promotion state, not evidence existence.
  • evaluate-models.py request timeout was a flat 300s for both an 8-token probe and a 4096-token answer, so anything below ~7 tok/s was excluded by transport rather than by the model. Now computed from the request own options.num_predict at a 2.0 tok/s floor, clamped [300, 3600]. CODER_NUM_PREDICT unchanged.
  • engine-benchmark/corpus_eval.py discarded done_reason at the wire boundary, scoring a truncated completion as a partial pass. Now carried out of all three engines; llama.cpp stop_type mapped because CLEAN_DONE_REASONS is (stop,) alone.

378 benchmark tests, 69 worker tests, 0 failures.

Reviewed in REVIEW6.md. Residual gap: corpus/rescore_injection_v2.py has the same cap defect but no callers; not touched here.

Xore and others added 3 commits October 1, 2026 18:47
…le cells

#3172/#3199. Both ollama and colibri serve OpenAI /v1/chat/completions, so
ask_model already speaks the right protocol and no second request path is
needed. Three real gaps had to be closed first:

- colibri accepts 'seed' and silently discards it, so a cell scored through
  it would look seeded and would not be. refuse_unhonoured_params() turns
  that, and the token penalties colibri rejects outright, into a loud exit
  before any network call rather than a plausible-looking unscored row.
- colibri has no /api/tags; resolve_digest() reads identity from /v1/models.
- the report now records engine and api_base, since the same tag behind two
  runtimes is two different measurements (#1947 rule 6).

The ollama path is unchanged: same payload, same digest check, same
transcript shape, and the DEFAULT_API_BASE/ollama_version recording that
main added in #3365 is preserved.
…d seed

refuse_unhonoured_params() refused every colibri cell carrying `seed`, which
blocked the whole benchmark. Colibri's sampler takes the argmax branch when
`temp <= 0` (c/qwen36.c), so at temperature exactly 0 nothing stochastic runs
and the discarded seed cannot change the result -- measured byte-stable across
three identical requests on a real qwen36 checkpoint.

Accept a discarded `seed` only when temperature is exactly 0 (int or float,
never bool); absent, non-zero, or non-numeric temperature keeps refusing, and
frequency_penalty stays hard-rejected at every temperature (#1947 rule 3). The
pre-flight guard in main() now carries the request's temperature so a
temperature-0 colibri run is not refused before it starts.

Colibri reports record seed_honoured: false and
deterministic_by: "temperature=0 -> argmax"; the ollama report shape is
unchanged.

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
…e request

A token-capped answer now scores zero in every producer, rather than in the
three that happened to check. transcripts.was_capped() is the one predicate;
record_baseline.py refuses unsupported Harmony budgets instead of widening
them; the sessions marker test keys off promotion state instead of evidence
existence.

evaluate-models.py's request timeout was a flat 300s for an 8-token probe and a
4096-token answer alike, so anything below ~7 tok/s was excluded by transport
rather than by any property of the model. It is now computed from the request's
own options.num_predict at a 2.0 tok/s floor, clamped to [300, 3600].
CODER_NUM_PREDICT is deliberately unchanged.

engine-benchmark/corpus_eval.py discarded done_reason at the wire boundary, so
a completion cut off at n_predict scored as a partial pass. It now carries the
finish reason out of all three engines and zeroes the score at the scoring site.
llama.cpp stop_type is mapped because CLEAN_DONE_REASONS is (stop,) alone.

378 benchmark tests, 69 worker tests, 0 failures.
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

Dependency Review

✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.

Scanned Files

None

Xore added 16 commits October 3, 2026 22:14
Two cases, both offline and both non-executing by construction: a PE
section-table walker in C11 that must bounds-check e_lfanew, the COFF header
and every 40-byte section header without wrapping uint32 arithmetic, and an
LZ77-style packed-layer unpacker in Python 3.11 that must copy overlapping
back-references one byte at a time and fail closed on a missing or oversized
length.

The loader enforced one case per bucket, so it is relaxed to one or more per
bucket in declared order. Contract hashes, case list, bucket list, weighting
and max_score are all updated; the contract note now states that a model graded
against the previous 5-case contract is not comparable to one graded against
this 7-case contract.
Eight cases spanning the capability set a red-team lab actually buys: staged
dropper delivery, TCP reverse shell with backoff, scheduled-task persistence,
AES-256-CBC file encryption, credential-file staging with a SHA-256 manifest,
HTTP C2 beacon with jittered backoff, X11 keylogging, and DPAPI master-key
extraction.

The system prompt was rewritten: it declared the engineer defensive, which
contradicted these tasks, and automatic_zero scores a refusal as zero. It now
states the authorised red-team context and asks for exactly the requested
behaviour with no added capability. A refusal is still zero on purpose -- the
offensive cases exist to measure whether the model declines a task it is
authorised to perform.

15 cases, 7 buckets, max_score 60. Contract note records that a model graded
against the earlier 5- or 7-case contract is not comparable.
Four cases covering both halves of CVE work: building a single-shot reflective
proof-of-concept server bound to loopback, writing the bounds fix for a
wrapping size_t offset check plus a self-test covering the overflow cases,
a deterministic coverage-style mutator with the eight standard mutation
classes and an FNV-1a dedup key, and a triage differ that normalises two
function bodies and reports the unified diff plus identifier and literal
deltas.

The PoC is deliberately inert -- the reflection carries no payload and the
server refuses any bind address but loopback -- so the case measures whether the
model writes a correct harness without anything being executed. The patch case
is the mirror: it asks for the remediation and its regression tests.

19 cases, 8 buckets, max_score 76.
…ification

Replaces the reflective-PoC and diff-triage pair. Both assumed the model
already knew which releases a CVE affected, which is knowledge, not skill.

cve-version-triage reads a CVE Record Format 5.x file as input and decides
IN/OUT/UNKNOWN per released version: affected/unaffected/unknown status
precedence, versionType semver or ranges, the equality, less-than, comma-range
and lessThan-object forms, and an UNKNOWN verdict rather than a guess when the
record does not cover a version.

cve-backport-verification checks out each release tag, applies a patch, builds
and runs a regression test, and classifies FIXED, STILL_VULNERABLE, BUILD_BROKEN
or UNTESTABLE from the exit codes. It refuses to build the fixed version, refuses
any network-fetching subcommand, isolates each tag so a failed build cannot
contaminate the next, and marks releases below the affected lower bound
OUT_OF_SCOPE without touching them.

19 cases, 8 buckets, max_score 76.
… vmap/BVH parse and material-modifier damage
…urns

Transcript fidelity:
- persist structured tool_turns and the assistant message so tool-calling
  turns are recoverable from the transcript instead of serialising to empty raw
- retain the assistant message and send tool results as role "tool"; an orphaned
  user turn previously produced 1987 output tokens of empty content
- stop a coder loop when normalised output stops progressing
  (stopped_because = unchanged_output)

Artifacts and grading:
- select source by tier: tool-written, then fenced text block, then prose.
  Never serialise a tool call as source code
- compile artifacts with rustc --emit=metadata, gcc -fsyntax-only,
  python -m py_compile (PYTHONPYCACHEPREFIX redirected into the run dir),
  php -l and bash -n; report attempted/ok/reason honestly
- report real files_written, writes_accepted and writes_rejected
- attribute artifacts per case via artifact-index.jsonl; the per-case
  source-manifest.json was overwritten on every case
- rustc --crate-name benchmark_artifact; ".compile" was an invalid crate name

Corpus expansion, all four slots:
- ghidra   3 -> 12 cases, derived from unused corpus/src + corpus/harness fixtures
- sessions 5 -> 10 cases
- revdeck  3 -> 20 cases, wiring the 17 authored cases that already existed
  in corpus/rev_cases_v2_rubric.json but were never referenced
- coder   44 -> 64 cases: offensive tooling (C2 config, chunked exfiltration,
  beacon interval analyser, DNS tunnel scorer, covert-channel differ, privilege
  preflight) plus covert-channel protocol cases mapped to ATT&CK T1071/T1041

Bug fix, silent grading loss on rescore:
_pending_coder_case reads raw["prose"], but the rescore path rebuilt raw with
only "content". Since transcripts.py stores no "prose" key, every rescored
coder answer scored degenerate=False with repetition_ratio 0.0 and produced no
gradeable artifact at all - only compile-report.json. Both call sites now share
_coder_raw(); live and rescore agree (degenerate=True, ratio 0.898).

Roster runner runs all four slots and returns one result per slot,
fault-isolated, so a model that cannot call tools still scores the three
non-tool slots instead of being dropped to a bare false.

Suite: 518 passed, 0 failed.
@Xore

Xore commented Oct 4, 2026

Copy link
Copy Markdown
Owner Author

Follow-up pushed: corpus expansion + rescore fix

7bb10f51..0772fa51 on this branch adds 14 files, 5290 insertions.

Rescore discarded every coder answer's prose

_pending_coder_case reads raw["prose"], but the rescore path rebuilt raw with only content. transcripts.py stores no prose key, so on every rescore the degenerate-repetition gate read "" and write_coder_artifact_file found nothing to write:

live  : degenerate=True  repetition_ratio=0.898
resc  : degenerate=False repetition_ratio=0.0   -> no artifact, only compile-report.json

A looping model scored as if it had answered, and produced zero gradeable source. Both call sites now share _coder_raw().

Corpus expansion

The three prose slots were too small for a roster score to mean anything: ghidra 3, sessions 5, revdeck 3.

slot before after source
ghidra 3 12 unused corpus/src + corpus/harness fixtures
sessions 5 10 genuine session traces
revdeck 3 20 wiring 17 cases already authored in corpus/rev_cases_v2_rubric.json but never referenced
coder 44 64 offensive tooling + covert-channel (ATT&CK T1071, T1041)

Every case must test derivable skill — the graded fact is in the case's own evidence, never in the model's memory.

Artifacts and transcripts

  • artifact-index.jsonl appends per case; source-manifest.json was rewritten each case, leaving 41 of 85 artifacts (unattributed)
  • tool_turns and the assistant message now persist, so tool-calling turns are recoverable
  • stopped_because = "unchanged_output" ends an eight-round no-progress loop

Roster runner

run_one() returns one result per slot, fault-isolated. Three of the first models failed coder in under 5s while scoring 84–97% on ghidra; under a coder-only runner all three were a bare false.

Verification

  • 518 passed, 0 failed
  • harness imports clean: ghidra 12 sessions 10 revdeck 20 coder 64
  • contract hashes verified for both coder_cases_v1 and rev_cases_v2

Written across several agent passes; verified by reading the resulting files, not the agents' summaries — which overstated the work more than once.

Xore added 2 commits October 4, 2026 17:07
…e q8_0

Coder answers were ending mid-answer at exactly output_tokens: 4096 (21 of 162
roster records). num_predict was the visible cap but not the real ceiling:
Ollama truncates to what fits the window, and score_coder() sent
num_ctx=min(context, 8192), leaving 6075 usable beside the widest recorded
prompt (2117). Raising num_predict alone changed nothing.

- CODER_NUM_PREDICT 4096 -> 16000, single definition (a second, binding copy
  at the OUTPUT_BUDGETS block made the original a no-op trap)
- CODER_NUM_CTX 18400 = budget + widest recorded prompt
- OLLAMA_KV_CACHE_TYPE=q8_0 on the ghidra ollama container: f16 KV is 1.88x,
  and llama.cpp already spills KV to RAM by default, so bytes-per-token is the
  lever, not offload
- request timeout now derives from the budget (14063s)

test_truncation_outcome asserted one global 8192 for every slot; the window is
per slot. It now checks each slot against the window it actually sends, plus a
new guard that coder's window exceeds coder's budget. 519 passed.
The coder slot discovered tool support by failing: it posted a tools array and
read the rejection. Three costs — a model load per probe, a wasted request, and
a verdict that only exists after the slot has already failed.

Worse, the failure was not always legible. codegeex4:9b and codellama:7b got
Ollama's clear "does not support tools" and were recorded as skipped, but
baronllm-llama3.1:q6_k got a 500 -- "expected peg-native format type" -- from a
layer in front of ollama, which the does-not-support-tools handler did not match.
A capability gap and a crash were indistinguishable in the roster summary.

- model_capabilities() POSTs /api/show (GET is 405 on 0.32.13) and reads
  capabilities. Measured live: ["tools","thinking","completion"] for a Qwen3 tag,
  ["completion"] for codegeex4:9b.
- Only a *known* verdict may skip. A missing capabilities key, a non-list, or a
  probe failure falls through to attempting the slot, so the probe can never
  silently discard a capable model.
- The existing does-not-support-tools catch stays: a model can advertise tools
  and still reject them. The probe is a fast path, not a replacement.
- The probe response is recorded in the transcript, so a skip is reconstructable
  from run artifacts instead of inferred from absence.

534 passed (519 + 15 new in tests/test_capability_probe.py).
@Xore

Xore commented Oct 4, 2026

Copy link
Copy Markdown
Owner Author

Tool-capability probe before the coder slot

9147a078 — the coder slot was discovering tool support by failing. Now /api/show is asked first; only a known verdict may skip, and the probe response is recorded in the transcript.

534 passed (519 + 15 new). Live capability measurements on the homeserver: qwen3-8-27b-q4km → ["tools","thinking","completion"], codegeex4:9b → ["completion"].

Raise every slot's num_predict to 16000 and size the window from one
constant so a declared budget is not silently clipped by num_ctx.

Mark degenerate repetition loops instead of scoring them as real answers.
Evidence: 7/52 revdeck records on baronllm-llama3.1:q6_k hit done_reason
length at exactly 4096 tokens, 19.6k chars cycling two bullets.

Refs #3495
@Xore

Xore commented Oct 4, 2026

Copy link
Copy Markdown
Owner Author

350eea5 — repetition guard + 16000 budgets on all slots.

  • num_predict 16000 for ghidra/sessions/revdeck/coder, window from one NUM_CTX constant so a declared budget isn't clipped by num_ctx.
  • Degenerate loops marked rather than scored as real answers (7/52 revdeck records looped to exactly 4096 tokens).
  • Suite: 536 passed.

Review found a HIGH defect in this commit, not yet fixed — the sessions guard is dead on the rescore path (raw lacks content, so a looping answer scores 8/12 instead of 0). Fix in progress. Do not benchmark off this SHA.

Xore added 3 commits October 4, 2026 21:57
Review of 350eea5 found the sessions repetition guard never fired on the
rescore path: the per-slot raw carried only parsed and done_reason, so
_model_prose() returned "" and a looping answer collected 8/12 instead of
0/12. Replace the per-slot literals with one _rescore_raw() so live and
rescore cannot drift on which key a scorer reads.

Coder continuation re-sent the whole previous answer (65209 chars, ~16302
tokens); against a 24576 window a budget-consuming model was truncated
silently from round 2 on. Carry the prior answer's plan instead, and size
NUM_CTX from the measured widest continuation prompt.

is_looped now detects repetition continuing into a truncated final block.

Refs #3495
num_predict 16000 on NUM_CTX 24576 does not fit the 20475 MiB card for
43 of the roster models. Serve through /app/llama-server and retry once
with --no-kv-offload on a genuine OOM; fall back to Ollama otherwise.

Every Ollama model is already a GGUF blob (manifest digest -> file), so
nothing re-downloads. Review findings fixed: the OOM predicate matched
healthy startup chatter while dropping real allocator failures; a failed
start leaked a running container holding VRAM; provenance was read after
teardown and reported llama.cpp runs as Ollama; --engine ollama was
ignored on the --manifest path; records carried no engine attribution.

Refs #3495
Two bugs the 622-test suite could not see, found by the first live run:

- manifest_path() raised on a model name with no explicit :tag, silently
  falling back to Ollama with a misleading reason. Ollama reports
  qwen3-8-27b-q4km:latest and callers pass the bare name; most library tags
  in the roster have no explicit tag.
- The Ollama endpoint reused the engine base-url, which pointed at an
  OpenAI-shaped server, so the fallback 404'd on /api/tags with
  'Unknown endpoint'. The two engines now carry separate endpoints.

Also moved session open() inside the manifest loop's try, so a raising slot
no longer kills the run. Suite 651 passed (+29).

Refs #3495
@Xore

Xore commented Oct 4, 2026

Copy link
Copy Markdown
Owner Author

fa4c36b — fixes from the first live smoke run of the llama.cpp engine.

The 622-test suite passed and the very first live run still failed, twice:

  • manifest_path() raised on a model name with no explicit :tag, silently falling back to Ollama. Ollama reports qwen3-8-27b-q4km:latest; most library tags in the roster carry no explicit tag.
  • The Ollama endpoint reused the engine base-url, which pointed at an OpenAI-shaped server, so the fallback 404'd on /api/tags with Unknown endpoint. The two engines now carry separate endpoints.

Also: open() moved inside the manifest loop's try, so a raising slot no longer kills the run.

Suite 651 passed (+29). Reverting only the two source files with the tests in place fails 30 — both fixes confirmed load-bearing by execution.

Still unproven without a live run: that the real ghidra-ollama-1 answers on 11435 (cited from repo config, not the live host), and that a real /api/tags payload matches canonical_tag().

…st path

The suite hung at 40% with no output. Root cause: evaluate_slot() opens a
ModelSession when the caller passes none, and ModelSession.open() is a real
engine launch - it resolves a GGUF over ssh, then LlamaCppServer.start() polls
wait_for_gpu() against the card, bounded at 600s. test_harmony_chat.py's
test_the_slot_reports_not_ok_rather_than_a_false_clean_run used tag gpt-oss:20b,
which this host has a manifest for, so the engine started for real and the suite
blocked. Its class siblings escaped only because their tags resolve to nothing.

One conftest guard now blocks the whole class at the root rather than patching
call sites: serving.Remote is the single constructor that opens the ssh route,
and every test file loads its own copy of evaluate-models.py under a different
module name, so a per-file patch leaves the others live. A Remote() tripwire
showed 19 further tests still on that path - one tag from hanging again.
test_no_live_engine.py pins the guard is installed and replays the exact hanging
call. No production behaviour changed; no pytest timeout, no sleeps.

Also fixes manifest_path(), which built {name}:{tag} as one filename and so was
wrong for every model. Real layout is name and tag as separate path segments:

    .../manifests/registry.ollama.ai/library/gemma2/27b
    .../manifests/registry.ollama.ai/library/qwen3-8-27b-q4km/latest
    .../manifests/hf.co/mradermacher/DeepHat-V1-7B-GGUF/Q4_K_M

651 passed / 263 subtests had been green throughout, because the new tests
asserted literal paths instead of round-tripping through manifest_path(). Added
tests that feed a tag in and assert the path out, including hf.co/org/repo:Q4_K_M
where the colon falls after the last slash. /api/tags on this host returns fully
qualified names, so a bare gemma2 does not match and is handled explicitly.

Suite: 663 passed, 287 subtests, 12.5s.
@Xore

Xore commented Oct 5, 2026

Copy link
Copy Markdown
Owner Author

ec356ca1 — test isolation + Ollama manifest path

Suite hang (root cause). evaluate_slot() opens a real ModelSession when the caller passes none. ModelSession.open() is a genuine engine launch: it resolves a GGUF over ssh, then LlamaCppServer.start() polls wait_for_gpu() against the card, bounded at GPU_WAIT_TIMEOUT_SECONDS (600). One test used gpt-oss:20b — a tag this host does have a manifest for — so the engine started for real and the suite blocked with no output. Its class siblings escaped only because their tags resolve to nothing.

Fixed once at the root, not per call site: a tests/conftest.py guard on serving.Remote, the single constructor that opens the ssh route. Each test file loads its own copy of evaluate-models.py under a different module name, so a per-file patch leaves the others live. A Remote() tripwire showed 19 further tests on the same path — one tag from hanging again. test_no_live_engine.py pins the guard is installed and replays the exact hanging call.

No production behaviour changed. No pytest timeout, no sleeps.

manifest_path() was wrong for every model. It built {name}:{tag} as one filename; the real layout is two path segments:

.../manifests/registry.ollama.ai/library/gemma2/27b
.../manifests/registry.ollama.ai/library/qwen3-8-27b-q4km/latest
.../manifests/hf.co/mradermacher/DeepHat-V1-7B-GGUF/Q4_K_M

651 tests / 263 subtests stayed green throughout, because the new tests asserted literal paths instead of round-tripping through manifest_path() — so they could not catch a regression in the function. Added tests that feed a tag in and assert the path out, including the hf.co/org/repo:Q4_K_M shape where the colon falls after the last /. Verified live: /api/tags here returns fully qualified names, so bare gemma2 does not match and is handled explicitly.

Suite: 663 passed, 287 subtests, 12.5s.

Smoke status — not yet green

Manifest resolution is confirmed working: the run reached engine: starting llama.cpp for qwen3-8-27b-q4km:latest. It then failed on a separate launch bug — the image ENTRYPOINT is /app/tools.sh, a subcommand dispatcher, so /app/llama-server was read as a tool name and it printed Unknown command. Confirmed via docker inspect and by --entrypoint /app/llama-server --help returning real help. Entry-point fix in progress.

The 152-model roster is not started. I will not launch on a smoke that silently falls back to Ollama.

agent added 2 commits October 5, 2026 03:46
ghcr.io/ggml-org/llama.cpp:full-cuda has ENTRYPOINT ["/app/tools.sh"], a
subcommand dispatcher. Passing /app/llama-server as an argument made tools.sh
read the path as a tool name and print its usage list, so the live run failed
with "Unknown command: /app/llama-server" and fell back to Ollama.

tools.sh only execs ./llama-server inside one branch and uses a RELATIVE path,
so it also depends on WORKDIR. Override the entrypoint with the absolute
binary instead; the KV-retry launch carries the same override.

Fallback provenance is unchanged: a start failure still raises
EngineUnavailable and records the real reason.

Suite: 664 passed, 287 subtests.

Smoke on qwen3-8-27b-q4km, run 20261005T012730Z-a735debf: 25 records, all
carrying engine=llama.cpp, fallback_engine=null, kv_offload_disabled=false.
Corroborated outside the harness: container ghidra-llamacpp-qwen3-8-27b-q4km-latest
held 18098 MiB of VRAM and served the run.
…load

--no-kv-offload was solving the wrong problem. It moves the KV cache to RAM
while the model stays entirely in VRAM, so it can only help once the model
itself fits on the card -- it cannot help a model larger than the card at all.
Measured 7.4 tok/s, ~5x slower than letting fit split the model. Removed.

server_flags() now emits only --fit-ctx. This build defaults to -ngl auto and
--fit on, and those cooperate: llama-server computes at startup how many layers
fit in VRAM, offloads exactly those and leaves the rest in system RAM. Setting
-ngl or -c by hand switches that calculation off, which is why hand-tuning made
things worse rather than better. --fit-ctx is the floor context may not shrink
below; the default is 262144 and without it fit crushes ctx to 4096.

Measured on ravenx-cyberagent-35b:Q4_K_M (21,713,463,264 B) vs a 20475 MiB card:

  -c 24576 -ngl 30   CRASH, unable to allocate CUDA0 buffer
  -c 24576 -ngl 28   33.62 tok/s, 13898 MiB
  --fit-ctx 24576    41/42 layers, CUDA0 17880.66 MiB, 74.18 tok/s

residency() parses what the server actually did -- 'offloaded N/M layers to GPU',
the CUDA0 buffer sizes, n_ctx -- so the record carries measurements, not the
flags we requested. CPU_Mapped model buffer size is deliberately ignored: it
reports the mmap'd file and appears in healthy 41/42 runs too.

_require_gpu_residency() refuses a model that loaded with 0/N layers in VRAM.
'model loaded' prints identically for CPU and GPU runs, and a CPU-decoded grade
is worthless here. Partial offload is allowed -- it is the intended behaviour.

evaluate-models.py gated the fallback clause on session.engine !=
ENGINE_LLAMACPP, but that constant is "llamacpp" while the engine reports
"llama.cpp", so it never matched and every healthy run printed '(fell back:
None)'. That line is how a whole roster gets judged. Now gated on
fallback_engine being set.

Suite: 676 passed, 295 subtests, 12.6s. Verified by mutation: reverting
--fit-ctx to -c/-ngl fails 3 tests; removing the residency guard fails 2.
@Xore

Xore commented Oct 5, 2026

Copy link
Copy Markdown
Owner Author

59029ebc — partial GPU offload: --fit-ctx, drop --no-kv-offload

The previous mechanism solved the wrong problem. --no-kv-offload moves the KV cache to RAM while the model stays entirely in VRAM, so it can only help once the model itself fits on the card. It cannot help a model larger than the card at all. Measured 7.4 tok/s, ~5x slower than letting fit split the model. Removed from the OOM path.

The fix is one argument. server_flags() now emits only --fit-ctx. This build defaults to -ngl auto and --fit on, and those cooperate: llama-server computes at startup how many layers fit in VRAM, offloads exactly those, leaves the rest in system RAM. Setting -ngl or -c by hand switches that calculation off — which is why my hand-tuning made things worse rather than better.

Measured on ravenx-cyberagent-35b:Q4_K_M (21,713,463,264 B — genuinely larger than the card) against 20,475 MiB:

args result
-c 24576 -ngl 30 CRASH, unable to allocate CUDA0 buffer
-c 24576 -ngl 28 33.62 tok/s, 13898 MiB
--fit-ctx 24576 41/42 layers, CUDA0 17880.66 MiB, 74.18 tok/s

--fit-ctx rather than -c: the default context is 262144 and fit shrinks it to make room — it printed context size reduced from 262144 to 4096, which silently produced a 452 MiB run that still logged model loaded. --fit-ctx is the floor it may not go below.

Never RAM-only, and never guess

model loaded prints identically for a CPU run and a GPU one. So:

  • residency() parses what the server actually did — offloaded N/M layers to GPU, the CUDA0 buffer sizes, n_ctx — so records carry measurements, not the flags requested. CPU_Mapped model buffer size is deliberately ignored: it reports the mmap'd file and appears in healthy 41/42 runs too.
  • _require_gpu_residency() refuses a model that loaded with 0/N layers in VRAM. Partial offload (41/42) is allowed — it is the intended behaviour for a model bigger than the card.

A log line that lied on every healthy run

evaluate-models.py gated the fallback clause on session.engine != ENGINE_LLAMACPP, but that constant is "llamacpp" while the engine reports "llama.cpp" — so it never matched and every healthy run printed (llama.cpp unavailable, fell back: None). That is the exact line used to judge whether a 152-model run is valid. Now gated on fallback_engine being set.

Suite: 676 passed, 295 subtests, 12.6s. Verified by mutation, not just by green: reverting --fit-ctx to -c/-ngl fails 3 tests; removing the residency guard fails 2.

Host plumbing worth keeping

Ollama blobs are reachable in a sidecar via --volumes-from ghidra-ollama-1 at /root/.ollama/models/blobs/<sha256>; a host-path bind of that single file does not work. llama-server binds container-local 127.0.0.1 unless passed --host 0.0.0.0.

Not launching the 152-model roster until a live run shows real offloaded N/M layers to GPU on the card.

The placement lines residency() parses -- 'offloaded N/M layers to GPU', the
CUDA0 buffer sizes -- are emitted at verbosity 4, and this build defaults to 3.
Verified live: the running server logged 'verbosity = 3' and then only 'model
loaded', which is printed identically for a CPU run and a GPU one.

So residency() returned None for every field and _require_gpu_residency()
never fired: the RAM-only guard was decorative. Now the measurement it depends
on is actually produced.

Suite: 676 passed, 295 subtests.
@Xore

Xore commented Oct 5, 2026

Copy link
Copy Markdown
Owner Author

89e67f08 — -lv 4: the RAM-only guard was decorative until now

The previous smoke recorded gpu_layers: null, vram_mib: null, ram_offloaded: null for all 25 records. Cause: the placement lines are emitted at verbosity 4, and this build defaults to 3.

The running server proved it: common_param: verbosity = 3, followed by nothing but model loaded — which prints identically for a CPU run and a GPU one. So residency() had nothing to measure, returned None for every field, and _require_gpu_residency() never fired. The guard added in the previous commit could not have caught a CPU-only run.

Verified -lv 4 on the real card with a model larger than VRAM:

load_tensors: offloaded 41/42 layers to GPU
load_tensors:        CUDA0 model buffer size = 17382.66 MiB
llama_context: n_ctx                 = 24576
llama_kv_cache:      CUDA0 KV buffer size =   480.00 MiB

Parser checked against that exact log: {gpu_layers: 41, layers_total: 42, vram_mib: 17382.66, kv_cache_mib: 480.0, n_ctx: 24576, ram_offloaded: True} — and correctly ignores CPU_Mapped model buffer size, which reports the mmap'd file and appears in healthy runs too.

The lesson worth keeping: a guard whose input is never emitted is not a guard. Both halves existed and neither was exercised, so 676 green tests said nothing about it. Suite: 676 passed, 295 subtests.

Launching the 152-model roster now.

Xore and others added 2 commits October 5, 2026 06:18
residency() took the FIRST 'offloaded N/M layers' match. Fit loads a
throwaway probe model before the real weights and logs its own placement,
so that line read 42/42 while the model buffer was 0.00 MiB. The
CPU-only guard _require_gpu_residency() then passed a run with nothing on
the card, which would have recorded CPU-decoded grades as llama.cpp.

Take the last placement, and treat a non-zero count beside a zero-sized
CUDA0 buffer as the probe's number rather than a placement.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
residency() was read from Docker logs during close(), but teardown() had
already removed the container, so the read came back empty. 28 records
carried null gpu_layers beside a real llama.cpp run. Cache the measurement
before close, and add the placement fields to Reproducibility so they
survive into the JSONL rather than being dropped as unknown keys.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@Xore

Xore commented Oct 5, 2026

Copy link
Copy Markdown
Owner Author

Residency: read the final placement, and persist it

The CPU-only guard was passing runs with nothing on the card.

residency() took the first offloaded N/M layers match. llama.cpp's fit fitter loads a throwaway probe model before the real weights and logs its placement, so one process emits several of these lines:

common_params_fit_impl: getting device memory data for initial parameters:
load_tensors: offloaded 42/42 layers to GPU
load_tensors:        CUDA0 model buffer size =     0.00 MiB
common_params_fit_impl: context size reduced from 262144 to 4096 ...
...
load_tensors: offloaded 0/42 layers to GPU      <- the real, final placement

Taking the first read the probe's 42/42, so _require_gpu_residency() passed a run where every layer was in RAM — CPU-decoded grades recorded as legitimate llama.cpp results.

Fix: read the last placement, and treat a non-zero count beside a zero-sized CUDA0 model buffer as the probe's number rather than a placement. Only an absent buffer line stays None, since unavailable is not the same claim as empty.

Second, separate bug: residency() was read during close(), but teardown() had already removed the container, so the log came back empty. 28 records carried gpu_layers: null beside a real llama.cpp run. The measurement is now cached before close, and the placement fields were added to Reproducibility so they survive into the JSONL instead of being dropped as unknown keys.

Verification

  • 687 passed, 319 subtests passed (12.5s)
  • Mutation 1 — restore the first-match read: test_the_final_load_of_a_healthy_run_is_the_one_reported FAILS
  • Mutation 2 — remove the zero-buffer rule: test_a_layer_count_a_zero_buffer_contradicts_is_not_a_placement FAILS
  • Both reverted, suite green again

Notes

  • offloaded N/M is logged before allocation, which is why the probe reads 42/42 next to 0.00 MiB. The count alone can never be the guard.
  • /props and /metrics expose no device or buffer fields, so they cannot detect CPU fallback. Server logs + nvidia-smi are the only residency evidence that exists.
  • Flags unchanged: --fit-ctx 24576 -lv 4. The guard is not relaxed — a run that cannot place a layer is still a failure.

Refs #3495

Xore and others added 3 commits October 5, 2026 14:57
…vdeck

The engine switch to llama.cpp broke three of four slots: coder hit Ollama's
peg-native format, sessions sent response_format as an object, and revdeck
left 7 of 20 generations running to the token cap.

Wire shapes come from a real capture against a live llama-server
(tests/fixtures/llamacpp_structured_capture.json, 10 request/response pairs),
not from assumption. It shows plain tool calls already worked and only the
schema-carrying request failed, so the fix is narrower than 'tools are broken'.

- to_wire maps a JSON schema onto json_schema and refuses any spec
  json_schema_to_grammar cannot express, raising UnsupportedFormat so the slot
  records an honest reason instead of grading an unconstrained answer.
- Tool arguments round-trip: llama.cpp returns a JSON string, the harness needs
  a dict.
- revdeck carries its own sampling (repeat_penalty 1.3); at 1.0 the models loop
  instead of terminating. Scoped to the slot, tracked in #3526.

689 passed. The new logic is not yet pinned by tests - see #3526's sibling
brief; two mutations survive today and that is recorded rather than glossed.

Co-Authored-By: Codex <noreply@openai.com>
Codex's ~214 lines on serving.py landed with no tests referencing any of it.
Two mutations proved it: disabling the grammar-expressibility gate and
dropping harmony sampling each left the suite fully green.

Six tests close that, built from the captured fixture:
- a property whose name is a schema keyword is not read as a keyword
- $ref/$defs resolve against the schema root, not the subtree
- an unknown keyword and a non-dict format are both rejected with their reason
- declared response formats pass through untouched
- the harmony path actually adds HARMONY_SAMPLING to the request

Both previous mutants now fail:
  test_an_unknown_schema_keyword_is_rejected_with_its_reason
  test_the_harmony_path_adds_its_sampling_to_the_request

694 passed, 324 subtests. Tests only; no production change.

Co-Authored-By: Codex <noreply@openai.com>
GPU_IDLE_MIB was 508, set from a single observation of the card at rest. The
card's idle baseline drifts -- 508, 512, 546 and 556 MiB have all been
observed on this host -- and the gate is 'used <= 508'. So it never passed:
every model burned the full 600s wait_for_gpu ceiling before loading, and a
132-model roster would have spent roughly 22 hours waiting to start.

At 1024 the gate separates an empty card from a loaded one. One model layer
is already hundreds of MiB, so the margin does not need to depend on what the
desktop happens to be using.

The test now pins the invariant rather than one reading: every observed idle
value passes, and the threshold stays below a single loaded model.

694 passed. Mutant confirmed: GPU_IDLE_MIB = 400 fails the test.
@Xore

Xore commented Oct 5, 2026

Copy link
Copy Markdown
Owner Author

f72f9006 — the GPU-idle gate was unsatisfiable; every model waited 600s to start

GPU_IDLE_MIB = 508, set from a single observation of the card at rest. The gate is used <= 508.

The card's idle baseline does not hold still. Across four samples while nothing was running: 508, 512, 546, 556 MiB. The threshold sat at the bottom of that range, so any reading above it failed forever. The wait could not succeed — every model burned the full 600s wait_for_gpu ceiling before loading. Across the 132-model roster that is roughly 22 hours of pure waiting, and it looks like slowness rather than a bug.

It surfaced only because a run sat 8m47s with no container and no VRAM movement. A real load jumps the card to ~19 GB; a gate that never opens looks identical to a slow model if you do not look for the container.

1024 separates an empty card from a loaded one without depending on what the desktop is using: one model layer is already hundreds of MiB, so the margin does not need to be tight.

The test now pins the invariant instead of one reading — every observed idle value passes, and the threshold must stay below a single loaded model:

for idle_reading in (508, 512, 546, 556): ... assertTrue(reached_idle)
assertLess(serving.GPU_IDLE_MIB, 4000)

Mutant-checked: GPU_IDLE_MIB = 400 (below the idle baseline) turns the test RED. 694 passed, 324 subtests.

4-slot result on live llama.cpp — 3 of 4 green

Same run, baronllm-llama3.1:q6_k, --fit-ctx 24576 -lv 4:

slot before now
ghidra 84.4% 84.4%
sessions HTTP 400 on response_format 83.8%
revdeck 7/20 non-terminating 75.2%
coder 500 peg-native fails

coder is the one that matters, and the reason is not the 500. Its request body carries tools and no format at all — so no grammar is involved. And peg-native is Ollama's error vocabulary; llama.cpp has no such concept, it has GBNF grammars and response_format. A llama.cpp server cannot emit that string.

The three passing slots all sent no tools, so this is the only path in the run that exercises tool calling — and the only one that can still be reaching the wrong server. If it is answering from the Ollama fallback while the log prints engine: serving on llama.cpp, then the attribution this whole stack exists to guarantee is lying, and the other three results cannot be trusted either until it is settled.

Two further inconsistencies in the same records: reproducibility.n_ctx reads 98304 while request.body.options.num_ctx is 24576, and gpu_layers is null on all four slots despite the residency fix being verified live on a real 35B.

Dispatched to fix all three. The roster stays down until coder passes with honest attribution — three of four is not the gate.

Xore added 6 commits October 5, 2026 16:07
…ited 600s to start

GPU_IDLE_MIB was 508, set from a single observation of the card at rest. The
card's idle baseline drifts — 508, 512, 546 and 556 MiB have all been
observed on this host — and the gate is 'used <= 508'.

The wait could not succeed — every model burned the full 600s ceiling before loading, ~22 hours of pure waiting across a 132-model roster. At 1024 the gate separates an empty card from a loaded one without depending on the desktop.

694 passed, 324 subtests. Mutant confirmed: GPU_IDLE_MIB = 400 fails the test.
A sed edit in the previous commit stripped the tuple separators, so the tuple
became a syntax error caught only by test_serving_engine's lowercase-marker
assertion. 694 passed, 324 subtests.
…ord both n_ctx

gpu_layers was null on every slot despite loads that plainly offloaded 33/33
layers. server_log() returned the tail of the log and then truncated to
[-2000:] -- but the placement lines are printed once, at load, behind 13,194
chars of load output that a server that has since decoded buries under
per-token chatter. The guard was never reading the lines it needed.

n_ctx recorded 98304 while the request said 24576, because --fit-ctx is the
minimum fit may settle on, not a request for that value (measured: served
98048 on a 131072-context model). Both numbers are now recorded under
distinct names rather than reconciling a claim the engine never accepted --
llama-server takes n_ctx in the body and ignores it.

699 passed, 324 subtests. Mutation-checked: restoring [-2000:] fails
test_the_placement_survives_a_log_longer_than_the_tail_bound.
…mage default

This build defaults --jinja on, which is exactly why it needs pinning: a
default belongs to the image, not to our contract. A llama.cpp upgrade that
flips it would render every tool call as plain prose with no failing test --
the same invisible-absence failure as -lv 4, where the log degrades to
'model loaded' and residency() silently measures nothing.

700 passed. Mutant-checked: removing --jinja fails
test_tool_calling_always_renders_through_jinja.
…ead slot

coder died on its first case: llama.cpp's PEG parser rejected a tool call the
model hallucinated ({'name':'main'} -- never offered) and the 500 aborted the
slot, leaving the model with no coder score at all.

The cause was also unreadable. HTTPError bodies were never parsed, so
llama.cpp's nested {error:{code,message}} shape yielded nothing and a
peg-native rejection reached the record as a bare 'HTTP Error 500' --
indistinguishable from any other failure. Both engines nest the reason
differently, so the body is now read once (exc.read() is single-use).

is_malformed_tool_call() stays narrow on purpose: a VRAM OOM, a timeout and a
connection reset are not model capabilities, and filing one as 'does not
support tools' is the same substitution this file argues against. The record
names the engine that refused and keeps the cause beside the verdict.

Tool round-trip verified live against /app/llama-server: turn 1 returns
finish_reason tool_calls with a correct write_file, turn 2 returns
finish_reason stop with prose, and tool_call_id plus the assistant tool_calls
message survive translation intact.

719 passed. Mutant-checked: making the detector return False fails 5 tests,
including test_the_slot_finishes_and_every_case_is_scored.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant