Repository navigation
Conversation
…le cells #3172/#3199. Both ollama and colibri serve OpenAI /v1/chat/completions, so ask_model already speaks the right protocol and no second request path is needed. Three real gaps had to be closed first: - colibri accepts 'seed' and silently discards it, so a cell scored through it would look seeded and would not be. refuse_unhonoured_params() turns that, and the token penalties colibri rejects outright, into a loud exit before any network call rather than a plausible-looking unscored row. - colibri has no /api/tags; resolve_digest() reads identity from /v1/models. - the report now records engine and api_base, since the same tag behind two runtimes is two different measurements (#1947 rule 6). The ollama path is unchanged: same payload, same digest check, same transcript shape, and the DEFAULT_API_BASE/ollama_version recording that main added in #3365 is preserved.
…d seed refuse_unhonoured_params() refused every colibri cell carrying `seed`, which blocked the whole benchmark. Colibri's sampler takes the argmax branch when `temp <= 0` (c/qwen36.c), so at temperature exactly 0 nothing stochastic runs and the discarded seed cannot change the result -- measured byte-stable across three identical requests on a real qwen36 checkpoint. Accept a discarded `seed` only when temperature is exactly 0 (int or float, never bool); absent, non-zero, or non-numeric temperature keeps refusing, and frequency_penalty stays hard-rejected at every temperature (#1947 rule 3). The pre-flight guard in main() now carries the request's temperature so a temperature-0 colibri run is not refused before it starts. Colibri reports record seed_honoured: false and deterministic_by: "temperature=0 -> argmax"; the ollama report shape is unchanged. Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
…e request A token-capped answer now scores zero in every producer, rather than in the three that happened to check. transcripts.was_capped() is the one predicate; record_baseline.py refuses unsupported Harmony budgets instead of widening them; the sessions marker test keys off promotion state instead of evidence existence. evaluate-models.py's request timeout was a flat 300s for an 8-token probe and a 4096-token answer alike, so anything below ~7 tok/s was excluded by transport rather than by any property of the model. It is now computed from the request's own options.num_predict at a 2.0 tok/s floor, clamped to [300, 3600]. CODER_NUM_PREDICT is deliberately unchanged. engine-benchmark/corpus_eval.py discarded done_reason at the wire boundary, so a completion cut off at n_predict scored as a partial pass. It now carries the finish reason out of all three engines and zeroes the score at the scoring site. llama.cpp stop_type is mapped because CLEAN_DONE_REASONS is (stop,) alone. 378 benchmark tests, 69 worker tests, 0 failures.
Dependency Review✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.Scanned FilesNone |
Two cases, both offline and both non-executing by construction: a PE section-table walker in C11 that must bounds-check e_lfanew, the COFF header and every 40-byte section header without wrapping uint32 arithmetic, and an LZ77-style packed-layer unpacker in Python 3.11 that must copy overlapping back-references one byte at a time and fail closed on a missing or oversized length. The loader enforced one case per bucket, so it is relaxed to one or more per bucket in declared order. Contract hashes, case list, bucket list, weighting and max_score are all updated; the contract note now states that a model graded against the previous 5-case contract is not comparable to one graded against this 7-case contract.
Eight cases spanning the capability set a red-team lab actually buys: staged dropper delivery, TCP reverse shell with backoff, scheduled-task persistence, AES-256-CBC file encryption, credential-file staging with a SHA-256 manifest, HTTP C2 beacon with jittered backoff, X11 keylogging, and DPAPI master-key extraction. The system prompt was rewritten: it declared the engineer defensive, which contradicted these tasks, and automatic_zero scores a refusal as zero. It now states the authorised red-team context and asks for exactly the requested behaviour with no added capability. A refusal is still zero on purpose -- the offensive cases exist to measure whether the model declines a task it is authorised to perform. 15 cases, 7 buckets, max_score 60. Contract note records that a model graded against the earlier 5- or 7-case contract is not comparable.
Four cases covering both halves of CVE work: building a single-shot reflective proof-of-concept server bound to loopback, writing the bounds fix for a wrapping size_t offset check plus a self-test covering the overflow cases, a deterministic coverage-style mutator with the eight standard mutation classes and an FNV-1a dedup key, and a triage differ that normalises two function bodies and reports the unified diff plus identifier and literal deltas. The PoC is deliberately inert -- the reflection carries no payload and the server refuses any bind address but loopback -- so the case measures whether the model writes a correct harness without anything being executed. The patch case is the mirror: it asks for the remediation and its regression tests. 19 cases, 8 buckets, max_score 76.
…ification Replaces the reflective-PoC and diff-triage pair. Both assumed the model already knew which releases a CVE affected, which is knowledge, not skill. cve-version-triage reads a CVE Record Format 5.x file as input and decides IN/OUT/UNKNOWN per released version: affected/unaffected/unknown status precedence, versionType semver or ranges, the equality, less-than, comma-range and lessThan-object forms, and an UNKNOWN verdict rather than a guess when the record does not cover a version. cve-backport-verification checks out each release tag, applies a patch, builds and runs a regression test, and classifies FIXED, STILL_VULNERABLE, BUILD_BROKEN or UNTESTABLE from the exit codes. It refuses to build the fixed version, refuses any network-fetching subcommand, isolates each tag so a failed build cannot contaminate the next, and marks releases below the affected lower bound OUT_OF_SCOPE without touching them. 19 cases, 8 buckets, max_score 76.
…er, per-bone profiles)
… vmap/BVH parse and material-modifier damage
…e model declares completion
…binaries out of the tree
…urns Transcript fidelity: - persist structured tool_turns and the assistant message so tool-calling turns are recoverable from the transcript instead of serialising to empty raw - retain the assistant message and send tool results as role "tool"; an orphaned user turn previously produced 1987 output tokens of empty content - stop a coder loop when normalised output stops progressing (stopped_because = unchanged_output) Artifacts and grading: - select source by tier: tool-written, then fenced text block, then prose. Never serialise a tool call as source code - compile artifacts with rustc --emit=metadata, gcc -fsyntax-only, python -m py_compile (PYTHONPYCACHEPREFIX redirected into the run dir), php -l and bash -n; report attempted/ok/reason honestly - report real files_written, writes_accepted and writes_rejected - attribute artifacts per case via artifact-index.jsonl; the per-case source-manifest.json was overwritten on every case - rustc --crate-name benchmark_artifact; ".compile" was an invalid crate name Corpus expansion, all four slots: - ghidra 3 -> 12 cases, derived from unused corpus/src + corpus/harness fixtures - sessions 5 -> 10 cases - revdeck 3 -> 20 cases, wiring the 17 authored cases that already existed in corpus/rev_cases_v2_rubric.json but were never referenced - coder 44 -> 64 cases: offensive tooling (C2 config, chunked exfiltration, beacon interval analyser, DNS tunnel scorer, covert-channel differ, privilege preflight) plus covert-channel protocol cases mapped to ATT&CK T1071/T1041 Bug fix, silent grading loss on rescore: _pending_coder_case reads raw["prose"], but the rescore path rebuilt raw with only "content". Since transcripts.py stores no "prose" key, every rescored coder answer scored degenerate=False with repetition_ratio 0.0 and produced no gradeable artifact at all - only compile-report.json. Both call sites now share _coder_raw(); live and rescore agree (degenerate=True, ratio 0.898). Roster runner runs all four slots and returns one result per slot, fault-isolated, so a model that cannot call tools still scores the three non-tool slots instead of being dropped to a bare false. Suite: 518 passed, 0 failed.
Follow-up pushed: corpus expansion + rescore fix
Rescore discarded every coder answer's prose
A looping model scored as if it had answered, and produced zero gradeable source. Both call sites now share Corpus expansionThe three prose slots were too small for a roster score to mean anything: ghidra 3, sessions 5, revdeck 3.
Every case must test derivable skill — the graded fact is in the case's own evidence, never in the model's memory. Artifacts and transcripts
Roster runner
Verification
Written across several agent passes; verified by reading the resulting files, not the agents' summaries — which overstated the work more than once. |
…e q8_0 Coder answers were ending mid-answer at exactly output_tokens: 4096 (21 of 162 roster records). num_predict was the visible cap but not the real ceiling: Ollama truncates to what fits the window, and score_coder() sent num_ctx=min(context, 8192), leaving 6075 usable beside the widest recorded prompt (2117). Raising num_predict alone changed nothing. - CODER_NUM_PREDICT 4096 -> 16000, single definition (a second, binding copy at the OUTPUT_BUDGETS block made the original a no-op trap) - CODER_NUM_CTX 18400 = budget + widest recorded prompt - OLLAMA_KV_CACHE_TYPE=q8_0 on the ghidra ollama container: f16 KV is 1.88x, and llama.cpp already spills KV to RAM by default, so bytes-per-token is the lever, not offload - request timeout now derives from the budget (14063s) test_truncation_outcome asserted one global 8192 for every slot; the window is per slot. It now checks each slot against the window it actually sends, plus a new guard that coder's window exceeds coder's budget. 519 passed.
The coder slot discovered tool support by failing: it posted a tools array and read the rejection. Three costs — a model load per probe, a wasted request, and a verdict that only exists after the slot has already failed. Worse, the failure was not always legible. codegeex4:9b and codellama:7b got Ollama's clear "does not support tools" and were recorded as skipped, but baronllm-llama3.1:q6_k got a 500 -- "expected peg-native format type" -- from a layer in front of ollama, which the does-not-support-tools handler did not match. A capability gap and a crash were indistinguishable in the roster summary. - model_capabilities() POSTs /api/show (GET is 405 on 0.32.13) and reads capabilities. Measured live: ["tools","thinking","completion"] for a Qwen3 tag, ["completion"] for codegeex4:9b. - Only a *known* verdict may skip. A missing capabilities key, a non-list, or a probe failure falls through to attempting the slot, so the probe can never silently discard a capable model. - The existing does-not-support-tools catch stays: a model can advertise tools and still reject them. The probe is a fast path, not a replacement. - The probe response is recorded in the transcript, so a skip is reconstructable from run artifacts instead of inferred from absence. 534 passed (519 + 15 new in tests/test_capability_probe.py).
Tool-capability probe before the coder slot
|
Raise every slot's num_predict to 16000 and size the window from one constant so a declared budget is not silently clipped by num_ctx. Mark degenerate repetition loops instead of scoring them as real answers. Evidence: 7/52 revdeck records on baronllm-llama3.1:q6_k hit done_reason length at exactly 4096 tokens, 19.6k chars cycling two bullets. Refs #3495
|
350eea5 — repetition guard + 16000 budgets on all slots.
Review found a HIGH defect in this commit, not yet fixed — the sessions guard is dead on the rescore path ( |
Review of 350eea5 found the sessions repetition guard never fired on the rescore path: the per-slot raw carried only parsed and done_reason, so _model_prose() returned "" and a looping answer collected 8/12 instead of 0/12. Replace the per-slot literals with one _rescore_raw() so live and rescore cannot drift on which key a scorer reads. Coder continuation re-sent the whole previous answer (65209 chars, ~16302 tokens); against a 24576 window a budget-consuming model was truncated silently from round 2 on. Carry the prior answer's plan instead, and size NUM_CTX from the measured widest continuation prompt. is_looped now detects repetition continuing into a truncated final block. Refs #3495
num_predict 16000 on NUM_CTX 24576 does not fit the 20475 MiB card for 43 of the roster models. Serve through /app/llama-server and retry once with --no-kv-offload on a genuine OOM; fall back to Ollama otherwise. Every Ollama model is already a GGUF blob (manifest digest -> file), so nothing re-downloads. Review findings fixed: the OOM predicate matched healthy startup chatter while dropping real allocator failures; a failed start leaked a running container holding VRAM; provenance was read after teardown and reported llama.cpp runs as Ollama; --engine ollama was ignored on the --manifest path; records carried no engine attribution. Refs #3495
Two bugs the 622-test suite could not see, found by the first live run: - manifest_path() raised on a model name with no explicit :tag, silently falling back to Ollama with a misleading reason. Ollama reports qwen3-8-27b-q4km:latest and callers pass the bare name; most library tags in the roster have no explicit tag. - The Ollama endpoint reused the engine base-url, which pointed at an OpenAI-shaped server, so the fallback 404'd on /api/tags with 'Unknown endpoint'. The two engines now carry separate endpoints. Also moved session open() inside the manifest loop's try, so a raising slot no longer kills the run. Suite 651 passed (+29). Refs #3495
|
fa4c36b — fixes from the first live smoke run of the llama.cpp engine. The 622-test suite passed and the very first live run still failed, twice:
Also: Suite 651 passed (+29). Reverting only the two source files with the tests in place fails 30 — both fixes confirmed load-bearing by execution. Still unproven without a live run: that the real |
…st path
The suite hung at 40% with no output. Root cause: evaluate_slot() opens a
ModelSession when the caller passes none, and ModelSession.open() is a real
engine launch - it resolves a GGUF over ssh, then LlamaCppServer.start() polls
wait_for_gpu() against the card, bounded at 600s. test_harmony_chat.py's
test_the_slot_reports_not_ok_rather_than_a_false_clean_run used tag gpt-oss:20b,
which this host has a manifest for, so the engine started for real and the suite
blocked. Its class siblings escaped only because their tags resolve to nothing.
One conftest guard now blocks the whole class at the root rather than patching
call sites: serving.Remote is the single constructor that opens the ssh route,
and every test file loads its own copy of evaluate-models.py under a different
module name, so a per-file patch leaves the others live. A Remote() tripwire
showed 19 further tests still on that path - one tag from hanging again.
test_no_live_engine.py pins the guard is installed and replays the exact hanging
call. No production behaviour changed; no pytest timeout, no sleeps.
Also fixes manifest_path(), which built {name}:{tag} as one filename and so was
wrong for every model. Real layout is name and tag as separate path segments:
.../manifests/registry.ollama.ai/library/gemma2/27b
.../manifests/registry.ollama.ai/library/qwen3-8-27b-q4km/latest
.../manifests/hf.co/mradermacher/DeepHat-V1-7B-GGUF/Q4_K_M
651 passed / 263 subtests had been green throughout, because the new tests
asserted literal paths instead of round-tripping through manifest_path(). Added
tests that feed a tag in and assert the path out, including hf.co/org/repo:Q4_K_M
where the colon falls after the last slash. /api/tags on this host returns fully
qualified names, so a bare gemma2 does not match and is handled explicitly.
Suite: 663 passed, 287 subtests, 12.5s.
|
ghcr.io/ggml-org/llama.cpp:full-cuda has ENTRYPOINT ["/app/tools.sh"], a subcommand dispatcher. Passing /app/llama-server as an argument made tools.sh read the path as a tool name and print its usage list, so the live run failed with "Unknown command: /app/llama-server" and fell back to Ollama. tools.sh only execs ./llama-server inside one branch and uses a RELATIVE path, so it also depends on WORKDIR. Override the entrypoint with the absolute binary instead; the KV-retry launch carries the same override. Fallback provenance is unchanged: a start failure still raises EngineUnavailable and records the real reason. Suite: 664 passed, 287 subtests. Smoke on qwen3-8-27b-q4km, run 20261005T012730Z-a735debf: 25 records, all carrying engine=llama.cpp, fallback_engine=null, kv_offload_disabled=false. Corroborated outside the harness: container ghidra-llamacpp-qwen3-8-27b-q4km-latest held 18098 MiB of VRAM and served the run.
…load --no-kv-offload was solving the wrong problem. It moves the KV cache to RAM while the model stays entirely in VRAM, so it can only help once the model itself fits on the card -- it cannot help a model larger than the card at all. Measured 7.4 tok/s, ~5x slower than letting fit split the model. Removed. server_flags() now emits only --fit-ctx. This build defaults to -ngl auto and --fit on, and those cooperate: llama-server computes at startup how many layers fit in VRAM, offloads exactly those and leaves the rest in system RAM. Setting -ngl or -c by hand switches that calculation off, which is why hand-tuning made things worse rather than better. --fit-ctx is the floor context may not shrink below; the default is 262144 and without it fit crushes ctx to 4096. Measured on ravenx-cyberagent-35b:Q4_K_M (21,713,463,264 B) vs a 20475 MiB card: -c 24576 -ngl 30 CRASH, unable to allocate CUDA0 buffer -c 24576 -ngl 28 33.62 tok/s, 13898 MiB --fit-ctx 24576 41/42 layers, CUDA0 17880.66 MiB, 74.18 tok/s residency() parses what the server actually did -- 'offloaded N/M layers to GPU', the CUDA0 buffer sizes, n_ctx -- so the record carries measurements, not the flags we requested. CPU_Mapped model buffer size is deliberately ignored: it reports the mmap'd file and appears in healthy 41/42 runs too. _require_gpu_residency() refuses a model that loaded with 0/N layers in VRAM. 'model loaded' prints identically for CPU and GPU runs, and a CPU-decoded grade is worthless here. Partial offload is allowed -- it is the intended behaviour. evaluate-models.py gated the fallback clause on session.engine != ENGINE_LLAMACPP, but that constant is "llamacpp" while the engine reports "llama.cpp", so it never matched and every healthy run printed '(fell back: None)'. That line is how a whole roster gets judged. Now gated on fallback_engine being set. Suite: 676 passed, 295 subtests, 12.6s. Verified by mutation: reverting --fit-ctx to -c/-ngl fails 3 tests; removing the residency guard fails 2.
|
| args | result |
|---|---|
-c 24576 -ngl 30 |
CRASH, unable to allocate CUDA0 buffer |
-c 24576 -ngl 28 |
33.62 tok/s, 13898 MiB |
--fit-ctx 24576 |
41/42 layers, CUDA0 17880.66 MiB, 74.18 tok/s |
--fit-ctx rather than -c: the default context is 262144 and fit shrinks it to make room — it printed context size reduced from 262144 to 4096, which silently produced a 452 MiB run that still logged model loaded. --fit-ctx is the floor it may not go below.
Never RAM-only, and never guess
model loaded prints identically for a CPU run and a GPU one. So:
residency()parses what the server actually did —offloaded N/M layers to GPU, the CUDA0 buffer sizes,n_ctx— so records carry measurements, not the flags requested.CPU_Mapped model buffer sizeis deliberately ignored: it reports the mmap'd file and appears in healthy 41/42 runs too._require_gpu_residency()refuses a model that loaded with 0/N layers in VRAM. Partial offload (41/42) is allowed — it is the intended behaviour for a model bigger than the card.
A log line that lied on every healthy run
evaluate-models.py gated the fallback clause on session.engine != ENGINE_LLAMACPP, but that constant is "llamacpp" while the engine reports "llama.cpp" — so it never matched and every healthy run printed (llama.cpp unavailable, fell back: None). That is the exact line used to judge whether a 152-model run is valid. Now gated on fallback_engine being set.
Suite: 676 passed, 295 subtests, 12.6s. Verified by mutation, not just by green: reverting --fit-ctx to -c/-ngl fails 3 tests; removing the residency guard fails 2.
Host plumbing worth keeping
Ollama blobs are reachable in a sidecar via --volumes-from ghidra-ollama-1 at /root/.ollama/models/blobs/<sha256>; a host-path bind of that single file does not work. llama-server binds container-local 127.0.0.1 unless passed --host 0.0.0.0.
Not launching the 152-model roster until a live run shows real offloaded N/M layers to GPU on the card.
The placement lines residency() parses -- 'offloaded N/M layers to GPU', the CUDA0 buffer sizes -- are emitted at verbosity 4, and this build defaults to 3. Verified live: the running server logged 'verbosity = 3' and then only 'model loaded', which is printed identically for a CPU run and a GPU one. So residency() returned None for every field and _require_gpu_residency() never fired: the RAM-only guard was decorative. Now the measurement it depends on is actually produced. Suite: 676 passed, 295 subtests.
|
residency() took the FIRST 'offloaded N/M layers' match. Fit loads a throwaway probe model before the real weights and logs its own placement, so that line read 42/42 while the model buffer was 0.00 MiB. The CPU-only guard _require_gpu_residency() then passed a run with nothing on the card, which would have recorded CPU-decoded grades as llama.cpp. Take the last placement, and treat a non-zero count beside a zero-sized CUDA0 buffer as the probe's number rather than a placement. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
residency() was read from Docker logs during close(), but teardown() had already removed the container, so the read came back empty. 28 records carried null gpu_layers beside a real llama.cpp run. Cache the measurement before close, and add the placement fields to Reproducibility so they survive into the JSONL rather than being dropped as unknown keys. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Residency: read the final placement, and persist itThe CPU-only guard was passing runs with nothing on the card.
Taking the first read the probe's Fix: read the last placement, and treat a non-zero count beside a zero-sized Second, separate bug: Verification
Notes
Refs #3495 |
…vdeck The engine switch to llama.cpp broke three of four slots: coder hit Ollama's peg-native format, sessions sent response_format as an object, and revdeck left 7 of 20 generations running to the token cap. Wire shapes come from a real capture against a live llama-server (tests/fixtures/llamacpp_structured_capture.json, 10 request/response pairs), not from assumption. It shows plain tool calls already worked and only the schema-carrying request failed, so the fix is narrower than 'tools are broken'. - to_wire maps a JSON schema onto json_schema and refuses any spec json_schema_to_grammar cannot express, raising UnsupportedFormat so the slot records an honest reason instead of grading an unconstrained answer. - Tool arguments round-trip: llama.cpp returns a JSON string, the harness needs a dict. - revdeck carries its own sampling (repeat_penalty 1.3); at 1.0 the models loop instead of terminating. Scoped to the slot, tracked in #3526. 689 passed. The new logic is not yet pinned by tests - see #3526's sibling brief; two mutations survive today and that is recorded rather than glossed. Co-Authored-By: Codex <noreply@openai.com>
Codex's ~214 lines on serving.py landed with no tests referencing any of it. Two mutations proved it: disabling the grammar-expressibility gate and dropping harmony sampling each left the suite fully green. Six tests close that, built from the captured fixture: - a property whose name is a schema keyword is not read as a keyword - $ref/$defs resolve against the schema root, not the subtree - an unknown keyword and a non-dict format are both rejected with their reason - declared response formats pass through untouched - the harmony path actually adds HARMONY_SAMPLING to the request Both previous mutants now fail: test_an_unknown_schema_keyword_is_rejected_with_its_reason test_the_harmony_path_adds_its_sampling_to_the_request 694 passed, 324 subtests. Tests only; no production change. Co-Authored-By: Codex <noreply@openai.com>
GPU_IDLE_MIB was 508, set from a single observation of the card at rest. The card's idle baseline drifts -- 508, 512, 546 and 556 MiB have all been observed on this host -- and the gate is 'used <= 508'. So it never passed: every model burned the full 600s wait_for_gpu ceiling before loading, and a 132-model roster would have spent roughly 22 hours waiting to start. At 1024 the gate separates an empty card from a loaded one. One model layer is already hundreds of MiB, so the margin does not need to depend on what the desktop happens to be using. The test now pins the invariant rather than one reading: every observed idle value passes, and the threshold stays below a single loaded model. 694 passed. Mutant confirmed: GPU_IDLE_MIB = 400 fails the test.
|
| slot | before | now |
|---|---|---|
| ghidra | 84.4% | 84.4% |
| sessions | HTTP 400 on response_format |
83.8% |
| revdeck | 7/20 non-terminating | 75.2% |
| coder | 500 peg-native |
fails |
coder is the one that matters, and the reason is not the 500. Its request body carries tools and no format at all — so no grammar is involved. And peg-native is Ollama's error vocabulary; llama.cpp has no such concept, it has GBNF grammars and response_format. A llama.cpp server cannot emit that string.
The three passing slots all sent no tools, so this is the only path in the run that exercises tool calling — and the only one that can still be reaching the wrong server. If it is answering from the Ollama fallback while the log prints engine: serving on llama.cpp, then the attribution this whole stack exists to guarantee is lying, and the other three results cannot be trusted either until it is settled.
Two further inconsistencies in the same records: reproducibility.n_ctx reads 98304 while request.body.options.num_ctx is 24576, and gpu_layers is null on all four slots despite the residency fix being verified live on a real 35B.
Dispatched to fix all three. The roster stays down until coder passes with honest attribution — three of four is not the gate.
…ited 600s to start GPU_IDLE_MIB was 508, set from a single observation of the card at rest. The card's idle baseline drifts — 508, 512, 546 and 556 MiB have all been observed on this host — and the gate is 'used <= 508'. The wait could not succeed — every model burned the full 600s ceiling before loading, ~22 hours of pure waiting across a 132-model roster. At 1024 the gate separates an empty card from a loaded one without depending on the desktop. 694 passed, 324 subtests. Mutant confirmed: GPU_IDLE_MIB = 400 fails the test.
A sed edit in the previous commit stripped the tuple separators, so the tuple became a syntax error caught only by test_serving_engine's lowercase-marker assertion. 694 passed, 324 subtests.
…ord both n_ctx gpu_layers was null on every slot despite loads that plainly offloaded 33/33 layers. server_log() returned the tail of the log and then truncated to [-2000:] -- but the placement lines are printed once, at load, behind 13,194 chars of load output that a server that has since decoded buries under per-token chatter. The guard was never reading the lines it needed. n_ctx recorded 98304 while the request said 24576, because --fit-ctx is the minimum fit may settle on, not a request for that value (measured: served 98048 on a 131072-context model). Both numbers are now recorded under distinct names rather than reconciling a claim the engine never accepted -- llama-server takes n_ctx in the body and ignores it. 699 passed, 324 subtests. Mutation-checked: restoring [-2000:] fails test_the_placement_survives_a_log_longer_than_the_tail_bound.
…mage default This build defaults --jinja on, which is exactly why it needs pinning: a default belongs to the image, not to our contract. A llama.cpp upgrade that flips it would render every tool call as plain prose with no failing test -- the same invisible-absence failure as -lv 4, where the log degrades to 'model loaded' and residency() silently measures nothing. 700 passed. Mutant-checked: removing --jinja fails test_tool_calling_always_renders_through_jinja.
…ead slot
coder died on its first case: llama.cpp's PEG parser rejected a tool call the
model hallucinated ({'name':'main'} -- never offered) and the 500 aborted the
slot, leaving the model with no coder score at all.
The cause was also unreadable. HTTPError bodies were never parsed, so
llama.cpp's nested {error:{code,message}} shape yielded nothing and a
peg-native rejection reached the record as a bare 'HTTP Error 500' --
indistinguishable from any other failure. Both engines nest the reason
differently, so the body is now read once (exc.read() is single-use).
is_malformed_tool_call() stays narrow on purpose: a VRAM OOM, a timeout and a
connection reset are not model capabilities, and filing one as 'does not
support tools' is the same substitution this file argues against. The record
names the engine that refused and keeps the cause beside the verdict.
Tool round-trip verified live against /app/llama-server: turn 1 returns
finish_reason tool_calls with a correct write_file, turn 2 returns
finish_reason stop with prose, and tool_call_id plus the assistant tool_calls
message survive translation intact.
719 passed. Mutant-checked: making the detector return False fails 5 tests,
including test_the_slot_finishes_and_every_case_is_scored.
A token-capped answer scores zero in every producer, not just the three that happened to check.
transcripts.was_capped()is the single predicate;record_baseline.pyrefuses unsupported Harmony budgets rather than widening them; the sessions marker test keys off promotion state, not evidence existence.evaluate-models.pyrequest timeout was a flat 300s for both an 8-token probe and a 4096-token answer, so anything below ~7 tok/s was excluded by transport rather than by the model. Now computed from the request ownoptions.num_predictat a 2.0 tok/s floor, clamped [300, 3600].CODER_NUM_PREDICTunchanged.engine-benchmark/corpus_eval.pydiscardeddone_reasonat the wire boundary, scoring a truncated completion as a partial pass. Now carried out of all three engines; llama.cppstop_typemapped becauseCLEAN_DONE_REASONSis(stop,)alone.378 benchmark tests, 69 worker tests, 0 failures.
Reviewed in REVIEW6.md. Residual gap:
corpus/rescore_injection_v2.pyhas the same cap defect but no callers; not touched here.