The problem
revdeck is the only slot whose models fail to terminate. On llama.cpp, 7 of 20 revdeck answers ran to the token cap instead of returning finish_reason: stop, and scored zero rather than as answers:
format-string-argument, heap-copy-error-handling, linked-list-sum-intent,
neutral-note-stack-overflow, process-injection,
use-after-free-path, witness-instruction-process-launch
This is not a slot in the ordinary sense — it behaves differently enough that the shared decode path is worth studying on its own. This issue is the research track; the current mitigation is noted below so the benchmark can proceed in the meantime.
What is already established
A real capture against baronllm-llama3.1:q6_k on llama.cpp (fixture:
analysis/ghidra/benchmarks/tests/fixtures/llamacpp_structured_capture.json) shows:
- At
repeat_penalty: 1.0, process-injection loops and never stops.
- At
repeat_penalty: 1.3, repeat_last_n: 4096 it terminates.
- The loop is not template-related:
gpt-oss (harmony) needs different sampling from Qwen-style models, which is why HARMONY_SAMPLING and REVDECK_SAMPLING are separate constants.
Current mitigation (decided, ship it)
REVDECK_SAMPLING = {"repeat_last_n": 4096, "repeat_penalty": 1.3}. Applied to the revdeck slot only. This is a deliberate, documented exception — the other three slots are unchanged. It is not the final answer; this issue tracks the research.
Open questions
- Is
repeat_penalty the right knob? It penalises repetition in correct answers too, not just runaway loops. A model that legitimately repeats a code pattern (format-string-argument, heap-copy-error-handling) is penalised for doing its job.
- Does the same penalty distort the other slots? If revdeck needs 1.3 to terminate, do sessions/ghidra/coder silently need it too?
- Does this reproduce on Ollama? The 7/20 figure is llama.cpp-only.
- Is there a slot-level fix? Explicit
--stop / <|im_end|> passed to llama-server, rather than per-slot sampling.
- Per-model vs per-slot? Some models may terminate cleanly at 1.0. The penalty is currently blanket for the slot.
Why this is research, not a bug
The non-stopping behaviour may be a property of the models themselves under constrained decoding, not a harness defect. That distinction matters: if it is the model, the correct fix is to record it and score it honestly; if it is the harness, the fix belongs in the decode path.
Acceptance
Refs #3495
The problem
revdeckis the only slot whose models fail to terminate. On llama.cpp, 7 of 20 revdeck answers ran to the token cap instead of returningfinish_reason: stop, and scored zero rather than as answers:This is not a slot in the ordinary sense — it behaves differently enough that the shared decode path is worth studying on its own. This issue is the research track; the current mitigation is noted below so the benchmark can proceed in the meantime.
What is already established
A real capture against
baronllm-llama3.1:q6_kon llama.cpp (fixture:analysis/ghidra/benchmarks/tests/fixtures/llamacpp_structured_capture.json) shows:repeat_penalty: 1.0,process-injectionloops and never stops.repeat_penalty: 1.3, repeat_last_n: 4096it terminates.gpt-oss(harmony) needs different sampling from Qwen-style models, which is whyHARMONY_SAMPLINGandREVDECK_SAMPLINGare separate constants.Current mitigation (decided, ship it)
REVDECK_SAMPLING = {"repeat_last_n": 4096, "repeat_penalty": 1.3}. Applied to the revdeck slot only. This is a deliberate, documented exception — the other three slots are unchanged. It is not the final answer; this issue tracks the research.Open questions
repeat_penaltythe right knob? It penalises repetition in correct answers too, not just runaway loops. A model that legitimately repeats a code pattern (format-string-argument, heap-copy-error-handling) is penalised for doing its job.--stop/<|im_end|>passed to llama-server, rather than per-slot sampling.Why this is research, not a bug
The non-stopping behaviour may be a property of the models themselves under constrained decoding, not a harness defect. That distinction matters: if it is the model, the correct fix is to record it and score it honestly; if it is the harness, the fix belongs in the decode path.
Acceptance
1.0, which need1.3, which loop at any setting.Refs #3495