Skip to content

revdeck decode research: non-stopping generations, repetition penalty as a slot-level exception #3526

Description

@Xore

The problem

revdeck is the only slot whose models fail to terminate. On llama.cpp, 7 of 20 revdeck answers ran to the token cap instead of returning finish_reason: stop, and scored zero rather than as answers:

format-string-argument, heap-copy-error-handling, linked-list-sum-intent,
neutral-note-stack-overflow, process-injection,
use-after-free-path, witness-instruction-process-launch

This is not a slot in the ordinary sense — it behaves differently enough that the shared decode path is worth studying on its own. This issue is the research track; the current mitigation is noted below so the benchmark can proceed in the meantime.

What is already established

A real capture against baronllm-llama3.1:q6_k on llama.cpp (fixture:
analysis/ghidra/benchmarks/tests/fixtures/llamacpp_structured_capture.json) shows:

  • At repeat_penalty: 1.0, process-injection loops and never stops.
  • At repeat_penalty: 1.3, repeat_last_n: 4096 it terminates.
  • The loop is not template-related: gpt-oss (harmony) needs different sampling from Qwen-style models, which is why HARMONY_SAMPLING and REVDECK_SAMPLING are separate constants.

Current mitigation (decided, ship it)

REVDECK_SAMPLING = {"repeat_last_n": 4096, "repeat_penalty": 1.3}. Applied to the revdeck slot only. This is a deliberate, documented exception — the other three slots are unchanged. It is not the final answer; this issue tracks the research.

Open questions

  1. Is repeat_penalty the right knob? It penalises repetition in correct answers too, not just runaway loops. A model that legitimately repeats a code pattern (format-string-argument, heap-copy-error-handling) is penalised for doing its job.
  2. Does the same penalty distort the other slots? If revdeck needs 1.3 to terminate, do sessions/ghidra/coder silently need it too?
  3. Does this reproduce on Ollama? The 7/20 figure is llama.cpp-only.
  4. Is there a slot-level fix? Explicit --stop / <|im_end|> passed to llama-server, rather than per-slot sampling.
  5. Per-model vs per-slot? Some models may terminate cleanly at 1.0. The penalty is currently blanket for the slot.

Why this is research, not a bug

The non-stopping behaviour may be a property of the models themselves under constrained decoding, not a harness defect. That distinction matters: if it is the model, the correct fix is to record it and score it honestly; if it is the harness, the fix belongs in the decode path.

Acceptance

Refs #3495

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions