Part of #3079 (round 7). Plan §7, experiments T4 (DPO) and T5 (GRPO). GPU work; queues behind the cold re-run and behind round7-5.
Why RL here, specifically
The benchmark already contains programmatic judges: injection_gate.classify_answer / gate_points / paired_verdict (v3, #2694), polarity.forbidden_hit, the pydantic contracts in llm-worker/contracts.py (Ollama structured output rejects a schema-invalid answer outright), and claims.py's pooled-claims coverage. Those are reward functions, as long as the prompts they score are training-side (corpus v1), never the 17 test cases. Injection resistance in particular is a property SFT teaches weakly and preference optimisation teaches directly — and the harness's own caveat is that the injection axis reads "coverage verified, resistance not measured" until the v3 positive control fires. This is the round that targets that axis on purpose.
T4 — DPO first (cheap, no vLLM)
- start from round7-5's best adapter per student
- pairs: corpus v1 S5 plus rejection-sampled pairs from the student itself (sample N answers per training prompt, keep (chosen, rejected) where the judges disagree: schema-valid vs invalid, instruction-resisted vs complied, polarity-right vs forbidden-term hit)
- Unsloth
DPOTrainer path, QLoRA, β and lr per Unsloth's DPO notebook defaults, fixed seed
T5 — GRPO (Unsloth + vLLM colocated)
- 14 B GRPO on 20 GB is feasible per Unsloth (vLLM standby:
UNSLOTH_VLLM_STANDBY=1 before import; keep max sequence modest; batch of generations ≥ 4)
- reward = weighted sum, weights recorded in the config: schema validity (hard 0 when invalid), injection gate verdict (hard 0 on
complied), polarity (forbidden_hit == 0), pooled-claims coverage against a training-side pool built with claims.py --pool /mnt-1/training/pools/train-v1.json from teacher answers on S2/S3 (not tier-a-v1.json), a mild length penalty
- loss variants worth one run each if time allows:
grpo, dr_grpo, gspo
- early stop on dev; export every checkpoint that beats its SFT parent on dev through
export_to_ollama.sh
Deliverables
Depends on round7-5, round7-3. Feeds round7-8.
Part of #3079 (round 7). Plan §7, experiments T4 (DPO) and T5 (GRPO). GPU work; queues behind the cold re-run and behind round7-5.
Why RL here, specifically
The benchmark already contains programmatic judges:
injection_gate.classify_answer/gate_points/paired_verdict(v3, #2694),polarity.forbidden_hit, the pydantic contracts inllm-worker/contracts.py(Ollama structured output rejects a schema-invalid answer outright), andclaims.py's pooled-claims coverage. Those are reward functions, as long as the prompts they score are training-side (corpus v1), never the 17 test cases. Injection resistance in particular is a property SFT teaches weakly and preference optimisation teaches directly — and the harness's own caveat is that the injection axis reads "coverage verified, resistance not measured" until the v3 positive control fires. This is the round that targets that axis on purpose.T4 — DPO first (cheap, no vLLM)
DPOTrainerpath, QLoRA, β and lr per Unsloth's DPO notebook defaults, fixed seedT5 — GRPO (Unsloth + vLLM colocated)
UNSLOTH_VLLM_STANDBY=1before import; keep max sequence modest; batch of generations ≥ 4)complied), polarity (forbidden_hit== 0), pooled-claims coverage against a training-side pool built withclaims.py --pool /mnt-1/training/pools/train-v1.jsonfrom teacher answers on S2/S3 (nottier-a-v1.json), a mild length penaltygrpo,dr_grpo,gspoexport_to_ollama.shDeliverables
reward.py(imports the harness's own modules — no reimplementation),dpo.py,grpo.py, configs, run cards (reward curves, KL, dev metrics, VRAM peak)docs/analysis/ghidra/benchmarks/injection-gate-protocol.md), paired verdicts, on the round-7 pin — the number this issue exists to moveDepends on round7-5, round7-3. Feeds round7-8.