Skip to content

round7-6: preference and RL round — DPO then GRPO with the benchmark's own judges as reward (schema, injection gate, polarity, pooled claims) #3085

Description

@Xore

Part of #3079 (round 7). Plan §7, experiments T4 (DPO) and T5 (GRPO). GPU work; queues behind the cold re-run and behind round7-5.

Why RL here, specifically

The benchmark already contains programmatic judges: injection_gate.classify_answer / gate_points / paired_verdict (v3, #2694), polarity.forbidden_hit, the pydantic contracts in llm-worker/contracts.py (Ollama structured output rejects a schema-invalid answer outright), and claims.py's pooled-claims coverage. Those are reward functions, as long as the prompts they score are training-side (corpus v1), never the 17 test cases. Injection resistance in particular is a property SFT teaches weakly and preference optimisation teaches directly — and the harness's own caveat is that the injection axis reads "coverage verified, resistance not measured" until the v3 positive control fires. This is the round that targets that axis on purpose.

T4 — DPO first (cheap, no vLLM)

  • start from round7-5's best adapter per student
  • pairs: corpus v1 S5 plus rejection-sampled pairs from the student itself (sample N answers per training prompt, keep (chosen, rejected) where the judges disagree: schema-valid vs invalid, instruction-resisted vs complied, polarity-right vs forbidden-term hit)
  • Unsloth DPOTrainer path, QLoRA, β and lr per Unsloth's DPO notebook defaults, fixed seed

T5 — GRPO (Unsloth + vLLM colocated)

  • 14 B GRPO on 20 GB is feasible per Unsloth (vLLM standby: UNSLOTH_VLLM_STANDBY=1 before import; keep max sequence modest; batch of generations ≥ 4)
  • reward = weighted sum, weights recorded in the config: schema validity (hard 0 when invalid), injection gate verdict (hard 0 on complied), polarity (forbidden_hit == 0), pooled-claims coverage against a training-side pool built with claims.py --pool /mnt-1/training/pools/train-v1.json from teacher answers on S2/S3 (not tier-a-v1.json), a mild length penalty
  • loss variants worth one run each if time allows: grpo, dr_grpo, gspo
  • early stop on dev; export every checkpoint that beats its SFT parent on dev through export_to_ollama.sh

Deliverables

  • reward.py (imports the harness's own modules — no reimplementation), dpo.py, grpo.py, configs, run cards (reward curves, KL, dev metrics, VRAM peak)
  • artefacts registered in round7-8 with the SFT parent as control
  • the injection axis reported per the v3 protocol (docs/analysis/ghidra/benchmarks/injection-gate-protocol.md), paired verdicts, on the round-7 pin — the number this issue exists to move

Depends on round7-5, round7-3. Feeds round7-8.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or requestllmLLM analysis workermlML worker and GPU scoring

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions