Part of #3079 (round 7). Plan §7, experiment T1. GPU work; queues behind the cold re-run via chain_round7.sh.
Why
Every base in the matrix learned decompiler output and honeypot transcripts only incidentally. Unsloth's continued-pretraining path (UnslothTrainer, embed_tokens + lm_head in target_modules, a separate embedding_learning_rate) is the cheapest way to move a base's language toward Ghidra pseudocode and attacker shell sessions before it is taught the jobs in round7-5. Whether it helps on this benchmark is the question; the answer is a row, either way.
Run
- base: the round7-5 students (
Qwen3-14B, gpt-oss-20b; Gemma-4-12B if VRAM allows the embedding modules) — same weights the SFT runs start from, so CPT-then-SFT vs SFT-only is a clean A/B
- data: corpus v1 S6 (round7-3): decompiler output, sanitised sessions, REx86 text, permissively-licensed RE corpora; target 50–200 M tokens, measured not guessed
- settings: QLoRA 4-bit, r = 64–128 on all linear + embeddings,
embedding_learning_rate ≈ 1/10 of learning_rate, use_gradient_checkpointing="unsloth" (offloads activations to system RAM — 93 GB today, 128 GB once the DIMMs land), sequence length 8 k, packing on; unsloth_tiled_mlp=True for any 32 k-context run
- throughput: record tokens/s on this card for each student in the run card — that number decides how much CPT is affordable, and nobody has measured it here
Deliverables
Depends on round7-1, round7-3. Feeds round7-5 and round7-8.
Part of #3079 (round 7). Plan §7, experiment T1. GPU work; queues behind the cold re-run via
chain_round7.sh.Why
Every base in the matrix learned decompiler output and honeypot transcripts only incidentally. Unsloth's continued-pretraining path (
UnslothTrainer,embed_tokens+lm_headintarget_modules, a separateembedding_learning_rate) is the cheapest way to move a base's language toward Ghidra pseudocode and attacker shell sessions before it is taught the jobs in round7-5. Whether it helps on this benchmark is the question; the answer is a row, either way.Run
Qwen3-14B,gpt-oss-20b;Gemma-4-12Bif VRAM allows the embedding modules) — same weights the SFT runs start from, so CPT-then-SFT vs SFT-only is a clean A/Bembedding_learning_rate≈ 1/10 oflearning_rate,use_gradient_checkpointing="unsloth"(offloads activations to system RAM — 93 GB today, 128 GB once the DIMMs land), sequence length 8 k, packing on;unsloth_tiled_mlp=Truefor any 32 k-context runDeliverables
analysis/ghidra/training/configs/cpt-*.yaml, run cards (loss curve, tokens, wall-clock, VRAM peak) underdocs/benchmarks/training/runs//mnt-1/training/runs/t1-cpt-<student>/, mirrored per the plan's durability ruleexport_to_ollama.shat Q4_K_M and scored raw (before SFT) by round7-8 — a CPT-only row tells whether the language shift alone moves the rubricDepends on round7-1, round7-3. Feeds round7-5 and round7-8.