Skip to content

round7-4: continued pretraining with Unsloth on domain text — decompiler output, sessions, RE corpora — before SFT #3083

Description

@Xore

Part of #3079 (round 7). Plan §7, experiment T1. GPU work; queues behind the cold re-run via chain_round7.sh.

Why

Every base in the matrix learned decompiler output and honeypot transcripts only incidentally. Unsloth's continued-pretraining path (UnslothTrainer, embed_tokens + lm_head in target_modules, a separate embedding_learning_rate) is the cheapest way to move a base's language toward Ghidra pseudocode and attacker shell sessions before it is taught the jobs in round7-5. Whether it helps on this benchmark is the question; the answer is a row, either way.

Run

  • base: the round7-5 students (Qwen3-14B, gpt-oss-20b; Gemma-4-12B if VRAM allows the embedding modules) — same weights the SFT runs start from, so CPT-then-SFT vs SFT-only is a clean A/B
  • data: corpus v1 S6 (round7-3): decompiler output, sanitised sessions, REx86 text, permissively-licensed RE corpora; target 50–200 M tokens, measured not guessed
  • settings: QLoRA 4-bit, r = 64–128 on all linear + embeddings, embedding_learning_rate ≈ 1/10 of learning_rate, use_gradient_checkpointing="unsloth" (offloads activations to system RAM — 93 GB today, 128 GB once the DIMMs land), sequence length 8 k, packing on; unsloth_tiled_mlp=True for any 32 k-context run
  • throughput: record tokens/s on this card for each student in the run card — that number decides how much CPT is affordable, and nobody has measured it here

Deliverables

  • config under analysis/ghidra/training/configs/cpt-*.yaml, run cards (loss curve, tokens, wall-clock, VRAM peak) under docs/benchmarks/training/runs/
  • adapter + merged 16-bit under /mnt-1/training/runs/t1-cpt-<student>/, mirrored per the plan's durability rule
  • each CPT checkpoint exported through round7-1's export_to_ollama.sh at Q4_K_M and scored raw (before SFT) by round7-8 — a CPT-only row tells whether the language shift alone moves the rubric
  • handed to round7-5 as an alternative starting point (CPT-then-SFT row)

Depends on round7-1, round7-3. Feeds round7-5 and round7-8.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or requestin-progressActively being worked onllmLLM analysis workermlML worker and GPU scoring

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions