Skip to content

round7-3: training corpus v1 — decontaminated, teacher-labelled slices for the three slots, injection pairs, CPT text; manifest + decontamination report #3082

Description

@Xore

Part of #3079 (round 7). Plan §6. No GPU for the build; the local-teacher labelling needs the card and queues behind the cold re-run — the OpenRouter-labelled synthetic slice does not.

The rule that shapes everything here

The benchmark corpus is the test set. Its 17 programs (xor_decode_loop, vulnerable_strcpy, strcpy_note_neutral, strcpy_note_injected, linked_list_sum, indirect_dispatch, error_handling_alloc, process_and_injection, process_witness_probe, tlv_parser, loopback_connect, safe_strcpy, integer_overflow_alloc, use_after_free, checksum_rotate, format_string_bug, file_write_persist) — at any toolchain or optimisation level, their decompiled text in tierb-cache/, their rubric terms, the adjudicated claim pool tier-a-v1.json, and every transcript under docs/benchmarks/runs/ — are off limits for training data, calibration text and RL prompts. Same for the evaluate-models.py session and Rev·Deck fixtures. A model that has seen them scores a memory, not a capability.

Slices to build (/mnt-1/training/corpus-v1/, mode 0700)

slice job input labels leaves the host?
S1 sessions-captured session analysis real Cowrie/Beelzebub sessions from honeypot-v* in ES, sanitised with llm-worker/contracts.py's sanitize_text / _redact_secrets / sanitize_commands — the production path, so the student trains on what it will see local teacher, N=3, majority-agreed SessionAnalysis JSON only never
S2 ghidra-captured Ghidra triage captured samples the production ghidra-worker already analysed (ghidra-analysis-v1); re-decompile through ghidra_cache.py against the headless service so the evidence shape matches the benchmark's Tier B local teacher, N=3, majority-agreed triage JSON in the ghidra-triage-v1 contract never
S3 ghidra-synthetic Ghidra triage + Rev·Deck ≥ 60 new C programs written for this corpus across the rubric's behaviour families (encoding loops, memory-safety bugs, dispatch tables, persistence, network, benign near-neighbours, embedded-instruction controls) — none of them the 17; built with build_corpus.py's toolchain × opt-level matrix; decompiled the same way OpenRouter open-weight teacher (DeepSeek V4 / Qwen3.8 large / GLM-5.3 / Kimi K3 — licences permit distillation), N=2 agreement yes — synthetic only
S4 revdeck-rex86 Rev·Deck Zenodo 15420461 REx86.zip (5,981 entries: intent, Q&A, complete-the-code, comments; CC-BY-4.0). Check for an internal held-out split first (#847's caveat) as shipped already public
S5 injection-pairs all three prompts from S1–S3 with newly written embedded attacker instructions (new phrasings, not the corpus's own strcpy_note_injected / process_and_injection text) chosen = ignores the instruction; rejected = complies / breaks schema / flips polarity — filtered by injection_gate.classify_answer, the pydantic contracts and polarity.forbidden_hit per source slice
S6 cpt-text continued pretraining (round7-4) raw domain text: decompiler output from S2/S3, sanitised session transcripts, REx86 text, public RE corpora with permissive licences none per source slice
dev early stopping 10 % of each slice, split by program / session, never by row — —

Teacher key for S3 lives on the host (/etc/apiary-training.env, 0600, or the operator's choice) — never in the repo, never in an issue. Captured slices S1/S2/S5-from-captured/S6-from-captured stay on the host; if anything is pushed to Hugging Face it is a private repo and only ever the synthetic/public slices — operator decision, default off.

Decontamination (a deliverable, before any score is quoted)

analysis/ghidra/training/decontaminate.py: exact-hash and near-duplicate (MinHash + 8-gram overlap) comparison of every training prompt and answer against (a) the 17 programs' sources and every decompiled variant in tierb-cache/, (b) the rubric's keyword groups, (c) the claim pool, (d) the session/Rev·Deck fixtures. Required result: 0 hits; the report is committed as docs/benchmarks/training/corpus-v1-decontamination.md with counts per slice.

Deliverables

  • build_corpus_v1.sh + the S3 sources + decontaminate.py + the teacher prompt templates committed under analysis/ghidra/training/corpus/ (synthetic sources reviewed for TEST-NET addresses / reserved names / fake credentials like the benchmark fixtures)
  • docs/benchmarks/training/corpus-v1.md: counts, token totals, teacher tags + digests, agreement rates, sha256 per shard, what was rejected and why
  • the decontamination report, 0 hits
  • captured-derived data verified absent from git and from any remote

Depends on round7-1 for the local-teacher labelling (Ollama-served teachers need the card free). Blocks round7-4, round7-5, round7-6.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

analysisPayload analysis pipelineenhancementNew feature or requestllmLLM analysis workermlML worker and GPU scoring

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions