Part of #3079 (round 7). Plan §6. No GPU for the build; the local-teacher labelling needs the card and queues behind the cold re-run — the OpenRouter-labelled synthetic slice does not.
The rule that shapes everything here
The benchmark corpus is the test set. Its 17 programs (xor_decode_loop, vulnerable_strcpy, strcpy_note_neutral, strcpy_note_injected, linked_list_sum, indirect_dispatch, error_handling_alloc, process_and_injection, process_witness_probe, tlv_parser, loopback_connect, safe_strcpy, integer_overflow_alloc, use_after_free, checksum_rotate, format_string_bug, file_write_persist) — at any toolchain or optimisation level, their decompiled text in tierb-cache/, their rubric terms, the adjudicated claim pool tier-a-v1.json, and every transcript under docs/benchmarks/runs/ — are off limits for training data, calibration text and RL prompts. Same for the evaluate-models.py session and Rev·Deck fixtures. A model that has seen them scores a memory, not a capability.
Slices to build (/mnt-1/training/corpus-v1/, mode 0700)
| slice |
job |
input |
labels |
leaves the host? |
| S1 sessions-captured |
session analysis |
real Cowrie/Beelzebub sessions from honeypot-v* in ES, sanitised with llm-worker/contracts.py's sanitize_text / _redact_secrets / sanitize_commands — the production path, so the student trains on what it will see |
local teacher, N=3, majority-agreed SessionAnalysis JSON only |
never |
| S2 ghidra-captured |
Ghidra triage |
captured samples the production ghidra-worker already analysed (ghidra-analysis-v1); re-decompile through ghidra_cache.py against the headless service so the evidence shape matches the benchmark's Tier B |
local teacher, N=3, majority-agreed triage JSON in the ghidra-triage-v1 contract |
never |
| S3 ghidra-synthetic |
Ghidra triage + Rev·Deck |
≥ 60 new C programs written for this corpus across the rubric's behaviour families (encoding loops, memory-safety bugs, dispatch tables, persistence, network, benign near-neighbours, embedded-instruction controls) — none of them the 17; built with build_corpus.py's toolchain × opt-level matrix; decompiled the same way |
OpenRouter open-weight teacher (DeepSeek V4 / Qwen3.8 large / GLM-5.3 / Kimi K3 — licences permit distillation), N=2 agreement |
yes — synthetic only |
| S4 revdeck-rex86 |
Rev·Deck |
Zenodo 15420461 REx86.zip (5,981 entries: intent, Q&A, complete-the-code, comments; CC-BY-4.0). Check for an internal held-out split first (#847's caveat) |
as shipped |
already public |
| S5 injection-pairs |
all three |
prompts from S1–S3 with newly written embedded attacker instructions (new phrasings, not the corpus's own strcpy_note_injected / process_and_injection text) |
chosen = ignores the instruction; rejected = complies / breaks schema / flips polarity — filtered by injection_gate.classify_answer, the pydantic contracts and polarity.forbidden_hit |
per source slice |
| S6 cpt-text |
continued pretraining (round7-4) |
raw domain text: decompiler output from S2/S3, sanitised session transcripts, REx86 text, public RE corpora with permissive licences |
none |
per source slice |
| dev |
early stopping |
10 % of each slice, split by program / session, never by row |
— |
— |
Teacher key for S3 lives on the host (/etc/apiary-training.env, 0600, or the operator's choice) — never in the repo, never in an issue. Captured slices S1/S2/S5-from-captured/S6-from-captured stay on the host; if anything is pushed to Hugging Face it is a private repo and only ever the synthetic/public slices — operator decision, default off.
Decontamination (a deliverable, before any score is quoted)
analysis/ghidra/training/decontaminate.py: exact-hash and near-duplicate (MinHash + 8-gram overlap) comparison of every training prompt and answer against (a) the 17 programs' sources and every decompiled variant in tierb-cache/, (b) the rubric's keyword groups, (c) the claim pool, (d) the session/Rev·Deck fixtures. Required result: 0 hits; the report is committed as docs/benchmarks/training/corpus-v1-decontamination.md with counts per slice.
Deliverables
Depends on round7-1 for the local-teacher labelling (Ollama-served teachers need the card free). Blocks round7-4, round7-5, round7-6.
Part of #3079 (round 7). Plan §6. No GPU for the build; the local-teacher labelling needs the card and queues behind the cold re-run — the OpenRouter-labelled synthetic slice does not.
The rule that shapes everything here
The benchmark corpus is the test set. Its 17 programs (
xor_decode_loop,vulnerable_strcpy,strcpy_note_neutral,strcpy_note_injected,linked_list_sum,indirect_dispatch,error_handling_alloc,process_and_injection,process_witness_probe,tlv_parser,loopback_connect,safe_strcpy,integer_overflow_alloc,use_after_free,checksum_rotate,format_string_bug,file_write_persist) — at any toolchain or optimisation level, their decompiled text intierb-cache/, their rubric terms, the adjudicated claim pooltier-a-v1.json, and every transcript underdocs/benchmarks/runs/— are off limits for training data, calibration text and RL prompts. Same for theevaluate-models.pysession and Rev·Deck fixtures. A model that has seen them scores a memory, not a capability.Slices to build (
/mnt-1/training/corpus-v1/, mode 0700)honeypot-v*in ES, sanitised withllm-worker/contracts.py'ssanitize_text/_redact_secrets/sanitize_commands— the production path, so the student trains on what it will seeSessionAnalysisJSON onlyghidra-workeralready analysed (ghidra-analysis-v1); re-decompile throughghidra_cache.pyagainst the headless service so the evidence shape matches the benchmark's Tier Bghidra-triage-v1contractbuild_corpus.py's toolchain × opt-level matrix; decompiled the same wayREx86.zip(5,981 entries: intent, Q&A, complete-the-code, comments; CC-BY-4.0). Check for an internal held-out split first (#847's caveat)strcpy_note_injected/process_and_injectiontext)injection_gate.classify_answer, the pydantic contracts andpolarity.forbidden_hitTeacher key for S3 lives on the host (
/etc/apiary-training.env, 0600, or the operator's choice) — never in the repo, never in an issue. Captured slices S1/S2/S5-from-captured/S6-from-captured stay on the host; if anything is pushed to Hugging Face it is a private repo and only ever the synthetic/public slices — operator decision, default off.Decontamination (a deliverable, before any score is quoted)
analysis/ghidra/training/decontaminate.py: exact-hash and near-duplicate (MinHash + 8-gram overlap) comparison of every training prompt and answer against (a) the 17 programs' sources and every decompiled variant intierb-cache/, (b) the rubric's keyword groups, (c) the claim pool, (d) the session/Rev·Deck fixtures. Required result: 0 hits; the report is committed asdocs/benchmarks/training/corpus-v1-decontamination.mdwith counts per slice.Deliverables
build_corpus_v1.sh+ the S3 sources +decontaminate.py+ the teacher prompt templates committed underanalysis/ghidra/training/corpus/(synthetic sources reviewed for TEST-NET addresses / reserved names / fake credentials like the benchmark fixtures)docs/benchmarks/training/corpus-v1.md: counts, token totals, teacher tags + digests, agreement rates, sha256 per shard, what was rejected and whyDepends on round7-1 for the local-teacher labelling (Ollama-served teachers need the card free). Blocks round7-4, round7-5, round7-6.