Skip to content

Fix manufactured confidence on ambiguous inputs with one-vs-all distillation - #6

Merged
ealmloff merged 2 commits into
mainfrom
issue5-ova-calibration
Jul 24, 2026
Merged

ealmloff merged 2 commits into
mainfrom
issue5-ova-calibration

Conversation

@ealmloff

@ealmloff ealmloff commented Jul 24, 2026 •

Copy link
Copy Markdown
Member

The fix retrains the shipped checkpoint on a rebuilt public corpus with a target scheme that never renormalizes away out-of-head mass (cf. "Revisiting One-vs-All Classifiers for Predictive Uncertainty and OOD Detection", Padhy et al. 2020):

  • Cache the teacher's raw per-class head marginals (--head-marginal-targets) and distill them with per-class sigmoid cross-entropy (--soft-loss-mode bce). Out-of-scope mass simply lowers every class target instead of being renormalized into false confidence.
  • Discount the hard argmax target by the teacher's in-head mass (--mass-discounted-hard-labels) and drop rows whose in-head mass is noise (--min-teacher-head-mass 0.1).
  • Self-distill from a ~2x larger one-vs-all parent (wordseq-b1536-k3-m2048-med-3conv-hidden, 0.9492 test teacher parity) whose sigmoid marginals are cached via the new scripts/cache_self_distill.py.
  • The runtime is unchanged: softmax over BCE-trained per-class log-odds is near-uniform when every class is unlikely and confident when one dominates.

The corpus is rebuilt from public sources with the original layout (scripts/build_finetune_corpus.py): bigcode/the-stack-smol-xl and bigcode/the-stack per-language samples, GitHub repos for labels absent from The Stack (objectivec, gradle, gemfile), and a small synthetic set covering the issue #5 ambiguity, including train-only bare dash lists whose teacher targets are genuinely split. Files from one repository always land in the same split.

Results on the exact issue reproduction: Markdown 0.86 / Yaml 0.06 (was Yaml 0.91 / Markdown 0.07). Genuinely balanced inputs now report split probabilities: a bare - first\n- second list reads Yaml 0.54 / Markdown 0.28 (was 0.99), name lists 0.43/0.39, and the median top-1 over 200 sampled bare lists drops 0.92 -> 0.51. Aggregates improve across the board on the rebuilt held-out split versus the previous
artifact: fs accuracy 0.9262 -> 0.9424, teacher parity 0.9221 ->
0.9441, macro recall 0.9299 -> 0.9397; tree-mode accuracy on local repos: dioxus 95.6% -> 97.1%, wasm-bindgen 90.2% -> 95.4%, docsite 91.2% -> 93.9%.

Also included:

  • Regression tests for the issue snippet, capitalized name lists, bare-list uncertainty (top-2 = {yaml, markdown}, top-1 < 0.9), and YAML guards.
  • scripts/build_fs_labels.py and scripts/cache_self_distill.py, both previously referenced by the tooling but missing from the repo, and scripts/make_pruned48_config.py to regenerate the pruned teacher config from the magika pip package.
  • Teacher-vetted synthetic hard-boundary generators (hard_gen_*.py) for the remaining student-error confusion cells, with negative A/B results documented: at this capacity the c/cpp and js/ts sets only relocate errors inside genuinely arbitrary boundaries, so the shipped recipe does not sample them.
  • Trainer support for all of the above plus --checkpoint-weights (npz save/resume for non-exportable parent architectures), with the legacy renormalizing recipe preserved as the default.
  • Docs: README/MODEL_CARD metrics and SHA-256 for the new artifact, an analysis of the remaining ambiguous pairs in the confusion matrix, the full reproducible fine-tune recipe in TRAINING.md (including A/B findings: FP-then-QAT collapses on short schedules; hard-loss 0.5 re-sharpens ambiguous rows on long ones), and regenerated confusion reports/images.

README confusion images are pinned to this branch name; re-pin to the merge commit after landing, matching the existing convention.

Fixes #5


Devin Review

Status Commit
⚪ Not started —

Run Devin Review

Open in Devin Review (Staging)

ealmloff and others added 2 commits July 23, 2026 18:29
…llation (#5)

Issue #5: `# Heading` + a single-word dash list scored Yaml 0.91 /
Markdown 0.07 even though the Magika teacher reads it as markdown 0.96.
Root cause: the teacher predicts over 214 labels, and the training
cache renormalized its probabilities over the 48 exported labels.
That conditions on "the input is one of the head labels" and
manufactures confident targets for inputs the teacher mostly places on
txt/unknown - a bare `- item` list (valid YAML and valid Markdown)
trained toward yaml 0.9+ while the teacher kept ~70% of its mass
out-of-head.

The fix retrains the shipped checkpoint on a rebuilt public corpus
with a target scheme that never renormalizes away out-of-head mass
(cf. "Revisiting One-vs-All Classifiers for Predictive Uncertainty and
OOD Detection", Padhy et al. 2020):

- Cache the teacher's raw per-class head marginals
  (--head-marginal-targets) and distill them with per-class sigmoid
  cross-entropy (--soft-loss-mode bce). Out-of-scope mass simply
  lowers every class target instead of being renormalized into false
  confidence.
- Discount the hard argmax target by the teacher's in-head mass
  (--mass-discounted-hard-labels) and drop rows whose in-head mass is
  noise (--min-teacher-head-mass 0.1).
- Self-distill from a ~2x larger one-vs-all parent
  (wordseq-b1536-k3-m2048-med-3conv-hidden, 0.9492 test teacher
  parity) whose sigmoid marginals are cached via the new
  scripts/cache_self_distill.py.
- The runtime is unchanged: softmax over BCE-trained per-class
  log-odds is near-uniform when every class is unlikely and confident
  when one dominates.

The corpus is rebuilt from public sources with the original layout
(scripts/build_finetune_corpus.py): bigcode/the-stack-smol-xl and
bigcode/the-stack per-language samples, GitHub repos for labels absent
from The Stack (objectivec, gradle, gemfile), and a small synthetic
set covering the issue #5 ambiguity, including train-only bare dash
lists whose teacher targets are genuinely split. Files from one
repository always land in the same split.

Results on the exact issue reproduction: Markdown 0.86 / Yaml 0.06
(was Yaml 0.91 / Markdown 0.07). Genuinely balanced inputs now report
split probabilities: a bare `- first\n- second` list reads Yaml 0.54 /
Markdown 0.28 (was 0.99), name lists 0.43/0.39, and the median top-1
over 200 sampled bare lists drops 0.92 -> 0.51. Aggregates improve
across the board on the rebuilt held-out split versus the previous
artifact: fs accuracy 0.9262 -> 0.9424, teacher parity 0.9221 ->
0.9441, macro recall 0.9299 -> 0.9397; tree-mode accuracy on local
repos: dioxus 95.6% -> 97.1%, wasm-bindgen 90.2% -> 95.4%, docsite
91.2% -> 93.9%.

Also included:
- Regression tests for the issue snippet, capitalized name lists,
  bare-list uncertainty (top-2 = {yaml, markdown}, top-1 < 0.9), and
  YAML guards.
- scripts/build_fs_labels.py and scripts/cache_self_distill.py, both
  previously referenced by the tooling but missing from the repo, and
  scripts/make_pruned48_config.py to regenerate the pruned teacher
  config from the magika pip package.
- Teacher-vetted synthetic hard-boundary generators (hard_gen_*.py)
  for the remaining student-error confusion cells, with negative A/B
  results documented: at this capacity the c/cpp and js/ts sets only
  relocate errors inside genuinely arbitrary boundaries, so the
  shipped recipe does not sample them.
- Trainer support for all of the above plus --checkpoint-weights (npz
  save/resume for non-exportable parent architectures), with the
  legacy renormalizing recipe preserved as the default.
- Docs: README/MODEL_CARD metrics and SHA-256 for the new artifact,
  an analysis of the remaining ambiguous pairs in the confusion
  matrix, the full reproducible fine-tune recipe in TRAINING.md
  (including A/B findings: FP-then-QAT collapses on short schedules;
  hard-loss 0.5 re-sharpens ambiguous rows on long ones), and
  regenerated confusion reports/images.

README confusion images are pinned to this branch name; re-pin to the
merge commit after landing, matching the existing convention.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <166158716+staging-devin-ai-integration[bot]@users.noreply.github.com>
The model inspects just the first and last 4096 bytes of a file, but
the example read every file in full, so tree walks over directories
with large artifacts spent their time streaming bytes the model never
looks at. Read the head and tail blocks and report the real file size
instead; output is byte-identical on full-tree runs.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <166158716+staging-devin-ai-integration[bot]@users.noreply.github.com>
@ealmloff
ealmloff merged commit ee77127 into main Jul 24, 2026
5 checks passed
@ealmloff ealmloff mentioned this pull request Jul 24, 2026
3 tasks done
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Very high YAML score for an ambiguous Markdown/YAML snippet

1 participant