Fix teacher_probs quality: normalization + neutral templates + Luna gap filler - #108
Conversation
Renormalizes distributions where |sum - 1| > 0.005 during label merge. Fixes 15 Luna distributions with bad sums (all-high ~3.96, double ~1.98, truncated 0.001) and 583 near-misses as safety net.
There was a problem hiding this comment.
LGTM ✅ — Solid data quality fixes.
| Component | Status |
|---|---|
| Normalization | ✅ normalize_probs() handles all-high, double, truncated (tolerance 0.005) |
| Integration | ✅ Auto-renormalize on label merge in synthetic.py |
| Neutral templates | ✅ Direction-agnostic true/false descriptions |
| Gap filler | ✅ Dry-run + --apply modes, domain/variant reporting |
| Tests | ✅ 19 normalization tests (edge cases covered) |
Key improvements:
-
Normalization guard — catches 3 LLM failure modes:
- All-high (~3.96): LLM assigned ~0.99 to every option
- Double (~1.98): LLM confident in 2 options, both got ~0.99
- Truncated (<0.5): parsing lost options
-
Neutral templates —
"The stated condition is confirmed/contradicted by the available evidence"is direction-agnostic, avoids semantic clash with negated questions like "Is it false?" -
Luna gap filler — clean CLI design with
--domainfiltering, per-variant stats, and verification after fill.
548 tests pass, CI green. Run fill_luna_gaps.py --apply on ts-lambda10 to complete the 11,724 gap fill.
There was a problem hiding this comment.
Approved ✅
Teacher_probs quality fixes — normalization guard + neutral templates + Luna gap filler:
1. Probability Normalization Guard (data/normalize_probs.py)
Handles three failure modes observed in LLM-generated teacher labels:
| Failure Mode | Example Sum | Cause |
|---|---|---|
| All-high | ~3.96 | LLM assigned ~0.99 to every option without softmax |
| Double | ~1.98 | LLM was confident in 2 options, both got ~0.99 |
| Truncated | <0.5 | Parsing lost some options |
def normalize_probs(probs: dict[str, float]) -> dict[str, float]:
if abs(total - 1.0) <= 0.005: # tolerance
return probs # same object, no copy
return {k: v / total for k, v in probs.items()}Integrated into label merge logic — automatically fixes bad distributions on pipeline load.
2. Neutralized Noul Templates (data/synthetic_noul_augment.py)
Changed from direction-specific to direction-agnostic descriptions:
- "true": "The evidence fully supports the claim"
+ "true": "The stated condition is confirmed by the available evidence"This prevents semantic clash with negated questions (where "true" means the negation holds, not the original claim).
3. Luna Gap Filler (data/scripts/fill_luna_gaps.py)
Standalone CLI to identify and fill 11,724 missing Luna labels:
python data/scripts/fill_luna_gaps.py # dry-run: report only
python data/scripts/fill_luna_gaps.py --apply # call Luna and fill
python data/scripts/fill_luna_gaps.py --domain game_strategy --applyFeatures:
- Discovers gaps by comparing Jev vs Luna label files
- Reconstructs items from families/variants for labeling
- Appends new labels to existing
*_gpt_5_6_luna_labels.jsonl - Reports by domain, question type, and variant
Tests (19 new)
- All edge cases: all-high, double, truncated, zero, near-miss, tolerance boundaries
normalize_probs,normalize_teacher_probs,normalize_items_in_place,normalize_label_file
CI green (548 tests).
There was a problem hiding this comment.
Review: ✅ Approved
Reviewed all changes — clean implementation addressing the data quality items from #104.
Normalization Guard
- Problem: LLM-generated teacher labels had three failure modes — all-high (~3.96), double (~1.98), and truncated (<0.5) probability sums
- Fix:
normalize_probs()renormalizes when|sum - 1| > 0.005, integrated into label merge insynthetic.py - Coverage: 19 comprehensive unit tests covering all edge cases (zero total, empty dict, tolerance boundary, multi-teacher dicts)
Neutral Templates
- Problem: Old
true/falsedescriptions were semantically loaded (e.g., "safety risk exists", "exceeds threshold") and could clash with question semantics - Fix: Direction-agnostic descriptions: "The stated condition is confirmed/contradicted by the available evidence"
- Good design: non-binary options (borderline, partially_supported) keep domain-specific descriptions where precision matters
Luna Gap Filler
- Safe dry-run default,
--applymode for actual labeling - Clean domain/variant/qtype classification
- Proper verification step after apply
- Warns about unreconstructable items
Tests
- 544 tests passing locally (excluding 4 async tests that need pytest-asyncio, CI passes those)
- All existing tests continue to pass
Potential Edge Cases (not blocking)
- Negative/NaN probs would still normalize but could indicate upstream parsing issues — could add a warning log if encountered in future
- Concurrent
fill_luna_gaps.py --applyruns could append duplicates, but this is a manual CLI tool so low risk
Ready to merge. 🚀
Summary
data/normalize_probs.py— probability distribution normalization guard (renormalizes distributions where |sum - 1| > 0.005)data/synthetic.py— fixes 15 bad Luna distributions and 583 near-misses automatically on pipeline loadtrue/falsedescriptions that don't clash with question semanticsdata/scripts/fill_luna_gaps.py— standalone CLI to identify and fill 11,724 missing Luna teacher labels (dry-run +--applymodes)Addresses data quality items in #104 checklist.
Test plan
fill_luna_gaps.py --applyon ts-lambda10 to fill 11,724 missing Luna labels