Skip to content

Fix teacher_probs quality: normalization + neutral templates + Luna gap filler - #108

Merged
Oaklight merged 3 commits into
mainfrom
worktree-fix+teacher-probs-quality
Sep 27, 2026
Merged

Oaklight merged 3 commits into
mainfrom
worktree-fix+teacher-probs-quality

Conversation

@Oaklight

Copy link
Copy Markdown
Owner

Summary

  • Add data/normalize_probs.py — probability distribution normalization guard (renormalizes distributions where |sum - 1| > 0.005)
  • Integrate normalization into label merge logic in data/synthetic.py — fixes 15 bad Luna distributions and 583 near-misses automatically on pipeline load
  • Neutralize noul augmentation template descriptions — direction-agnostic true/false descriptions that don't clash with question semantics
  • Add data/scripts/fill_luna_gaps.py — standalone CLI to identify and fill 11,724 missing Luna teacher labels (dry-run + --apply modes)

Addresses data quality items in #104 checklist.

Test plan

  • 19 normalization unit tests (all edge cases: all-high, double, truncated, zero, near-miss)
  • 35 noul augment tests still pass with neutralized templates
  • Full suite: 548 passed, 3 skipped
  • Run fill_luna_gaps.py --apply on ts-lambda10 to fill 11,724 missing Luna labels
  • Re-run probability sum verification after Luna gap fill

Renormalizes distributions where |sum - 1| > 0.005 during label merge.
Fixes 15 Luna distributions with bad sums (all-high ~3.96, double ~1.98,
truncated 0.001) and 583 near-misses as safety net.

@clementine-oaklight clementine-oaklight Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM ✅ — Solid data quality fixes.

Component Status
Normalization ✅ normalize_probs() handles all-high, double, truncated (tolerance 0.005)
Integration ✅ Auto-renormalize on label merge in synthetic.py
Neutral templates ✅ Direction-agnostic true/false descriptions
Gap filler ✅ Dry-run + --apply modes, domain/variant reporting
Tests ✅ 19 normalization tests (edge cases covered)

Key improvements:

  1. Normalization guard — catches 3 LLM failure modes:

    • All-high (~3.96): LLM assigned ~0.99 to every option
    • Double (~1.98): LLM confident in 2 options, both got ~0.99
    • Truncated (<0.5): parsing lost options
  2. Neutral templates — "The stated condition is confirmed/contradicted by the available evidence" is direction-agnostic, avoids semantic clash with negated questions like "Is it false?"

  3. Luna gap filler — clean CLI design with --domain filtering, per-variant stats, and verification after fill.

548 tests pass, CI green. Run fill_luna_gaps.py --apply on ts-lambda10 to complete the 11,724 gap fill.

@milo-oaklight milo-oaklight Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved ✅

Teacher_probs quality fixes — normalization guard + neutral templates + Luna gap filler:

1. Probability Normalization Guard (data/normalize_probs.py)

Handles three failure modes observed in LLM-generated teacher labels:

Failure Mode Example Sum Cause
All-high ~3.96 LLM assigned ~0.99 to every option without softmax
Double ~1.98 LLM was confident in 2 options, both got ~0.99
Truncated <0.5 Parsing lost some options
def normalize_probs(probs: dict[str, float]) -> dict[str, float]:
    if abs(total - 1.0) <= 0.005:  # tolerance
        return probs  # same object, no copy
    return {k: v / total for k, v in probs.items()}

Integrated into label merge logic — automatically fixes bad distributions on pipeline load.

2. Neutralized Noul Templates (data/synthetic_noul_augment.py)

Changed from direction-specific to direction-agnostic descriptions:

- "true": "The evidence fully supports the claim"
+ "true": "The stated condition is confirmed by the available evidence"

This prevents semantic clash with negated questions (where "true" means the negation holds, not the original claim).

3. Luna Gap Filler (data/scripts/fill_luna_gaps.py)

Standalone CLI to identify and fill 11,724 missing Luna labels:

python data/scripts/fill_luna_gaps.py              # dry-run: report only
python data/scripts/fill_luna_gaps.py --apply       # call Luna and fill
python data/scripts/fill_luna_gaps.py --domain game_strategy --apply

Features:

  • Discovers gaps by comparing Jev vs Luna label files
  • Reconstructs items from families/variants for labeling
  • Appends new labels to existing *_gpt_5_6_luna_labels.jsonl
  • Reports by domain, question type, and variant

Tests (19 new)

  • All edge cases: all-high, double, truncated, zero, near-miss, tolerance boundaries
  • normalize_probs, normalize_teacher_probs, normalize_items_in_place, normalize_label_file

CI green (548 tests).

@Oaklight
Oaklight merged commit 7c7eec2 into main Sep 27, 2026
4 checks passed
@Oaklight
Oaklight deleted the worktree-fix+teacher-probs-quality branch September 27, 2026 04:22

@elena-oaklight elena-oaklight Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: ✅ Approved

Reviewed all changes — clean implementation addressing the data quality items from #104.

Normalization Guard

  • Problem: LLM-generated teacher labels had three failure modes — all-high (~3.96), double (~1.98), and truncated (<0.5) probability sums
  • Fix: normalize_probs() renormalizes when |sum - 1| > 0.005, integrated into label merge in synthetic.py
  • Coverage: 19 comprehensive unit tests covering all edge cases (zero total, empty dict, tolerance boundary, multi-teacher dicts)

Neutral Templates

  • Problem: Old true/false descriptions were semantically loaded (e.g., "safety risk exists", "exceeds threshold") and could clash with question semantics
  • Fix: Direction-agnostic descriptions: "The stated condition is confirmed/contradicted by the available evidence"
  • Good design: non-binary options (borderline, partially_supported) keep domain-specific descriptions where precision matters

Luna Gap Filler

  • Safe dry-run default, --apply mode for actual labeling
  • Clean domain/variant/qtype classification
  • Proper verification step after apply
  • Warns about unreconstructable items

Tests

  • 544 tests passing locally (excluding 4 async tests that need pytest-asyncio, CI passes those)
  • All existing tests continue to pass

Potential Edge Cases (not blocking)

  • Negative/NaN probs would still normalize but could indicate upstream parsing issues — could add a warning log if encountered in future
  • Concurrent fill_luna_gaps.py --apply runs could append duplicates, but this is a manual CLI tool so low risk

Ready to merge. 🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant