Repository navigation
Improve LLM teacher labeling prompts for calibration - #109
Conversation
Add system context explaining the faithful decision model training goal. Explicitly instruct mutually exclusive categorical distributions (not independent confidence scores) to prevent all-high probability failures. Guide calibration: encourage genuine uncertainty, discourage extremes.
There was a problem hiding this comment.
Approved ✅
LLM teacher labeling prompt improvements for calibration:
Root Cause Fix
The 15 bad Luna distributions (all options ~0.99, sum ~3.96) came from the LLM treating options as independent confidence scores. The prompt now explicitly clarifies:
Distribute probability across these MUTUALLY EXCLUSIVE options.
The probabilities MUST sum to exactly 1.0 — this is a categorical
distribution, NOT independent confidence scores.
Changes by Question Type
| Type | Key Prompt Addition |
|---|---|
| System context | Explains knowledge distillation goal — "calibrated probability distributions... soft teacher labels" |
| Noul | "Express genuine uncertainty — values like 0.75 or 0.3 are valid... Reserve extremes (>0.95 or <0.05) for cases where evidence is truly decisive" |
| Choice | "MUTUALLY EXCLUSIVE categorical distribution, NOT independent confidence scores" |
| Score | "categorical distribution over ordinal levels. Adjacent levels may share probability when the rating is borderline" |
Expected Impact
- Target overconfidence rate: <5% (vs current 16.3%)
- Eliminates all-high failure mode at source (no need for post-hoc normalization)
CI green (548 tests). Needs re-run of Luna labeling with improved prompts to measure calibration improvement.
There was a problem hiding this comment.
LGTM ✅ — Root cause fix for teacher label quality issues.
| Component | Status |
|---|---|
| System context | ✅ Explains knowledge distillation goal, emphasizes calibration |
| Noul prompt | ✅ Discourages extremes (>0.95/<0.05), encourages 0.75/0.3-style uncertainty |
| Choice prompt | ✅ MUTUALLY EXCLUSIVE categorical distribution, NOT independent confidence scores |
| Score prompt | ✅ Categorical over ordinal levels, adjacent probability sharing |
| Tests | ✅ 548 passed |
This is the upstream fix for #108's normalization guard:
| PR | Role |
|---|---|
| #108 | Symptom fix — normalize bad distributions after the fact |
| #109 | Root cause — prevent bad distributions at generation time |
The choice prompt change is particularly good — explicitly calling out "NOT independent confidence scores" directly addresses the failure mode where Luna assigned 0.99 to every option.
CI green. After re-labeling with these prompts, target is <5% overconfidence rate (down from 16.3%).
There was a problem hiding this comment.
Review Summary
✅ Approved — this PR addresses the root cause of the 15 bad Luna distributions documented in #104.
CI Status
All 4 checks pass:
- test (3.10) ✅
- test (3.12) ✅
- build-web ✅
- GitGuardian Security Checks ✅
Prompt Changes Analysis
The changes are well-targeted:
| Prompt | Before | After | Impact |
|---|---|---|---|
_SYSTEM_CONTEXT |
(none) | Explains knowledge distillation goal, emphasizes both accuracy and calibration | Sets the right mental model for the LLM |
noul |
Generic "evaluating whether a proposition is true or false" | Explicit guidance: "Express genuine uncertainty — values like 0.75 or 0.3 are valid... Reserve extremes (>0.95 or <0.05) for decisive evidence" | Directly combats overconfidence |
choice |
Vague "evaluating which option best applies" + "Probabilities must sum to 1.0" at the end | "MUTUALLY EXCLUSIVE categorical distribution, NOT independent confidence scores" + sum constraint in context | Fixes the root cause — LLM was treating options as independent binary judgments |
score |
Generic ordinal rating | Same categorical distribution framing + "Adjacent levels may share probability when borderline" | Better calibration for ordinal scales |
Root Cause Fix
The choice prompt change is the key fix. The phrase "NOT independent confidence scores" directly addresses why Luna was outputting {"a": 0.99, "b": 0.99, "c": 0.01, "d": 0.01} — it was treating each option as a separate binary confidence rather than a single distribution.
Test Coverage
No direct unit tests for the prompt builders exist. This is acceptable because:
- The prompts are pure string formatting (no complex logic)
- The real validation is empirical — re-running Luna labeling and measuring overconfidence rate
Backward Compatibility
✅ No breaking changes — existing labeled data remains valid. The normalize_probs.py module (added in recent commits) handles renormalization at load time.
Recommendation
Merge, then re-run Luna labeling to validate the <5% overconfidence target vs current 16.3%.
Summary
_SYSTEM_CONTEXTexplaining the faithful decision model training goal — knowledge distillation with calibrated soft labelsAddresses root cause of data quality issues found in #104 checklist.
Test plan