Skip to content

Improve LLM teacher labeling prompts for calibration - #109

Merged
Oaklight merged 1 commit into
mainfrom
worktree-fix+llm-label-prompts
Sep 27, 2026
Merged

Oaklight merged 1 commit into
mainfrom
worktree-fix+llm-label-prompts

Conversation

@Oaklight

Copy link
Copy Markdown
Owner

Summary

  • Add _SYSTEM_CONTEXT explaining the faithful decision model training goal — knowledge distillation with calibrated soft labels
  • Noul prompt: guide calibration, discourage extremes (>0.95 / <0.05) unless evidence is decisive
  • Choice prompt: explicitly state "MUTUALLY EXCLUSIVE categorical distribution, NOT independent confidence scores" — fixes the root cause of 15 bad Luna distributions where all options got 0.99
  • Score prompt: emphasize categorical distribution over ordinal levels, allow adjacent level probability sharing

Addresses root cause of data quality issues found in #104 checklist.

Test plan

  • Full test suite: 548 passed, 3 skipped
  • Re-run Luna labeling with improved prompts, compare overconfidence rate (target: <5% vs current 16.3%)
  • Verify probability sums after re-labeling

Add system context explaining the faithful decision model training goal.
Explicitly instruct mutually exclusive categorical distributions (not
independent confidence scores) to prevent all-high probability failures.
Guide calibration: encourage genuine uncertainty, discourage extremes.

@milo-oaklight milo-oaklight Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved ✅

LLM teacher labeling prompt improvements for calibration:

Root Cause Fix

The 15 bad Luna distributions (all options ~0.99, sum ~3.96) came from the LLM treating options as independent confidence scores. The prompt now explicitly clarifies:

Distribute probability across these MUTUALLY EXCLUSIVE options.
The probabilities MUST sum to exactly 1.0 — this is a categorical
distribution, NOT independent confidence scores.

Changes by Question Type

Type Key Prompt Addition
System context Explains knowledge distillation goal — "calibrated probability distributions... soft teacher labels"
Noul "Express genuine uncertainty — values like 0.75 or 0.3 are valid... Reserve extremes (>0.95 or <0.05) for cases where evidence is truly decisive"
Choice "MUTUALLY EXCLUSIVE categorical distribution, NOT independent confidence scores"
Score "categorical distribution over ordinal levels. Adjacent levels may share probability when the rating is borderline"

Expected Impact

  • Target overconfidence rate: <5% (vs current 16.3%)
  • Eliminates all-high failure mode at source (no need for post-hoc normalization)

CI green (548 tests). Needs re-run of Luna labeling with improved prompts to measure calibration improvement.

@clementine-oaklight clementine-oaklight Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM ✅ — Root cause fix for teacher label quality issues.

Component Status
System context ✅ Explains knowledge distillation goal, emphasizes calibration
Noul prompt ✅ Discourages extremes (>0.95/<0.05), encourages 0.75/0.3-style uncertainty
Choice prompt ✅ MUTUALLY EXCLUSIVE categorical distribution, NOT independent confidence scores
Score prompt ✅ Categorical over ordinal levels, adjacent probability sharing
Tests ✅ 548 passed

This is the upstream fix for #108's normalization guard:

PR Role
#108 Symptom fix — normalize bad distributions after the fact
#109 Root cause — prevent bad distributions at generation time

The choice prompt change is particularly good — explicitly calling out "NOT independent confidence scores" directly addresses the failure mode where Luna assigned 0.99 to every option.

CI green. After re-labeling with these prompts, target is <5% overconfidence rate (down from 16.3%).

@elena-oaklight elena-oaklight Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review Summary

✅ Approved — this PR addresses the root cause of the 15 bad Luna distributions documented in #104.

CI Status

All 4 checks pass:

  • test (3.10) ✅
  • test (3.12) ✅
  • build-web ✅
  • GitGuardian Security Checks ✅

Prompt Changes Analysis

The changes are well-targeted:

Prompt Before After Impact
_SYSTEM_CONTEXT (none) Explains knowledge distillation goal, emphasizes both accuracy and calibration Sets the right mental model for the LLM
noul Generic "evaluating whether a proposition is true or false" Explicit guidance: "Express genuine uncertainty — values like 0.75 or 0.3 are valid... Reserve extremes (>0.95 or <0.05) for decisive evidence" Directly combats overconfidence
choice Vague "evaluating which option best applies" + "Probabilities must sum to 1.0" at the end "MUTUALLY EXCLUSIVE categorical distribution, NOT independent confidence scores" + sum constraint in context Fixes the root cause — LLM was treating options as independent binary judgments
score Generic ordinal rating Same categorical distribution framing + "Adjacent levels may share probability when borderline" Better calibration for ordinal scales

Root Cause Fix

The choice prompt change is the key fix. The phrase "NOT independent confidence scores" directly addresses why Luna was outputting {"a": 0.99, "b": 0.99, "c": 0.01, "d": 0.01} — it was treating each option as a separate binary confidence rather than a single distribution.

Test Coverage

No direct unit tests for the prompt builders exist. This is acceptable because:

  1. The prompts are pure string formatting (no complex logic)
  2. The real validation is empirical — re-running Luna labeling and measuring overconfidence rate

Backward Compatibility

✅ No breaking changes — existing labeled data remains valid. The normalize_probs.py module (added in recent commits) handles renormalization at load time.

Recommendation

Merge, then re-run Luna labeling to validate the <5% overconfidence target vs current 16.3%.

@Oaklight
Oaklight merged commit 85d6bfa into main Sep 27, 2026
4 checks passed
@Oaklight
Oaklight deleted the worktree-fix+llm-label-prompts branch September 27, 2026 06:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant