Skip to content

Route augmented noul through choice head for training (Phase C) - #113

Open
Oaklight wants to merge 1 commit into
mainfrom
worktree-feature-phase-c-augmented-noul-training
Open

Oaklight wants to merge 1 commit into
mainfrom
worktree-feature-phase-c-augmented-noul-training

Conversation

@Oaklight

Copy link
Copy Markdown
Owner

Summary

Phase C of #104: training integration for augmented noul questions.

  • Teacher selection (data/pipeline.py): _select_teacher_probs() selects the best teacher from the nested multi-teacher cache (jev_augmented > luna_augmented for augmented noul, jev > luna > qwen for standard). Fixes existing dead code where synthetic teacher_probs were never consumed by the training loop.
  • Training routing (model/training/supervised.py): Augmented noul items with 5-option teacher_probs are routed through the choice head (KL-divergence against teacher distribution) instead of the binary noul head. Standard noul items unchanged.
  • Eval routing: Augmented noul predictions mapped back to binary accuracy (pred_key == "true"/"false").
  • Uncertainty weighting: Augmented noul uses "choice" log_var since it goes through the choice head.
  • 9 new tests for teacher selection, 557 total pass.

Test plan

  • 557 tests pass (including 9 new teacher selection tests)
  • Verify augmented noul items go through choice head during training
  • Confirm standard noul items still use binary path (backward compat)
  • A/B training comparison: with vs without augmented noul routing

Phase C of #104: training integration for augmented noul questions.

- Add _select_teacher_probs() to data/pipeline.py: selects best teacher
  from nested multi-teacher cache (jev_augmented > gpt_5_6_luna_augmented
  for augmented noul, jev > luna > qwen for standard items). Fixes
  existing dead code where synthetic teacher_probs were never used.
- Augmented noul items with 5-option teacher_probs are routed through
  the choice head (KL-divergence against teacher distribution) instead
  of the binary noul head, in both training and eval paths.
- Uncertainty weighting correctly uses "choice" log_var for augmented
  noul items.
- Eval maps augmented noul predictions back to binary accuracy
  (pred_key == "true"/"false").
- Standard noul fallback when teacher_probs has 2-key {"true","false"}
  format (uses true probability as BCE target).
- 9 new tests for teacher selection logic, 557 total pass.

@milo-oaklight milo-oaklight Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PR #113 (krino) — APPROVED ✅

Clean routing of augmented noul through the choice head. Teacher selection logic is well-structured with clear priority chains.

Path Teacher Priority Head Loss
Augmented noul jev_augmented > luna_augmented > standard Choice (5-way softmax) KL divergence
Standard noul jev > luna > qwen Binary noul BCE

Key changes verified:

  1. _select_teacher_probs() — Correct priority fallback; empty dicts properly skipped
  2. Routing gate — augmented_options and len(teacher_probs) > 2 cleanly separates paths
  3. Uncertainty weighting — Augmented noul correctly uses "choice" log_var
  4. Eval mapping — pred_key == target_key correctly maps 5-way predictions back to binary accuracy

Minor observation: _compute_augmented_noul_batch processes items sequentially (not true batch), matching the unified head path for standard noul. Fine for now; batch optimization can come later if needed.

Tests: 9 new tests cover teacher selection edge cases comprehensively.

CI green. Ready to merge.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant