Add unified decision head: all types as score problems (issue #98 Option A) - #105
Conversation
Implement UnifiedHead that uses a single AttentionHead for noul, choice, and score via cross-attention over options. Noul is treated as 2-option softmax over "false"/"true" with logit-difference trick for BCE loss compatibility (sigmoid(b-a) = softmax([a,b])[1]). - Add UnifiedHead class to model/src/heads.py with noul_logit(), predict_noul(), predict_choice(), predict_score() methods - Add unified_head parameter to DecisionModel; when enabled, choice_head and score_head alias the unified instance, noul goes through forward_noul with "false"/"true" option encoding - Update _compute_noul_batch in supervised.py to fall back to per-item forward_noul when noul_head is None (unified mode) - Fix optimizer param group deduplication for unified/shared heads - Add --unified-head CLI flag and unified_head YAML config key - Add 14 tests covering shapes, sigmoid-softmax equivalence, param count, drop-in compatibility, gradient flow, and masking Closes #98 (Option A)
There was a problem hiding this comment.
LGTM ✅ — Clean unified head architecture.
| Component | Status |
|---|---|
| UnifiedHead | ✅ Single AttentionHead for all types via cross-attention |
| Noul as 2-option | ✅ sigmoid(b-a) = softmax([a,b])[1] identity verified |
| Head aliasing | ✅ choice_head/score_head → same UnifiedHead |
| Param dedup | ✅ Optimizer groups correctly handle aliased heads |
| Tests | ✅ 14 new tests (shapes, equivalence, gradients, drop-in) |
Note: With unified head, _compute_noul_batch falls back to per-item forward_noul calls since noul now requires cross-attention with option embeddings. This loses batch efficiency for noul training — acceptable tradeoff for architectural simplicity, but worth monitoring if noul dominates training time.
563 tests pass, CI green.
There was a problem hiding this comment.
Approved ✅
Unified decision head — all question types as score problems (issue #98 Option A):
Core Design
- Single
UnifiedHeadwraps oneAttentionHeadfor all three question types - Noul treated as 2-option softmax over "false"/"true" embeddings
- Key identity:
sigmoid(b - a) = softmax([a, b])[1]— BCE loss remains compatible
Architecture
UnifiedHead
└─ AttentionHead(Q, K, V)
├─ forward(ctx, opts, mask) → logits [batch, n_options]
├─ noul_logit(ctx, 2_opts) → [batch, 1] via logit_true - logit_false
├─ predict_choice() → softmax distribution
└─ predict_score() → (probs, expected_value)
Integration
choice_headandscore_headalias the sameUnifiedHeadinstance — existing training code unchangednoul_head = None;forward_noul()encodes "false"/"true" and calls unified cross-attention_compute_noul_batch()falls back to per-item forward when unified (cross-attention requires option embeddings)- Optimizer param group deduplication handles aliased heads correctly
Wiring
- CLI:
--unified-head - YAML:
unified_head: true - Param savings printed on init
Tests (14 new)
- Output shapes match ChoiceHead/ScoreHead
noul_logit()returns [batch, 1] compatible with BCE- Sigmoid-softmax equivalence proven numerically (atol=1e-6)
predict_*methods return valid distributions- Param count: unified < separate (NoulHead + ChoiceHead + ScoreHead)
- Drop-in replacement: accepts same args including option masking
- Gradient flows from all three types to shared attention weights
rival_awareflag propagates
CI green (563 tests).
There was a problem hiding this comment.
✅ Approved
Summary: Clean implementation of the unified decision head (Issue #98 Option A).
What I checked:
- CI: All 4 checks passing
- Local tests: 563/563 passing (ran full suite)
- UnifiedHead tests: 14 new tests covering shapes, sigmoid-softmax equivalence, gradient flow, param savings, masking, and backward compat
Architecture review:
-
UnifiedHead design — Single
AttentionHeadfor all three types with clean interfaces:forward()matches ChoiceHead/ScoreHead signature (drop-in replacement)noul_logit()uses the mathematically correctsigmoid(b-a) = softmax([a,b])[1]identity- Prediction methods mirror the type-specific heads
-
Integration — Properly handled:
choice_headandscore_headalias the unified instancenoul_head = Nonewith routing handled inforward_noul- Optimizer param group deduplication via
id()prevents double-counting
-
Backward compatibility —
unified_head=Falsepreserves existing behavior exactly -
Training code — Per-item noul forward in
_compute_noul_batchis slower than batched linear head, but necessary for cross-attention and documented in PR description. Acceptable trade-off for param savings.
LGTM 🚀
Summary
UnifiedHeadclass that uses a singleAttentionHeadfor all three question types (noul, choice, score) via cross-attention over optionssigmoid(b-a) = softmax([a,b])[1]) for BCE loss compatibilityunified_head=True,choice_headandscore_headalias the sameUnifiedHeadinstance, andnoul_headis set toNone(noul routing handled inforward_noul)_compute_noul_batch) falls back to per-itemforward_noulcalls when unified head is enabled since noul now requires cross-attention with option embeddingschoice_headandscore_headpoint to the same module--unified-headCLI flag andunified_headYAML config key fortrain_multitask.pyCloses #98 (Option A)
Test plan
python -m pytest tests/ -x -q— 563 passed)TestUnifiedHeadtests pass covering:noul_logit()returns[batch, 1]shape compatible with BCE losssigmoid(logit_true - logit_false) == softmax(logits)[1]predict_noulreturns valid probabilities in[0, 1]predict_choicereturns distribution summing to 1predict_scorereturns valid probabilities and expected value in rangerival_awareflag propagates to inner AttentionHeadtrain_multitask.py --unified-headon training server)