Skip to content

Add unified decision head: all types as score problems (issue #98 Option A) - #105

Merged
Oaklight merged 1 commit into
mainfrom
feature/unified-head
Sep 27, 2026
Merged

Oaklight merged 1 commit into
mainfrom
feature/unified-head

Conversation

@Oaklight

Copy link
Copy Markdown
Owner

Summary

  • Add UnifiedHead class that uses a single AttentionHead for all three question types (noul, choice, score) via cross-attention over options
  • Noul is treated as 2-option softmax over "false"/"true" with the logit-difference trick (sigmoid(b-a) = softmax([a,b])[1]) for BCE loss compatibility
  • When unified_head=True, choice_head and score_head alias the same UnifiedHead instance, and noul_head is set to None (noul routing handled in forward_noul)
  • Training code (_compute_noul_batch) falls back to per-item forward_noul calls when unified head is enabled since noul now requires cross-attention with option embeddings
  • Optimizer param group deduplication handles the case where choice_head and score_head point to the same module
  • New --unified-head CLI flag and unified_head YAML config key for train_multitask.py
  • 14 new tests covering output shapes, sigmoid-softmax equivalence, param count savings, drop-in compatibility with ChoiceHead/ScoreHead, gradient flow through all three types, and option masking

Closes #98 (Option A)

Test plan

  • All 563 existing tests pass (python -m pytest tests/ -x -q — 563 passed)
  • 14 new TestUnifiedHead tests pass covering:
    • Output shapes match ChoiceHead/ScoreHead for choice and score types
    • noul_logit() returns [batch, 1] shape compatible with BCE loss
    • Mathematical equivalence: sigmoid(logit_true - logit_false) == softmax(logits)[1]
    • predict_noul returns valid probabilities in [0, 1]
    • predict_choice returns distribution summing to 1
    • predict_score returns valid probabilities and expected value in range
    • Param count: unified < separate (NoulHead + ChoiceHead + ScoreHead)
    • Drop-in replacement: accepts same args as ChoiceHead/ScoreHead including option masking
    • Gradient flows from all three types to shared attention weights
    • rival_aware flag propagates to inner AttentionHead
  • Integration test with real backbone (requires GPU — run train_multitask.py --unified-head on training server)

Implement UnifiedHead that uses a single AttentionHead for noul, choice,
and score via cross-attention over options.  Noul is treated as 2-option
softmax over "false"/"true" with logit-difference trick for BCE loss
compatibility (sigmoid(b-a) = softmax([a,b])[1]).

- Add UnifiedHead class to model/src/heads.py with noul_logit(),
  predict_noul(), predict_choice(), predict_score() methods
- Add unified_head parameter to DecisionModel; when enabled, choice_head
  and score_head alias the unified instance, noul goes through
  forward_noul with "false"/"true" option encoding
- Update _compute_noul_batch in supervised.py to fall back to per-item
  forward_noul when noul_head is None (unified mode)
- Fix optimizer param group deduplication for unified/shared heads
- Add --unified-head CLI flag and unified_head YAML config key
- Add 14 tests covering shapes, sigmoid-softmax equivalence, param count,
  drop-in compatibility, gradient flow, and masking

Closes #98 (Option A)

@clementine-oaklight clementine-oaklight Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM ✅ — Clean unified head architecture.

Component Status
UnifiedHead ✅ Single AttentionHead for all types via cross-attention
Noul as 2-option ✅ sigmoid(b-a) = softmax([a,b])[1] identity verified
Head aliasing ✅ choice_head/score_head → same UnifiedHead
Param dedup ✅ Optimizer groups correctly handle aliased heads
Tests ✅ 14 new tests (shapes, equivalence, gradients, drop-in)

Note: With unified head, _compute_noul_batch falls back to per-item forward_noul calls since noul now requires cross-attention with option embeddings. This loses batch efficiency for noul training — acceptable tradeoff for architectural simplicity, but worth monitoring if noul dominates training time.

563 tests pass, CI green.

@milo-oaklight milo-oaklight Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved ✅

Unified decision head — all question types as score problems (issue #98 Option A):

Core Design

  • Single UnifiedHead wraps one AttentionHead for all three question types
  • Noul treated as 2-option softmax over "false"/"true" embeddings
  • Key identity: sigmoid(b - a) = softmax([a, b])[1] — BCE loss remains compatible

Architecture

UnifiedHead
  └─ AttentionHead(Q, K, V)
        ├─ forward(ctx, opts, mask) → logits [batch, n_options]
        ├─ noul_logit(ctx, 2_opts) → [batch, 1] via logit_true - logit_false
        ├─ predict_choice() → softmax distribution
        └─ predict_score() → (probs, expected_value)

Integration

  • choice_head and score_head alias the same UnifiedHead instance — existing training code unchanged
  • noul_head = None; forward_noul() encodes "false"/"true" and calls unified cross-attention
  • _compute_noul_batch() falls back to per-item forward when unified (cross-attention requires option embeddings)
  • Optimizer param group deduplication handles aliased heads correctly

Wiring

  • CLI: --unified-head
  • YAML: unified_head: true
  • Param savings printed on init

Tests (14 new)

  • Output shapes match ChoiceHead/ScoreHead
  • noul_logit() returns [batch, 1] compatible with BCE
  • Sigmoid-softmax equivalence proven numerically (atol=1e-6)
  • predict_* methods return valid distributions
  • Param count: unified < separate (NoulHead + ChoiceHead + ScoreHead)
  • Drop-in replacement: accepts same args including option masking
  • Gradient flows from all three types to shared attention weights
  • rival_aware flag propagates

CI green (563 tests).

@Oaklight
Oaklight merged commit eb8c15d into main Sep 27, 2026
4 checks passed
@Oaklight
Oaklight deleted the feature/unified-head branch September 27, 2026 02:45

@elena-oaklight elena-oaklight Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Approved

Summary: Clean implementation of the unified decision head (Issue #98 Option A).

What I checked:

  • CI: All 4 checks passing
  • Local tests: 563/563 passing (ran full suite)
  • UnifiedHead tests: 14 new tests covering shapes, sigmoid-softmax equivalence, gradient flow, param savings, masking, and backward compat

Architecture review:

  1. UnifiedHead design — Single AttentionHead for all three types with clean interfaces:

    • forward() matches ChoiceHead/ScoreHead signature (drop-in replacement)
    • noul_logit() uses the mathematically correct sigmoid(b-a) = softmax([a,b])[1] identity
    • Prediction methods mirror the type-specific heads
  2. Integration — Properly handled:

    • choice_head and score_head alias the unified instance
    • noul_head = None with routing handled in forward_noul
    • Optimizer param group deduplication via id() prevents double-counting
  3. Backward compatibility — unified_head=False preserves existing behavior exactly

  4. Training code — Per-item noul forward in _compute_noul_batch is slower than batched linear head, but necessary for cross-attention and documented in PR description. Acceptable trade-off for param savings.

LGTM 🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Architecture: unified decision head with shared attention layers

1 participant