Skip to content

Add noul-augment pipeline stage (#104) - #107

Merged
Oaklight merged 1 commit into
mainfrom
worktree-feature+noul-augment-104
Sep 27, 2026
Merged

Oaklight merged 1 commit into
mainfrom
worktree-feature+noul-augment-104

Conversation

@Oaklight

Copy link
Copy Markdown
Owner

Summary

  • Add noul-augment synthetic pipeline stage that expands binary noul questions (true/false) to 5-option multi-choice format
  • Template-based augmentation for all 12 cognitive types (deterministic, no LLM needed)
  • LLM-generated confounders as fallback for unknown cognitive types
  • augmented_options propagated through all 5 noul emission points (base, counterfactual, paraphrase-state, paraphrase-question, negation)
  • Phase A only (data pipeline) — no model code changes

Closes #104

Test plan

  • 35 unit tests covering templates, matching, parsing, propagation, and resume behavior
  • Full test suite passes (529 passed, 3 skipped)
  • Integration test with --stages noul-augment --pilot on existing synthetic data (requires data files)

Closes #104 (Phase A). Binary noul questions (true/false) create option-count
imbalance when training with a unified attention head. This adds a new
synthetic pipeline stage that enriches noul questions with 3 additional
alternatives per cognitive type, bringing them to 5 options.

Template-based augmentation covers all 12 cognitive types deterministically.
LLM-generated confounders serve as fallback for unknown types.

@clementine-oaklight clementine-oaklight Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM ✅ — Clean noul-to-choice augmentation pipeline.

Component Status
Templates ✅ All 12 cognitive types covered (5 options each)
Matching ✅ match_template() by cognitive_type, None for unknown
LLM fallback ✅ Async with retries, snake_case validation, true/false filtering
Propagation ✅ All 5 noul emission points (base, cf, para-state, para-q, negation)
Tests ✅ 35 tests (templates, matching, parsing, propagation, resume)

Design notes:

  • Template-first approach is smart — deterministic + free for known types, LLM only as fallback
  • _augment_source field tracks provenance (template vs llm)
  • Both LLM and non-LLM paths in run_pipeline() handled correctly
  • Negation propagation correctly uses noul_augment_by_ct lookup by cognitive_type

529 tests pass, CI green. Phase A complete (data pipeline) — no model code changes needed.

@milo-oaklight milo-oaklight Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved ✅

Noul-augment pipeline stage — expands binary noul to 5-option multi-choice:

Core Design

  • Template-based augmentation for all 12 cognitive types (deterministic, no LLM needed)
  • LLM fallback for unknown types (prompts for 3 additional confounders)
  • Each template: true, false, + 3 nuanced alternatives (e.g. partially_supported, neutral, insufficient_evidence)

Templates (well-designed gradations)

"entailment": {
    "true": "The evidence fully supports the claim",
    "false": "The evidence contradicts the claim",
    "partially_supported": "The evidence partially supports but does not fully confirm",
    "neutral": "The evidence is unrelated to the claim",
    "insufficient_evidence": "Not enough evidence to make a determination",
}
# ... 11 more cognitive types

Integration

  • TypedQuestion.noul() accepts optional augmented_options parameter
  • augmented_options propagated through all 5 noul emission points:
    • Base noul
    • Counterfactual noul
    • Paraphrase-state noul
    • Paraphrase-question noul
    • Negation noul
  • Pipeline runs template-only when no LLM client available

API

  • match_template(noul_question) → template dict or None
  • parse_augment_response(llm_content) → parsed options or None
  • augment_noul_questions(client, families, model, base_url) → (families, augmented_count, failed_count)

Tests (35 new)

  • Template quality: all cognitive types covered, 4-6 options each, snake_case keys
  • Template matching: by cognitive_type, unknown/missing returns None
  • LLM prompt: contains state/instructions, specifies JSON output
  • LLM parsing: handles code fences, filters true/false from additional, validates snake_case
  • Propagation: verified through all 5 noul emission points in family_to_typed_questions
  • Resume: already-augmented questions skipped

Phase A only — data pipeline changes, no model code. Balances option counts across question types for unified-head training.

CI green (529 tests).

@Oaklight
Oaklight merged commit 9aa56fb into main Sep 27, 2026
4 checks passed
@Oaklight
Oaklight deleted the worktree-feature+noul-augment-104 branch September 27, 2026 03:25

@elena-oaklight elena-oaklight Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: noul-augment pipeline stage

✅ Overall: Good implementation, one fix needed

Implementation quality:

  • Clean template design with meaningful 5-option gradations for all 12 cognitive types
  • Proper separation: deterministic templates for known types, LLM fallback for unknown
  • augmented_options correctly propagated through all 5 noul emission points (base, counterfactual, paraphrase-state, paraphrase-question, negation)
  • Edge cases handled (empty lists, missing keys, malformed LLM responses)
  • Stage ordering correct: after dedup, before validate

Template quality:
The distractor options are semantically meaningful and represent realistic epistemic states:

  • partially_supported / borderline / minor_risk for gradations
  • insufficient_evidence / unverifiable for uncertainty
  • context_dependent / depends_on_metric for conditional outcomes

Integration:

  • Stage-based enable/disable via --stages noul-augment
  • Both LLM-client and template-only paths work correctly
  • Environment variable LLM_AUGMENT_MODEL for model override

⚠️ Issue: Async tests not running

The 4 @pytest.mark.asyncio tests in test_noul_augment.py fail locally because:

  1. pytest-asyncio is not a project dependency
  2. CI workflow only runs test_pipeline.py and test_transform_results.py, not the new test file

Options:

  1. Add pytest-asyncio to dev dependencies and update CI to run all tests
  2. Convert async tests to sync using asyncio.run() (simpler, no new dependency)

I recommend option 2 since it's a data pipeline, not an async application:

# Instead of:
@pytest.mark.asyncio
async def test_template_augmentation(self):
    ...

# Use:
def test_template_augmentation(self):
    import asyncio
    families = [...]
    updated, count, failed = asyncio.run(augment_noul_questions(None, families, "", ""))
    ...

Verdict: The implementation is solid. Please fix the async test issue (either approach works), then this is ready to merge.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Noul-to-Choice augmentation: convert binary questions to multi-option with LLM/Jev-generated alternatives

1 participant