Repository navigation
Add noul-augment pipeline stage (#104) - #107
Conversation
Closes #104 (Phase A). Binary noul questions (true/false) create option-count imbalance when training with a unified attention head. This adds a new synthetic pipeline stage that enriches noul questions with 3 additional alternatives per cognitive type, bringing them to 5 options. Template-based augmentation covers all 12 cognitive types deterministically. LLM-generated confounders serve as fallback for unknown types.
There was a problem hiding this comment.
LGTM ✅ — Clean noul-to-choice augmentation pipeline.
| Component | Status |
|---|---|
| Templates | ✅ All 12 cognitive types covered (5 options each) |
| Matching | ✅ match_template() by cognitive_type, None for unknown |
| LLM fallback | ✅ Async with retries, snake_case validation, true/false filtering |
| Propagation | ✅ All 5 noul emission points (base, cf, para-state, para-q, negation) |
| Tests | ✅ 35 tests (templates, matching, parsing, propagation, resume) |
Design notes:
- Template-first approach is smart — deterministic + free for known types, LLM only as fallback
_augment_sourcefield tracks provenance (template vs llm)- Both LLM and non-LLM paths in
run_pipeline()handled correctly - Negation propagation correctly uses
noul_augment_by_ctlookup by cognitive_type
529 tests pass, CI green. Phase A complete (data pipeline) — no model code changes needed.
There was a problem hiding this comment.
Approved ✅
Noul-augment pipeline stage — expands binary noul to 5-option multi-choice:
Core Design
- Template-based augmentation for all 12 cognitive types (deterministic, no LLM needed)
- LLM fallback for unknown types (prompts for 3 additional confounders)
- Each template:
true,false, + 3 nuanced alternatives (e.g.partially_supported,neutral,insufficient_evidence)
Templates (well-designed gradations)
"entailment": {
"true": "The evidence fully supports the claim",
"false": "The evidence contradicts the claim",
"partially_supported": "The evidence partially supports but does not fully confirm",
"neutral": "The evidence is unrelated to the claim",
"insufficient_evidence": "Not enough evidence to make a determination",
}
# ... 11 more cognitive typesIntegration
TypedQuestion.noul()accepts optionalaugmented_optionsparameteraugmented_optionspropagated through all 5 noul emission points:- Base noul
- Counterfactual noul
- Paraphrase-state noul
- Paraphrase-question noul
- Negation noul
- Pipeline runs template-only when no LLM client available
API
match_template(noul_question)→ template dict or Noneparse_augment_response(llm_content)→ parsed options or Noneaugment_noul_questions(client, families, model, base_url)→ (families, augmented_count, failed_count)
Tests (35 new)
- Template quality: all cognitive types covered, 4-6 options each, snake_case keys
- Template matching: by cognitive_type, unknown/missing returns None
- LLM prompt: contains state/instructions, specifies JSON output
- LLM parsing: handles code fences, filters true/false from additional, validates snake_case
- Propagation: verified through all 5 noul emission points in
family_to_typed_questions - Resume: already-augmented questions skipped
Phase A only — data pipeline changes, no model code. Balances option counts across question types for unified-head training.
CI green (529 tests).
There was a problem hiding this comment.
Review: noul-augment pipeline stage
✅ Overall: Good implementation, one fix needed
Implementation quality:
- Clean template design with meaningful 5-option gradations for all 12 cognitive types
- Proper separation: deterministic templates for known types, LLM fallback for unknown
augmented_optionscorrectly propagated through all 5 noul emission points (base, counterfactual, paraphrase-state, paraphrase-question, negation)- Edge cases handled (empty lists, missing keys, malformed LLM responses)
- Stage ordering correct: after dedup, before validate
Template quality:
The distractor options are semantically meaningful and represent realistic epistemic states:
partially_supported/borderline/minor_riskfor gradationsinsufficient_evidence/unverifiablefor uncertaintycontext_dependent/depends_on_metricfor conditional outcomes
Integration:
- Stage-based enable/disable via
--stages noul-augment - Both LLM-client and template-only paths work correctly
- Environment variable
LLM_AUGMENT_MODELfor model override
⚠️ Issue: Async tests not running
The 4 @pytest.mark.asyncio tests in test_noul_augment.py fail locally because:
pytest-asynciois not a project dependency- CI workflow only runs
test_pipeline.pyandtest_transform_results.py, not the new test file
Options:
- Add
pytest-asyncioto dev dependencies and update CI to run all tests - Convert async tests to sync using
asyncio.run()(simpler, no new dependency)
I recommend option 2 since it's a data pipeline, not an async application:
# Instead of:
@pytest.mark.asyncio
async def test_template_augmentation(self):
...
# Use:
def test_template_augmentation(self):
import asyncio
families = [...]
updated, count, failed = asyncio.run(augment_noul_questions(None, families, "", ""))
...Verdict: The implementation is solid. Please fix the async test issue (either approach works), then this is ready to merge.
Summary
noul-augmentsynthetic pipeline stage that expands binary noul questions (true/false) to 5-option multi-choice formataugmented_optionspropagated through all 5 noul emission points (base, counterfactual, paraphrase-state, paraphrase-question, negation)Closes #104
Test plan
--stages noul-augment --piloton existing synthetic data (requires data files)