feat: source MAST phase 3 pilot corpus for 3 of 12 semantic-only modes - #16
Merged
Merged
Conversation
Follows through on docs/concepts/multi-agent-failures.md's "Phase 3 scoping" recommendation to pilot a small subset of MAST modes before committing to sourcing all twelve. scripts/gen_mast_pilot_corpus.py generates tests/data/mast_pilot_corpus.json (6 entries) for 2.5 Ignored Other Agent's Input, 2.4 Information Withholding, and 1.4 Loss of Conversation History — every entry transcribed from a real, cited GitHub issue, same discipline as tests/data/error_corpus_*.json. scripts/mast_mode_pilot_accuracy.py scores an experimental prompt against it, deliberately outside LLMClassifier's stable contract: its prompt and label set live only in the script, _SYSTEM_PROMPT/_parse_response() in triage/classifier/llm.py are untouched, and classify() still only ever returns one of the 9 stable FailureType members. Sourcing results: - 2.5 and 1.4 each got two clean, independently-sourced entries (different frameworks): FlowiseAI/Flowise#3512 + microsoft/autogen#6891 for 2.5; langchain-ai/langgraph#2395 + github/copilot-cli#1180 for 1.4. - 2.4 got only one clean entry (LibreChat-AI/LibreChat#10569). Its second real candidate, openai/openai-agents-python#348, turned out genuinely ambiguous between 2.4 and 2.5 on inspection — the reporter's own account doesn't establish which side dropped the information, and neither does the trace. Recorded as such (no single mast_mode, ambiguous_with instead, excluded from scoring) rather than forced into either label. That ambiguity is a real finding, not a corpus-building shortfall: it suggests 2.4/2.5 can be hard to tell apart even from a complete, real trace, which bears directly on whether either could ever be a reliable stable FailureType regardless of prompt quality — a classifier can't resolve an ambiguity a human reading the same trace can't resolve either. Neither script has been run against a real model — this environment has no LLM API key, same blocked status as scripts/hybrid_ambiguity_accuracy.py. No accuracy number is claimed. tests/test_mast_pilot_corpus.py adds 8 zero-API-call structural guards on the corpus itself (every entry cites a real source, every scored entry uses a piloted mode code, the 2/2/1 source-count split is pinned so it can't silently drift without the docstrings being updated too). Updated files: - scripts/gen_mast_pilot_corpus.py (new), scripts/mast_mode_pilot_accuracy.py (new), tests/data/mast_pilot_corpus.json (new, generated), tests/test_mast_pilot_corpus.py (new) - docs/concepts/multi-agent-failures.md — new "Pilot corpus sourced" subsection, phasing item 3 and top Status line updated - scripts/README.md — new Benchmarks row + writeup paragraph - CHANGELOG.md — records the pilot under [Unreleased] ### Added No triage/ core code changed. Verified: ruff check, ruff format --check, pytest (816 passed, up from 808), mypy --strict, mkdocs build --strict all clean. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M4WNEkbnKSx9mTg5Q1jX39
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Ownership. This PR was opened directly by an automated Claude Code session working in
this repo — see the "Exception" paragraph in this template. The
Co-Authored-By:trailer andsession-link footer on the commit and this description are required by the platform this
session runs on and can't be suppressed from inside the PR; nothing else here is exempt, and a
human still decides whether to merge it.
What does this change and why?
Follows through on
docs/concepts/multi-agent-failures.md's "Phase 3 scoping" recommendationto pilot a small subset of MAST modes with real evidence before committing to sourcing all
twelve.
scripts/gen_mast_pilot_corpus.pygeneratestests/data/mast_pilot_corpus.json(6 entries)for 2.5 Ignored Other Agent's Input, 2.4 Information Withholding, and 1.4 Loss of Conversation
History — every entry transcribed from a real, cited GitHub issue, same discipline as
tests/data/error_corpus_*.json.scripts/mast_mode_pilot_accuracy.pyscores an experimentalprompt against it, deliberately outside
LLMClassifier's stable contract: its prompt andlabel set live only in the script,
_SYSTEM_PROMPT/_parse_response()intriage/classifier/llm.pyare untouched, andclassify()still only ever returns one of the 9stable
FailureTypemembers.Sourcing results:
time): FlowiseAI/Flowise#3512 +
microsoft/autogen#6891 for 2.5;
langchain-ai/langgraph#2395 +
github/copilot-cli#1180 for 1.4.
(danny-avila/LibreChat#10569). Its
second real candidate,
openai/openai-agents-python#348,
turned out genuinely ambiguous between 2.4 and 2.5 on inspection — the reporter's own
account doesn't establish which side dropped the information, and neither does the trace.
Recorded as such (no single
mast_mode,ambiguous_withinstead, excluded from scoring)rather than forced into either label.
That ambiguity is a real finding, not a corpus-building shortfall: it suggests 2.4/2.5 can be
hard to tell apart even from a complete, real trace, which bears directly on whether either
could ever be a reliable stable
FailureTyperegardless of prompt quality — a classifier can'tresolve an ambiguity a human reading the same trace can't resolve either.
Neither script has been run against a real model — this environment has no LLM API key, same
blocked status as
scripts/hybrid_ambiguity_accuracy.py. No accuracy number is claimed.tests/test_mast_pilot_corpus.pyadds 8 zero-API-call structural guards on the corpus itself(every entry cites a real source, every scored entry uses a piloted mode code, the 2/2/1
source-count split is pinned so it can't silently drift without the docstrings being updated
too).
Related issue
None
Type of change
triage/core changes)Checklist
description are clean. I am the author and I am responsible for this change.
(N/A — see the "Ownership" exception above.)
pytest tests/ -x --tb=shortpasses locally (816 passed, up from 808 — 8 new tests)ruff check .,ruff format --check ., andmypy triage/ --strictare all cleanopenai/anthropic/langchain/langgraph/opentelemetry/etc. insidetriage/core — notriage/files touched at all; the new scripts importanthropic/openailazily inside functions, same convention as the existing accuracy scriptsdocs/concepts/multi-agent-failures.md(new "Pilot corpus sourced"subsection, phasing item 3, top Status line),
scripts/README.md(new Benchmarks row +writeup),
CHANGELOG.md([Unreleased]→### Added)FailureTypeorRecoveryActionadded (the whole point of this pilot ismeasuring before considering that)
If this touches
rules.pyor an error corpus — N/A, this PR touches neitherrules.pynorany
FailureTypeaccuracy corpus;mast_pilot_corpus.jsonis a new, separate corpus for adifferent measurement (MAST-mode labeling, not
FailureTyperouting accuracy).Anything reviewers should look at closely?
The ambiguous entry (
openai-agents-python#348, ingen_mast_pilot_corpus.py) is the partworth double-checking — it's a judgment call that the trace genuinely doesn't disambiguate 2.4
from 2.5, not a cop-out from labeling something harder to source. If a reviewer disagrees and
thinks it should carry a single label, that's a legitimate pushback worth having before this
lands.
🤖 Generated with Claude Code
https://claude.ai/code/session_01M4WNEkbnKSx9mTg5Q1jX39
Generated by Claude Code