Skip to content

feat: source MAST phase 3 pilot corpus for 3 of 12 semantic-only modes - #16

Merged
mattekudacy merged 1 commit into
mainfrom
claude/serene-ritchie-g11dsf
Sep 12, 2026
Merged

mattekudacy merged 1 commit into
mainfrom
claude/serene-ritchie-g11dsf

Conversation

@mattekudacy

Copy link
Copy Markdown
Owner

Ownership. This PR was opened directly by an automated Claude Code session working in
this repo — see the "Exception" paragraph in this template. The Co-Authored-By: trailer and
session-link footer on the commit and this description are required by the platform this
session runs on and can't be suppressed from inside the PR; nothing else here is exempt, and a
human still decides whether to merge it.

What does this change and why?

Follows through on docs/concepts/multi-agent-failures.md's "Phase 3 scoping" recommendation
to pilot a small subset of MAST modes with real evidence before committing to sourcing all
twelve.

scripts/gen_mast_pilot_corpus.py generates tests/data/mast_pilot_corpus.json (6 entries)
for 2.5 Ignored Other Agent's Input, 2.4 Information Withholding, and 1.4 Loss of Conversation
History — every entry transcribed from a real, cited GitHub issue, same discipline as
tests/data/error_corpus_*.json. scripts/mast_mode_pilot_accuracy.py scores an experimental
prompt against it, deliberately outside LLMClassifier's stable contract: its prompt and
label set live only in the script, _SYSTEM_PROMPT/_parse_response() in
triage/classifier/llm.py are untouched, and classify() still only ever returns one of the 9
stable FailureType members.

Sourcing results:

That ambiguity is a real finding, not a corpus-building shortfall: it suggests 2.4/2.5 can be
hard to tell apart even from a complete, real trace, which bears directly on whether either
could ever be a reliable stable FailureType regardless of prompt quality — a classifier can't
resolve an ambiguity a human reading the same trace can't resolve either.

Neither script has been run against a real model — this environment has no LLM API key, same
blocked status as scripts/hybrid_ambiguity_accuracy.py. No accuracy number is claimed.
tests/test_mast_pilot_corpus.py adds 8 zero-API-call structural guards on the corpus itself
(every entry cites a real source, every scored entry uses a piloted mode code, the 2/2/1
source-count split is pinned so it can't silently drift without the docstrings being updated
too).

Related issue

None


Type of change

  • New feature (a pilot corpus + experimental scoring script; no triage/ core changes)
  • Documentation

Checklist

  • No AI-attribution trailers or badges anywhere in this PR — commit messages and this
    description are clean. I am the author and I am responsible for this change.
    (N/A — see the "Ownership" exception above.)
  • pytest tests/ -x --tb=short passes locally (816 passed, up from 808 — 8 new tests)
  • ruff check ., ruff format --check ., and mypy triage/ --strict are all clean
  • No new imports of openai/anthropic/langchain/langgraph/opentelemetry/etc. inside
    triage/ core — no triage/ files touched at all; the new scripts import anthropic/
    openai lazily inside functions, same convention as the existing accuracy scripts
  • Docs updated: docs/concepts/multi-agent-failures.md (new "Pilot corpus sourced"
    subsection, phasing item 3, top Status line), scripts/README.md (new Benchmarks row +
    writeup), CHANGELOG.md ([Unreleased] → ### Added)
  • N/A — no new FailureType or RecoveryAction added (the whole point of this pilot is
    measuring before considering that)

If this touches rules.py or an error corpus — N/A, this PR touches neither rules.py nor
any FailureType accuracy corpus; mast_pilot_corpus.json is a new, separate corpus for a
different measurement (MAST-mode labeling, not FailureType routing accuracy).


Anything reviewers should look at closely?

The ambiguous entry (openai-agents-python#348, in gen_mast_pilot_corpus.py) is the part
worth double-checking — it's a judgment call that the trace genuinely doesn't disambiguate 2.4
from 2.5, not a cop-out from labeling something harder to source. If a reviewer disagrees and
thinks it should carry a single label, that's a legitimate pushback worth having before this
lands.

🤖 Generated with Claude Code

https://claude.ai/code/session_01M4WNEkbnKSx9mTg5Q1jX39


Generated by Claude Code

Follows through on docs/concepts/multi-agent-failures.md's "Phase 3
scoping" recommendation to pilot a small subset of MAST modes before
committing to sourcing all twelve.

scripts/gen_mast_pilot_corpus.py generates tests/data/mast_pilot_corpus.json
(6 entries) for 2.5 Ignored Other Agent's Input, 2.4 Information Withholding,
and 1.4 Loss of Conversation History — every entry transcribed from a real,
cited GitHub issue, same discipline as tests/data/error_corpus_*.json.
scripts/mast_mode_pilot_accuracy.py scores an experimental prompt against it,
deliberately outside LLMClassifier's stable contract: its prompt and label
set live only in the script, _SYSTEM_PROMPT/_parse_response() in
triage/classifier/llm.py are untouched, and classify() still only ever
returns one of the 9 stable FailureType members.

Sourcing results:
- 2.5 and 1.4 each got two clean, independently-sourced entries (different
  frameworks): FlowiseAI/Flowise#3512 + microsoft/autogen#6891 for 2.5;
  langchain-ai/langgraph#2395 + github/copilot-cli#1180 for 1.4.
- 2.4 got only one clean entry (LibreChat-AI/LibreChat#10569). Its second
  real candidate, openai/openai-agents-python#348, turned out genuinely
  ambiguous between 2.4 and 2.5 on inspection — the reporter's own account
  doesn't establish which side dropped the information, and neither does
  the trace. Recorded as such (no single mast_mode, ambiguous_with instead,
  excluded from scoring) rather than forced into either label.

That ambiguity is a real finding, not a corpus-building shortfall: it
suggests 2.4/2.5 can be hard to tell apart even from a complete, real trace,
which bears directly on whether either could ever be a reliable stable
FailureType regardless of prompt quality — a classifier can't resolve an
ambiguity a human reading the same trace can't resolve either.

Neither script has been run against a real model — this environment has no
LLM API key, same blocked status as scripts/hybrid_ambiguity_accuracy.py.
No accuracy number is claimed. tests/test_mast_pilot_corpus.py adds 8
zero-API-call structural guards on the corpus itself (every entry cites a
real source, every scored entry uses a piloted mode code, the 2/2/1
source-count split is pinned so it can't silently drift without the
docstrings being updated too).

Updated files:
- scripts/gen_mast_pilot_corpus.py (new), scripts/mast_mode_pilot_accuracy.py
  (new), tests/data/mast_pilot_corpus.json (new, generated),
  tests/test_mast_pilot_corpus.py (new)
- docs/concepts/multi-agent-failures.md — new "Pilot corpus sourced"
  subsection, phasing item 3 and top Status line updated
- scripts/README.md — new Benchmarks row + writeup paragraph
- CHANGELOG.md — records the pilot under [Unreleased] ### Added

No triage/ core code changed. Verified: ruff check, ruff format --check,
pytest (816 passed, up from 808), mypy --strict, mkdocs build --strict all
clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M4WNEkbnKSx9mTg5Q1jX39
@mattekudacy
mattekudacy merged commit 678533c into main Sep 12, 2026
7 checks passed
@mattekudacy
mattekudacy deleted the claude/serene-ritchie-g11dsf branch September 12, 2026 07:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants