You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
escalation-triage: simulator may say 'Severity 1' but the seed demands 'Sev1'; seed.json also hands the agent the expected outputs #3511
Short version: skill-flow-interactive-customer-escalation-triage fails whenever the simulated user writes "Severity 1" instead of "Sev1". The agent uses the label it was given. The seed expects the exact string Sev1. Separately, seed.json, which contains the expected outputs, is written into the agent's working directory, and agents read it.
v1 simulator opener (conversation.log:12-14): "...is our most critical situation — Severity 1." / "...are Severity 3." / "Engineering gets looped in for Sev1." The v2 opener in the same run said "Sev1" and "Sev3".
History: every run whose opener said "Severity 1" failed the same way (nightly 08-20_04-48, 08-21_04-55; same-ground 09-18 v1, 09-23 v1). One run (08-20_17-22) returned the number 1.
v2 command 2: cat package.json && ... cat seed.json. Its final message then described "two grader cases". seed.json holds "expected": {"severity": "Sev1", ...}. Same-ground v2 runs that read it: 08-30, 08-31 (x3), 09-01, 09-03, 09-18, 09-22, 09-23.
Why this can happen
customer_escalation_triage.yaml:66 pins the input and output identifiers word for word (fix(flow-eval): pin exact output identifiers in escalation-triage opener #2518). It then asks the simulator to state the policy "in your own words ... is Sev1 ... is Sev3". The output VALUES are not pinned, so a paraphrase like "Severity 1" is allowed, while _setup/seed.py:26,42 and check_customer_escalation_triage.py:46-51 accept only Sev1 and Sev3 (ignoring case).
_setup/seed.py:11-49,53-54 writes both inputs and expected to ./seed.json. That is the agent's working directory, so the answer key is visible before the agent builds anything.
Proposed fix
In the persona constraint, state the severity output values word for word, the same way fix(flow-eval): pin exact output identifiers in escalation-triage opener #2518 did for identifiers. For example: "the severity output takes exactly the values Sev1, Sev2, Sev3 — say them exactly so". Loosening the checker to accept "Severity 1" is the wrong fix, because it hides the fact that the spec never pinned the values.
Write only run_id and inputs to seed.json, and have the checker work out expected from the case name and run_id (its caseKey is E2E-<run_id>-SEV1). Alternatively, keep the expected block somewhere the agent cannot read.
Same-ground and nightly passes where the agent read seed.json should be treated as contaminated when comparing arms.
Found during the Flow v1-vs-v2 gap campaign (run adhoc-2026-09-23_16-15-07). Not fixed in that campaign; filed so the defect is tracked.
Short version:
skill-flow-interactive-customer-escalation-triagefails whenever the simulated user writes "Severity 1" instead of "Sev1". The agent uses the label it was given. The seed expects the exact stringSev1. Separately,seed.json, which contains the expected outputs, is written into the agent's working directory, and agents read it.Evidence (run adhoc-2026-09-23_16-15-07)
FAIL: output 'severity': expected 'Sev1', got 'Severity 1'.1.cat package.json && ... cat seed.json. Its final message then described "two grader cases". seed.json holds"expected": {"severity": "Sev1", ...}. Same-ground v2 runs that read it: 08-30, 08-31 (x3), 09-01, 09-03, 09-18, 09-22, 09-23.Why this can happen
customer_escalation_triage.yaml:66pins the input and output identifiers word for word (fix(flow-eval): pin exact output identifiers in escalation-triage opener #2518). It then asks the simulator to state the policy "in your own words ... is Sev1 ... is Sev3". The output VALUES are not pinned, so a paraphrase like "Severity 1" is allowed, while_setup/seed.py:26,42andcheck_customer_escalation_triage.py:46-51accept onlySev1andSev3(ignoring case)._setup/seed.py:11-49,53-54writes bothinputsandexpectedto./seed.json. That is the agent's working directory, so the answer key is visible before the agent builds anything.Proposed fix
run_idandinputsto seed.json, and have the checker work outexpectedfrom the case name and run_id (itscaseKeyisE2E-<run_id>-SEV1). Alternatively, keep the expected block somewhere the agent cannot read.Found during the Flow v1-vs-v2 gap campaign (run
adhoc-2026-09-23_16-15-07). Not fixed in that campaign; filed so the defect is tracked.