A proactive QA capability for VC scripting — finds unknown/undocumented behaviors before customers do, and classifies customer escalations before engineering time is spent in the wrong direction.
This repo supports two distinct use cases, both powered by the same SKILL.md:
When to run: Now, and after every release.
Goal: Generate a structured list of capability combinations that will break, produce wrong state, or have undocumented behavior — before a customer hits them in production.
Output: A domain test case markdown in proactive/investigations/. A one-line entry added to proactive/CHANGELOG.md.
How to invoke in claude.ai:
"Using vc-capability-qa, generate test cases for the [domain] domain. Act as QA. Start from what a customer building on VC would try to do, not from what the code supports."
How to invoke in Claude Desktop (with TFS code access):
"Using vc-capability-qa and investigate-issue, generate test cases for the [domain] domain. For each [Guessing] claim, read the actual imposter code and replace with [Certain] plus source line reference."
Cadence: Run for each new domain. Re-run existing domains after major releases to catch regressions. Log every run in proactive/CHANGELOG.md.
Domains to cover next (in priority order):
| Domain | Trigger |
|---|---|
| ACW + State Management (MCH) | Claro ORC-52396 — Category 3 audit needed |
| Cross-BU / Multi-BU | Walmart COM-00003890 — identity gap at BU boundary |
| Outbound + Consult/Transfer | Adjacent to voice domain gaps already found |
| Digital + Voice Blending | Agent handling both simultaneously |
| AI Agent + Voice Handoff | Cognigy June 2026 — segment model unknown |
When to run: When a customer escalation arrives.
Goal: Classify the escalation before spending engineering time investigating.
Output: Classification with evidence, adjacent risks identified, customer README saved to customers/{name}/. Investigations updated with any new combinations the escalation reveals.
How to invoke in claude.ai:
"Using vc-capability-qa, classify this escalation: [paste escalation description or ticket]. Is this a product gap or an understanding gap? What adjacent combinations are also at risk? Update the investigations if this reveals new test cases."
Classification decision tree:
Is the failure deterministic (same flow always fails)?
YES → Category 1 — check imposter script emissions (product gap)
NO → Is it MCH-specific or load-dependent?
YES → Category 3 — audit shared global variables (product gap)
NO → Is it about correlating identities across a boundary?
YES → Category 2 — escalate to architect (product gap)
NO → Likely understanding gap — check test case library
for documented behavior before raising engineering ticket
Real escalations classified using this flow:
| Customer | Classification | Status |
|---|---|---|
| Citi | Category 1 + 2 — recording correlation + consult identity | Open — architect engagement needed |
| Services Australia | Category 1 — missing state emissions (5 items) | Committed NiCE 27.3 Jul-Sep 2027 |
| Ford ORC-48255 | Category 1 — ITO blocked during WAIT | Open — onAgentHangup scoped, XL/XXL |
| Claro ORC-52396 | Category 3 — race condition on shared global variable | Resolved CU6 May 2026 |
| Walmart COM-00003890 | Category 2 — identity gap at BU boundary | Open — pending product evaluation, ETA 7/31/2026 |
Every new escalation that reveals a previously unknown combination must update the relevant domain investigation file. This is the core mechanism that makes the test case library grow over time.
When Flow 2 is run on a new escalation:
- If the combination is not in the test cases → add it as a new TC tagged ✅ Customer-Confirmed
- If the combination is in the test cases but was 🔴 Unknown/Untested → update status to ✅ Customer-Confirmed with customer evidence
- If the escalation corrects a [Guessing] claim → update to [Certain] with the customer evidence as source
The five customers already in this repo each contributed test cases this way — Services Australia confirmed 5 test cases, Citi confirmed 2, Ford confirmed 1. Every future escalation adds to this. Over time, 🔴 Unknown/Untested test cases get confirmed or refuted, and the library becomes an increasingly accurate map of what works and what doesn't.
vc-capability-qa/
├── README.md ← This file
├── SKILL.md ← Full skill definition — load this in Claude
│
├── proactive/ ← Flow 1 outputs
│ ├── CHANGELOG.md ← Log of every domain sweep (release, date, findings)
│ └── investigations/
│ ├── vc-voice-domain-test-cases-v2.md ← Voice domain (complete — 15 test cases)
│ └── ... (add new domains here)
│
└── customers/ ← Flow 2 inputs and outputs
├── citi/
│ ├── README.md ← Ask, classification, status, adjacent risks, TC contributions
│ └── artifacts/ ← Raw docs, email extracts, Jira exports as .md
├── service-australia/
│ ├── README.md
│ └── artifacts/
├── ford/
│ ├── README.md
│ └── artifacts/
├── claro/
│ ├── README.md
│ └── artifacts/
└── walmart/
├── README.md
└── artifacts/
When you have the artifact (PDF, docx, email, screenshot): upload it to Claude, it reads and extracts key facts, writes the customer README, and saves a cleaned .md version to artifacts/.
When you don't have the artifact: Claude pulls from email, Teams, and Jira directly and generates the README from what it finds. The artifacts/ folder contains the synthesized summary.
Why this matters: When Flow 2 runs on a new escalation, Claude reads all customer READMEs first. This enables immediate pattern-matching against known cases without re-investigating from scratch. The wider the customer folder grows, the faster and more accurate the classification becomes.
Every time Flow 1 is run, add one line to proactive/CHANGELOG.md:
| Date | Domain | Release baseline | Run by | TCs generated | Customer-confirmed | Notes |
Why it matters: When a post-release bug is reported, the changelog tells you whether the affected domain was swept before or after that release. If before — regression introduced by the release. If after — already in test cases, nobody acted on it.
Sources are read in this order — repo context always first:
| Priority | Source | What It Provides |
|---|---|---|
| 1 | customers/ folder |
Primary context — read before anything else |
| 2 | proactive/investigations/ |
Existing test cases — check before generating new ones |
| 3 | claude-vc/.claude/skills/ |
VC architecture, scripting, contact/agent lifecycle |
| 4 | Confluence MCR space | LLDs, Known Problems (cloudId: 0e508bed-9911-4fa0-9106-53d761fb5715) |
| 5 | Jira ORC / CXREC | Confirmed bugs, commitments, refinement notes |
| 6 | Microsoft 365 (email + Teams) | Escalation patterns, customer-confirmed failures |
| 7 | TFS source code | Imposter scripts, Agent.cs, MCHAgent.cs (Claude Desktop + semantic index only) |
Flow 1 (proactive) → populates "What UCs are not supported" → informs VC scripting scope on the roadmap.
Flow 2 (reactive) → classifies "Product gap vs understanding gap" → routes escalations correctly and surfaces product gaps with enough evidence to scope and commit.
The feedback loop between Flow 2 and proactive/investigations/ is what makes the system compound over time — each customer escalation makes the next classification faster and the next proactive sweep more accurate.
Siddharatha Joshi | VC Scripting (IVR, Routing, State Management)
Created: 2026-06-03 | Skill version: 1.1