Find what broke in your agent calls. Pin it so it never ships again.
pip install hotato
hotato vapi health # analyze your last 100 Vapi calls
hotato autopsy ./call.wav # or analyze one local fileZero config. Works with Vapi, Retell, Bland, Synthflow, Millis, or local audio. No judges. No cloud. No bill. MIT.
- Barge-in → Say-do gaps: Caller interrupts to cancel; agent says "canceled" but the booking tool still fires: a bug that fires actions the caller canceled. (Timing from the audio; the tool-fire check reads your call's tool log: hotato ingests Vapi/OTel traces.)
- Latency spikes: 800ms → 5s unpredictability that makes users hang up.
- Dead air: Long silences that kill conversation flow.
- Talk-over: Agent speaks over the caller; never yields.
pip install hotato
export VAPI_API_KEY=...
hotato vapi health --last 7d --output report.htmlOpen report.html. See your Voice Stability Score and every critical incident.
export RETELL_API_KEY=...
hotato retell health --call-id CALL_ID--call-id is required and repeatable: Retell has no verified
list-recent-calls endpoint, so hotato never guesses one. hotato bland health,
hotato synthflow health, and hotato millis health follow the Vapi shape;
those stacks export one mixed channel, so their reports carry the
measured-confidence mono observations block.
hotato autopsy ./call.wavWrites a detailed, self-contained HTML incident report under hotato-output/;
open it in your browser.
autopsy finds bugs. scan tracks trends across a folder of calls. When you
are ready, pin incidents to your CI so they never ship again: hotato pin
turns one incident into a portable failure check, and hotato prove is the CI
check that re-runs every stored piece of evidence and fails closed. Every
verdict carries its evidence across five dimensions (outcome, policy,
conversation, speech, reliability).
For continuous use: run hotato vapi health on a schedule, and open
hotato console --production-db DB to inspect stored runs locally.
The step's exit code is the verdict: 0 pass, 1 fail, 2 refuse.
# .github/workflows/voice-qa.yml
on: [pull_request]
jobs:
hotato:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: attenlabs/hotato@v1.16.0
with:
contracts: contracts/
hotato-version: 1.16.0Copy-paste workflow with a commit-SHA pin: docs/CI.md.
Point Claude Code, Cursor, or any coding agent at this repo: it reads
AGENTS.md and runs the loop end to end, offline, no key. The MCP
server exposes the scorer plus read/verify/propose tools over local stdio:
uvx --from "hotato[mcp]" hotato-mcp (docs/MCP.md).
hotato runs offline, on the machine that invokes it. The core is stdlib-only Python: no account, no key, no network call of its own. Your traces, prompts, and audio stay local, and the local-judge lane is opt-in and quality-gated, separate from the deterministic core.
The whole loop, command by command: docs/LIFECYCLE.md.
First touch to a CI gate: docs/GETTING-STARTED.md.
Feed it what you already have: docs/CONNECT.md ·
docs/TRACE.md · docs/SIMULATE.md.
What every verdict stands on: docs/EVIDENCE-CONTRACT.md.
Next to the hosted alternatives: docs/COMPARE.md.
The deep toolkit -- capture, simulation, load, benchmarking, the fix ladder,
the fleet control plane -- lives under hotato lab (hotato lab --help).
The public commands are durable; hotato lab evolves faster; every pre-1.17
top-level spelling keeps working unchanged.
| Property | Value |
|---|---|
| Footprint | ~10 MiB installed, 0 runtime dependencies (stdlib-only) |
| Reproducibility | byte-for-byte, content-addressed checks |
| Exit codes | 0 pass · 1 fail · 2 refuse |
| Release integrity | OIDC Trusted Publishing + build-provenance attested |
| Runtime | offline, off the production data path |
Verify the measurement yourself
PYTHONPATH=src python3 -m hotato.benchmark \
--scenarios corpus/real/scenarios --audio corpus/real/audioOn 13 recorded AMI Meeting Corpus clips, the median error between measured caller-onset and the human word-alignment label is 20 ms. Provenance: corpus/real/README.md · method: METHODOLOGY.md.
Timing is measurable only when the two voices arrive on separate channels; a mono or mixed export is marked NOT SCORABLE and refused (hotato trust --stereo call.wav). The full four-tier evidence policy (what each verdict stands on, per input) is docs/EVIDENCE-CONTRACT.md.
Issues and PRs welcome: CONTRIBUTING.md · SECURITY.md · CHANGELOG · docs/
MIT (LICENSE)
mcp-name: io.github.attenlabs/hotato