Skip to content

attenlabs/hotato

Use this GitHub action with your project
Add this Action to an existing workflow or create a new one
View on Marketplace

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

577 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

hotato

PyPI version Downloads per month Python versions CI status MIT license

hotato

Find what broke in your agent calls. Pin it so it never ships again.

pip install hotato
hotato vapi health              # analyze your last 100 Vapi calls
hotato autopsy ./call.wav       # or analyze one local file

Zero config. Works with Vapi, Retell, Bland, Synthflow, Millis, or local audio. No judges. No cloud. No bill. MIT.

hotato.dev

What it finds

  • Barge-in → Say-do gaps: Caller interrupts to cancel; agent says "canceled" but the booking tool still fires: a bug that fires actions the caller canceled. (Timing from the audio; the tool-fire check reads your call's tool log: hotato ingests Vapi/OTel traces.)
  • Latency spikes: 800ms → 5s unpredictability that makes users hang up.
  • Dead air: Long silences that kill conversation flow.
  • Talk-over: Agent speaks over the caller; never yields.

Quickstart

Vapi

pip install hotato
export VAPI_API_KEY=...
hotato vapi health --last 7d --output report.html

Open report.html. See your Voice Stability Score and every critical incident.

Retell

export RETELL_API_KEY=...
hotato retell health --call-id CALL_ID

--call-id is required and repeatable: Retell has no verified list-recent-calls endpoint, so hotato never guesses one. hotato bland health, hotato synthflow health, and hotato millis health follow the Vapi shape; those stacks export one mixed channel, so their reports carry the measured-confidence mono observations block.

Local audio

hotato autopsy ./call.wav

Writes a detailed, self-contained HTML incident report under hotato-output/; open it in your browser.

From finding bugs to preventing them

autopsy finds bugs. scan tracks trends across a folder of calls. When you are ready, pin incidents to your CI so they never ship again: hotato pin turns one incident into a portable failure check, and hotato prove is the CI check that re-runs every stored piece of evidence and fails closed. Every verdict carries its evidence across five dimensions (outcome, policy, conversation, speech, reliability).

Read more →

For continuous use: run hotato vapi health on a schedule, and open hotato console --production-db DB to inspect stored runs locally.

Wire it into CI

The step's exit code is the verdict: 0 pass, 1 fail, 2 refuse.

# .github/workflows/voice-qa.yml
on: [pull_request]
jobs:
  hotato:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: attenlabs/hotato@v1.16.0
        with:
          contracts: contracts/
          hotato-version: 1.16.0

Copy-paste workflow with a commit-SHA pin: docs/CI.md.

Point your agent at it

Point Claude Code, Cursor, or any coding agent at this repo: it reads AGENTS.md and runs the loop end to end, offline, no key. The MCP server exposes the scorer plus read/verify/propose tools over local stdio: uvx --from "hotato[mcp]" hotato-mcp (docs/MCP.md).

Nothing leaves your machine

hotato runs offline, on the machine that invokes it. The core is stdlib-only Python: no account, no key, no network call of its own. Your traces, prompts, and audio stay local, and the local-judge lane is opt-in and quality-gated, separate from the deterministic core.

Go deeper

The whole loop, command by command: docs/LIFECYCLE.md. First touch to a CI gate: docs/GETTING-STARTED.md. Feed it what you already have: docs/CONNECT.md · docs/TRACE.md · docs/SIMULATE.md. What every verdict stands on: docs/EVIDENCE-CONTRACT.md. Next to the hosted alternatives: docs/COMPARE.md.

The deep toolkit -- capture, simulation, load, benchmarking, the fix ladder, the fleet control plane -- lives under hotato lab (hotato lab --help). The public commands are durable; hotato lab evolves faster; every pre-1.17 top-level spelling keeps working unchanged.

Specifications

Property Value
Footprint ~10 MiB installed, 0 runtime dependencies (stdlib-only)
Reproducibility byte-for-byte, content-addressed checks
Exit codes 0 pass · 1 fail · 2 refuse
Release integrity OIDC Trusted Publishing + build-provenance attested
Runtime offline, off the production data path
Verify the measurement yourself
PYTHONPATH=src python3 -m hotato.benchmark \
  --scenarios corpus/real/scenarios --audio corpus/real/audio

On 13 recorded AMI Meeting Corpus clips, the median error between measured caller-onset and the human word-alignment label is 20 ms. Provenance: corpus/real/README.md · method: METHODOLOGY.md.

Timing is measurable only when the two voices arrive on separate channels; a mono or mixed export is marked NOT SCORABLE and refused (hotato trust --stereo call.wav). The full four-tier evidence policy (what each verdict stands on, per input) is docs/EVIDENCE-CONTRACT.md.

Contribute

Issues and PRs welcome: CONTRIBUTING.md · SECURITY.md · CHANGELOG · docs/

License

MIT (LICENSE)

Know when to pass it on.

mcp-name: io.github.attenlabs/hotato

About

Find what broke in your agent calls. Pin it so it never ships again. Local voice-agent call forensics and regression guards.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages