Skip to content

Roadmap: improve realtime latency p95 evidence #7

Description

@ykshv

Goal

Bring first-audio and barge-in p95 closer to the release targets under a long-running local GPU profile.

Scope

  • Separate cached deterministic responses from normal LLM/TTS responses.
  • Measure first audio p50/p95 and underruns over a longer soak.
  • Review speculative turn-taking thresholds and TTS first-clause budget.
  • Track GPU headroom and thermal behavior.

Acceptance

  • latency_report.py output clearly separates response classes.
  • p95 regressions are visible in CI/manual release evidence.
  • Any changed threshold is justified by measured behavior.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    latencyRealtime latency and p95 performance workroadmapPlanned public roadmap work

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions