🪜 Voice Latency Gap Simulator

Where the milliseconds go in one voice-agent turn — user stops talking → first bot audio. Teaching approximation, not a benchmark.

E2E: — ms

Turn pipeline stages

fixed tunable

Model: frame = one 32 ms VAD window; tail = the fixed 160 ms of speech after the last voiced frame (kept for naturalness); hangover + min-silence gate the endpoint; then ASR commits, packets travel, the LLM thinks, audio comes back. Real stacks differ (streaming codecs, speculative TTS) — the shape, not the exact ms, is the lesson.

Controls

Presets (click to load · ms)