Where the milliseconds go in one voice-agent turn — user stops talking → first bot audio. Teaching approximation, not a benchmark.
Bot audio wasted before cutoff ≈ — ms — grows with hangover: the echo-suppressor needs the tail gone before the user can take the floor.
Model: frame = one 32 ms VAD window; tail = the fixed 160 ms of speech after the last voiced frame (kept for naturalness); hangover + min-silence gate the endpoint; then ASR commits, packets travel, the LLM thinks, audio comes back. Real stacks differ (streaming codecs, speculative TTS) — the shape, not the exact ms, is the lesson.