Skip to main content
← Back to glossary
AI & automation (AIaaS)

Turn latency

What is Turn latency?

Turn latency is the time between a caller finishing speaking and an AI voice agent's spoken response starting to play — the round trip through speech-to-text, the language model's reasoning, and text-to-speech generation. Turn latency is one of the most important quality metrics for a conversational voice AI, since a delay much longer than a human's natural response pause makes the interaction feel robotic or broken, even if the eventual answer is accurate.

More detail

A production voice AI pipeline typically streams each stage — transcribing speech as it arrives, starting to generate a response before the full transcript is even final, and beginning speech playback before the whole reply is generated — to shrink perceived turn latency.

Turn latency is usually measured and tracked separately from overall call duration, since a technically short call can still feel poor if each individual turn has a noticeable, unnatural pause.

Frequently asked

Why does turn latency matter more for voice AI than chat AI?
A chat conversation tolerates a visible typing delay far better than a voice conversation tolerates silence — callers expect a response within roughly the pause a human would take, so voice AI pipelines are engineered specifically to minimize turn latency.
How do voice AI systems reduce turn latency?
By streaming each stage — transcription, language model reasoning, and speech generation — instead of waiting for one stage to fully finish before starting the next, so processing overlaps rather than happening strictly in sequence.

Build it on Orbit

Voice, messaging, email, video, and AI agents on one platform and one pay-as-you-go bill. Start free — no credit card required.