TTS (Text-to-Speech) — Text-to-Speech
What is TTS (Text-to-Speech)?
TTS (Text-to-Speech) converts written text into spoken audio, letting an IVR or AI voice agent read a dynamic response — an account balance, an appointment time — aloud without a human ever recording it. Modern neural TTS voices sound close to natural human speech, with adjustable tone, pace, and language, and can be generated in real time so an AI voice agent's text response is heard by the caller within a fraction of a second.
More detail
Older, concatenative TTS stitched together pre-recorded speech fragments and often sounded robotic; modern neural TTS models generate audio waveforms directly, producing far more natural intonation and pacing.
Real-time TTS latency matters directly for conversational AI voice agents — a caller expects a response to start within roughly the same pause a human would take before replying.
Frequently asked
- Does modern TTS still sound robotic?
- Generally no — neural TTS models used today produce speech with natural-sounding intonation and pacing that's often difficult to distinguish from a human recording, a significant improvement over older concatenative TTS systems.
- Why does TTS speed matter for an AI voice agent?
- A caller expects a response within roughly the same pause a human would take, so how quickly TTS can generate audio from text directly affects how natural a conversation with an AI voice agent feels — this is part of what's measured as turn latency.
See also
Build it on Orbit
Voice, messaging, email, video, and AI agents on one platform and one pay-as-you-go bill. Start free — no credit card required.