Skip to main content
← Back to glossary
AI & automation (AIaaS)

TTS (Text-to-Speech) — Text-to-Speech

What is TTS (Text-to-Speech)?

TTS (Text-to-Speech) converts written text into spoken audio, letting an IVR or AI voice agent read a dynamic response — an account balance, an appointment time — aloud without a human ever recording it. Modern neural TTS voices sound close to natural human speech, with adjustable tone, pace, and language, and can be generated in real time so an AI voice agent's text response is heard by the caller within a fraction of a second.

More detail

Older, concatenative TTS stitched together pre-recorded speech fragments and often sounded robotic; modern neural TTS models generate audio waveforms directly, producing far more natural intonation and pacing.

Real-time TTS latency matters directly for conversational AI voice agents — a caller expects a response to start within roughly the same pause a human would take before replying.

Frequently asked

Does modern TTS still sound robotic?
Generally no — neural TTS models used today produce speech with natural-sounding intonation and pacing that's often difficult to distinguish from a human recording, a significant improvement over older concatenative TTS systems.
Why does TTS speed matter for an AI voice agent?
A caller expects a response within roughly the same pause a human would take, so how quickly TTS can generate audio from text directly affects how natural a conversation with an AI voice agent feels — this is part of what's measured as turn latency.

Build it on Orbit

Voice, messaging, email, video, and AI agents on one platform and one pay-as-you-go bill. Start free — no credit card required.