STT (Speech-to-Text) — Speech-to-Text
What is STT (Speech-to-Text)?
STT (Speech-to-Text), also called automatic speech recognition (ASR), converts a caller's spoken audio into text a system can process — feeding an IVR's natural-language menu, a voicemail transcript, or an AI voice agent's understanding of what a caller just said. STT accuracy varies with audio quality, background noise, accent, and vocabulary, which is why a good voice AI pipeline is tuned for the specific domain (product names, account terminology) it needs to recognize.
More detail
Modern STT systems process audio in a streaming fashion, transcribing speech as it arrives rather than waiting for the caller to finish talking, which is what lets an AI voice agent respond with low turn latency.
Domain-specific vocabulary — product names, account numbers, industry jargon — often needs custom tuning or a supplied word list, since general-purpose STT models can misrecognize terms they weren't trained on.
Frequently asked
- Is STT the same thing as voicemail transcription?
- Voicemail transcription is one application of STT — the same underlying speech-recognition technology also powers real-time IVR voice menus, live call transcripts, and the listening half of an AI voice agent's conversation loop.
- Why does an AI voice agent sometimes mishear a name or account number?
- STT accuracy drops for unusual vocabulary, background noise, and strong accents, so uncommon names or long alphanumeric codes are especially error-prone unless the system has been tuned with domain-specific vocabulary for that use case.
See also
Build it on Orbit
Voice, messaging, email, video, and AI agents on one platform and one pay-as-you-go bill. Start free — no credit card required.