Skip to main content
← Back to glossary
AI & automation (AIaaS)

STT (Speech-to-Text) — Speech-to-Text

What is STT (Speech-to-Text)?

STT (Speech-to-Text), also called automatic speech recognition (ASR), converts a caller's spoken audio into text a system can process — feeding an IVR's natural-language menu, a voicemail transcript, or an AI voice agent's understanding of what a caller just said. STT accuracy varies with audio quality, background noise, accent, and vocabulary, which is why a good voice AI pipeline is tuned for the specific domain (product names, account terminology) it needs to recognize.

More detail

Modern STT systems process audio in a streaming fashion, transcribing speech as it arrives rather than waiting for the caller to finish talking, which is what lets an AI voice agent respond with low turn latency.

Domain-specific vocabulary — product names, account numbers, industry jargon — often needs custom tuning or a supplied word list, since general-purpose STT models can misrecognize terms they weren't trained on.

Frequently asked

Is STT the same thing as voicemail transcription?
Voicemail transcription is one application of STT — the same underlying speech-recognition technology also powers real-time IVR voice menus, live call transcripts, and the listening half of an AI voice agent's conversation loop.
Why does an AI voice agent sometimes mishear a name or account number?
STT accuracy drops for unusual vocabulary, background noise, and strong accents, so uncommon names or long alphanumeric codes are especially error-prone unless the system has been tuned with domain-specific vocabulary for that use case.

Build it on Orbit

Voice, messaging, email, video, and AI agents on one platform and one pay-as-you-go bill. Start free — no credit card required.