An AI voice agent in 2026 has no single sticker price — it costs the sum of four per-minute parts (speech-to-text, the language model, text-to-speech, and the telephony minute) plus whatever platform fee sits on top. Configured for a typical support or booking line, an all-in cost commonly lands somewhere between roughly $0.05 and $0.30+ per talk-minute, and the spread is that wide because a premium voice, a frontier model, and a hard-to-reach destination each move the number on their own. This guide breaks the cost into its parts, shows why one quoted per-minute figure means little in isolation, compares the three ways vendors bill for it, and gives you a way to reason about return against the cost of a human-handled call.
The short version: "how much does an AI voice agent cost" is the wrong question until you know which four parts you are paying for and who is marking each one up. Price the stack, not the sticker.
The four parts of an AI voice agent's per-minute cost
Every real-time voice agent runs the same pipeline on every turn: it transcribes what the caller said (speech-to-text), decides what to say (a language model), speaks the reply (text-to-speech), and does all of that over a live phone call (telephony). Each stage is metered separately, and each is where a vendor can add margin.
| Cost component | What you pay for | Directional 2026 range | What moves it |
|---|---|---|---|
| Speech-to-text (STT) | Transcribing caller audio in real time, per minute of audio | ~$0.005–$0.02 / min | Streaming vs batch, language, custom vocabulary |
| Language model (LLM) | Tokens in and out per turn — the reasoning and the reply | Highly variable, per 1M tokens | Model tier, prompt length, tool calls, context reuse |
| Text-to-speech (TTS) | Synthesizing the agent's spoken reply, per character or minute | ~$0.01–$0.30 / min | Standard vs premium/cloned voices, language count |
| Telephony (carrier minutes) | Carrying the actual call over the phone network | ~$0.004–$0.02 / min (US) | Destination country, inbound vs outbound, carrier path |
| Platform / orchestration | The pipeline, tools, testing, dashboards, and support | Per-minute uplift or monthly plan | Bundled vs à la carte, volume tier, SLA |
Two things fall out of this table. First, the LLM is the least predictable line: a chatty prompt, a long system message repeated every turn, or a tool call that pulls in a large context can quietly double the token bill, while prompt caching and per-agent model tiering pull it back down. Second, telephony is the part most people forget to count — a voice agent has to ride a real phone network, and the per-minute rate plus the number of carrier hops between the agent and the handset is a cost line, not a rounding error.
Why a single per-minute quote misleads
A headline "$0.09 per minute" tells you almost nothing without the configuration behind it. The number a vendor markets is usually a best case — a standard voice, a mid-tier model, a short prompt, a domestic call, and a p50 (median) turn that hides the p95 spikes. Change any one input and the real bill moves:
- The voice. A premium or cloned voice can cost 10–20× a standard one. If a demo used the cheap voice and production uses the premium one, the per-minute cost is not the same number.
- The model. Routing every turn through a frontier model is dramatically more expensive than tiering — a small model for routing and simple turns, a larger one only when the conversation needs it.
- The destination. A call to a mobile number in a hard-to-reach country can cost many times a domestic landline minute. A per-minute quote assumes a destination.
- The minimums and overage. Monthly platform minimums, per-seat fees, and overage rates above a bundled allotment often matter more than the marginal per-minute rate once you are at volume.
- The p95, not the p50. An agent that averages 700 ms per turn but spikes to 1.5 seconds on every tool call burns more minutes per resolved conversation — latency is a cost lever, not just a quality one.
> "The honest answer to 'what does an AI voice agent cost' is that it depends on four moving parts and who is marking each one up — so compare the whole stack, on your own configuration, not a marketed per-minute headline."
Ask any vendor — us included — for the per-minute cost on your voice, your model, and your destinations, and for the p95 under load, not just the marketed p50. See the per-stage methodology behind those turn numbers on the latency benchmark page.
The three ways vendors bill for it
Under all the marketing, AI voice agent pricing comes in three shapes, and the right one depends on how much of the stack you want to own.
| Pricing model | You assemble and pay for | Best for | The trade-off |
|---|---|---|---|
| À la carte specialist | Each vendor separately — STT, LLM, TTS, and your own telephony | Engineering teams that want to tune every layer | Lowest marginal control cost, highest integration and reconciliation overhead |
| Bundled voice-agent platform | One per-minute rate that wraps the pipeline; bring your own numbers | Teams that want a phone agent live fast | Simple bill, but telephony and the rest of your stack live elsewhere |
| All-in-one CPaaS/CCaaS | One account for voice agents, messaging, and carrier termination | Teams running voice next to SMS, WhatsApp, email, and a contact center | One bill and one vendor, in exchange for less à la carte model choice |
The à la carte route can show the lowest marginal per-minute number and still cost the most to run, because you are now operating four contracts, four dashboards, and four invoices — and the integration time is real engineering spend that never shows up on any vendor's pricing page.
How to calculate ROI, not just cost
Cost per minute is an input; the number that decides the investment is cost per resolved conversation measured against the fully-loaded cost of a human handling the same call. The framework:
- Fully-loaded human cost per minute. Take agent wages plus benefits, tooling, management, and idle time, divided by talk-minutes. In many support markets this lands well above $0.50 per talk-minute — often $0.75–$1.50+ once overhead is counted.
- All-in AI cost per minute. The four components above plus the platform fee, on your real configuration.
- Deflection / resolution rate. The share of conversations the agent finishes without a human. An agent that resolves 60% of calls at a fifth of the per-minute cost is a very different return from one that resolves 20% and escalates the rest.
- The math.
Monthly saving ≈ (calls handled by AI) × (avg minutes) × (human cost/min − AI cost/min), minus one-time build and tuning time.
An illustrative example — your numbers will differ: a line taking 10,000 calls a month at 4 minutes each, where the agent resolves 60% and the human-handled minute costs $0.90 against an all-in AI cost of $0.12, saves on the order of 10,000 × 0.60 × 4 × ($0.90 − $0.12) ≈ $18,700 per month before build time. Treat that as a shape, not a promise: the sensitivity is almost entirely in the resolution rate and the human cost baseline, so measure both on a pilot before you extrapolate.
The other half of ROI rarely shows up in a spreadsheet: after-hours and overflow coverage the business could not staff at all, faster time-to-answer that lifts conversion, and consistent handling that a stretched queue cannot match. Those are real returns; just do not double-count them against the per-minute saving.
How Orbit prices AI voice agents
Orbit by Devotel bills AI voice agents pay-as-you-go, with voice agents, SMS, WhatsApp, RCS, email, embeddable video, and a built-in contact center on one account and one bill — so you are not reconciling a voice-AI vendor, a messaging vendor, and a telephony provider as three invoices. Published per-country voice rates are on the voice pricing page, and the full plan structure is on the pricing page.
Two things keep the cost side honest. The pipeline runs on Orbit's own streamed speech-to-text → model → text-to-speech stack with per-agent model tiering, prompt caching, and speculative execution against a published ~1.1-second p50 turn target — a conservative internal SLO, with the per-stage methodology open on the latency benchmark page. And the telephony line rides Devotel's own wholesale softswitch, which connects to 500+ global carriers directly rather than reselling an aggregator hop — one fewer intermediary marking up the per-minute rate, and outbound calls terminate over that same softswitch rather than a carrier you bring or a resold path. See the AI voice agents overview, the cloud phone system, and the CPaaS overview.
For a platform-by-platform view of where the agents themselves differ, see Best AI voice agents in 2026; for a broader provider shortlist, the best Twilio alternatives guide; and for the latency-as-cost angle, how to reduce AI voice agent latency.
Frequently asked questions
How much does an AI voice agent cost per minute in 2026?
There is no single rate — it is the sum of speech-to-text, the language model, text-to-speech, and the telephony minute, plus a platform fee. For a typical support or booking configuration the all-in cost commonly lands between roughly $0.05 and $0.30+ per talk-minute, and it moves with your voice choice, model tier, and call destinations. Get a quote on your own configuration rather than trusting a marketed headline.
What makes up the cost of an AI voice agent?
Four metered parts. Speech-to-text transcribes the caller (~$0.005–$0.02/min), the language model reasons and replies (priced per token and highly variable), text-to-speech speaks the answer (~$0.01–$0.30/min depending on the voice), and telephony carries the call (~$0.004–$0.02/min in the US, more to hard-to-reach destinations). A platform or orchestration fee sits on top of those four.
Are AI voice agents cheaper than human agents?
Usually, once resolution rate is factored in. The fully-loaded cost of a human talk-minute often exceeds $0.50 and can reach $1.50+ with overhead, versus an all-in AI cost frequently in the $0.05–$0.30 range. The return depends on how many conversations the agent resolves without escalating — a high deflection rate is what turns a lower per-minute cost into real savings, so measure it on a pilot before extrapolating.
Why do AI voice agent prices vary so much between vendors?
Because vendors bundle the four components differently and mark up different layers. A marketed per-minute figure usually assumes a standard voice, a mid-tier model, a short prompt, and a domestic call; premium voices, frontier models, long prompts, and international destinations each raise it independently. À la carte specialists can show a low marginal rate but shift integration and telephony cost onto you.
How do I estimate ROI for an AI voice agent?
Compare cost per resolved conversation, not cost per minute. Multiply the calls the agent handles by average minutes by the gap between the fully-loaded human cost per minute and the all-in AI cost per minute, then subtract one-time build and tuning time. The result is most sensitive to the resolution rate and the human cost baseline, so pin both from a pilot before scaling.
Does Orbit charge separately for speech-to-text, the model, and text-to-speech?
No — Orbit bills AI voice agents pay-as-you-go on one account alongside every other channel, so you are not stitching together and reconciling separate STT, model, TTS, and telephony invoices. Published per-country voice rates are on the voice pricing page, and outbound calls terminate over Devotel's own wholesale softswitch rather than a resold aggregator hop.
The takeaway
AI voice agent pricing only looks confusing until you separate it into its four parts — speech-to-text, the language model, text-to-speech, and telephony — and ask who is marking each one up. Price the whole stack on your own voice, model, and destinations; judge the p95 under load, not the marketed p50; and measure ROI as cost per resolved conversation against a fully-loaded human minute, not as a per-minute number in isolation. Do that, and the "$0.09 a minute" headline stops mattering and the real economics come into focus.
Published 20 July 2026.