Short answer: an AI voice agent's all-in cost is the sum of four per-minute components (speech-to-text, the language model, text-to-speech, and the telephony minute) plus a platform fee, and a typical support configuration runs roughly $0.05–$0.30+ per talk-minute all-in. Its ROI comes from cost per resolved conversation, not cost per minute: weigh that all-in rate against a fully-loaded human talk-minute of roughly $0.50–$1.50+, scaled by the share of calls the agent finishes without escalating. This page breaks the cost into its parts, then walks the ROI framework step by step, so the marketed "$0.09-a-minute" headline stops being the number you sign on.
The four components of the per-minute cost
Every real-time voice agent runs the same pipeline on every turn of every call: it transcribes what the caller said, decides what to reply, synthesizes the reply aloud, and carries it all over the phone network. Each stage is metered separately, and each is a line a vendor can mark up.
| Component | What it pays for | Directional range | What moves it |
|---|---|---|---|
| Speech-to-text (STT) | Transcribing caller audio, per minute of audio | ~$0.005–$0.02/min | Streaming vs batch, language, custom vocabulary |
| Language model (LLM) | Tokens in and out per turn (the reasoning and the reply) | Highly variable, per 1M tokens | Model tier, prompt length, tool calls, prompt caching |
| Text-to-speech (TTS) | Synthesizing the agent's spoken reply | ~$0.01–$0.30/min | Standard vs premium or cloned voices |
| Telephony | Carrying the call over the phone network | ~$0.004–$0.02/min (US) | Destination country, carrier hops, inbound vs outbound |
| Platform / orchestration | Pipelines, tools, dashboards, support | Per-minute uplift or flat plan | Bundled vs à la carte, volume tier, SLA |
The LLM line is the least predictable: a long system prompt repeated every turn, or a tool call that hauls a large context in, quietly doubles the token bill; prompt caching and per-agent model tiering pull it back down. Telephony is the line most buyers forget — a voice agent rides a real phone network, and an international destination multiplies the per-minute rate that a domestic headline quote assumes.
The full component math, with the per-stage time budget that ties latency to burned minutes, is in the AI voice component TCO guide. For why a single quoted rate misleads regardless of these parts, the AI voice agent pricing guide runs the three vendor billing models side by side.
Why the marketed per-minute headline misleads
A headline "$0.09 per minute" is the best case of a five-line bill wearing one line's clothes. It assumes a standard voice, a mid-tier model, a short prompt, a domestic call, and a p50 (median) turn. Change one input and the bill moves; change two and it doubles. The fix is procedural: ask every vendor, Orbit included, for an all-in per-minute figure on your voice, your model, your destinations, and the p95 turn under load; then compare whole stacks, not marketings.
Independent analyses of the unbundled ("bring your own AI stack") shape, where the vendor headlines only its platform fee and passes STT, LLM, TTS, and telephony through from other vendors, put that headline-vs-real gap at roughly 3–8×. The true all-in cost of voice AI and SMS explainer collects those published analyses and gives you a calculator to rerun the math on your own configuration.
Calculating ROI, step by step
Cost per minute is an input. The number that decides the investment is cost per resolved conversation, measured against the fully-loaded cost of a human handling the same call.
1. Price the human alternative honestly
A human talk-minute is not wages divided by spoken minutes. Take wages plus payroll burden, tooling, management, and shrinkage (breaks, training, idle queue time; industry contact-center shrinkage runs 30–35%), then divide by the minutes a seat actually spends talking. That lands well above $0.50 in most support markets, often $0.75–$1.50+ once overhead is counted. The break-even math, tier by interaction complexity, is worked through in the TCO-vs-human-agent post.
2. Price the all-in AI minute on your configuration
Sum the four components above plus the platform fee, using your real voice, model tier, and destinations, not a marketed headline. Orbit's own published per-minute figure is worked as an example below so each model has one concrete anchor.
3. Estimate resolution rate conservatively
The deflection/resolution rate, meaning the share of conversations the agent finishes without a human, is the single most over-promised number in a vendor's demo. Model the low end; even a modest pilot beats a projection.
4. Run the math
Monthly saving ≈ (calls handled by AI) × (average minutes) ×
(human cost per minute − AI cost per minute)
− one-time build and tuningWorked example, a shape not a promise: a line takes 10,000 calls a month at 4 minutes each, the agent resolves 60%, the human minute costs $0.90, and the all-in AI minute costs $0.12. That is 10,000 × 0.60 × 4 × (0.90 − 0.12) ≈ $18,700 per month. The interactive cost calculator runs the same arithmetic on your own stack. Gartner's $80-billion-2026 labor-cost forecast and its rise from an estimated 1.6% automated interactions in 2022 to 10% by 2026 is the aggregation view of that same per-line math (source below).
The levers that move the ROI either way
- Resolution rate is the dominant term, and it compounds against fragmentation, not for it. An agent with no sight of a customer's prior SMS or WhatsApp thread re-asks answered questions and escalates, dragging realized resolution below the demo number. Voice plus messaging on one account, with one contact record, is the shared-context condition the math above assumes.
- Telephony path. Every carrier hop between the agent and the handset adds margin to the fourth component. Orbit terminates outbound calls over Devotel's own wholesale softswitch, 500+ global carriers direct, rather than reselling an aggregator hop; one intermediary fewer marking the minute.
- Model routing. Per-agent model tiering (a small model for routing and simple turns, a frontier model only when judgment is needed) with prompt caching halves the LLM line that otherwise dominates the spread; the model-cost explainer maps which turns belong on which tier.
Where Orbit prices and sits
Orbit by Devotel bills AI voice agents pay-as-you-go, at $0.014 per voice minute an agent is on a call, $0.007 per AI agent message, and an optional $0.005 per resolved outcome if you opt into outcome-based billing, with no platform fee, no seat fee, and no monthly minimum, all published per-country on the pricing page and the per-destination voice rate cards like voice pricing in the US. Voice runs alongside SMS, WhatsApp, RCS, email, video, and the contact center on one account and one bill, so the consolidation cost the four-component breakdown warns about collapses into one rate card. See the AI voice agents overview for the agent surface and the best AI voice agents in 2026 round-up for how the major platforms bill.
Frequently asked questions
How much does an AI voice agent cost per minute in 2026?
There is no single rate. It is the sum of speech-to-text, the language model, text-to-speech, and the telephony minute, plus a platform fee. A typical support or booking configuration commonly lands between roughly $0.05 and $0.30+ per talk-minute all-in, and it moves with voice choice, model tier, and destinations. Get an all-in quote on your own configuration.
What makes up the cost of an AI voice agent?
Four metered parts. Speech-to-text transcribes the caller (~$0.005–$0.02/min), the language model reasons and replies (priced per token and highly variable), text-to-speech speaks the answer (~$0.01–$0.30/min depending on the voice), and telephony carries the call (~$0.004–$0.02/min in the US). A platform or orchestration fee sits on top.
What's the ROI formula for an AI voice agent?
Monthly saving ≈ (calls handled by AI) × (average minutes) × (human cost per minute − AI cost per minute), minus one-time build and tuning. The result is most sensitive to resolution rate and the human cost baseline, so pin both from a pilot before scaling.
Are AI voice agents cheaper than human agents?
Usually, once resolution rate is factored in. A fully-loaded human talk-minute commonly exceeds $0.50, versus an all-in AI minute frequently in the $0.05–$0.30 band. The return depends on how many conversations the agent resolves without escalating, so measure that on a pilot.
How does Orbit bill for AI voice agents?
Pay-as-you-go: $0.014 per voice minute an agent is on a call, $0.007 per AI agent message, and an optional $0.005 per resolved outcome under outcome-based billing, with no platform or seat fee; voice, messaging, and the contact center share one account and one bill, so there is no four-vendor reconciliation tax on top.
Why do AI voice agent prices vary so much between vendors?
Because vendors bundle the four components differently and mark up different layers. A marketed figure usually assumes a standard voice, a mid-tier model, and a domestic destination; premium voices, frontier models, and international destinations each raise it independently, and à la carte specialists shift integration and telephony cost onto you.
Sources and further reading
- AI voice agent pricing in 2026: cost breakdown, comparisons & ROI: the time-ordered field-note version of this guide, with the three vendor billing models side by side.
- AI voice component TCO: time-and-cost math across STT, LLM, and TTS: the per-stage time budget that ties latency to burned minutes.
- AI voice agent TCO vs human agent TCO: break-even math: the hybrid escalation model behind step 1.
- The true all-in cost of voice AI and SMS (interactive calculator): the published independent analyses behind the 3–8× headline-vs-real gap.
- Which AI models make voice agents cost-effective, and why: the tier map that resolves the model-routing lever.
- Gartner Newsroom: Gartner Predicts Conversational AI Will Reduce Contact Center Agent Labor Costs by $80 Billion in 2026: the $80 billion labor-cost projection and the forecast that automated agent interactions rise from an estimated 1.6% in 2022 to 10% by 2026.
Published 25 September 2026. Part of the Orbit resources library: foundational guides for teams building on communications infrastructure.