Skip to main content
Back to resources

AI Voice Agent Cost and ROI: A Complete Breakdown for Buyers

An evergreen breakdown of what an AI voice agent actually costs — the four per-minute components plus the platform fee — and how to calculate its ROI against the fully-loaded cost of a human-handled call, with the levers that move the math either way.

Orbit Editorial Team

Short answer: an AI voice agent's all-in cost is the sum of four per-minute components (speech-to-text, the language model, text-to-speech, and the telephony minute) plus a platform fee, and a typical support configuration runs roughly $0.05–$0.30+ per talk-minute all-in. Its ROI comes from cost per resolved conversation, not cost per minute: weigh that all-in rate against a fully-loaded human talk-minute of roughly $0.50–$1.50+, scaled by the share of calls the agent finishes without escalating. This page breaks the cost into its parts, then walks the ROI framework step by step, so the marketed "$0.09-a-minute" headline stops being the number you sign on.

The four components of the per-minute cost

Every real-time voice agent runs the same pipeline on every turn of every call: it transcribes what the caller said, decides what to reply, synthesizes the reply aloud, and carries it all over the phone network. Each stage is metered separately, and each is a line a vendor can mark up.

ComponentWhat it pays forDirectional rangeWhat moves it
Speech-to-text (STT)Transcribing caller audio, per minute of audio~$0.005–$0.02/minStreaming vs batch, language, custom vocabulary
Language model (LLM)Tokens in and out per turn (the reasoning and the reply)Highly variable, per 1M tokensModel tier, prompt length, tool calls, prompt caching
Text-to-speech (TTS)Synthesizing the agent's spoken reply~$0.01–$0.30/minStandard vs premium or cloned voices
TelephonyCarrying the call over the phone network~$0.004–$0.02/min (US)Destination country, carrier hops, inbound vs outbound
Platform / orchestrationPipelines, tools, dashboards, supportPer-minute uplift or flat planBundled vs à la carte, volume tier, SLA

The LLM line is the least predictable: a long system prompt repeated every turn, or a tool call that hauls a large context in, quietly doubles the token bill; prompt caching and per-agent model tiering pull it back down. Telephony is the line most buyers forget — a voice agent rides a real phone network, and an international destination multiplies the per-minute rate that a domestic headline quote assumes.

The full component math, with the per-stage time budget that ties latency to burned minutes, is in the AI voice component TCO guide. For why a single quoted rate misleads regardless of these parts, the AI voice agent pricing guide runs the three vendor billing models side by side.

Why the marketed per-minute headline misleads

A headline "$0.09 per minute" is the best case of a five-line bill wearing one line's clothes. It assumes a standard voice, a mid-tier model, a short prompt, a domestic call, and a p50 (median) turn. Change one input and the bill moves; change two and it doubles. The fix is procedural: ask every vendor, Orbit included, for an all-in per-minute figure on your voice, your model, your destinations, and the p95 turn under load; then compare whole stacks, not marketings.

Independent analyses of the unbundled ("bring your own AI stack") shape, where the vendor headlines only its platform fee and passes STT, LLM, TTS, and telephony through from other vendors, put that headline-vs-real gap at roughly 3–8×. The true all-in cost of voice AI and SMS explainer collects those published analyses and gives you a calculator to rerun the math on your own configuration.

Calculating ROI, step by step

Cost per minute is an input. The number that decides the investment is cost per resolved conversation, measured against the fully-loaded cost of a human handling the same call.

1. Price the human alternative honestly

A human talk-minute is not wages divided by spoken minutes. Take wages plus payroll burden, tooling, management, and shrinkage (breaks, training, idle queue time; industry contact-center shrinkage runs 30–35%), then divide by the minutes a seat actually spends talking. That lands well above $0.50 in most support markets, often $0.75–$1.50+ once overhead is counted. The break-even math, tier by interaction complexity, is worked through in the TCO-vs-human-agent post.

2. Price the all-in AI minute on your configuration

Sum the four components above plus the platform fee, using your real voice, model tier, and destinations, not a marketed headline. Orbit's own published per-minute figure is worked as an example below so each model has one concrete anchor.

3. Estimate resolution rate conservatively

The deflection/resolution rate, meaning the share of conversations the agent finishes without a human, is the single most over-promised number in a vendor's demo. Model the low end; even a modest pilot beats a projection.

4. Run the math

Monthly saving ≈ (calls handled by AI) × (average minutes) ×
                 (human cost per minute − AI cost per minute)
                 − one-time build and tuning

Worked example, a shape not a promise: a line takes 10,000 calls a month at 4 minutes each, the agent resolves 60%, the human minute costs $0.90, and the all-in AI minute costs $0.12. That is 10,000 × 0.60 × 4 × (0.90 − 0.12) ≈ $18,700 per month. The interactive cost calculator runs the same arithmetic on your own stack. Gartner's $80-billion-2026 labor-cost forecast and its rise from an estimated 1.6% automated interactions in 2022 to 10% by 2026 is the aggregation view of that same per-line math (source below).

The levers that move the ROI either way

  • Resolution rate is the dominant term, and it compounds against fragmentation, not for it. An agent with no sight of a customer's prior SMS or WhatsApp thread re-asks answered questions and escalates, dragging realized resolution below the demo number. Voice plus messaging on one account, with one contact record, is the shared-context condition the math above assumes.
  • Telephony path. Every carrier hop between the agent and the handset adds margin to the fourth component. Orbit terminates outbound calls over Devotel's own wholesale softswitch, 500+ global carriers direct, rather than reselling an aggregator hop; one intermediary fewer marking the minute.
  • Model routing. Per-agent model tiering (a small model for routing and simple turns, a frontier model only when judgment is needed) with prompt caching halves the LLM line that otherwise dominates the spread; the model-cost explainer maps which turns belong on which tier.

Where Orbit prices and sits

Orbit by Devotel bills AI voice agents pay-as-you-go, at $0.014 per voice minute an agent is on a call, $0.007 per AI agent message, and an optional $0.005 per resolved outcome if you opt into outcome-based billing, with no platform fee, no seat fee, and no monthly minimum, all published per-country on the pricing page and the per-destination voice rate cards like voice pricing in the US. Voice runs alongside SMS, WhatsApp, RCS, email, video, and the contact center on one account and one bill, so the consolidation cost the four-component breakdown warns about collapses into one rate card. See the AI voice agents overview for the agent surface and the best AI voice agents in 2026 round-up for how the major platforms bill.

Frequently asked questions

How much does an AI voice agent cost per minute in 2026?

There is no single rate. It is the sum of speech-to-text, the language model, text-to-speech, and the telephony minute, plus a platform fee. A typical support or booking configuration commonly lands between roughly $0.05 and $0.30+ per talk-minute all-in, and it moves with voice choice, model tier, and destinations. Get an all-in quote on your own configuration.

What makes up the cost of an AI voice agent?

Four metered parts. Speech-to-text transcribes the caller (~$0.005–$0.02/min), the language model reasons and replies (priced per token and highly variable), text-to-speech speaks the answer (~$0.01–$0.30/min depending on the voice), and telephony carries the call (~$0.004–$0.02/min in the US). A platform or orchestration fee sits on top.

What's the ROI formula for an AI voice agent?

Monthly saving ≈ (calls handled by AI) × (average minutes) × (human cost per minute − AI cost per minute), minus one-time build and tuning. The result is most sensitive to resolution rate and the human cost baseline, so pin both from a pilot before scaling.

Are AI voice agents cheaper than human agents?

Usually, once resolution rate is factored in. A fully-loaded human talk-minute commonly exceeds $0.50, versus an all-in AI minute frequently in the $0.05–$0.30 band. The return depends on how many conversations the agent resolves without escalating, so measure that on a pilot.

How does Orbit bill for AI voice agents?

Pay-as-you-go: $0.014 per voice minute an agent is on a call, $0.007 per AI agent message, and an optional $0.005 per resolved outcome under outcome-based billing, with no platform or seat fee; voice, messaging, and the contact center share one account and one bill, so there is no four-vendor reconciliation tax on top.

Why do AI voice agent prices vary so much between vendors?

Because vendors bundle the four components differently and mark up different layers. A marketed figure usually assumes a standard voice, a mid-tier model, and a domestic destination; premium voices, frontier models, and international destinations each raise it independently, and à la carte specialists shift integration and telephony cost onto you.

Sources and further reading

Published 25 September 2026. Part of the Orbit resources library: foundational guides for teams building on communications infrastructure.

Ready to build?

Orbit puts voice, messaging, and AI agents on one platform with one pay-as-you-go bill. Start free — no credit card required.

AI Voice Agent Cost and ROI: A Complete Breakdown for Buyers — Orbit by Devotel