Skip to main content
Back to blog

How to benchmark an AI voice agent — the four metrics, where to ground them, and a benchmark sheet that survives procurement

Human QA samples calls and scores behavior; benchmarking an AI voice agent scores every conversation against outcome metrics. The four that decide vendors — deflection, first-contact resolution, fallback-to-human rate, and hallucination containment — plus where each is grounded in Orbit's shipped surfaces and a benchmark sheet format you can run in a two-week pilot.

Orbit Editorial Team

A vendor benchmark for an AI voice agent is an instrumented pilot, not a demo. The four numbers that separate vendors are deflection, first-contact resolution, fallback-to-human rate, and hallucination containment, and each one is measurable on your own call mix within two weeks — provided the platform exposes per-intent outcome data rather than an aggregate "handled calls" counter. This post lays out why the borrowed human-QA playbook mis-measures a machine, the four metrics with working definitions, where each is grounded in Devotel Orbit's shipped product, and a benchmark sheet format you can copy. The other two posts in this series — AI voice agent pricing and how to QA an AI voice agent — cover cost shape and the rubric/practice loop; this one covers the measurement layer that sits between them during a buying decision.

Why support teams benchmark AI voice agents differently from human QA

Human-agent QA solves a listening problem. Listeners are expensive, so the program samples two calls per agent per week, scores behavior on a calibrated form, and coaches the misses away. Every design decision in that playbook answers one constraint: a person must listen to each scored call.

An AI voice agent has no such constraint — every conversation is a transcript — and the borrowed playbook mis-measures as a result:

  • Behavior scoring is the wrong unit for a machine. A human agent drifts in behavior; an AI agent drifts in configuration. Benchmarking the machine means measuring outcomes (did the caller's task complete) rather than behaviors (was the greeting warm), because a prompt edit can move one intent's completion rate without changing anything you would hear as "behavior."
  • Sampling hides exactly what a benchmark must find. A 2% sample of human calls is acceptable because human performance is roughly uniform across the week. AI performance is uniform across nothing: one intent regresses while three hold. A benchmark that samples 2% of conversations will miss a broken booking flow behind healthy support scores — the whole point of the exercise is per-intent denominators.
  • Aggregate numbers launder vendor claims. A marketed "80% deflection" is an aggregate. Your benchmark has to decompose it by intent before it predicts anything about your deployment, because the vendor's aggregate was computed on someone else's call mix and possibly on a phantom demo bot, not on production traffic.

The practical consequence: a benchmark closer to regression testing than to coaching — score everything, split by intent, and treat the vendor's aggregate as a claim to verify rather than a number to adopt.

The four metrics that decide vendors

Every buyer query on this topic eventually collapses into four numbers. Define them once, in writing, before the pilot starts — vendors who cannot break their platform's data into these four shapes are telling you something.

Deflection rate

Definition: conversations the agent completes end-to-end, divided by all conversations that entered. The economic metric — it is the multiplier that turns a lower per-minute cost into savings in the ROI arithmetic, which is why it is also the number vendors are most tempted to inflate.

What to watch: deflection counts only completed tasks. A conversation where the agent talked for four minutes and then transferred anyway is not deflected; it is a four-minute toll on the queue. Normalize it per intent before comparing vendors — a vendor strong on FAQ deflection may be weak on the transactional intents (booking, order status, account changes) that carry your volume.

First-contact resolution (FCR)

Definition: conversations resolved on the first contact with no repeat contact and no human touch within your reopen window (seven days is a reasonable default for support queues). Deflection says the agent finished the task; FCR says it stayed finished.

What to watch: FCR is where overstated deflection gets caught. An agent that books an appointment badly — wrong location, no recitation, wrong timezone — shows deflection on the vendor's dashboard and a repeat call on yours. Measure FCR by joining voice outcomes against repeat contacts in the same unified inbox rather than trusting the voice channel's self-report.

Fallback-to-human rate

Definition: the share of conversations the agent hands to a person, split by why the fallback fired — a configured escalation rule, a low-confidence threshold, a dead tool call, or the caller demanding a human. The composition matters more than the headline: a 15% fallback rate that is mostly "caller asked for a person" is healthy; the same rate that is mostly empty tool responses is a misconfigured agent.

What to watch: the handoff itself is part of the metric. A fallback that transfers cold — the human picks up with no summary and the caller repeats everything — punishes the caller twice, so score the handoff summary as part of fallback quality, and break the rate by intent rather than reading one aggregate.

Hallucination containment

Definition: the rate at which the agent states policy, prices, or facts that your knowledge sources do not support, and the containment behavior when it does not know — a clean escalation or refusal is contained; inventing an answer is not. This is the newest of the four metrics and the one purchasing teams most often skip because it is the hardest to source from a platform that does not score it natively.

What to watch: containment is scored per conversation, not per claim — one invented policy answer in a twelve-turn conversation contaminates the whole outcome. Ground it in your own knowledge base during the pilot: seed the agent with your policy documents, then run a fixed question set (covered in the benchmark sheet below) and count unsupported claims by hand if the platform does not score them for you. Below 1% unsupported-claim rate per conversation is a sane pilot gate; the acceptable ceiling is yours to set, not the vendor's.

Where each metric is grounded in shipped product

A benchmark is only as repeatable as the surfaces it reads from. Orbit grounds each metric in a live dashboard surface, so the numbers you collect during a pilot are the same numbers you operate against in production — not a vendor-side export you cannot regenerate:

  • Deflection and resolution — the dashboard's inbox and voice conversation analytics break resolution and escalation by intent, and the Quality hub rolls LLM-judged outcomes into a per-agent scorecard across voice and inbox together, on every paid tier with no add-on SKU. Intent-level denominators are the first-class view, so the per-intent split the section above demands is the default, not a custom report.
  • First-contact resolution — because Orbit runs voice, SMS, WhatsApp, and email on one phone system and one unified inbox, repeat contacts are visible within the same account rather than in a second vendor's system. You measure the agent against the same seven-day reopen window you already use for human queues; sorting the agent scorecards by score shows which intent is generating the repeats.
  • Fallback-to-human — the per-agent scorecards and the Evaluations surface record every scored conversation one click from the aggregate, so a fallback spike decomposes into the individual handoffs that caused it; the rubric set you run in pilot can include an explicit pass/fail check on handoff-summary quality. The broader rubric design and how the Agents/Evaluations/Practice surfaces close the tuning loop are covered in the QA post — do not re-read them here; the benchmark's use is narrower: verify that per-conversation verdicts exist before you trust the aggregate.
  • Hallucination containment — the agent answers only from the knowledge sources you configure, and the same rubric mechanism scores unsupported claims as an explicit fail-condition check rather than folding them into a tone score. During the pilot, configure the knowledge documents that actually govern your policies, then let the rubric flag the unsupported-answer cases while you review the flagged conversations.

Two cross-cutting checks belong on this list too, because a benchmark that skips them under-prices the agent:

  • Latency under load. An agent that resolves well at 700 ms per turn in a demo can regress past awkward dead-air under production concurrency. Orbit publishes a per-stage methodology and a ~1.1-second p50 internal SLO on the latency benchmark page — treat any vendor that will not show a p95 turn time under load as unmeasured. The cost side of that same number is why the pricing post calls latency a cost lever, not just a quality one.
  • Telephony path. The benchmark should ride the production carrier path you will actually use — Orbit's voice pipeline terminates over Devotel's own wholesale softswitch rather than a resold aggregator hop, which is one fewer unmeasured intermediary between your agent and the handset. A pilot run on a demo path is not a benchmark of the production system.

How to set up the benchmark sheet

The sheet has four regions and runs in a two-week pilot on live or replayed traffic. Keep it to one screen per vendor; a benchmark you cannot scan in two minutes will not survive a procurement meeting.

Region 1 — denominator. Total conversations, split by intent, with the replays/production mix marked. Every metric divides by this; a vendor deck that reports deflection without intent counts is withholding the denominator.

Region 2 — the four metrics, per intent. One row per intent, four columns (deflection, FCR, fallback rate, hallucination containment), plus the aggregate in the last row. Add two annotation columns: the dominant fallback reason, and the reopen-window FCR deviation from the agent's self-reported completion. This region answers "which vendor" on your call mix, not on theirs.

Region 3 — fixed adversarial question set. Twenty to forty questions you never change across vendors: half answerable from your knowledge sources, a quarter deliberately out of scope, a quarter intentionally ambiguous. Score it manually once, then use it as the regression suite that survives the pilot — when a vendor's prompt update lands, rerun the fixed set before the change touches production traffic.

Region 4 — operability gates. Binary, non-negotiable rows: per-intent export available, p95 latency under load published or measurable, rubric customization supported, handoff summary configurable, knowledge sources versioned. A vendor fails gates, not metrics, when it cannot decompose the aggregate — which is also why the QA post's practice-loop detail reads as a buying criterion here: the remediation path has to exist before the benchmark's findings can change anything.

Run the sheet for two weeks per vendor on the same intent mix. The comparison that matters is not vendor A's aggregate against vendor B's aggregate but each vendor's worst intent — that is the number that moves first in production.

Frequently asked questions

What are the most important metrics for benchmarking an AI voice agent?

Four: deflection rate (tasks the agent completes end-to-end), first-contact resolution (it stays resolved within your reopen window), fallback-to-human rate split by reason, and hallucination containment (unsupported claims per conversation). Track all four per intent — aggregate numbers hide a broken intent behind healthy ones.

How long should an AI voice agent benchmark pilot run?

Two weeks per vendor on live production intent mix is the working answer — long enough to capture the weekly variance in call mix and any prompt or knowledge updates the vendor ships mid-pilot, short enough to hold the comparison honest. Run every vendor against the same fixed adversarial question set so the comparison survives personnel and platform changes.

Why not just use human QA methods to benchmark AI agents?

Because the constraint they were designed around — that listening is expensive — no longer exists, and because sampling, behavior scoring, and weekly aggregates each hide the failure modes a machine actually produces: per-intent regressions and data-independent claims. Benchmarking a machine is regression testing: score everything, split by intent, gate changes on a fixed question set.

What is a good fallback-to-human rate for an AI voice agent?

The composition matters more than the number. Fallback driven by caller preference is healthy at any level; fallback driven by dead tool calls or confidence thresholds is configuration debt. Decompose it by intent and reason first — a 15% fallback that is mostly "caller asked for a person" beats an 8% fallback that is mostly empty tool responses.

How does Orbit expose benchmark data during a pilot?

The dashboard's Quality hub, per-agent scorecards, and Evaluations surface report per-intent resolution, escalation, and LLM-judged outcomes across voice and inbox — the same views the pilot uses are the views production operations eventually run against, with no add-on SKU for the scoring. Latency is published with a per-stage methodology on the latency benchmark page.

Where to go next

The numbers this post measures are the inputs to the two adjacent decisions: the AI voice agent pricing guide turns a verified deflection rate into cost per resolved conversation, and the QA post turns the benchmark's fixed question set into the regression gate that keeps production stable after you pick a vendor. The voice AI agents overview covers the pillar itself. Run the two-week pilot, fill the sheet's four regions, and let the worst intent decide.

How to benchmark an AI voice agent — the four metrics, where to ground them, and a benchmark sheet that survives procurement — Orbit by Devotel