Pick up a buying guide for AI voice agents and it will hand you a list of vendors to compare head-to-head. That skips the decision that actually gates the shortlist: the vendors are not interchangeable members of one category. They cluster into four distinct families, each with a different build model, a different latency posture, and a different answer to where telephony lives. Two vendors from different families compare on paper but behave unlike one another in production, and a head-to-head across families is a category error that costs the buyer the first two months of a pilot. This post names the four families, maps the five evaluation checks to each, and tells you which family to shortlist before you open a vendor page.
Devotel Orbit ships the fourth family, so we have a position here, and the rest of this post argues the buyer-side case rather than ours where we can help it. The named long-form comparison across eight vendors sits in how to evaluate an AI voice agent platform; the measurement layer that decides inside a family is in how to benchmark an AI voice agent. What follows is the family sort both of those assume you have already run.
The four families
Every AI voice-agent vendor we track sits in one of four shapes. The family's borders are drawn by build model and telephony ownership, not by the marketing copy on the vendor's front page:
- Packaged outbound. The vendor is built for scripted, high-volume outbound with a deterministic flow builder — the agent stays on script, the model does not free-associate. Bland AI's Pathways is the canonical example. Latency is reported but secondary; the selling point is predictable cost at campaign scale.
- Managed inbound. A no-code or low-code builder gets a non-engineering team live in an afternoon, with inbound reception, booking, and routing templates ready made. Synthflow and the everyday Retell AI self-serve tier live here. The family trades deep customization for time-to-first-call, and telephony is a bring-your-own problem the vendor resells or leaves to you.
- Developer-first assembly. A full-control API exposes every knob — model, voice provider, telephony, latency tuning — and you assemble and operate the stack yourself. Vapi and Twilio's ConversationRelay plus a bring-your-own-LLM pattern are the two canonical shapes. The family is the right answer when the build model is "an engineering team owns this," and the wrong answer when it is not.
- Full communication platform. The voice agent runs beside the rest of the communications stack — SMS, WhatsApp, RCS, email, video, and a contact center — on one account and one bill, with owned carrier termination under it. Orbit by Devotel sits here, with Sinch and a few incumbent CPaaS/CCaaS suites at the acquired-product end of the same shape. The family is the answer when voice is one of several channels your customers use, not the only one.
The sort is not a quality judgment. A packaged-outbound vendor is a worse inbound receptionist; a developer-first platform is a worse "live by Friday" answer. Each family is excellent at exactly one buying problem and quietly bad at the other three.
Why a cross-family head-to-head misleads
The failure mode of the vendor-first guide is that it scores these families on the same table and reads the ranking as a verdict. Three specific distortions follow:
- Latency numbers cross indexes. A packaged-outbound vendor reports latency on its scripted flow; a managed-inbound vendor reports it on a canned template; a developer-first vendor reports a tuned p50 your own integration may never reach. Comparing them as one axis rewards the shape, not the vendor.
- Build model is priced as a feature, not a constraint. "Drag-and-drop builder" and "full API access" are answers to who builds, not checkboxes. A head-to-head that treats both as line items lets a marketing page outscore your team's actual staffing.
- Telephony is dismissed as an integration detail. Some families assume you bring your own carrier; one family terminates over its own wholesale softswitch. When the family compares "telephony included" against "telephony excluded" as if both were neutral facts, the lifetime cost and reliability difference between them is erased from the decision.
The fix is cheap: sort the list into families first, and run the head-to-head only inside the family you picked.
Which family fits which buyer
The family decision reduces to four questions, in order:
- Is the agent primarily outbound on a script? If yes, start in the packaged-outbound family; every other shape carries complexity you will not use.
- Does a non-engineering team own the build? If yes, the managed-inbound family gets you live without a sprint. If the team is engineering-led, managed builders become a ceiling rather than a shortcut.
- Do you need every layer under your own control? If yes, developer-first assembly is the price of admission, and it is worth paying when the latency, model, or telephony constraint is genuinely yours to tune.
- Will voice sit beside messaging, a contact center, or other channels your customers use? If yes, the full-platform family collapses four contracts into one, and the specialist families reintroduce the stitching this question was trying to avoid.
Most buyers answering "yes" to exactly one of these can skip the other three families entirely. The head-to-head that remains is short, and it runs on comparable vendors at last.
The five checks, run inside one family
The evaluation framework from how to evaluate an AI voice agent platform — build model, latency, model and voice support, telephony path, breadth beyond voice — is family-agnostic as a checklist. It yields different answers family to family:
- In packaged outbound, latency under load and per-minute cost at campaign scale dominate; breadth beyond voice is nearly irrelevant.
- In managed inbound, time-to-first-call and template coverage dominate; model bring-your-own is a secondary concern.
- In developer-first assembly, the API surface, tuning headroom, and your own integration bill dominate; a managed default is a warning sign, not a comfort.
- In the full-platform family, owned termination and channel breadth dominate; raw per-layer control is deliberately traded away.
Run those checks once the family is picked and the shortlist inside it is usually three vendors long. That is a head-to-head you can actually finish.
How to run the evaluation
- Answer the four family questions above, in order, and name the family you are buying from.
- Shortlist only vendors in that family. Cross out the rest of the guide's list; the vendor pages it omits are not a loss.
- Run the five checks within the family, weighting them as the family section above describes.
- Benchmark with a real call mix. The four-metric pilot format in how to benchmark an AI voice agent decides inside the family on your own traffic, not a demo script.
- Only then open a vendor head-to-head. The comparison is valid only when every vendor in it aims at the same buying problem.
Frequently asked questions
What are the four AI voice-agent families?
Packaged outbound (scripted flow builders for high-volume campaigns), managed inbound (no-code builders with canned templates), developer-first assembly (full-control APIs where you assemble the stack), and full communication platforms (voice agents run beside messaging and a contact center, on owned carrier termination). Every vendor in the category sits in one of the four.
Why not compare all vendors in one table?
Because the families answer different buying problems. A single table weights latency, build model, and telephony the same way for every row, and rewards whichever vendor happens to match the table's hidden assumptions. The sort-then-compare sequence removes the distortion.
Can a vendor span more than one family?
Vendor marketing will claim it. In practice the build model and the telephony answer place a vendor in exactly one family; the other three families' strengths do not transfer. Claims to the contrary are a roadmap slide, not a shipping shape.
Which family does Orbit by Devotel belong to?
The full communication platform family: AI voice agents run next to SMS, WhatsApp, RCS, email, video, and a contact center on one account, and outbound calls terminate over Devotel's own wholesale softswitch across 500+ carriers. If your evaluation fits that family's shape, the AI voice agents overview and voice pricing are the starting pages.
The takeaway
The vendor head-to-head is a real document, but it is a tool for the second half of the buying process. The first half is a family sort with exactly four options — packaged outbound, managed inbound, developer-first assembly, full platform — and the correct answer follows from four questions about script, staffing, control, and channel spread. Pick the family before you open a vendor page, and the head-to-head that remains is short, comparable, and finished in days rather than months. For the named eight-vendor framework that assumes you have done the sort, see how to evaluate an AI voice agent platform; for the benchmark that decides inside the family, see how to benchmark an AI voice agent.
Published 15 September 2026.