Skip to main content
Back to blog

The Six Families of AI Voice Agents: How to Choose an Archetype Before You Compare Vendors

Most AI-voice-agent shortlists fail before the first demo because the vendors on them belong to different families. This essay names the six archetypes: pipeline constructors, wrapper constructors, no-code builders, speech-model vendors, full-stack communications platforms, and scripted IVR. It shows how to pick your family first, then let the head-to-heads rank the vendors inside it.

Orbit Editorial Team

Quick answer: An AI voice agent falls into one of six families. Pipeline constructors give you separately addressable speech stages you compose yourself. Wrapper constructors bundle the same composition as one phone-call surface. No-code builders present staged pipelines to operators. Speech-model vendors ship an agents surface on their own synthesis and recognition models. Full-stack communications platforms carry the agent plus telephony, messaging, and the contact center on one account. Scripted IVR is the non-AI baseline you route out of. Pick the family first, because a vendor that does not fit your build model loses nothing in a head-to-head: it was never on your shortlist. Once the family is chosen, the vendor-by-vendor comparisons and the five evaluation checks rank the members.

Orbit by Devotel publishes this essay and is itself one of the vendors a reader will eventually compare. It sits in the full-stack-platform family, and that membership is stated rather than buried. The essay's job is the step before the comparison table: the blog already holds the vendor-by-vendor answers, but none of them tells a reader which kind of platform belongs on the list. This one does, and it deliberately includes the scripted-IVR baseline, because some traffic never needs a model at all.

Why archetype-first beats vendor-first

Most buyers start with names: someone posts "Vapi or Retell?" and the replies form a ranking. The ranking is unsound, because the two vendors belong to different families answering different questions. One asks "who composes the stages," the other asks "who should be able to assemble one." Scoring them against each other produces a winner that solves a problem the buyer may not have.

The vendor-first failure shows up two ways in practice. A team evaluates a pipeline constructor against a no-code builder, picks the pipeline for its control surface, then discovers the operators who own the production agent cannot read an API. The wrong family won, and the vendor choice inside it was irrelevant. Or a team evaluates two members of the same family, splits hairs on features both share, and never asks the question that would actually separate them from the neighbouring family: who owns the phone network, the messaging channels, and the human handback once the agent is live. Archetype-first inverts the order: choose the family that matches who builds, who operates, and what the agent must reach beyond the call; then the existing comparisons rank the vendors inside it.

The six archetypes

Pipeline constructors

Speech-to-speech as separately addressable stages (speech recognition, language model, synthesis) chosen and composed by the buyer, with the platform orchestrating the turns between them. The pitch is control: every stage is a knob, every stage vendor is swappable, and the pipeline is the product. The cost is surface area: the buyer governs every stage, every stage invoice, and every degraded-vendor failover. Choose this family when an engineering team owns the agent and open third-party model choice is the core requirement. Vapi and Synthflow sit here on the staged-rails side; the four-way framework comparison grades them against the wrapper family on the same criteria.

Wrapper constructors

The same composition presented as one bundled phone-call surface. The stages are still there; the product wraps them so the buyer manages the conversation, not the plumbing. Less per-stage control, less per-stage observability, and the same dependence on third-party speech suppliers underneath. It is a wrap, not a change of physics. Choose this family when a team wants conversation design without stage-level governance and accepts the bundle's defaults. Retell AI and Bland AI are the named members; their individual head-to-heads, Orbit vs Retell and Orbit vs Bland, grade them cell by cell.

No-code builders

A staged pipeline presented to operators through a visual builder, templates, and clone-to-deploy flows. The buyer composes the same stages, but the product assumes the person assembling them is not an engineer: sales ops, a front-office team, an agency reselling agents to local businesses. Choose this family when the people who will run the agent day to day have to be able to change it without filing a ticket to engineering. Check the template library and the handover model, not the builder's demo shine, and confirm there is an escape hatch when a flow needs something the canvas cannot express.

Speech-model vendors

The supplier of synthesis, recognition, and voice models, plus an agents surface built on those models. It belongs in the same evaluation because the other families depend on suppliers in this position: when the vendor also sells the agents surface, a buyer should know the pipeline constructor they picked may be one model contract away from this vendor occupying two rows in the same table. The family concession applies in reverse, too. A speech-model vendor's agents surface inherits its models' strengths and weaknesses, and open stage choice narrows to its own catalogue.

Full-stack communications platforms

The agent plus everything the agent eventually has to touch: telephony with an owned termination path, SMS, WhatsApp, RCS, and email beside it, handback to a human agent or queue inside the same contact center, and one usage bill across all of it. Devotel Orbit is this family, and it concedes the row the pipeline constructors lead on, open third-party stage choice, in exchange for owning the stack the agent runs on: a published per-stage latency budget and methodology, model-spend visibility per feature, and no second vendor between the agent and the phone network. Choose this family when the agent is going into production operations and the channels, handback, and bill consolidation matter more than a model marketplace.

Scripted IVR (the baseline to route out of)

The non-AI archetype, listed because every buyer measures AI against it. A scripted menu with pre-recorded, fixed answers resolves hours, locations, and payment replay cheaper and faster than any model, and it still earns its slot in a routing design. The question is never "AI or nothing." It is which intents justify a model and which a menu reads back in seven seconds with zero model spend. The IVR vs AI voice agents routing guide grades call types into menu, model, or human; most production deployments run all three on one call path, and on a full-stack platform that is a routing configuration, not a repurchase.

Choosing your family: the four questions that decide it

Four questions sort a buyer into a family before any vendor name matters.

Who builds and who operates? Engineers with an appetite for stage-level control point to pipeline constructors; operators assembling flows in a canvas point to no-code builders; a team that wants conversations without plumbing points to wrapper constructors. If the people who run the agent cannot change it, the family is wrong regardless of vendor quality inside it.

How open must the model choice be? If swapping speech-to-text or language-model vendors per stage is a stated requirement, mandated by procurement rather than a nice-to-have, pipeline constructors are the only family that satisfies it, and every other family concedes that row. If a managed, governed provider set is acceptable, the concession stops mattering and the other families' operational gains dominate.

What must the agent reach beyond the call? An agent that only talks is one product; an agent that sends the confirmation SMS, moves to WhatsApp when the caller prefers it, and hands the hard call to a human in the same queue is a different product. The more the answer includes channels and handback, the more it selects the full-stack family, or a deliberate multi-vendor assembly the buyer accepts owning.

Is the traffic even AI-shaped? Fixed answers, regulatory disclosures, and simple navigation belong to scripted IVR; natural-language intent resolution belongs to a model. Routing the wrong traffic into the AI families wastes the model spend the pricing guide itemizes; routing model-shaped traffic into menus wastes the callers.

From family to vendor: using the existing comparisons

Once the family is chosen, the essay hands off, and deliberately so, because the blog already answers the inside-the-family question in three registers, a layer each.

  • One vendor at a time. The head-to-heads apply one evaluation grid to one opponent: Orbit vs Vapi, Orbit vs Retell, Orbit vs Bland, Orbit vs Synthflow. Use them to deep-dive the single finalist the family choice left standing.
  • The family-internal four-way. The Vapi vs Retell vs Bland vs Synthflow comparison holds the four framework builders to the same grid at once (latency budget, model-cost governance, handback and outbound tooling) for buyers whose family choice landed on the constructor families.
  • The cross-family checks. The evaluation guide runs five checks (build model, latency, model and voice support, telephony path, and platform breadth) across eight providers. It is the sanity pass that confirms a family-level winner still survives cross-family scrutiny.

Two companion pieces then price and measure what these families ship: the AI voice agent pricing guide itemizes the four per-minute cost components every family pays, and the QA guide shows what production discipline looks like once the chosen vendor's agent goes live.

Frequently asked questions

Which archetype is Devotel Orbit?

The full-stack-communications-platform family: the AI voice agent shipped beside the telephony, the messaging channels, and the contact center on one account and one pay-as-you-go bill. Inside that family Orbit concedes one row openly, open third-party stage choice, where pipeline constructors lead, and leads on the operational rows: the published per-stage latency budget, the per-call quality index, handback to human and video agents, and one bill across voice, SMS, WhatsApp, RCS, email, and video.

Can I mix archetypes?

Mixed routing is the norm, not the exception: scripted IVR for fixed answers, an AI agent for intent resolution, a human queue for high-stakes callers. On a full-stack platform the mix is a routing configuration. Across families the mix is a procurement decision: a pipeline constructor for the agent assembled with a separate carrier and contact center is a legitimate architecture, priced and governed as the multi-vendor assembly it is.

Where do pipeline constructors beat platforms, honestly?

On open third-party model choice and stage-level control: a mandated bring-your-own-vendor environment, a research team tuning stage providers, or a buyer whose procurement requires per-stage substitutability. The four-way comparison credits those rows to the constructors in plain cells rather than shading them, because a conceded row is what makes the won rows legible.

Do I even need an AI archetype?

Not for fixed answers. Hours, locations, balances, and payment replay resolve in a scripted menu with no model spend; the routing guide grades call types into menu, model, or human before any AI archetype enters. AI earns its slot where callers bring natural-language intent that must be resolved, not navigated.

The Six Families of AI Voice Agents: How to Choose an Archetype Before You Compare Vendors — Orbit by Devotel