Skip to main content
Back to blog

OpenAI's voice moves, decoded for the AI-voice-agent lane

What the 2026 OpenAI API changes actually shut down for buyers whose agent talks on the phone — what voice-as-product from a model vendor enabled and disabled, the honest case where a direct OpenAI realtime pipeline still fits, and where a production telephony layer (SIP termination, STIR/SHAKEN attestation, A2P routing, E911, published per-minute pricing) picks the work up.

Orbit Editorial Team

Quick answer: OpenAI's 2026 API cycle removed two surfaces a "ChatGPT voice agent" could be prototyped on and re-priced the ones it kept — the Assistants API and hosted Agent Builder shut down, Realtime moved from beta to GA with revised per-minute pricing, and the remaining realtime and audio model snapshots carry retirement dates. What none of the changes produced is a phone call: a model vendor voice API still terminates nowhere on the public telephone network on its own. A demo that plays audio in a browser tab is a pipeline stage; the production gap — SIP termination, STIR/SHAKEN attestation, A2P messaging rules, E911 address registration, and a published per-minute telephony rate — is exactly where a communications platform fills the lane. This post is the industry-news decoding for that lane; the vendor-neutral timeline itself is owned by the OpenAI EOL announcement explainer, and the tenant response cycle is the model-vendor shift playbook.

If you are re-evaluating an AI-voice stack because an OpenAI announcement landed in your budget review, the test to run is the same one the Twilio Signal decoding ran on the incumbent's news: which row of the comparison actually moved, and which row was never a vendor event in the first place.

1. The trigger — what OpenAI changed, and when

OpenAI publishes a deprecation schedule, and the 2026 entries touched every surface a voice prototype could hang off. The dated scope, sorted by class:

  • Surface sunsets — the Assistants API and hosted Agent Builder. The Assistants API (threads, runs, server-managed vector stores) shut down in August 2026, with the Responses API paired with the Conversations API as the migration target. The hosted visual Agent Builder retires at end of November 2026, with the Agents SDK as the path. Both were quick prototypes' homes for "an assistant that responds" — buyers standing one up again now author it themselves.
  • Realtime moved from beta to GA — with new published pricing. The Realtime API went generally available in the August 2026 cycle with a revised per-minute price for the realtime models, replacing the beta pricing voice prototypes had been built against. A voice budget written during beta now re-opens at the GA rate.
  • Model snapshot retirements — including the audio families. The legacy GPT-3.5-turbo, GPT-4, and reasoning snapshots retire October through December 2026, and the audio and realtime families carry retirements rolling into early 2027. A pinned snapshot id is a dated event, not a version preference.
  • Policy restrictions — fine-tuning jobs. Creating fine-tuning jobs closed to new organizations in May 2026 and closes to existing customers in January 2027; inference on already-trained models continues until the base model retires.

The full back-compat reading of each class — which one is a re-author, which one is a re-pin, and why Chat Completions is explicitly not on the schedule — is the OpenAI EOL announcement explainer. The point for the voice lane is narrower: every change above governs the model pipeline, and none of them is a phone call.

2. What ChatGPT-voice-as-a-product actually enabled — and what it never did

"Voice from a model vendor" is a complete pipeline only in the narrow sense: audio in, audio out. For the AI-agent buyers naming a ChatGPT replacement in the compare funnel, it is worth stating both columns plainly.

What the vendor surface enabled:

  • Streaming speech-to-speech inside one vendor's models. The Realtime API carries audio in and audio out with mid-call tool use and interruptions handled at the model level — a genuine capability for in-app, browser, or sandbox voice.
  • A fast demo loop. Transcription, reasoning, and a spoken reply in one API call is the cheapest way to hear a voice agent work, which is why so many category proposals start there.
  • A priced-by-minute model hop. The GA realtime pricing is a published per-minute number; prototyping teams could estimate demo cost against it.

What it disabled, or never offered, for a production AI-voice agent:

  • PSTN reach. A Realtime session terminates in your application. The phone network — numbering, SIP termination, carrier routing — is not on the API surface, so the "voice agent" is not a phone number a customer can call.
  • Attestation and carrier trust. STIR/SHAKEN caller-ID attestation, the DNO/robocall posture, and phone-company routing obligations do not exist at the model layer; they attach where the call becomes a telephone call.
  • Messaging rules. The agent that follows up an inbound call with an SMS enters the A2P regime — registration, opt-out handling, throughput rules — none of which a model API answers.
  • Emergency obligations. E911 address registration and dispatchable-location handling attach to a voice service that reaches the phone network; a model vendor's audio surface is out of that scope entirely.

The framing worth retiring is "OpenAI used to sell voice and now it doesn't." OpenAI sells the model side of voice; the production half of the lane never moved, and the 2026 announcements did not change it.

3. The honest carve-out — where a direct OpenAI realtime pipeline still fits

The comparison above does not mean a direct pipeline is wrong for everyone. There are two configurations where it is the correct architecture, stated plainly so the rest of this post reads as a decision, not a sales pitch:

  • In-app voice, no phone number involved. A voice assistant inside your own product — a mobile app, a kiosk, a web workspace — where the audio never touches the public telephone network. The carrier and regulatory columns above do not apply, and the vendor's per-minute realtime pricing is the whole cost surface.
  • A bespoke, single-purpose bot with engineered latency. A narrow scripted flow (one intent, a fixed vocabulary, a controlled acoustic environment) where you hold the model, the prompts, and the pipeline yourself and the latency budget is tuned end-to-end by your own team. The 2026 deprecations are then a scheduled re-pin, and nothing else in this post is your work to do.

Where even those teams hit the ceiling is the moment the assistant has to reach a person: a caller dials a number, the agent has to call back on a real phone call, the campaign has to send the follow-up text. At that seam the model pipeline is one component and the production telephony layer is the other.

4. Where Devotel Orbit fills the production-grade lane

Devotel Orbit builds the half of the lane the model vendor never had. Concretely, for the buyer whose agent must work on the phone:

  • SIP termination and carrier-grade routing. Numbers and trunks on Orbit terminate calls and route them to your agent tier with the carrier posture a phone call needs, and the programmable voice API is the control surface for it.
  • STIR/SHAKEN attestation. Caller-ID attestation on outbound calls is the carrier-trust column — covered in the AI-voice compliance audit explainer and the DNO caller-ID guard explainer — and it lives on the voice platform, not in your prompt.
  • A2P messaging rules. The follow-up text inherits 10DLC registration and campaign obligations the platform already models as tenant-owned controls, so the conversational hand-off between a call and a message stays inside one compliance surface.
  • E911 obligations. US voice lines carry registered addresses and dispatchable-location handling as a documented flow, not an unowned risk.
  • Published per-minute pricing. The telephony leg is a priced, public number — the full destination voice pricing page is per country — and the four components of an AI voice minute (transcription, model, speech, telephony) are broken out in the AI voice agent pricing guide so an OpenAI-side estimate and the telephony estimate land on the same ledger.

The buyer comparison rows that matter on this lane — which incumbent's news actually changed them, which stayed — are kept on the Twilio Signal decoding and the AI-voice-agent evaluation framework; the Orbit-vs-incumbent head-to-heads stay the row-level reference, and this post is the news-side note that the OpenAI row is a pipeline event, not a telephony event.

5. The operator-resolution loop

A vendor API change is only finished when the operator's stack absorbs it. On Orbit that loop is a calendar, not a fire drill:

  • This week in Orbit. The vendor schedule re-pins and any tenant-facing change on our side land in the weekly installment — the current cadence lives at /blog/this-week-in-orbit-2026-10-08 — so release windows are reviewed on the same rhythm as the shipped record.
  • The standing playbook. The model-vendor shift playbook names the two moves each release window triggers — a migration re-pin and a cost-smoothing pass — and stands as the template for the next OpenAI, Anthropic, or Google window the same way.
  • The timeline of record. The OpenAI EOL announcement explainer stays the dated record of what shut down and when, so a buyer can re-check the timeline against their own pin dates.

Use the loop in this order: re-read the dated timeline, apply the playbook's two moves, confirm the resolution in the weekly installment. The news event closes when the third step runs to no-op.

Frequently asked questions

Did OpenAI deprecate voice itself?

No. The 2026 deprecations removed the Assistants API and hosted Agent Builder surfaces and cycled model snapshots; the realtime and audio model families remain, re-priced at GA rates with dated retirements per snapshot. The production carrier side of voice — termination, attestation, messaging rules — was never part of the OpenAI surface, so nothing in that layer changed.

When does a direct OpenAI realtime pipeline still make sense?

When the audio never touches the telephone network — in-app voice, a kiosk, a web workspace — or for a bespoke single-purpose bot where your team owns the model, prompts, and latency budget end to end. The moment the assistant must place or receive a real phone call, the telephony layer becomes load-bearing.

What should an AI-voice buyer actually re-check after this announcement?

Two things: which model snapshot your agent resolves to (a re-pin, on the deprecation calendar), and what the GA realtime pricing does to your per-minute estimate. The telephony comparison rows — termination, attestation, A2P, E911, published minutes — are unchanged by the announcement.

Does this change the Orbit comparison rows against the incumbents?

No. The OpenAI cycle is a model-pipeline event; the rows that distinguish a communications platform (SIP termination, STIR/SHAKEN, A2P compliance, E911, pricing) are the same rows the Twilio Signal decoding re-checks on the incumbent's news cycle. A model-vendor announcement re-opens your model pin, not the registry.

OpenAI's voice moves, decoded for the AI-voice-agent lane — Orbit by Devotel