Quick answer: OpenAI's public API surface now rotates on a published schedule, and each release window lands on an AI-voice tenant as exactly two work items — a migration re-pin (which model id the agent tier resolves to) and a cost-smoothing pass (what the per-token price step does to the budget). The 2026 cycle already announced the dates: the Assistants surface shut down in August, the legacy GPT-3.5/GPT-4/o1/o3/o4 snapshots retire October through December, hosted Evals and Agent Builder close at the end of November, and Chat Completions is explicitly not deprecated. Tenants whose model choice is a governed preset and whose spend is behind a governor treat each window as a scheduled maintenance block; tenants with hard-pinned model strings treat it as a call-path rewrite. This playbook is the standing template for the first posture.
The OpenAI EOL announcement explainer owns the dated timeline itself; this post owns the response cycle — what a CPaaS-side tenant does each time a model vendor ships a release window, and which controls on Devotel Orbit absorb the work. It is an industry explainer with a tenant-owned playbook attached, not an Orbit product announcement.
1. What changed in OpenAI's public API this quarter
The deprecation calendar is public and now spans more than model strings. The dated entries a tenant planning around release windows needs, sorted:
- Surface sunsets (re-author events). The Assistants API shut down in August 2026, with the Responses API paired with the Conversations API as the migration target. The hosted Evals platform went read-only in October and shuts down end of November 2026. The hosted Agent Builder retires end of November 2026, with the Agents SDK as the migration path.
- Model retirements (re-pin events). The legacy GPT-3.5-turbo, GPT-4, o1, o3-mini, and o4-mini snapshots retire October through December 2026, with audio and realtime families rolling into early 2027. A pinned snapshot id is a termination date.
- Policy restrictions (monitor events). New fine-tuning job creation closed to new organizations in May 2026 and closes to existing customers in January 2027; inference on already-trained models continues until the base model retires.
- Explicitly stable. Chat Completions does not appear on the deprecation schedule, and the schedule's own migration rows still recommend
/v1/chat/completionsas the replacement for retiring surfaces.
The full back-compat reading — which class hits which integration shape, and why "OpenAI is deprecating Chat Completions" is a mis-sort — is the OpenAI EOL announcement explainer. The point here is narrower: the vendor now publishes a standing rotation, so the tenant needs a standing response, not a one-off fire drill per window.
2. What a vendor change does to an AI voice agent on a CPaaS
A voice agent consumes the LLM at exactly one point in the pipeline: after transcript capture and grounding, before TTS render. A vendor release window lands on that choke-point through three conduits, and two of them have direct Devotel Orbit surfaces that take the impact:
- The model id itself. When a snapshot family retires, any integration resolving to that id errors at generation time, and the call dies mid-turn. On Orbit the model id lives behind the model-presets catalog — four curated bundles (Balanced, High Intelligence, Ultra Fast, Cost Saver) that pre-wire the LLM, transcriber, and voice as one reviewed decision. A vendor retirement reaches the catalog, not the call path: the preset's underlying model is re-pinned once, and every agent instantiated from it inherits the bump.
- The per-token price step. A re-pin across model generations is rarely price-neutral — the replacement model carries its own token tariff, and the voice-agent LLM line is the one that scales with conversation depth, not just traffic. That is the variance the LLM spend governor exists to bound: budgets, spend telemetry, and alert rules per agent, so the price step shows up as a governor data point the same day the re-pin lands, not on the invoice.
- The evaluation gate. Any model swap invalidates the previous quality baseline. The gate that survives a vendor window is the one owned by the tenant — authored against saved failing conversations — not one hosted on the vendor's own platform, which the 2026 Evals shutdown demonstrated concretely.
The per-step cost chain a re-pin touches — STT, LLM, TTS on one budget row — is walked in the time-and-cost budget explainer. The latency half of that row is why a re-pin is also a quality decision, not only a finance one: the preset catalog's per-bundle latency annotation is the comparison surface for it.
3. The tactical response playbook
Three standing moves, run per release window, in this order:
- Retain on the shared-backbone model. Stay on the vendor's current supported generation — the Chat Completions primitive and the current model family — rather than chasing every intermediate snapshot. The snapshots are the volatility; the supported generation is where the vendor's own migration rows point. A tenant on the backbone absorbs a release window as a preset bump; a tenant pinned to an intermediate snapshot absorbs it as an emergency.
- Re-pin versions on a cadence, not on a deadline. Treat every pinned model id as a deprecation with a date attached — because the vendor now publishes the date. On Orbit the re-pin is one reviewed catalog update: edit the preset, re-run the evaluation gate against saved conversations, and the instantiated agents move together. Off Orbit the same move is a sweep for hard-coded model strings; either way, the cadence (monthly review against the vendor's published calendar) is what keeps the work in the maintenance column.
- Smooth the billing side on the finance ledger. A re-pin changes the token tariff before it changes anything else visible. Set the governor budget and alert threshold against the replacement model's per-token price before the swap ships, not after the first invoice. The LLM spend governor walkthrough covers the alert-rule shape; the goal of this move is that a vendor's pricing re-pin never reaches the tenant's finance review as a surprise.
4. What Devotel Orbit does differently
The posture on Orbit is tenant-owned controls, not a platform mandate — the platform ships the surfaces; the tenant decides the cadence and the thresholds:
- Model presets as the migration surface. The vendor window reaches one catalog entry instead of N agent configs. Re-pinning is a single deliberate action, and a wrong pick is cheap to reverse — re-instantiate a different preset and discard the draft.
- The LLM spend governor as the cost-smoothing surface. Budgets, per-agent spend telemetry, and alert rules sit between the vendor's tariff and the tenant's invoice, so a price step is a governor event, not a reconciliation finding.
- An audit trail the tenant owns. Model-tier decisions — which preset serves which agent, when the re-pin happened, what the gate re-run returned — are recorded in the tenant's own audit posture, alongside the rest of the tenant-owned control surface (quiet hours, the opt-out ledger, credit assignments). The compliant reading of a vendor window is one the tenant can evidence afterward, and the evidence lives with the tenant, not the vendor's hosted log.
The same tenant-owned posture governs how this content itself is built for answer engines: the AI-search visibility explainer covers why a standing, dated explainer per release cycle — rather than a catch-all template — is what keeps a vendor-shift category readable at all.
5. Frequently asked questions
Is the completion API dead?
No. Chat Completions does not appear on OpenAI's deprecation schedule, and the schedule's own migration rows recommend /v1/chat/completions as the replacement for the retiring surfaces. The dated events are the Assistants API, the hosted Evals and Agent Builder surfaces, the legacy model snapshots, and fine-tuning job creation. The full sort is in the EOL announcement explainer.
What happens to stored context?
Context survives a vendor window in proportion to where it lives. Conversation state held in the platform's own store — transcripts, saved conversations, the grounding knowledge base — is untouched by any vendor rotation; the model swap changes the generation step, not the record. State held inside the vendor's own objects (threads, runs, hosted vector stores on the Assistants surface) is the exposure, and the August 2026 shutdown is the dated lesson: a tenant-owned store is the durable answer, a vendor-side store is borrowed time.
How do I re-pin a model version?
On Orbit, at the preset, not per agent: open the Model Presets page, re-instantiate the bundle whose underlying model the window retires, and re-run the evaluation gate against your saved failing conversations before discarding the old draft. The model-presets walkthrough covers the instantiate flow; the spend governor covers the budget and alert reset that should ship with the re-pin.
6. Where to go next
- OpenAI's deprecation cycle, decoded — the dated timeline and the back-compat sort this playbook responds to.
- Model presets — latency, cost, and quality trade-offs — the catalog surface a re-pin lands on.
- LLM spend management — cost governance for AI agents — the governor surface the price step flows through.
- The time and cost model behind every AI voice turn — the per-step budget a model swap moves.
- How to win AI search visibility as a CPaaS in 2026 — why the response playbook needs a standing surface, not a catch-all.
What this does not say (disclaimers)
This is an industry explainer with a tenant-owned playbook attached, not a Devotel Orbit product announcement. Nothing here announces a shipped capability, a licensing change, or a pricing change; the migration cadence and the governor thresholds remain the tenant's own decisions, and compliance posture maps to tenant-owned controls only. The OpenAI deprecation calendar is industry news a CPaaS buyer reads for risk — the integration inventory, not this blog, decides what any given release window touches.