Skip to main content
Back to blog

Voice clones on Orbit: custom brand TTS for AI agents

A voice clone is a custom TTS voice your organization owns — trained from a recorded call with documented consent, or designed from a text description. This guide walks the Voice Clones page on Devotel Orbit — the two creation modes, the Synthetic-voice watermark credential workbench, how a clone plugs into AI voice agents as the speech voice, the consent rules that apply before you train one, and how a brand voice carries into multilingual messaging.

Orbit Editorial Team

Quick answer: A voice clone is a custom text-to-speech voice your organization owns — the difference between your AI agent reading as a generic assistant and answering in the voice your customers already associate with your brand. On Devotel Orbit you train one under Voice → Voice Clones, either from a recorded call held to documented speaker consent, or by describing the voice you want in text and synthesizing a fresh one. The first clone per organization is free; further clones are billed from the credit wallet with the price disclosed on the form, and a failed training run refunds automatically. A finished clone becomes the speech voice of any AI voice agent you assign it to. This guide covers the page end to end — including the consent you must hold before cloning a real voice and the watermark workbench that marks the clone's synthesized output.

What a voice clone is — and is not

A voice clone on Orbit is a speaking voice registered against your organization, trained from audio you supply or designed from a written description. It answers one question for callers — who does this sound like — with an answer you chose deliberately instead of a vendor default. Three properties define it:

  1. It is yours alone. Clones are per-organization, created from your recorded calls or your description, and visible only on your account. Deleting a clone also removes it from the upstream provider, so the asset does not outlive your use of it.
  2. It is a speech (TTS) voice, not an identity. Assigning a clone to an AI voice agent changes how the agent sounds, never who can call or what the agent may say — routing, prompts, and policies stay where they already live.
  3. Its output can be proven. Every clone carries a Synthetic-voice watermark workbench that mints a tamper-evident content credential over the exact audio the clone produced and can verify a credential later. That matters because the same cloning technology that gives you a brand voice is also the year's fraud story — the sibling post on cloned-voice fraud and voice security covers the inbound risk angle (attackers imitating your staff and customers) and the enrollment-plus-liveness defense, while the watermark is the outbound half: provable, disclosed provenance for the audio your tenant ships.

Creating a clone under Voice → Voice Clones

The page lists every clone for the organization with a status — Training, Ready, or Failed — and offers Train new clone with two source modes. Both are gated to owner, admin, and developer roles, because creation charges the wallet and deletion is destructive.

  • From recording. Pick one of your recorded calls as the training sample — the picker only lists calls that actually have a recording, and at least 30 seconds of clean single-speaker audio gives the best result. Before the clone is ever used, the Listen to sample control on each row plays the training source through a lazy signed URL, so you can audit exactly what a voice was built from.
  • From description. Write a free-text description — accent, age, tone, pace — and Orbit synthesizes a brand-new voice that belongs to no one. This path needs no recorded audio and therefore no speaker consent, which makes it the fastest safe option when you cannot obtain consent from a specific speaker.

The first clone per organization is free; the form states the price before you commit, and a training run that fails refunds the charge automatically. Creation modes hide rather than fail when unavailable — if the page shows that clone creation is unavailable on your plan, that verdict comes from the server, not from a form that can only fail. Clones land as Training and typically finish within a couple of minutes; the list refreshes on its own while one is mid-training.

Deletion closes the loop: the row's trash action asks for confirmation, then removes the clone from your account and from the upstream provider. If any agent still uses the voice, retire the assignment first.

How a clone reaches an AI voice agent

The clone lifecycle completes where the voice is spent: Voice → Agents, where each AI voice agent takes a speech voice for the synthetic half of the conversation. Assigning a ready clone there replaces the stock TTS voice across that agent's calls — the same agent logic, prompts, and numbers, spoken in your brand voice. The full configuring walk-through of the agent side is in the AI voice agents omnichannel guide.

Voice clones are not limited to live agent conversations. A clone can also power a voicemail-drop message, so a campaign's prerecorded voice note carries the same brand voice as the agent that answers callbacks. Between the two surfaces, one clone covers outbound campaigns, inbound support, and anything in between that speaks.

Watermark the output. Once a clone is Ready, its row opens the Synthetic-voice watermark workbench — create a credential against the clip the clone produced, or verify one in hand. Minting returns a credential bound to the exact audio bytes, so any re-encode or edit breaks the match; verifying a pasted credential distinguishes a genuine mark from a forged one or from a real credential re-stapled onto different audio. The workbench also resolves the audible disclaimer text to play before the clip, mapped from your AI-disclosure switch under Settings → Compliance, and carries the EU AI Act Article 50 and FCC synthetic-voice disclosure flags. Both halves are tenant-owned: you choose when to mark output and what reaches the listener; the platform ships the marking and verification capability. Nothing here originates a call or message — marking and verifying are pure provenance operations over audio you already have.

Consent before you clone a real voice

A voice is a person's biometric identifier, so the first rule is consent: only train a clone from a recording whose speaker consented to that use, on a recorded call made under the applicable recording-consent rule. Those obligations vary by jurisdiction — some regions require every party's consent, others one — and the obligations remain yours regardless of where the platform is hosted. The explainer on call recording consent rules lays out the two regimes and how to record the rule that applies to you.

The page supports that posture mechanically rather than by mandate. The source picker restricts training to calls your organization chose to record, and the Listen to sample player gives every colleague reviewing the clone a way to check the origin audio. When consent for a specific speaker is unavailable — or when cloning a voice was never the right call — switch to From description: a synthesized voice that impersonates no one and carries a clean provenance story by construction. The decision is tenant-owned; the platform gives you both modes rather than pushing one default.

Brand voice beyond calls — multilingual consistency

The reason to own a voice rather than rent a vendor default is continuity: the greeting a caller hears from your support line, the message a campaign drops to voicemail, and the answer an AI agent gives should all sound like the same brand. Text makes the same decision per language rather than once per channel — the explainer on per-contact language preference and localized template variants shows how preference-driven variant resolution keeps your tone consistent at send time instead of guessing at blast time. Together the two posts describe one practice: decide the voice once, apply it everywhere, and let the platform resolve the per-contact language.

Frequently asked questions

Which creation mode should I start with?

From recording if a specific, consented voice is the brand asset — a founder, a dedicated support persona, a voice talent in contract. From description if you are still choosing the voice, cannot secure consent, or want a persona that belongs to no one and cannot be mistaken for a real employee.

Can a failed training run cost me?

No — the charge refunds automatically when training fails, and failed clones do not consume your free-clone quota. Only a clone that reaches Ready counts.

What does the watermark actually prove?

That a specific credential matches specific audio, tamper-evidently. Minting signs the hash of the exact bytes; verification re-checks both the signature and, when you attach audio, whether the audio still matches. A credential copied onto edited audio fails the audio-match, and a forged credential fails the signature — the two failure reasons are reported distinctly.

Is a cloned voice allowed on outbound AI calls?

Yes, subject to your consent posture — the same wire the outbound FCC AI-voice rules draw: synthesized and cloned voices on outbound AI calls sit on the recorded written consent of the called party, and Orbit's compliance classification treats cloned voices as their own class. The deepfake explainer above covers the regulatory direction; the controls and records you need are the tenant-owned ones described here.

The takeaway

A voice clone is a small asset with an outsized brand effect — one recording or one description becomes every spoken interaction you run. Train it under Voice → Voice Clones on held consent, assign it to an AI voice agent or a voicemail drop, and mark its output with the Synthetic-voice watermark so provenance travels with the audio. The sibling posts cover the fraud-defense side of the same technology and the consent regimes behind recording; this one is the how-to for the side that belongs to you.

Voice clones on Orbit: custom brand TTS for AI agents — Orbit by Devotel