Skip to main content
An agent can speak in a voice cloned from your own recordings: yours, a colleague’s, or a voice artist’s who has agreed to it. You send the recordings once, get back a voice id, and put that id in an agent’s voice.tts.voice. Four steps:
  1. Make the clone with POST /v1/voices.
  2. Listen to it with GET /v1/voices/{id}/preview.
  3. Put its provider_voice_id in an agent’s voice.
  4. Place a call, or a test call, with that agent.
The console does the same under Voices, at voice.wixzel.com/voices, where you can record yourself in the browser or upload a clip.

Which engines can use a clone

A clone is made at one provider, and only that provider’s engines can speak in it. Choose the provider with provider when you make the clone.
gemini_live cannot speak in a cloned voice. Gemini Live only offers its own prebuilt voices.
To use the same voice on both kinds of engine, clone it twice, once on each provider. Each clone counts toward the limit and is charged separately.

1. Make the clone

Send the recordings as files in a multipart request:
Or, from JSON, as links the API downloads once:
Use one or the other in a request, not both. Links must be public https or http URLs. Private and local addresses are refused. The response is the voice:
id is the voice’s id in this API, for reading, renaming and deleting it. provider_voice_id is the id an agent uses. engines lists where it can speak. The request waits while the provider processes the audio, which usually takes a few seconds and can take up to a minute.

2. Listen to it

The response is an mp3 of the voice saying a short line. language is optional. Without it, the preview is in the recording’s language, or English for an ElevenLabs clone with no language set. A Sarvam clone can preview in any Sarvam language, which is a quick way to hear it speak a language it was not recorded in. Previews are free. The first one in each language is made and then kept, so playing it again is instant.

3. Give it to an agent

Put provider_voice_id in voice.tts.voice, on an agent whose engine is in the voice’s engines:
The same works on POST /v1/agents. Your clones are also listed first in GET /v1/engines/{engine}/voices, with "category": "cloned", so a voice picker built on that endpoint shows them without extra work. A clone used on an engine that cannot speak it is refused with unsupported_voice, and the message says which engines it does speak on.

4. Call

Place a call with the agent as usual, or check the voice first with a test call, which rings a number and speaks one phrase:
consent: true confirms that the voice is yours, or that the speaker has given you permission to clone it and to use it on calls. A request without it is refused with 400 invalid_parameter. The confirmation is stored with the time, your account and the IP address the request came from. Do not clone a voice you do not have permission to use. See the Acceptable use policy.

Who can see and use a clone

Only the account that made it.
  • GET /v1/voices and GET /v1/engines/{engine}/voices show an account its own clones and nobody else’s.
  • An agent create or update that names another account’s clone is refused with unsupported_voice, the same answer as for a voice that does not exist.
  • The recordings are passed to the provider to make the clone. Wixzel Voice does not store them.

Cost and limits

The $1.00 is charged when the clone is made. It appears in /v1/usage/events as a voice_clone line. If your balance cannot cover it, the request is refused with 402 insufficient_credits before anything is sent to the provider. Deleting a clone does not refund the fee. Speech in an ElevenLabs clone is billed like any ElevenLabs voice on the same model. Speech in a Sarvam clone is billed like Bulbul v3: Sarvam prices cloned speech at ₹30 per 10,000 characters, the same as Bulbul v3. The per-minute prices on Voice engines apply unchanged. An account can hold 3 clones. A fourth is refused with 409 account_limit_reached. Delete one you no longer need, or contact support to raise the limit.
On a self-hosted install the clones are made on your own provider accounts and the 1.00feeis1.00 fee is 0.

Recording tips

The clone sounds like the recording, so the recording matters more than anything else.
  • Length. For ElevenLabs, 30 seconds to a few minutes in total. For Sarvam, one file of 10 to 30 seconds.
  • One speaker. No other voices, and no music.
  • A quiet room. No echo, fans or traffic. A phone held close in a small, furnished room is better than a good microphone in an empty one.
  • Speak the way the agent will. Talk naturally, at the pace and tone of a phone call. A clone of someone reading in a flat voice will sound flat on calls.
  • Format. WAV or FLAC is best. mp3, m4a, ogg and webm are accepted. Up to 10 MB per file.

Managing voices

PATCH changes name and description only. The recordings and the provider cannot be changed; to change them, make a new clone. DELETE deletes the clone at the provider too, and frees one of your 3 places. It is refused with 409 voice_in_use while an agent uses the voice, and the message names the agents. Give them another voice first. The list is cursor-paginated like every other list. See Pagination.

Scopes

Reading voices and previews needs voices:read. Making, renaming and deleting them needs voices:write.

SDKs and the MCP server

The SDKs have the same operations:
client.voices.list, retrieve, update, delete and preview cover the rest. The MCP server has the tools list_voices, get_voice, preview_voice, create_voice_clone, update_voice and delete_voice, so you can make a clone from Claude Code, Codex or Cursor. create_voice_clone takes audio_urls, since an MCP client cannot upload a file.

Errors

On agent writes (POST /v1/agents, PATCH /v1/agents/{id}): See Errors for the error format.