> ## Documentation Index
> Fetch the complete documentation index at: https://docs.voice.wixzel.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Clone a voice for your AI phone agent

> Make a cloned voice from your own recordings with POST /v1/voices, preview it, and give it to an agent. ElevenLabs clones speak on the classic and Deepgram engines; Sarvam clones speak Indian languages, including Malayalam.

An agent can speak in a voice cloned from your own recordings: yours, a
colleague's, or a voice artist's who has agreed to it. You send the recordings
once, get back a voice id, and put that id in an agent's `voice.tts.voice`.

Four steps:

1. Make the clone with `POST /v1/voices`.
2. Listen to it with `GET /v1/voices/{id}/preview`.
3. Put its `provider_voice_id` in an agent's voice.
4. Place a call, or a test call, with that agent.

The console does the same under **Voices**, at
[voice.wixzel.com/voices](https://voice.wixzel.com/voices), where you can
record yourself in the browser or upload a clip.

## Which engines can use a clone

A clone is made at one provider, and only that provider's engines can speak in
it. Choose the provider with `provider` when you make the clone.

| `provider` | Speaks on | Recordings | Languages |
| - | - | - | - |
| `elevenlabs` (default) | `classic`, `deepgram_agent` | 1 to 5 files; 30 seconds to a few minutes in total | The languages of the agent's ElevenLabs model |
| `sarvam` | `sarvam` | 1 file; 10 to 30 seconds is enough | Every language Sarvam supports, including Malayalam, whatever language the recording is in |

<Note>
  `gemini_live` cannot speak in a cloned voice. Gemini Live only offers its own
  prebuilt voices.
</Note>

To use the same voice on both kinds of engine, clone it twice, once on each
provider. Each clone counts toward the [limit](#cost-and-limits) and is charged
separately.

## 1. Make the clone

Send the recordings as files in a multipart request:

```bash theme={null}
curl https://api.voice.wixzel.com/v1/voices \
  -H "Authorization: Bearer $WIXZEL_API_KEY" \
  -F name="Meera, front desk" \
  -F provider=elevenlabs \
  -F consent=true \
  -F files=@meera-1.wav \
  -F files=@meera-2.wav
```

Or, from JSON, as links the API downloads once:

```bash theme={null}
curl https://api.voice.wixzel.com/v1/voices \
  -H "Authorization: Bearer $WIXZEL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Meera, Malayalam",
    "provider": "sarvam",
    "language": "ml-IN",
    "consent": true,
    "audio_urls": ["https://example.com/recordings/meera-ml.wav"]
  }'
```

Use one or the other in a request, not both. Links must be public `https` or
`http` URLs. Private and local addresses are refused.

| Field | Notes |
| - | - |
| `name` | Required. Up to 100 characters. |
| `description` | Optional. Up to 500 characters. |
| `provider` | `elevenlabs` (default) or `sarvam`. Decides the engines, as above. |
| `language` | The language spoken in the recording, such as `ml-IN` or `hi-IN`. Required for `sarvam`. |
| `consent` | Required, and must be `true`. See [Consent](#consent). |
| `remove_background_noise` | ElevenLabs only. Cleans up a noisy recording. It can make a clean one sound worse, so leave it off unless you need it. |
| `files` | Multipart only. The recordings. |
| `audio_urls` | JSON only. Links to the recordings, up to 5. |

The response is the voice:

```json theme={null}
{
  "id": "6703f1c2a9e4b1d2c3f4a5b6",
  "object": "voice",
  "name": "Meera, front desk",
  "description": null,
  "provider": "elevenlabs",
  "provider_voice_id": "c38kUX8pkfYO2kHyqfFy",
  "engines": ["classic", "deepgram_agent"],
  "language": null,
  "created_at": "2026-10-08T09:12:44.000Z",
  "updated_at": "2026-10-08T09:12:44.000Z"
}
```

`id` is the voice's id in this API, for reading, renaming and deleting it.
`provider_voice_id` is the id an agent uses. `engines` lists where it can speak.

The request waits while the provider processes the audio, which usually takes a
few seconds and can take up to a minute.

## 2. Listen to it

```bash theme={null}
curl "https://api.voice.wixzel.com/v1/voices/$VOICE_ID/preview?language=ml-IN" \
  -H "Authorization: Bearer $WIXZEL_API_KEY" \
  -o preview.mp3
```

The response is an mp3 of the voice saying a short line. `language` is
optional. Without it, the preview is in the recording's language, or English
for an ElevenLabs clone with no `language` set. A Sarvam clone can preview in
any Sarvam language, which is a quick way to hear it speak a language it was
not recorded in.

Previews are free. The first one in each language is made and then kept, so
playing it again is instant.

## 3. Give it to an agent

Put `provider_voice_id` in `voice.tts.voice`, on an agent whose engine is in
the voice's `engines`:

<CodeGroup>
  ```bash ElevenLabs clone, classic engine theme={null}
  curl -X PATCH https://api.voice.wixzel.com/v1/agents/$AGENT_ID \
    -H "Authorization: Bearer $WIXZEL_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "voice": {
        "stt": { "model": "deepgram/nova-3", "language": "en-US" },
        "llm": { "model": "openrouter/gpt-4o-mini" },
        "tts": { "model": "elevenlabs/eleven_flash_v2_5", "voice": "c38kUX8pkfYO2kHyqfFy" }
      }
    }'
  ```

  ```bash Sarvam clone, sarvam engine theme={null}
  curl -X PATCH https://api.voice.wixzel.com/v1/agents/$AGENT_ID \
    -H "Authorization: Bearer $WIXZEL_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "language": "ml-IN",
      "voice": {
        "stt": { "model": "sarvam/saaras-v3-realtime", "language": "ml-IN" },
        "llm": { "model": "sarvam/sarvam-105b" },
        "tts": { "model": "sarvam/bulbul-v3", "voice": "svc-..." }
      }
    }'
  ```
</CodeGroup>

The same works on `POST /v1/agents`. Your clones are also listed first in
`GET /v1/engines/{engine}/voices`, with `"category": "cloned"`, so a voice
picker built on that endpoint shows them without extra work.

A clone used on an engine that cannot speak it is refused with
`unsupported_voice`, and the message says which engines it does speak on.

## 4. Call

Place a call with the agent as usual, or check the voice first with a test
call, which rings a number and speaks one phrase:

```bash theme={null}
curl https://api.voice.wixzel.com/v1/agents/$AGENT_ID/test-call \
  -H "Authorization: Bearer $WIXZEL_API_KEY" \
  -H "Idempotency-Key: $(uuidgen)" \
  -H "Content-Type: application/json" \
  -d '{ "to": "+919876543210" }'
```

## Consent

`consent: true` confirms that the voice is yours, or that the speaker has given
you permission to clone it and to use it on calls. A request without it is
refused with `400 invalid_parameter`. The confirmation is stored with the time,
your account and the IP address the request came from.

Do not clone a voice you do not have permission to use. See the
[Acceptable use policy](https://voice.wixzel.com/acceptable-use).

## Who can see and use a clone

Only the account that made it.

* `GET /v1/voices` and `GET /v1/engines/{engine}/voices` show an account its own
  clones and nobody else's.
* An agent create or update that names another account's clone is refused with
  `unsupported_voice`, the same answer as for a voice that does not exist.
* The recordings are passed to the provider to make the clone. Wixzel Voice does
  not store them.

## Cost and limits

| | |
| - | - |
| Making a clone | **\$1.00**, once, from your credit balance |
| Previews | Free |
| Speech in a clone | The same per-character price as the provider's other voices |
| Clones per account | **3** |

The \$1.00 is charged when the clone is made. It appears in
[`/v1/usage/events`](/billing#tracing-a-charge) as a `voice_clone` line. If your
balance cannot cover it, the request is refused with `402
insufficient_credits` before anything is sent to the provider. Deleting a clone
does not refund the fee.

Speech in an ElevenLabs clone is billed like any ElevenLabs voice on the same
model. Speech in a Sarvam clone is billed like Bulbul v3: Sarvam prices cloned
speech at ₹30 per 10,000 characters, the same as Bulbul v3. The per-minute
prices on [Voice engines](/voice-engines#what-a-call-costs) apply unchanged.

An account can hold 3 clones. A fourth is refused with `409
account_limit_reached`. Delete one you no longer need, or contact support to
raise the limit.

<Note>
  On a [self-hosted](/self-hosting) install the clones are made on your own
  provider accounts and the $1.00 fee is $0.
</Note>

## Recording tips

The clone sounds like the recording, so the recording matters more than
anything else.

* **Length.** For ElevenLabs, 30 seconds to a few minutes in total. For Sarvam,
  one file of 10 to 30 seconds.
* **One speaker.** No other voices, and no music.
* **A quiet room.** No echo, fans or traffic. A phone held close in a small,
  furnished room is better than a good microphone in an empty one.
* **Speak the way the agent will.** Talk naturally, at the pace and tone of a
  phone call. A clone of someone reading in a flat voice will sound flat on
  calls.
* **Format.** WAV or FLAC is best. mp3, m4a, ogg and webm are accepted. Up to
  10 MB per file.

## Managing voices

```bash theme={null}
# List your voices
curl https://api.voice.wixzel.com/v1/voices \
  -H "Authorization: Bearer $WIXZEL_API_KEY"

# Read one
curl https://api.voice.wixzel.com/v1/voices/$VOICE_ID \
  -H "Authorization: Bearer $WIXZEL_API_KEY"

# Rename it
curl -X PATCH https://api.voice.wixzel.com/v1/voices/$VOICE_ID \
  -H "Authorization: Bearer $WIXZEL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "name": "Meera, reception" }'

# Delete it
curl -X DELETE https://api.voice.wixzel.com/v1/voices/$VOICE_ID \
  -H "Authorization: Bearer $WIXZEL_API_KEY"
```

`PATCH` changes `name` and `description` only. The recordings and the
provider cannot be changed; to change them, make a new clone.

`DELETE` deletes the clone at the provider too, and frees one of your 3 places.
It is refused with `409 voice_in_use` while an agent uses the voice, and the
message names the agents. Give them another voice first.

The list is cursor-paginated like every other list. See
[Pagination](/pagination).

### Scopes

Reading voices and previews needs `voices:read`. Making, renaming and deleting
them needs `voices:write`.

### SDKs and the MCP server

The [SDKs](/sdks) have the same operations:

```ts theme={null}
const voice = await client.voices.create({
  name: 'Meera, Malayalam',
  provider: 'sarvam',
  language: 'ml-IN',
  consent: true,
  audio_urls: ['https://example.com/recordings/meera-ml.wav'],
});
```

`client.voices.list`, `retrieve`, `update`, `delete` and `preview` cover the
rest.

The [MCP server](/mcp-server) has the tools `list_voices`, `get_voice`,
`preview_voice`, `create_voice_clone`, `update_voice` and `delete_voice`, so you can make a
clone from Claude Code, Codex or Cursor. `create_voice_clone` takes
`audio_urls`, since an MCP client cannot upload a file.

## Errors

| Status | `code` | Meaning |
| - | - | - |
| `400` | `missing_audio` | No recordings. Send `files` in a multipart request or `audio_urls` in JSON. |
| `400` | `ambiguous_audio` | Both `files` and `audio_urls` were sent. Send one. |
| `400` | `unsupported_audio` | A file is not audio. Send mp3, wav, m4a, ogg, webm or flac. |
| `400` | `too_many_files` | More recordings than the provider takes: 5 for `elevenlabs`, 1 for `sarvam`. |
| `400` | `missing_language` | `provider` is `sarvam` and `language` is missing. |
| `400` | `invalid_parameter` | A field is invalid. With `param: "consent"`, `consent` was not `true`. |
| `400` | `unsafe_audio_url` | An `audio_urls` link points to a private or local address. |
| `400` | `audio_url_unreachable` | An `audio_urls` link could not be downloaded. |
| `400` | `voice_clone_rejected` | The provider refused the recording. The message gives its reason. Usually too short, too noisy, or more than one speaker. |
| `402` | `insufficient_credits` | Your balance cannot cover the \$1.00 fee. Top up and try again. |
| `409` | `account_limit_reached` | The account already has 3 clones. Delete one first. |
| `409` | `voice_in_use` | On `DELETE`: an agent uses this voice. Give it another voice first. |
| `413` | `file_too_large` | A recording is larger than 10 MB. |
| `502` | `provider_unavailable` | The provider failed. Nothing was charged. Try again. |
| `503` | `voice_capacity` | Cloning is unavailable on the platform for now. Nothing was charged. Try again later. |

On agent writes (`POST /v1/agents`, `PATCH /v1/agents/{id}`):

| Status | `code` | Meaning |
| - | - | - |
| `400` | `unsupported_voice` | The voice is another account's clone, does not exist, or is a clone on a provider this engine cannot speak. |

See [Errors](/errors) for the error format.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.