> ## Documentation Index
> Fetch the complete documentation index at: https://docs.voice.wixzel.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Clone a voice

> Makes a cloned voice from recordings of one person speaking, and charges the one-off fee in the price book (`voice_clone`). `consent` must be `true`. Send the recordings as `files` in a multipart request, or as `audio_urls` in JSON. The recordings are passed to the provider and not stored. Put the returned `provider_voice_id` in an agent’s `tts.voice`; only engines listed in `engines` can speak in it. Requires `voices:write`.



## OpenAPI

````yaml /openapi.json post /v1/voices
openapi: 3.1.0
info:
  title: Wixzel Voice API
  version: '2026-09-01'
  description: >-
    APIs for agentic telephony. One API key, one balance, every voice engine.


    ## Authentication

    Send your key as `Authorization: Bearer wv_live_...`. Keys are scoped: grant
    only what an integration needs. There is no admin scope.


    ## Versioning

    The `/v1` prefix covers additive changes. Behavioural changes ship behind a
    dated `Wixzel-Version` header, and existing keys keep the behaviour they
    were created with.


    ## Money

    Amounts are integer **micro-USD** (1,000,000 = $1.00). Voice usage is billed
    per second, per token and per character, so a float dollar figure cannot
    represent it without disagreeing with the ledger.


    ## Errors

    Every error carries a stable `code` and a `request_id`. Match on `code`; the
    `message` is for humans and may change.
  contact:
    name: Wixzel Voice
    url: https://voice.wixzel.com
servers:
  - url: https://api.voice.wixzel.com
    description: Production
security: []
tags:
  - name: Calls
    description: Place calls and read what happened on them.
  - name: Realtime
    description: Talk to an agent from a browser or app, with no phone line.
  - name: Agents
    description: The prompt, voice and behaviour of a caller.
  - name: Leads
    description: People to call, and data to merge into prompts.
  - name: Campaigns
    description: Call a list of leads with one agent.
  - name: Knowledge bases
    description: Facts an agent can draw on mid-call.
  - name: Phone numbers
    description: Numbers on your SIP trunks.
  - name: SIP trunks
    description: Your carrier connections.
  - name: Appointments
    description: Bookings, including ones agents make on calls.
  - name: Usage
    description: Itemised billing lines.
  - name: Billing
    description: Balance, ledger and top-ups.
  - name: API keys
    description: Create, scope and rotate keys.
  - name: Engines
    description: What the platform can serve, and what it costs.
  - name: Voices
    description: Cloned voices, made from your own recordings, for your agents to speak in.
  - name: Webhooks
    description: Where events are sent, and what happened when they got there.
  - name: MCP servers
    description: Other apps’ tools, for agents to use while they are on a call.
paths:
  /v1/voices:
    post:
      tags:
        - Voices
      summary: Clone a voice
      description: >-
        Makes a cloned voice from recordings of one person speaking, and charges
        the one-off fee in the price book (`voice_clone`). `consent` must be
        `true`. Send the recordings as `files` in a multipart request, or as
        `audio_urls` in JSON. The recordings are passed to the provider and not
        stored. Put the returned `provider_voice_id` in an agent’s `tts.voice`;
        only engines listed in `engines` can speak in it. Requires
        `voices:write`.
      requestBody:
        content:
          multipart/form-data:
            schema:
              $ref: '#/components/schemas/CreateVoiceMultipart'
          application/json:
            schema:
              $ref: '#/components/schemas/CreateVoice'
      responses:
        '201':
          description: The new voice
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Voice'
        '400':
          description: Invalid request
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Error'
        '401':
          description: Missing or invalid API key
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Error'
        '402':
          description: Not enough credit for the cloning fee
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Error'
        '403':
          description: Key lacks the required scope
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Error'
        '409':
          description: The account already has its maximum number of voices
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Error'
        '413':
          description: A recording is larger than 10 MB
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Error'
        '429':
          description: Rate limited. Retry after the interval in `Retry-After`.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Error'
        '503':
          description: Voice cloning is temporarily unavailable. Nothing was charged.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/Error'
      security:
        - bearerAuth: []
components:
  schemas:
    CreateVoiceMultipart:
      type: object
      properties:
        name:
          type: string
          minLength: 1
          maxLength: 100
          example: Dr. Meera, front desk
        description:
          type: string
          maxLength: 500
        provider:
          type: string
          enum:
            - elevenlabs
            - sarvam
          default: elevenlabs
          description: >-
            Where the clone is made, which decides the engines it speaks on.
            `elevenlabs` for the `classic` and `deepgram_agent` engines;
            `sarvam` for the `sarvam` engine, which is the one to use for Indian
            languages.
        language:
          type: string
          pattern: ^[a-z]{2}-[A-Z]{2}$
          description: The language spoken in the recording. Required for `sarvam`.
          example: ml-IN
        remove_background_noise:
          type: boolean
          description: >-
            ElevenLabs only. Cleans up a noisy recording; it can make a clean
            one worse.
        consent:
          type: boolean
          description: >-
            Must be `true`. You confirm that the voice is yours, or that the
            speaker has given you permission to clone it and use it on calls.
            Recorded with the time and your account.
        files:
          type: array
          items:
            type: string
            format: binary
          minItems: 1
          maxItems: 5
          description: >-
            The recordings. 30 seconds to a few minutes of one person speaking
            clearly, with no music or other voices. mp3, wav, m4a, ogg, webm or
            flac, up to 10 MB each. ElevenLabs takes up to 5 files; Sarvam takes
            1, and 10 to 30 seconds is enough.
      required:
        - name
        - consent
        - files
    CreateVoice:
      type: object
      properties:
        name:
          type: string
          minLength: 1
          maxLength: 100
          example: Dr. Meera, front desk
        description:
          type: string
          maxLength: 500
        provider:
          type: string
          enum:
            - elevenlabs
            - sarvam
          default: elevenlabs
          description: >-
            Where the clone is made, which decides the engines it speaks on.
            `elevenlabs` for the `classic` and `deepgram_agent` engines;
            `sarvam` for the `sarvam` engine, which is the one to use for Indian
            languages.
        language:
          type: string
          pattern: ^[a-z]{2}-[A-Z]{2}$
          description: The language spoken in the recording. Required for `sarvam`.
          example: ml-IN
        remove_background_noise:
          type: boolean
          description: >-
            ElevenLabs only. Cleans up a noisy recording; it can make a clean
            one worse.
        consent:
          type: boolean
          description: >-
            Must be `true`. You confirm that the voice is yours, or that the
            speaker has given you permission to clone it and use it on calls.
            Recorded with the time and your account.
        audio_urls:
          type: array
          items:
            type: string
            maxLength: 2048
            format: uri
          minItems: 1
          maxItems: 5
          description: >-
            JSON requests only: public https links to the recordings, fetched
            once and not stored. Multipart requests send the files themselves as
            `files` instead.
      required:
        - name
        - consent
    Voice:
      type: object
      properties:
        id:
          type: string
          pattern: ^[0-9a-f]{24}$
          example: 6a96a3ead6e886d42462dd3e
        object:
          type: string
          enum:
            - voice
        name:
          type: string
        description:
          type:
            - string
            - 'null'
        provider:
          type: string
          enum:
            - elevenlabs
            - sarvam
          description: >-
            Where the clone is made, which decides the engines it speaks on.
            `elevenlabs` for the `classic` and `deepgram_agent` engines;
            `sarvam` for the `sarvam` engine, which is the one to use for Indian
            languages.
        provider_voice_id:
          type: string
          description: Put this in an agent’s `tts.voice` to make it speak in this voice.
          example: c38kUX8pkfYO2kHyqfFy
        engines:
          type: array
          items:
            type: string
          description: The engines that can speak in this voice.
          example:
            - classic
            - deepgram_agent
        language:
          type:
            - string
            - 'null'
          description: >-
            The language of the recording. Sarvam clones can still speak every
            language Sarvam supports.
        created_at:
          type: string
          format: date-time
          description: ISO 8601, always UTC.
          example: '2026-09-01T12:00:00.000Z'
        updated_at:
          type: string
          format: date-time
          description: ISO 8601, always UTC.
          example: '2026-09-01T12:00:00.000Z'
      required:
        - id
        - object
        - name
        - description
        - provider
        - provider_voice_id
        - engines
        - language
        - created_at
        - updated_at
    Error:
      type: object
      properties:
        error:
          type: object
          properties:
            type:
              type: string
              enum:
                - invalid_request_error
                - authentication_error
                - permission_error
                - rate_limit_error
                - insufficient_credits
                - not_found_error
                - conflict_error
                - api_error
            code:
              type: string
              description: Stable machine-readable code.
              example: agent_not_found
            message:
              type: string
              description: >-
                Human-readable explanation. Do not match on this — match on
                code.
            param:
              type: string
              description: Which field caused the failure, when applicable.
            doc_url:
              type: string
            request_id:
              type: string
              description: Quote this when asking for help.
              example: req_01HXYZ...
          required:
            - type
            - code
            - message
            - request_id
      required:
        - error
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      description: >-
        Your Wixzel Voice API key: `Authorization: Bearer wv_live_...`. Keys are
        scoped; grant only what the integration needs.

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.