> ## Documentation Index
> Fetch the complete documentation index at: https://docs.voice.wixzel.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Compose a voice AI pipeline: speech-to-text, LLM and text-to-speech

> Name the speech-to-text, language model and text-to-speech models yourself, from Deepgram, OpenRouter, ElevenLabs or Sarvam, and see what each costs.

An agent's `voice` takes one of two shapes: a **composed** pipeline of three
named models, or one **realtime** model that does the whole conversation. This
page is about the first. For the shapes themselves, see
[Voice engines](/voice-engines).

## The shape

```json theme={null}
"voice": {
  "stt": { "model": "deepgram/nova-3", "language": "en-US" },
  "llm": { "model": "openrouter/gpt-4o-mini", "temperature": 0.7 },
  "tts": { "model": "elevenlabs/eleven_turbo_v2_5", "voice": "21m00Tcm4TlvDq8ikWAM" }
}
```

Every model is a `provider/model` pair. All three stages are required:
a pipeline missing one cannot hold a conversation, and defaulting it would
silently bill you for a model you never chose.

## Which models you can name

Only models in the price book. The book is the meter's source of truth, so
"no price, no call" is enforced when the agent is created, not on the first
call, where the failure would land on your customer. Name something unpriced
and the answer is a `400` that lists what is available for that stage:

```json theme={null}
{
  "error": {
    "type": "invalid_request_error",
    "code": "unsupported_model",
    "message": "No pricing for acme/whisper-9. Available for stt: deepgram/nova-3, deepgram/nova-3-multilingual, sarvam/saaras-v3-realtime.",
    "param": "voice"
  }
}
```

The live list is [`GET /v1/engines`](/api-reference/engines): each engine
carries its `models`, with the unit each is billed in and the price per unit.
Prefer it to hardcoding ids: a model whose provider is degraded disappears
from that list before your calls start failing.

Each entry carries a `default` flag, and the distinction matters for what a call
costs you:

* **`default: true`** is what the engine runs when your agent names no model,
  and it is the model the engine's headline per-minute price is quoted from.
* **Everything else is selectable.** Name it in the agent's `voice` and the call
  is priced from *your* choice. Choosing a dearer model does not raise the
  credit anyone else needs; reservations are computed from each agent's own
  configuration, not from the most expensive model in the catalogue.

Language-model ids from OpenRouter are vendor-qualified, and only the first
slash separates provider from model:

```
openrouter/anthropic/claude-sonnet-4.5
└────┬───┘ └──────────┬─────────────┘
 provider          model id
```

<Note>
  `GET /v1/engines` lists the models an engine **uses by default**, which is what
  its published per-minute price is calculated from. The language models below are
  all selectable on the classic pipeline even though only `openrouter/gpt-4o-mini`
  appears in that list. Naming one of the others is enough; you are billed at its
  own rate.
</Note>

### Language models

Any OpenRouter model in the price book can be named. Their ids are
**vendor-qualified**: `openrouter/` names the provider, and everything after
the first slash is OpenRouter's own model id:

```json theme={null}
"llm": { "model": "openrouter/anthropic/claude-sonnet-4.5" }
```

`openrouter/gpt-4o-mini` is the one exception: it predates that convention and
is kept working as written, so an agent that already names it needs no change.

Which to pick is mostly a latency and cost question, and on a phone call latency
is the one your caller notices. The cheap fast tiers (`gpt-4o-mini`,
`google/gemini-2.5-flash-lite`, `mistralai/mistral-small-3.2-24b-instruct`)
answer quickly enough that the pause after someone stops speaking feels natural.
A frontier model reasons better and takes longer doing it, which on a voice call
reads as hesitation rather than intelligence. Prices per million tokens are
below; a talkative minute is roughly 2,400 input and 500 output tokens.

| Stage | Model | Family | Price | ≈ per minute |
| - | - | - | - | - |
| Language model | `openrouter/anthropic/claude-haiku-4.5` | classic | \$0.0014 per 1k tokens in | \$0.0034 |
| Language model | `openrouter/anthropic/claude-haiku-4.5` | classic | \$0.0070 per 1k tokens out | \$0.0034 |
| Language model | `openrouter/anthropic/claude-sonnet-4.5` | classic | \$0.0042 per 1k tokens in | \$0.0101 |
| Language model | `openrouter/anthropic/claude-sonnet-4.5` | classic | \$0.0210 per 1k tokens out | \$0.0101 |
| Language model | `openrouter/deepseek/deepseek-chat-v3.1` | classic | \$0.0003 per 1k tokens in | \$0.0008 |
| Language model | `openrouter/deepseek/deepseek-chat-v3.1` | classic | \$0.0013 per 1k tokens out | \$0.0006 |
| Language model | `openrouter/google/gemini-2.5-flash` | classic | \$0.0004 per 1k tokens in | \$0.0010 |
| Language model | `openrouter/google/gemini-2.5-flash` | classic | \$0.0035 per 1k tokens out | \$0.0017 |
| Language model | `openrouter/google/gemini-2.5-flash-lite` | classic | \$0.0001 per 1k tokens in | \$0.0003 |
| Language model | `openrouter/google/gemini-2.5-flash-lite` | classic | \$0.0006 per 1k tokens out | \$0.0003 |
| Language model | `openrouter/gpt-4o-mini` | classic | \$0.0002 per 1k tokens in | \$0.0005 |
| Language model | `openrouter/gpt-4o-mini` | classic | \$0.0008 per 1k tokens out | \$0.0004 |
| Language model | `openrouter/meta-llama/llama-3.3-70b-instruct` | classic | \$0.0001 per 1k tokens in | \$0.0003 |
| Language model | `openrouter/meta-llama/llama-3.3-70b-instruct` | classic | \$0.0004 per 1k tokens out | \$0.0002 |
| Language model | `openrouter/mistralai/mistral-small-3.2-24b-instruct` | classic | \$0.0001 per 1k tokens in | \$0.0003 |
| Language model | `openrouter/mistralai/mistral-small-3.2-24b-instruct` | classic | \$0.0003 per 1k tokens out | \$0.0001 |
| Language model | `openrouter/openai/gpt-4.1-mini` | classic | \$0.0006 per 1k tokens in | \$0.0013 |
| Language model | `openrouter/openai/gpt-4.1-mini` | classic | \$0.0022 per 1k tokens out | \$0.0011 |
| Language model | `openrouter/openai/gpt-4o` | classic | \$0.0035 per 1k tokens in | \$0.0084 |
| Language model | `openrouter/openai/gpt-4o` | classic | \$0.0140 per 1k tokens out | \$0.0067 |
| Language model | `sarvam/sarvam-105b` | sarvam | \$0.0004 per 1k tokens in | \$0.0011 |
| Language model | `sarvam/sarvam-105b` | sarvam | \$0.0011 per 1k tokens out | \$0.0005 |
| Realtime | `deepgram/agent` | realtime | \$0.0018 per second | \$0.1050 |
| Realtime | `google/gemini-live` | realtime | \$0.0042 per 1k audio tokens in | \$0.0063 |
| Realtime | `google/gemini-live` | realtime | \$0.0168 per 1k audio tokens out | \$0.0252 |
| Realtime | `google/gemini-live` | realtime | \$0.0007 per 1k tokens in | \$0.0017 |
| Realtime | `google/gemini-live` | realtime | \$0.0028 per 1k tokens out | \$0.0013 |
| Speech-to-text | `deepgram/nova-3` | classic | \$0.0002 per second | \$0.0108 |
| Speech-to-text | `deepgram/nova-3-multilingual` | classic | \$0.0002 per second | \$0.0129 |
| Speech-to-text | `sarvam/saaras-v3-realtime` | sarvam | \$0.0001 per second | \$0.0076 |
| Text-to-speech | `elevenlabs/eleven_flash_v2_5` | classic | \$0.0700 per 1k characters | \$0.0630 |
| Text-to-speech | `elevenlabs/eleven_multilingual_v2` | classic | \$0.1400 per 1k characters | \$0.1260 |
| Text-to-speech | `elevenlabs/eleven_turbo_v2_5` | classic | \$0.0700 per 1k characters | \$0.0630 |
| Text-to-speech | `sarvam/bulbul-v3` | sarvam | \$0.0457 per 1k characters | \$0.0411 |

Per-minute figures assume 15 synthesised characters, 40 input tokens and 8 output tokens per second, and exclude the flat \$0.012 per minute platform fee. A language model is billed in two units, so its line appears twice.

## Choosing the voice

`tts.voice` is the speaker your caller hears. It is a different kind of value
from `tts.model`: the model is the synthesiser, the voice is who it sounds like.

```json theme={null}
"tts": { "model": "elevenlabs/eleven_turbo_v2_5", "voice": "21m00Tcm4TlvDq8ikWAM" }
```

Each family draws from its own roster (ElevenLabs voice ids for `classic`,
Bulbul speaker names for `sarvam`), so ask the engine rather than guessing:

```bash theme={null}
curl https://api.voice.wixzel.com/v1/engines/classic/voices \
  -H "Authorization: Bearer $WIXZEL_API_KEY"
```

Omit `voice` and the pipeline uses its default, which is rarely what you want
past a first test. A voice from the wrong roster is refused when the agent is
created, listing the valid ones.

## Choosing the language

`stt.language` is a code from the engine's own list:

```bash theme={null}
curl https://api.voice.wixzel.com/v1/engines/classic/languages \
  -H "Authorization: Bearer $WIXZEL_API_KEY"
```

A composed engine only lists a language when **both** its speech-to-text and its
text-to-speech serve it. That intersection is narrower than either provider's
own list, and it is the honest number: a language you can transcribe but not
speak back gives an agent that understands the caller and cannot answer them.

Set `"language": "multi"` to let Deepgram detect and code-switch instead of
pinning one. See [Voice engines](/voice-engines#discovering-what-is-available)
for the response shape.

## One family per pipeline

The three stages must come from the same **family**, because each family runs
on its own pipeline:

| Family | Speech-to-text | Language model | Text-to-speech |
| - | - | - | - |
| `classic` | `deepgram/…` | `openrouter/…` | `elevenlabs/…` |
| `sarvam` | `sarvam/…` | `sarvam/…` | `sarvam/…` |

The `stt` provider decides which pipeline runs: a Deepgram model selects the
classic pipeline, a Sarvam model selects the Sarvam one. That pipeline then
runs its own family end to end. A `tts` from another family passes pricing
validation but is not what that pipeline speaks with, so keep all three
together.

## Variants and language

Within a family the pipeline picks the exact variant for the agent's
`language`: the classic pipeline uses ElevenLabs Turbo for English and Flash
for everything else, and Nova-3 in the matching language mode. Naming a
`tts.model` explicitly overrides that choice. What you are billed is the model
that actually ran, at its own rate: the estimate at admission reserves for the
dearest variant in the family, and the settle charges for the one used.

## What it costs

A composed minute is the sum of its stages plus the flat platform fee, at
typical speech rates: 15 characters of synthesised speech, 40 input tokens
and 8 output tokens per second. Each stage is metered in its own unit
(seconds of audio, tokens, characters) every few seconds while the call is
live, so a quiet call costs less than the estimate and a talkative one more.

The [pricing page](https://voice.wixzel.com/pricing#compose) has a builder
that prices any mix and writes the JSON for it.

Prices are frozen when a call is admitted, so a change to the book never
moves a running meter. See [Credits](/credits) for how the balance works.

## Or hand the whole turn to one model

```json theme={null}
"voice": {
  "realtime": { "model": "google/gemini-live", "voice": "Charon", "language": "auto" }
}
```

A realtime model replaces all three stages and is billed per token or per
second, depending on the provider. It cannot be combined with `stt` or `llm`;
the API refuses a config that names both.

`voice` and `language` work the same way here, against the realtime engine's own
rosters: `GET /v1/engines/gemini_live/voices` lists Gemini's prebuilt voices.
`"language": "auto"` is Gemini's default and lets the model follow the caller,
switching mid-conversation if they do.
