Skip to main content
An agent’s voice takes one of two shapes: a composed pipeline of three named models, or one realtime model that does the whole conversation. This page is about the first. For the shapes themselves, see Voice engines.

The shape

Every model is a provider/model pair. All three stages are required: a pipeline missing one cannot hold a conversation, and defaulting it would silently bill you for a model you never chose.

Which models you can name

Only models in the price book. The book is the meter’s source of truth, so “no price, no call” is enforced when the agent is created, not on the first call, where the failure would land on your customer. Name something unpriced and the answer is a 400 that lists what is available for that stage:
The live list is GET /v1/engines: each engine carries its models, with the unit each is billed in and the price per unit. Prefer it to hardcoding ids: a model whose provider is degraded disappears from that list before your calls start failing. Each entry carries a default flag, and the distinction matters for what a call costs you:
  • default: true is what the engine runs when your agent names no model, and it is the model the engine’s headline per-minute price is quoted from.
  • Everything else is selectable. Name it in the agent’s voice and the call is priced from your choice. Choosing a dearer model does not raise the credit anyone else needs; reservations are computed from each agent’s own configuration, not from the most expensive model in the catalogue.
Language-model ids from OpenRouter are vendor-qualified, and only the first slash separates provider from model:
GET /v1/engines lists the models an engine uses by default, which is what its published per-minute price is calculated from. The language models below are all selectable on the classic pipeline even though only openrouter/gpt-4o-mini appears in that list. Naming one of the others is enough; you are billed at its own rate.

Language models

Any OpenRouter model in the price book can be named. Their ids are vendor-qualified: openrouter/ names the provider, and everything after the first slash is OpenRouter’s own model id:
openrouter/gpt-4o-mini is the one exception: it predates that convention and is kept working as written, so an agent that already names it needs no change. Which to pick is mostly a latency and cost question, and on a phone call latency is the one your caller notices. The cheap fast tiers (gpt-4o-mini, google/gemini-2.5-flash-lite, mistralai/mistral-small-3.2-24b-instruct) answer quickly enough that the pause after someone stops speaking feels natural. A frontier model reasons better and takes longer doing it, which on a voice call reads as hesitation rather than intelligence. Prices per million tokens are below; a talkative minute is roughly 2,400 input and 500 output tokens. Per-minute figures assume 15 synthesised characters, 40 input tokens and 8 output tokens per second, and exclude the flat $0.012 per minute platform fee. A language model is billed in two units, so its line appears twice.

Choosing the voice

tts.voice is the speaker your caller hears. It is a different kind of value from tts.model: the model is the synthesiser, the voice is who it sounds like.
Each family draws from its own roster (ElevenLabs voice ids for classic, Bulbul speaker names for sarvam), so ask the engine rather than guessing:
Omit voice and the pipeline uses its default, which is rarely what you want past a first test. A voice from the wrong roster is refused when the agent is created, listing the valid ones.

Choosing the language

stt.language is a code from the engine’s own list:
A composed engine only lists a language when both its speech-to-text and its text-to-speech serve it. That intersection is narrower than either provider’s own list, and it is the honest number: a language you can transcribe but not speak back gives an agent that understands the caller and cannot answer them. Set "language": "multi" to let Deepgram detect and code-switch instead of pinning one. See Voice engines for the response shape.

One family per pipeline

The three stages must come from the same family, because each family runs on its own pipeline: The stt provider decides which pipeline runs: a Deepgram model selects the classic pipeline, a Sarvam model selects the Sarvam one. That pipeline then runs its own family end to end. A tts from another family passes pricing validation but is not what that pipeline speaks with, so keep all three together.

Variants and language

Within a family the pipeline picks the exact variant for the agent’s language: the classic pipeline uses ElevenLabs Turbo for English and Flash for everything else, and Nova-3 in the matching language mode. Naming a tts.model explicitly overrides that choice. What you are billed is the model that actually ran, at its own rate: the estimate at admission reserves for the dearest variant in the family, and the settle charges for the one used.

What it costs

A composed minute is the sum of its stages plus the flat platform fee, at typical speech rates: 15 characters of synthesised speech, 40 input tokens and 8 output tokens per second. Each stage is metered in its own unit (seconds of audio, tokens, characters) every few seconds while the call is live, so a quiet call costs less than the estimate and a talkative one more. The pricing page has a builder that prices any mix and writes the JSON for it. Prices are frozen when a call is admitted, so a change to the book never moves a running meter. See Credits for how the balance works.

Or hand the whole turn to one model

A realtime model replaces all three stages and is billed per token or per second, depending on the provider. It cannot be combined with stt or llm; the API refuses a config that names both. voice and language work the same way here, against the realtime engine’s own rosters: GET /v1/engines/gemini_live/voices lists Gemini’s prebuilt voices. "language": "auto" is Gemini’s default and lets the model follow the caller, switching mid-conversation if they do.