voice takes one of two shapes: a composed pipeline of three
named models, or one realtime model that does the whole conversation. This
page is about the first. For the shapes themselves, see
Voice engines.
The shape
provider/model pair. All three stages are required:
a pipeline missing one cannot hold a conversation, and defaulting it would
silently bill you for a model you never chose.
Which models you can name
Only models in the price book. The book is the meter’s source of truth, so “no price, no call” is enforced when the agent is created, not on the first call, where the failure would land on your customer. Name something unpriced and the answer is a400 that lists what is available for that stage:
GET /v1/engines: each engine
carries its models, with the unit each is billed in and the price per unit.
Prefer it to hardcoding ids: a model whose provider is degraded disappears
from that list before your calls start failing.
Each entry carries a default flag, and the distinction matters for what a call
costs you:
default: trueis what the engine runs when your agent names no model, and it is the model the engine’s headline per-minute price is quoted from.- Everything else is selectable. Name it in the agent’s
voiceand the call is priced from your choice. Choosing a dearer model does not raise the credit anyone else needs; reservations are computed from each agent’s own configuration, not from the most expensive model in the catalogue.
GET /v1/engines lists the models an engine uses by default, which is what
its published per-minute price is calculated from. The language models below are
all selectable on the classic pipeline even though only openrouter/gpt-4o-mini
appears in that list. Naming one of the others is enough; you are billed at its
own rate.Language models
Any OpenRouter model in the price book can be named. Their ids are vendor-qualified:openrouter/ names the provider, and everything after
the first slash is OpenRouter’s own model id:
openrouter/gpt-4o-mini is the one exception: it predates that convention and
is kept working as written, so an agent that already names it needs no change.
Which to pick is mostly a latency and cost question, and on a phone call latency
is the one your caller notices. The cheap fast tiers (gpt-4o-mini,
google/gemini-2.5-flash-lite, mistralai/mistral-small-3.2-24b-instruct)
answer quickly enough that the pause after someone stops speaking feels natural.
A frontier model reasons better and takes longer doing it, which on a voice call
reads as hesitation rather than intelligence. Prices per million tokens are
below; a talkative minute is roughly 2,400 input and 500 output tokens.
Per-minute figures assume 15 synthesised characters, 40 input tokens and 8 output tokens per second, and exclude the flat $0.012 per minute platform fee. A language model is billed in two units, so its line appears twice.
Choosing the voice
tts.voice is the speaker your caller hears. It is a different kind of value
from tts.model: the model is the synthesiser, the voice is who it sounds like.
classic,
Bulbul speaker names for sarvam), so ask the engine rather than guessing:
voice and the pipeline uses its default, which is rarely what you want
past a first test. A voice from the wrong roster is refused when the agent is
created, listing the valid ones.
Choosing the language
stt.language is a code from the engine’s own list:
"language": "multi" to let Deepgram detect and code-switch instead of
pinning one. See Voice engines
for the response shape.
One family per pipeline
The three stages must come from the same family, because each family runs on its own pipeline:
The
stt provider decides which pipeline runs: a Deepgram model selects the
classic pipeline, a Sarvam model selects the Sarvam one. That pipeline then
runs its own family end to end. A tts from another family passes pricing
validation but is not what that pipeline speaks with, so keep all three
together.
Variants and language
Within a family the pipeline picks the exact variant for the agent’slanguage: the classic pipeline uses ElevenLabs Turbo for English and Flash
for everything else, and Nova-3 in the matching language mode. Naming a
tts.model explicitly overrides that choice. What you are billed is the model
that actually ran, at its own rate: the estimate at admission reserves for the
dearest variant in the family, and the settle charges for the one used.
What it costs
A composed minute is the sum of its stages plus the flat platform fee, at typical speech rates: 15 characters of synthesised speech, 40 input tokens and 8 output tokens per second. Each stage is metered in its own unit (seconds of audio, tokens, characters) every few seconds while the call is live, so a quiet call costs less than the estimate and a talkative one more. The pricing page has a builder that prices any mix and writes the JSON for it. Prices are frozen when a call is admitted, so a change to the book never moves a running meter. See Credits for how the balance works.Or hand the whole turn to one model
stt or llm;
the API refuses a config that names both.
voice and language work the same way here, against the realtime engine’s own
rosters: GET /v1/engines/gemini_live/voices lists Gemini’s prebuilt voices.
"language": "auto" is Gemini’s default and lets the model follow the caller,
switching mid-conversation if they do.
