Pipeline or speech-to-speech? Run the numbers.

Find the best voice AI architecture for your use case. Paste your system prompt, set your call duration and the language your agent speaks, and compare Gemini 3.8 Live, Gemini 3.1 Flash Live, GPT-Realtime-mini, Grok Voice 2.0 and GPT-Live-1 with an orchestrated STT + LLM + TTS pipeline, at public list prices.

  • 6 architectures
  • Turn-by-turn pricing
  • Response-time chains
  • Nothing leaves your browser
01 Your use case

Four inputs. The whole page updates as you type.

Paste your prompt, set the call length, say who calls whom and in which language. Prices are public list prices as of 15 September 2026, model costs only.

configure-your-callsimple-inbound
Preset
1System prompt≈ 4 characters per token
or pick a size
2Call durationaverage, ring excluded

≈ 11 dialogue turns at 5.3 turns per minute

3Direction
4Agent languageweighs the voices
Advanced
▸Tools, conversation shape, pipeline components, backend, EU residency
Tools
Tools the agent can call

Each tool adds about 180 tokens of definitions to every turn

Tool calls per call

One extra model round-trip each. Auto: about 1.5 per minute

Conversation
Turns per minute

Speech-to-speech re-bills the context on every turn

Share of the call with speech

The rest is silence and thinking time

Agent share of the speech

Drives TTS and audio output

Pipeline stackpick a component or set your price
Speech-to-text

Billed on the whole call

LLM

Input / output per 1M; the repeated prefix bills at the cached rate (10%)

Text-to-speech

Billed on the minutes of generated speech

GPT-Live-1
Backend model

The model GPT-Live-1 delegates reasoning and tools to

Options
best-for-your-callinbound · 2 min · 3,000 tok
Best value for moneySpeech-to-speechinbound · feel first

Gemini 3.1 Flash Live

$0.047/mincost per minute
$0.094per 2 min call
~0.8 sresponse · 1 hop
feel control voice · en p95 1.4 s
Why

+187% over the cheapest (GPT-Realtime-2.1-mini), for a smoother conversation and a faster answer (~0.8 s vs ~0.9 s): one hop, no STT → LLM → TTS hand-offs. That is $0.061 more per call. It buys the best voice of the field: Gemini 3.1’s voice family leads the Artificial Analysis Pronunciation Robustness benchmark (88.1%), so names, amounts, dates and emails come out right. The top feel earns +100% of extra budget on inbound calls. Trade-off: control 3.5/5. Speech-to-speech behaves like a fixed temperature of 1; keep the pipeline for deterministic tool use.

Other interesting architectures
Cheapest · GPT-Realtime-2.1-mini$0.016/min−65%~0.9 s
Pipeline$0.020/min−57%~1.4 s
Gemini 3.8 Live$0.026/min−44%~0.8 s

Price delta versus the best value; hover the cheapest for why it wins on price.

Free, 15 min. Bring your prompt, leave with a stack.

Sign up
02 Response time

Why speech-to-speech feels faster.

From the caller’s last word to the agent’s first. A pipeline chains four vendors, each a network hop with its own queue: the delays add up, and the slowest hop of the moment sets the tail. Speech-to-speech is one model, one hop.

Grok Voice 2.01 hop~0.7 s typical · p95 1.3 s

End of speech, in-model 200 msFirst audio 500 ms

Gemini 3.1 Flash Live1 hop~0.8 s typical · p95 1.4 s

End of speech, in-model 250 msFirst audio 500 ms

Gemini 3.8 Live1 hop~0.8 s typical · p95 1.4 s

End of speech, in-model 250 msFirst audio 500 ms

GPT-Realtime-2.1-mini1 hop~0.9 s typical · p95 1.6 s

End of speech, in-model 300 msFirst audio 600 ms

GPT-Live-1 + backend1 hop~0.9 s typical · p95 1.8 s

Full duplex, no explicit turn detection 200 msFirst audio (backend thinks in parallel) 700 ms

Pipeline4 hops~1.4 s typical · p95 2.6 s

VAD: silence wait + endpointing 400 msSTT final transcript 250 msLLM first token 350 msTTS first audio 200 msHop overhead (network, queues) 200 ms

Typical figures from public benchmarks and field measurements, not vendor numbers. The sum of vendor-quoted latencies is a floor: turn detection has to wait for silence, the transcript has to be final before the LLM starts, the first token has to arrive before the TTS can speak, and every hop adds serialisation, network and queueing time. A speech-to-speech model hears the end of the sentence and starts answering in the same stream, which is why the conversation feels smoother than the numbers suggest.

03 Cost per minute

What a minute costs, by architecture.

Model costs only, list prices 15 September 2026. Hover a segment for its share.

GPT-Realtime-2.1-miniSpeech-to-speechcheapest$0.016 / min

Prompt + tools (cached after turn 1) $0.002 (15%)Audio input (fresh + cached history) $0.003 (19%)Audio output $0.011 (66%)

PipelinePipeline$0.020 / min

Speech-to-text (Deepgram Nova-3) $0.006 (29%)LLM (GPT-5.6 Luna) $0.001 (5%)Text-to-speech (Cartesia Sonic) $0.014 (67%)

Gemini 3.8 LiveSpeech-to-speech$0.026 / min

Prompt + tools re-billed every turn $0.009 (36%)Audio history re-read every turn $0.012 (45%)Audio output $0.005 (17%)Transcription (text output) $0.000 (1%)

Gemini 3.1 Flash LiveSpeech-to-speech$0.047 / min

Prompt + tools re-billed every turn $0.017 (36%)Audio history re-read every turn $0.021 (45%)Audio output $0.008 (17%)Transcription (text output) $0.001 (1%)

GPT-Live-1 + backendSpeech-to-speech + LLM$0.051 / min

GPT-Live-1 voice layer ($0.05/min) $0.050 (98%)Backend (GPT-5.6 Luna, 4 delegations) $0.001 (2%)

Grok Voice 2.0Speech-to-speech$0.082 / min

Session audio (flat per minute) $0.080 (98%)Text items ($0.004 each) $0.002 (2%)

04 Calculation details

The call, turn by turn.

The totals above are the sum of these turns. Each turn: the caller speaks, the architecture answers after its response-time chain (VAD included for the pipeline), the agent speaks, then a pause. Open an architecture to see what every turn bills and how long it waits.

05 Sensitivity

Where the answer flips.

Cost per minute as one input moves and the others stay at your values. The vertical marker is your use case.

PipelineGemini 3.1 Flash LiveGemini 3.8 LiveGPT-Realtime-2.1-miniGrok Voice 2.0GPT-Live-1 + backend
By system prompt size
$0$0.050$0.10$0.1510010.1k20kprompt tokens

you · 2.6k prompt tokens

Realtime-mini$0.016

Pipeline$0.020

Gemini 3.8$0.025

Gemini 3.1$0.045

GPT-Live-1$0.051

Grok 2.0$0.082

By call duration
$0$0.050$0.10$0.1515 s5.1 min10 mincall duration

you · 1.9 min call duration

Realtime-mini$0.016

Pipeline$0.020

Gemini 3.8$0.025

Gemini 3.1$0.045

GPT-Live-1$0.051

Grok 2.0$0.082

By number of tools
$0$0.025$0.050$0.075$0.1002040tools

you · 2 tools

Realtime-mini$0.016

Pipeline$0.020

Gemini 3.8$0.026

Gemini 3.1$0.047

GPT-Live-1$0.051

Grok 2.0$0.082

The pipeline and Realtime-mini stay flat because their prompt is cached. Gemini Live climbs with every token you add to the prompt, every extra tool and every minute of call: the Live API re-bills the whole session context on each of the ~5.3 turns per minute. GPT-Live-1 is a flat line at $0.05/min plus a few backend tokens; Grok Voice 2.0 is flat at $0.08/min whatever the prompt, the tools or the minutes.

06 Side by side

The main options at a glance.

Price for your use case; feel, control and response time are what you get whatever the price. Hover a cell for the detail.

Architecture$ / min$ / callPrompt sensitivityFeelVoice · enControlResponse timeAA indexBest for

GPT-Realtime-2.1-minicheapest

Native audio in and out, prompt caching on the session

$0.016$0.033low~0.9 s1 hop · p95 1.6 snot scoredSimple flows where one vendor and a native feel beat fine control

Pipeline

Deepgram Nova-3 → GPT-5.6 Luna → Cartesia Sonic

$0.020$0.041flat~1.4 s4 hops · p95 2.6 sper componentVolume, long calls, heavy prompts, many tools, strict control

Gemini 3.8 Live

Native audio in, native audio out, one model; ~44% below Gemini 3.1 Flash Live

$0.026$0.052high~0.8 s1 hop · p95 1.4 s76.0Short calls with a compact prompt: the cheapest native feel

Gemini 3.1 Flash Livebest value

Native audio in, native audio out, one model

$0.047$0.094high~0.8 s1 hop · p95 1.4 s71.5Short calls with a compact prompt where the feel matters most

GPT-Live-1 + backend

Full-duplex voice layer, reasoning delegated to a text model

$0.051$0.102none~0.9 s1 hop · p95 1.8 s81.5Complex tasks needing reasoning, with a predictable bill

Grok Voice 2.0

Native audio in and out, flat $0.08 per audio minute (Think Fast 2.0)

$0.082$0.164flat~0.7 s1 hop · p95 1.3 s82.9Long calls or heavy prompts at a flat, predictable price; fastest answer of the field

Feel: how smooth and natural the conversation sounds; the pipeline scores 3 with a premium STT and TTS pair, 2 otherwise, and every score moves with the voice fit in the agent’s language (a native-grade voice earns half a dot, a 3 loses a full dot, a 2 loses two: a bad voice is a bad experience). Voice: reputation of each model’s voices in that language, 1 to 5. Control: 5 for the deterministic pipeline (temperature, structured output, any component); speech-to-speech models are rated in proportion to their Artificial Analysis Speech to Speech Index. Response time: typical delay between the caller’s last word and the agent’s first, and the 95th percentile; hops are the vendors in series. AA index: Artificial Analysis Speech to Speech Index (pipelines are scored per component).

both keep your exact settings

best value

Gemini 3.1 · $0.047/min · ~0.8 s

07 When is what best

Four situations, four answers.

What moves the answer: Gemini context compression, the number of tools, the dialogue pace and EU residency. All of them are in the advanced settings.

Short calls, compact prompt

under 1 min · under 3,000 tokens · a couple of tools

Under a minute, a prompt under 3,000 tokens, a couple of tools: speech-to-speech is in the same price band as the pipeline. Realtime-mini is usually the cheapest of all. Take the smoother conversation.

1 hoptake the feel

Long calls, heavy prompt, many tools

over 4 min · over 6,000 tokens · tool-heavy

Past 4 minutes or 6,000 tokens the pipeline pulls away: the prompt is cached, the transcript is small text. Gemini Live costs 2 to 5 times more here; GPT-Live-1 stays flat but never wins on price.

2 to 5×Gemini Live premium

Reasoning and multi-step actions

booking · CRM look-ups · identity checks

Booking, CRM look-ups, identity checks: GPT-Live-1 with a strong backend keeps the voice fluid while the backend thinks, at a predictable $0.05 per minute floor. A pipeline with a strong LLM does the same with full control over the tool calls.

$0.05per min floor, flat

Custom or cloned voice, brand voice

the voice is the product

Speech-to-speech voices are the model's own: Gemini 3.1 and 3.8 Live now speak French or Spanish natively, so language alone no longer forces a pipeline (set the agent language above and the scores follow). A cloned voice, a brand voice or a sovereign model still does: only a pipeline lets you plug ElevenLabs, Gradium, Cartesia or your own TTS.

any TTSpipeline only
08 Claude skill

Run CallShift from Claude.

Drop our skill into Claude and drive your whole environment in plain language: list and duplicate agents, review a campaign's ROI, add contacts, assign SIP numbers, wire up webhooks. It talks to our API for you.

  • Manages your CallShift agents, campaigns and calls through the API, with a disposable key
  • Reviews your agent against 140 voice AI best practices: prompt structure, turn-taking, STT and TTS, tools, outbound compliance, cost control
  • Architecture advice: which family fits your prompt size, with the simulator for the numbers
Free · works in Claude and Claude Code · EN, FR, ES

Add it in Claude under Settings → Capabilities → Skills, or drop it in your Claude Code skills folder.

09 Method

How the simulator computes a minute.

One call is modelled as a duration, a number of dialogue turns and a share of speech: 5.3 dialogue turns per minute by default, 75% of the call with someone speaking, 60% of that speech by the agent, 3.36 text tokens per second of speech (about 150 words per minute) and 180 tokens per tool definition. Every architecture is priced on that same call, turn by turn: the caller speaks, the architecture answers after its response-time chain (VAD, STT, LLM, TTS and hops for the pipeline; one in-model hop for speech-to-speech), the agent speaks, then a pause. The totals are the sum of the turns, and the turn-by-turn section shows every one of them.

  • Pipeline = STT on the whole call (Deepgram Nova-3 by default, $0.0058 per minute) + TTS on the agent's speech (Cartesia Sonic by default, $0.03 per minute) + LLM tokens: the prompt and tool definitions fresh on the first turn then cached at 10%, the transcript history cached, the newest utterance fresh, one extra round-trip per tool call, the agent's words as output.
  • Gemini 3.1 Flash Live = on every turn the prompt and tools as text input ($0.75 per 1M) plus all audio so far as audio input ($3 per 1M, 25 tokens per second), the agent's audio as output ($12 per 1M) and the transcript as text output ($4.50 per 1M). No caching; optional sliding-window compression.
  • Gemini 3.8 Live = the same billing model at 0.56 times every Gemini 3.1 Flash Live rate, the ratio of the Artificial Analysis cost per hour reported in Google DeepMind's evaluation ($0.84 against $1.50).
  • GPT-Realtime-2.1-mini = the prompt fresh once ($0.60 per 1M) then cached ($0.06), each user utterance once as fresh audio ($10 per 1M, 10 tokens per second), the audio history at the cached rate ($0.30), the agent's audio as output ($20 per 1M, 20 tokens per second).
  • GPT-Live-1 = $0.05 per minute for the voice layer plus one backend delegation per minute and per tool call, each re-reading the cached prompt and the transcript so far and producing about 216 output tokens at the backend's rates (GPT-5.6 Luna: $0.20 in, $0.02 cached, $1.20 out per 1M).

Model costs only, in every architecture. Quality and latency come from artificialanalysis.ai: the Speech to Speech Index (GPT-Live-1 with GPT-6 Astra 81.5, GPT-Realtime-2.1 High 79.1, Gemini 3.8 Live 76.0, Gemini 3.1 Flash Live High 71.5) and the time to first audio measured on reasoning questions (about 1.3 s, 2.3 s and 3.0 s with thinking on; conversational turns without thinking are well under a second). The response-time chains are typical figures from public benchmarks and field measurements, not vendor numbers: a pipeline's response time is turn detection + STT + LLM + TTS + one network hop per vendor, and its tail is set by the slowest hop of the moment. Public list prices checked on 15 September 2026; volume and enterprise discounts are not modelled.

10 Questions

Frequently asked questions.

Is a speech-to-speech model cheaper than an STT + LLM + TTS pipeline?

It depends on the model. At September 2026 list prices, a pipeline (Deepgram Nova-3 at $0.0058 per minute, a cached cheap LLM, Cartesia Sonic at $0.03 per minute of speech) costs about $0.02 per minute of call, whatever the prompt size. GPT-Realtime-2.1-mini lands in the same range. Gemini 3.1 Flash Live starts higher and grows with the prompt size and the call length because it re-bills the whole context on every turn. GPT-Live-1 is a flat $0.05 per minute plus backend tokens.

Why does the size of my system prompt matter so much for Gemini Live?

The Gemini Live API has no prompt caching: on every dialogue turn it bills the system prompt, the tool definitions and all the audio exchanged so far as input tokens. A 3,000-token prompt on a 2-minute call with 11 turns is billed 11 times, 33,000 tokens, before any audio. A pipeline sends the same prompt to its LLM on every turn too, but prompt caching bills the repeated prefix at 10% of the input price.

Why does call duration change the answer?

Speech-to-speech models keep the whole conversation in context. Each new turn re-reads the audio history, so the cost of a turn grows with the minutes already spoken: a long call costs more per minute than a short one. Gemini Live's context window compression caps that history with a sliding window. A pipeline is flat: STT is billed on the call duration, TTS on the words the agent says, and the transcript history is small text that caching makes almost free.

What is prompt caching and why does it matter for a voice agent?

A voice agent sends a new request to its language model on every dialogue turn, and each request carries the whole context again: the system prompt, the tool definitions, the transcript so far and the new utterance. On an 11-turn call a 3,000-token prompt is sent 11 times. Prompt caching means the provider recognises that the beginning of the request is identical to the previous one and bills that repeated part at a fraction of the price: 10% on OpenAI and Gemini text models, about 3% on the Realtime audio history. It works automatically when the stable content (prompt, tools) comes first and the changing content (date, caller data, new turn) last. The Gemini Live API has no cache, which is why its cost climbs with the prompt size. The simulator applies each provider's native caching as is: on for the pipeline LLMs, GPT-Realtime-mini and the GPT-Live-1 backend, never for Gemini Live.

How many tokens is my system prompt?

Paste it into the simulator: it counts about 4 characters per token, in your browser, without sending anything. Gemini and OpenAI tokenizers are within about 10% of each other on prose. Each tool definition adds roughly 180 tokens (name, description, parameter schema) that ride along on every turn.

If the price is the same, which architecture should I pick?

The speech-to-speech one, for the feel: full-duplex listening, natural prosody, interruptions and backchannels handled by the model, and no STT to LLM to TTS hand-offs. The trade-off is control. A speech-to-speech model behaves like a fixed temperature of 1 and offers little structured output, while a pipeline gives you temperature 0, deterministic tool calls and the freedom to swap any component or voice.

What is GPT-Live-1 and how is it billed?

GPT-Live-1 is OpenAI's full-duplex speech-to-speech model released in September 2026. The voice layer costs a flat $0.05 per minute billed per second, silences included, and it delegates reasoning and tool calls to a backend text model (GPT-5.6 Luna, Terra, Sol or GPT-6 Astra) billed per token with prompt caching. It ranks first on the Artificial Analysis Speech to Speech Index with Astra as backend, with about 1.3 seconds to first audio.

How is Gemini 3.8 Live priced in the simulator, and where does the 44% come from?

Gemini 3.8 Live is billed like Gemini 3.1 Flash Live: no prompt caching, the prompt, tools and audio history re-billed on every turn, at about 44% less per token. The ratio comes from Google DeepMind's model evaluation, which reports the Artificial Analysis cost per hour of input audio on the Big Bench Audio subset: $0.84 per hour for Gemini 3.8 Live against $1.50 for Gemini 3.1 Flash Live (Minimal) and $1.75 for Gemini 3.1 Flash Live (High). The simulator applies that 0.56 ratio to every Gemini 3.1 Flash Live token rate. The same document gives Gemini 3.8 Live a 76.0 score on the Artificial Analysis Speech to Speech Index (82.6 with extended thinking) against 71.5 for Gemini 3.1 Flash Live (High), and the top experience score on ServiceNow's EVA-Bench. In our listening it is slightly disappointing on voice and feel next to Gemini 3.1 Flash Live, whose voice family leads the Artificial Analysis Pronunciation Robustness benchmark, so the simulator rates its feel and voice below 3.1's. No separate latency figure is published, so the response-time chain of Gemini 3.1 Flash Live is reused. Source: Gemini 3.8 Audio (Live, Live Extended Thinking) model evaluation, Google DeepMind, September 2026.

How is Grok Voice 2.0 priced, and why is it flat?

Grok Voice Think Fast 2.0, xAI's speech-to-speech model launched on 29 July 2026, bills $0.08 per minute of session audio (up from $0.05 for version 1.0), plus $0.004 per text item injected into the conversation; function call outputs are not billed. Nothing depends on the prompt size, the history or the number of turns, so the simulator prices it as a flat line: it is never the cheapest on a short call with a compact prompt, and it wins once minutes or tokens grow. Artificial Analysis measures 0.70 s average time to first audio, the only top-5 model under a second, and an 82.9 Speech to Speech Index for the High reasoning variant, which sets its control score. No EU data residency option is published. Source: Artificial Analysis, Grok Voice Think Fast 2.0 on the Speech to Speech Index, July 2026.

How does the agent language change the recommendation?

Each model and each pipeline component carries a voice score, 1 to 5, for English, French, Spanish and Dutch: reputation from blind listening benchmarks, vendor language lists and what we hear in production, not a measurement, and editable in the code. Gemini 3.1 Flash Live has the best voice and feel of the field: its voice family, Gemini 3.1 Flash TTS, leads the Artificial Analysis Pronunciation Robustness benchmark at 88.1% (ahead of SpaceXAI TTS at 87.6% and ElevenLabs Eleven v3 at 85.6%), the test of whether names, amounts, dates, codes and emails are said correctly. It scores 5 in every language and the top feel, and every other option, Gemini 3.8 Live included, sits half a point below its reputation on both. In French, Gradium, ElevenLabs and Cartesia are the strong pipeline voices; OpenAI's voices keep an accent; Grok has no French reputation yet. In Spanish, Gradium and ElevenLabs lead the pipeline, Gemini 3.8 Live and Grok's regional voices follow. In Dutch and Flemish, ElevenLabs has native Flemish voices, while Cartesia and Deepgram support it with an accent and Gradium does not ship it. The score moves the feel (a 5 earns half a dot, a 3 loses a full dot, a 2 loses two: a bad voice in the caller's language is a bad experience whatever the model does elsewhere), an option scoring under 3 is not proposed as best value while a better-sounding one fits the budget, the best-value pick follows, and the pipeline selects show the score of every STT and TTS in that language. Belgium is covered by French and Dutch. Source: Artificial Analysis Pronunciation Robustness benchmark, September 2026.

Can I keep the data in the EU with each architecture, and what does it cost?

Yes, with all of them except Grok Voice 2.0, which has no published EU option, at different prices. Gemini 3.1 Flash Live and the Gemini text models run in an EU region through Vertex AI at regional pricing, 10% above the global endpoint. OpenAI offers EU data residency projects: GPT-Realtime-mini and the text models keep their list price, GPT-Live-1's voice layer costs 10% more. A pipeline is the easiest to keep in the EU because residency is decided component by component: STT and TTS on their EU endpoints, the LLM on an EU project or an EU region, and any vendor that cannot comply is swapped without touching the rest. Tick EU data residency in the simulator to apply these prices and adapt the recommendation.

Which prices does the simulator use and how fresh are they?

Public list prices checked on 15 September 2026 from the Gemini API and Vertex AI pricing pages, the OpenAI API pricing page, xAI's pricing for Grok Voice, the Deepgram, Cartesia and Gradium pricing pages (Gradium's per-character credits converted at about 900 characters per spoken minute), and artificialanalysis.ai for quality and latency. Volume discounts, committed-use contracts and enterprise rates are not modelled. Actual bills differ with the conversation shape: adjust the turns per minute, speech share and tool calls in the advanced settings.

Still have questions?

Contact us

Contact
Next step

See it on your own calls.

Book a call with a voice AI engineer: we start from your real call flows and show you the agent, the architecture and the numbers that fit them.

1:1with a voice AI engineer, no slides
2architectures we ship in production: speech-to-speech and pipelines
1rollout plan to leave with: volume, pricing, telephony