Case study · I build voice AI agents

About 20% better than their old pipeline.

uMinds runs the sales funnels of more than ten creators. In a head-to-head A/B test live in production, their Spanish speech-to-speech agents beat the Cartesia + GPT-4.1 mini + Deepgram pipeline they ran before on the metric that pays: performance. Same campaign, one architecture swapped.

uMindsCreator sales funnels · 10+ creators

  • Gemini 3.1 Flash Live
  • Speech-to-speech
  • Custom SIP at scale
  • Spanish
01 Results

What changed, in figures.

~20%better performance than their orchestrated pipeline, A/B tested in production

Performance on the same campaign (pipeline = 100)
Pipeline100
Speech-to-speech~120
  • 50,000calls completed in under 3 hours
  • 45concurrent lines, live in production
  • 5 CPSon a single number, custom SIP

Figures reflect one customer's production configuration and results, not a contractual commitment.

02 The need

Where they started.

uMinds runs the sales funnels of more than ten creators, at volume: campaigns of tens of thousands of calls, on their own SIP telephony.

They already had voice agents in production on an orchestrated pipeline: Deepgram to transcribe, GPT-4.1 mini to reason, Cartesia to speak. The question was whether native speech-to-speech would do better on the metric that pays.

03 What we deployed

What runs in production.

Test
A/B, live in production, on the same campaign
Before
Pipeline: Deepgram (STT) + GPT-4.1 mini (LLM) + Cartesia (TTS)
After
Speech-to-speech: Gemini 3.1 Flash Live
Language
Spanish
Telephony
Their own SIP: 5 calls per second on a single number
Concurrency
45 lines
Accounts
3 sub-accounts, run from one place
04 How it works

Same campaign, one architecture swapped.

  1. Both architectures, live

    The same campaign ran on the pipeline and on speech-to-speech, in production.

  2. Measured on performance

    Speech-to-speech came out about 20% ahead of the pipeline.

  3. Scaled on the winner

    And it held up on the very same campaign: 50,000 calls completed in under 3 hours, on 45 lines at 5 calls per second.

05 In their words

From the people who run it.

Same campaign, one architecture swapped

Before: Cartesia + GPT-4.1 mini + Deepgram, orchestrated. After: Gemini 3.1 Flash Live, speech-to-speech, on their own SIP telephony, across 3 sub-accounts run from one place.

06 This case fits you if

Recognise yourself?

  • You run voice agents on a pipeline (STT, LLM, TTS) today.
  • You want to measure what speech-to-speech changes on your own campaign.
  • You need volume: concurrent lines, calls per second, your own SIP.
Next step

See it on your own calls.

Book a call with a voice AI engineer: we start from your real call flows and show you the agent, the architecture and the numbers that fit them.

1:1with a voice AI engineer, no slides
2architectures we ship in production: speech-to-speech and pipelines
1rollout plan to leave with: volume, pricing, telephony