About 20% better than their old pipeline.
uMinds runs the sales funnels of more than ten creators. In a head-to-head A/B test live in production, their Spanish speech-to-speech agents beat the Cartesia + GPT-4.1 mini + Deepgram pipeline they ran before on the metric that pays: performance. Same campaign, one architecture swapped.
uMindsCreator sales funnels · 10+ creators
What changed, in figures.
~20%better performance than their orchestrated pipeline, A/B tested in production
- 50,000calls completed in under 3 hours
- 45concurrent lines, live in production
- 5 CPSon a single number, custom SIP
Figures reflect one customer's production configuration and results, not a contractual commitment.
Where they started.
uMinds runs the sales funnels of more than ten creators, at volume: campaigns of tens of thousands of calls, on their own SIP telephony.
They already had voice agents in production on an orchestrated pipeline: Deepgram to transcribe, GPT-4.1 mini to reason, Cartesia to speak. The question was whether native speech-to-speech would do better on the metric that pays.
What runs in production.
- Test
- A/B, live in production, on the same campaign
- Before
- Pipeline: Deepgram (STT) + GPT-4.1 mini (LLM) + Cartesia (TTS)
- After
- Speech-to-speech: Gemini 3.1 Flash Live
- Language
- Spanish
- Telephony
- Their own SIP: 5 calls per second on a single number
- Concurrency
- 45 lines
- Accounts
- 3 sub-accounts, run from one place
Same campaign, one architecture swapped.
Both architectures, live
The same campaign ran on the pipeline and on speech-to-speech, in production.
Measured on performance
Speech-to-speech came out about 20% ahead of the pipeline.
Scaled on the winner
And it held up on the very same campaign: 50,000 calls completed in under 3 hours, on 45 lines at 5 calls per second.
From the people who run it.

I thought the biggest voice providers would be impossible to compete with. But CallShift focused on what others were missing. With strong bets on models like Gemini, they’ve delivered the quality and pricing that made us move away from other providers. For me, CallShift is the future of voice infrastructure.
Same campaign, one architecture swapped
Before: Cartesia + GPT-4.1 mini + Deepgram, orchestrated. After: Gemini 3.1 Flash Live, speech-to-speech, on their own SIP telephony, across 3 sub-accounts run from one place.
Recognise yourself?
- You run voice agents on a pipeline (STT, LLM, TTS) today.
- You want to measure what speech-to-speech changes on your own campaign.
- You need volume: concurrent lines, calls per second, your own SIP.
Another need? Another case.
See it on your own calls.
Book a call with a voice AI engineer: we start from your real call flows and show you the agent, the architecture and the numbers that fit them.
