A multi-turn prompt attack convinced one of the biggest providers’ voice agents to issue a $150 discount code.
Tarush Agarwal is the co-founder and CEO of Cekura.ai, which builds testing and verification infrastructure for voice agents. Before founding Cekura, he studied computer science at IIT Bombay and worked on low-latency quantitative trading systems in London and Chicago, where teams optimized performance at roughly seven to nine nanoseconds. Cekura entered Y Combinator after pivoting from a legal voice-agent product, later raised a $2.5 million seed round, and now works with more than 200 customers while running millions of simulations.
Voice agents can perform well in controlled tests and still fail during real conversations. Interruptions, background noise, mixed languages, transcription errors, emotional callers, and multi-turn manipulation can expose problems that never appear in a text-based evaluation.
Tarush’s most counterintuitive claim is that better models do not automatically produce better voice agents. Around half of Cekura’s customers still use GPT-4.1 because newer reasoning-heavy models can introduce delays that do not work in live calls. Production performance depends on the full system, including latency, transcription, turn detection, interruption handling, speech quality, instruction following, and the infrastructure connecting each component.
In Today’s Episode We Discuss:
00:01 Introducing Tarush Agarwal and Cekura.ai
00:27 From IIT Bombay to quantitative trading
02:45 Founder life versus low-latency engineering
04:26 Building voice agents for personal injury law firms
06:27 Pivoting during the first week of Y Combinator
07:11 Early growth, the $2.5 million seed round, and customer focus
09:49 The current state of voice AI
12:49 The metrics that determine voice-agent quality
15:17 Compliance, healthcare, and high-stakes conversations
18:02 How multi-turn prompt attacks exploit voice agents
19:17 The quiet problem with how companies run evals
22:33 Why testing voice agents through text is insufficient
24:06 Cascading systems versus speech-to-speech models
25:36 Building realistic simulation environments
27:12 What changed in voice AI over two years
29:29 Public benchmarks, latency gains, and accuracy limits
31:14 Cekura’s long-term vision beyond voice
32:16 Moving from founder-led sales to a dedicated GTM team
33:57 The product metric Tarush watches every day
35:37 Why voice AI could become larger than software
Cekura began after Tarush and his co-founders spent three hours after dinner manually calling their own legal voice agent. He explains why healthcare teams must simulate distressed patients, how multi-turn testing exposed the $150 discount exploit, and why his team sometimes shipped a bug fix before the customer reporting it had finished the call.
The episode returns to an old engineering principle: reliability begins when reality is allowed to break the system.
Pull Quotes
“Everyone talks about evals. I don’t think most people know how to do it correctly.”
“You need to build your own evals. You need to own your evals.”
Subscribe on Spotify: Spotify
Subscribe on Apple Podcasts: Apple Podcasts
Follow Tarush Agarwal on LinkedIn: LinkedIn
Follow Tarush Agarwal on X: X
Follow Brian on Linkedin: LinkedIn
Visit Our Website: Website
Subscribe to Our Newsletter: Newsletter
👂🎧 Watch, listen, and follow on your favorite platform: Link
🙏 Join the conversation on your favorite social network: Link

