The next generation of voice AI with Google DeepMind and Sierra AI
Share

Post Content

 

 [[{“value”:”Valeria Wu (Google DeepMind) and Soham Ray (Sierra AI) explore the latest advances in native audio models and the challenges of building better real-time voice experiences. They discuss conversational latency, multilingual code-switching, agentic voice tasks, and how new benchmarks can better measure the quality of real-world dialogue.

Watch along and learn:
*How native speech-to-speech models preserve the tone and prosody traditional pipelines lose.
*The engineering tradeoffs behind sub-second response times and seamless, human-like interruption handling.
*How models smoothly navigate mixed-language phrasing and dialect shifts without dropping context.
*How to orchestrate API calls and stateful tasks while maintaining an uninterrupted vocal flow.
*Moving beyond Word Error Rate toward dynamic metrics that measure conversational flow and turn-taking.

Subscribe to Google for Developers → https://goo.gle/developers

Products Mentioned: Google DeepMind
Speakers: Valeria Wu, Soham Ray

Chapters:
00:00 – Introduction & The State of Real-Time Voice AI
0:50 – Understanding Sierra & TAU
01:48 – Latency
03:52 – Fluid Multilinguality
04:22 – What’s Next?”}]] Read More Google for Developers 

By ali

Leave a Reply