Staging environment
Free Lesson

Evals for Voice AI: Learnings from Google Evals Team

Part of The AI Evaluation Handbook

30 min
Feb 10, 2026 7:00 PM
Virtual (Zoom)

In this video

What you'll learn

Multimodal Evals Challenges

Gain practical insights for building, benchmarking, or researching evals in voice and other multimodal contexts

Design Voice AI Evals

Learn from practical experience of evals for real-world voice-based conversational AI product like NotebookLM

Evaluation Driven Development

Use evals for real-world conversational quality, setting evals metrics and guardrails in voice systems.

Why this topic matters

How to do AI evals for voice-based products? How do evaluation frameworks shape the development, reliability, and user experience of emerging voice AI systems. Using Google’s NotebookLM as a case study, we’ll dive into practical insights on designing evals that capture real-world conversational quality, reasoning depth, and responsiveness in multimodal and voice-based contexts.

You'll learn from

Ravin Kumar

Ravin Kumar

Sr. Researcher at Google Deepmind

Google DeepMind
Google
sweetgreen
SpaceX
See all products from AI Evals and Analytics