Staging environment
Free Lesson

Part 3: Building Robust Evaluations for AI Agents

Part of From Automation to Multi-Agent Architectures

60 min
Feb 26, 2026 3:00 PM
Virtual (Zoom)

In this video

What you'll learn

Define What “Good” Looks Like for AI Agents

Identify success metrics for agent reasoning, actions, and outcomes—going far beyond simple accuracy scores.

Design Practical, Real-World Agent Evaluations

Build task-based, behavioral, and regression-style evals that reflect how agents actually operate in production.

Use Evaluations to Ship with Confidence

Apply eval results to debug failures, compare agent versions, and iterate safely without breaking existing behavior.

Why this topic matters

AI agents often seem to work, until real users, edge cases, and scale expose their failures. This matters because without proper evaluation, teams ship systems they can’t trust or improve. Robust evals turn agent systems from impressive demos into reliable, measurable, and production-ready products.

You'll learn from

Hamza Farooq

Hamza Farooq

Founder & Adjunct Professor | 15+ years | Google | Stanford | UCLA

Gabriela de Queiroz

Gabriela de Queiroz

Ex-Microsoft & IBM AI leader | AI Advisor for Startups

Previous attendees from:

Google
Stanford University
UCLA
Walmart
University of Minnesota
See all products from Hamza