Free Lesson
Part 3: Building Robust Evaluations for AI Agents
60 min
Feb 26, 2026 3:00 PM
Virtual (Zoom)
In this video
What you'll learn
Define What “Good” Looks Like for AI Agents
Identify success metrics for agent reasoning, actions, and outcomes—going far beyond simple accuracy scores.
Design Practical, Real-World Agent Evaluations
Build task-based, behavioral, and regression-style evals that reflect how agents actually operate in production.
Use Evaluations to Ship with Confidence
Apply eval results to debug failures, compare agent versions, and iterate safely without breaking existing behavior.
Why this topic matters
AI agents often seem to work, until real users, edge cases, and scale expose their failures. This matters because without proper evaluation, teams ship systems they can’t trust or improve. Robust evals turn agent systems from impressive demos into reliable, measurable, and production-ready products.
You'll learn from

Hamza Farooq
Founder & Adjunct Professor | 15+ years | Google | Stanford | UCLA

Gabriela de Queiroz
Ex-Microsoft & IBM AI leader | AI Advisor for Startups
Previous attendees from:
