Staging environment
Free Lesson

Practical Evaluation Strategies for AI Agents

Part of The AI Evaluation Handbook

45 min
Jan 22, 2026 12:00 PM
Virtual (Zoom)

In this video

What you'll learn

Identify what “good” looks like for an AI agent

Define success metrics for agent reasoning, actions, and outcomes, beyond simple accuracy.

Design practical evals for agent workflows

Build task-based, behavioral, and regression-style evaluations that reflect real-world usage.

Use evals to iterate and improve agent systems

Apply evaluation results to debug failures, compare agent versions, and confidently ship changes.

Why this topic matters

AI agents often appear to work until they’re exposed to real users, edge cases, and scale. Without proper evaluation, teams ship systems they can’t trust or improve. This topic matters because evals turn agents from impressive demos into reliable products by making behavior measurable, debuggable, and safe to deploy in production.

You'll learn from

Hamza Farooq

Hamza Farooq

Founder | Ex-Google | Adjunct UCLA & UMN, SCU | Venture Partner

Gabriela de Queiroz

Gabriela de Queiroz

Ex-Microsoft & IBM AI leader | AI Advisor for Startups

Previous attendees from:

Google
Apple
Airbnb
Amazon
Microsoft
See all products from Hamza