Free Lesson
Evaluating AI Agents before Users Break Them
Part of The AI Evaluation Handbook
60 min
Feb 10, 2026 1:00 PM
Virtual (Zoom)
In this video
What you'll learn
How to tell if an AI agent is actually helping users
Learn what “working” actually means from a product perspective.
What signals show problems early
Spot early signs of confusion, inconsistency, or user trust issues.
What to check before shipping or expanding an agent
Use a simple checklist to decide ship, fix, or stop.
How to talk about agent performance with your team
Ask the right questions without needing deep technical detail.
How to iterate toward production using Langfuse insights
Turn observed behavior into concrete improvements and safer deployments.
Why this topic matters
AI agents rarely fail loudly. They slowly degrade, behave inconsistently, or erode user trust. Many product teams lack clear ways to evaluate agent behavior before users are affected. This session covers practical evaluation frameworks PMs can use to understand agent behavior, spot risks early, and make confident product decisions without relying only on engineering intuition.
You'll learn from

Aki Wijesundara, PhD
AI Founder | Educator | Google AI Accelerator Alum

Marc Klingen
Co-Founder & CEO of Langfuse

Lotte Verheyden
Developer Relations at Langfuse
.png&w=1536&q=75)