Free Lesson
Evaluating AI Agents
Part of Building Production AI Systems
45 min
Oct 31, 2025 12:00 PM
Virtual (Zoom)
In this video
What you'll learn
Measuring AI in Uncertain Domains
Evaluating AI in markets is tough since no single "right" answer exists. Success requires creative metrics.
Going From Accuracy to Market Relevance
Beyond F1 scores, using market-centric metrics like correlation helps align AI behavior with financial realities.
How to Use Meta-Evaluation and Practical Tools
Defining "good enough" using tools like MLflow and Langchain/ Langgraph matters; metrics must stand up to scrutiny.
Why this topic matters
An AI agent was created to surface the most relevant companies for any given theme, but evaluating its performance turned out to be tricky. In this talk, the challenge of designing a solid evaluation framework is explored. Early benchmarks using F1 scores and MSE often punished good picks, so the approach was refined, leading to stronger evaluations and more confidence in the results.
You'll learn from

Amir Feizpour
Founder @ Aggregate Intellect

Samuel Dion-Girardeau
CTO at Tilt
