Staging environment
Free Lesson

Evaluating AI Agents

Part of Building Production AI Systems

45 min
Oct 31, 2025 12:00 PM
Virtual (Zoom)

In this video

What you'll learn

Measuring AI in Uncertain Domains

Evaluating AI in markets is tough since no single "right" answer exists. Success requires creative metrics.

Going From Accuracy to Market Relevance

Beyond F1 scores, using market-centric metrics like correlation helps align AI behavior with financial realities.

How to Use Meta-Evaluation and Practical Tools

Defining "good enough" using tools like MLflow and Langchain/ Langgraph matters; metrics must stand up to scrutiny.

Why this topic matters

An AI agent was created to surface the most relevant companies for any given theme, but evaluating its performance turned out to be tricky. In this talk, the challenge of designing a solid evaluation framework is explored. Early benchmarks using F1 scores and MSE often punished good picks, so the approach was refined, leading to stronger evaluations and more confidence in the results.

You'll learn from

Amir Feizpour

Amir Feizpour

Founder @ Aggregate Intellect

Samuel Dion-Girardeau

Samuel Dion-Girardeau

CTO at Tilt

See all products from aggregate intellect