Free Lesson
Redesign Your Product Metrics for AI Evals
Part of The AI Evaluation Handbook
30 min
Dec 17, 2025 4:00 PM
Virtual (Zoom)
In this video
What you'll learn
Understand How Product Data Science Traditionally Evaluates
Learn how teams use metrics, funnels, and experiments to measure feature impact in deterministic products.
Identify Why Traditional Methods Break for AI Features
See how probabilistic outputs, ambiguous correctness, and multi-step pipelines undermine traditional analytics.
Apply a New Simple Mental Model for Evaluating AI Features
Learn a practical model for evaluating AI across inputs, context, model behavior, output quality, and user value.
Why this topic matters
AI features break the assumptions behind traditional product metrics. Funnels, experiments, and success metrics often give misleading signals because AI outputs are probabilistic and multi-step. Understanding these emerging gaps in the evaluation of AI quality helps PMs, engineers, and data teams avoid misreads, design better evaluation workflows, and make more confident product decisions.
You'll learn from
.png&w=384&q=75)
Shane Butler
Principal Data Scientist, AI Evaluations at Ontra
Previously at Stripe, Nextdoor, PwC
