Staging environment
Free Lesson

Redesign Your Product Metrics for AI Evals

Part of The AI Evaluation Handbook

30 min
Dec 17, 2025 4:00 PM
Virtual (Zoom)

In this video

What you'll learn

Understand How Product Data Science Traditionally Evaluates

Learn how teams use metrics, funnels, and experiments to measure feature impact in deterministic products.

Identify Why Traditional Methods Break for AI Features

See how probabilistic outputs, ambiguous correctness, and multi-step pipelines undermine traditional analytics.

Apply a New Simple Mental Model for Evaluating AI Features

Learn a practical model for evaluating AI across inputs, context, model behavior, output quality, and user value.

Why this topic matters

AI features break the assumptions behind traditional product metrics. Funnels, experiments, and success metrics often give misleading signals because AI outputs are probabilistic and multi-step. Understanding these emerging gaps in the evaluation of AI quality helps PMs, engineers, and data teams avoid misreads, design better evaluation workflows, and make more confident product decisions.

You'll learn from

Shane Butler

Shane Butler

Principal Data Scientist, AI Evaluations at Ontra

Previously at Stripe, Nextdoor, PwC

Stripe
Nextdoor
Ontra
PwC UK
AppFolio
See all products from AI Analyst Lab