Staging environment
Free Lesson

How Evals Made GitHub Copilot Happen

Part of The AI Evaluation Handbook

30 min
May 12, 2025 6:00 PM
Virtual (Zoom)

In this video

What you'll learn

Build more reliable LLM-as-judge systems

See how the Copilot team improved their automated evaluation by validating the judges

Learn the three-part eval taxonomy that drove success

Understand the differences between algorithmic, subjective, and verifiable evaluation approaches

Avoid the "ratchet effect" trap in A/B testing

Learn how GitHub's team hit local maximums with their metrics and the techniques they developed to overcome them.

Why this topic matters

GitHub Copilot stands as one of the first commercially successful generative AI applications. To make the product work, the copilot team had to invent evaluation methodologies with no existing blueprint. John shares candid insights from these pioneering efforts, helping you avoid repeating their mistakes and improve your own evaluation practices.

You'll learn from

John Berryman

John Berryman

ML Researcher, Software Engineer, and Author; Worked on GitHub Copilot

Shawn Simister

Shawn Simister

Member of Technical Staff, Factory

Hamel Husain

Hamel Husain

ML Engineer with 20 years of experience

Previously at

GitHub
Eventbrite
See all products from Hamel Husain & Shreya Shankar