Free Lesson
How Evals Made GitHub Copilot Happen
Part of The AI Evaluation Handbook
30 min
May 12, 2025 6:00 PM
Virtual (Zoom)
In this video
What you'll learn
Build more reliable LLM-as-judge systems
See how the Copilot team improved their automated evaluation by validating the judges
Learn the three-part eval taxonomy that drove success
Understand the differences between algorithmic, subjective, and verifiable evaluation approaches
Avoid the "ratchet effect" trap in A/B testing
Learn how GitHub's team hit local maximums with their metrics and the techniques they developed to overcome them.
Why this topic matters
GitHub Copilot stands as one of the first commercially successful generative AI applications. To make the product work, the copilot team had to invent evaluation methodologies with no existing blueprint. John shares candid insights from these pioneering efforts, helping you avoid repeating their mistakes and improve your own evaluation practices.
You'll learn from

John Berryman
ML Researcher, Software Engineer, and Author; Worked on GitHub Copilot

Shawn Simister
Member of Technical Staff, Factory

Hamel Husain
ML Engineer with 20 years of experience
Previously at
.png&w=1536&q=75)