Free Lesson
Setting up your first AI eval with a LLM-as-judge
Part of The AI Evaluation Handbook
45 min
Feb 27, 2026 9:00 AM
Virtual (Zoom)
In this video
What you'll learn
Most common mistakes to avoid when building an LLM-as-judge
Understand why most teams' LLM judges don't work and the specific mistakes that make them unreliable.
How to write your judge instructions
Learn how to identify what to check through LLM-as-judge, define specific rules, and build the judge prompt
How to evaluate your LLM-as-judge
Know how to evaluate the evaluator by comparing judge scores to human labels and decide if you can trust the results.
Why this topic matters
Most teams building an LLM-as-a-judge make the same mistakes. They skip error analysis and ask the judge to look for hypothetical errors instead of real ones. They pack multiple criteria into one judge, creating high cognitive load that produces unreliable scores. They never validate the judge against human labels, so they don't know if it's accurate or noise.
You'll learn from

Madalina Turlea
Co-founder @Lovelaice, 10+ years in Product
Catalina Turlea
Founder @Lovelaice
