Staging environment
Free Lesson

Setting up your first AI eval with a LLM-as-judge

Part of The AI Evaluation Handbook

45 min
Feb 27, 2026 9:00 AM
Virtual (Zoom)

In this video

What you'll learn

Most common mistakes to avoid when building an LLM-as-judge

Understand why most teams' LLM judges don't work and the specific mistakes that make them unreliable.

How to write your judge instructions

Learn how to identify what to check through LLM-as-judge, define specific rules, and build the judge prompt

How to evaluate your LLM-as-judge

Know how to evaluate the evaluator by comparing judge scores to human labels and decide if you can trust the results.

Why this topic matters

Most teams building an LLM-as-a-judge make the same mistakes. They skip error analysis and ask the judge to look for hypothetical errors instead of real ones. They pack multiple criteria into one judge, creating high cognitive load that produces unreliable scores. They never validate the judge against human labels, so they don't know if it's accurate or noise.

You'll learn from

Madalina Turlea

Madalina Turlea

Co-founder @Lovelaice, 10+ years in Product

Catalina Turlea

Catalina Turlea

Founder @Lovelaice

See all products from Madalina