Free Lesson
Scaling Judge-Time Compute for Robust Auto LLM Evaluation
Part of The AI Evaluation Handbook
60 min
Jul 17, 2025 2:00 PM
Virtual (Zoom)
In this video
What you'll learn
LLM Judge Reliability Issues
Identify key failure modes in automated LLM evaluation including bias and inconsistent outputs
Judge-Time Compute Scaling
Apply reasoning model training techniques to improve evaluation reliability and accuracy
RL-Powered Evaluation Systems
Implement reinforcement learning methods to build more robust automated assessment tools
Why this topic matters
Reliable LLM evaluation is crucial for AI safety and quality in production systems. Poor judges waste resources and create unsafe deployments. Mastering robust evaluation techniques positions you as essential for companies deploying AI at scale - a rapidly growing field where quality assurance expertise is highly valued.
You'll learn from

Jason Liu
Consultant at the intersection of Information Retrieval and AI

Leonard Tang
Co-Founder & CEO @ Haize Labs
worked with
.png&w=1536&q=75)