Staging environment
Free Lesson

Scaling Judge-Time Compute for Robust Auto LLM Evaluation

Part of The AI Evaluation Handbook

60 min
Jul 17, 2025 2:00 PM
Virtual (Zoom)

In this video

What you'll learn

LLM Judge Reliability Issues

Identify key failure modes in automated LLM evaluation including bias and inconsistent outputs

Judge-Time Compute Scaling

Apply reasoning model training techniques to improve evaluation reliability and accuracy

RL-Powered Evaluation Systems

Implement reinforcement learning methods to build more robust automated assessment tools

Why this topic matters

Reliable LLM evaluation is crucial for AI safety and quality in production systems. Poor judges waste resources and create unsafe deployments. Mastering robust evaluation techniques positions you as essential for companies deploying AI at scale - a rapidly growing field where quality assurance expertise is highly valued.

You'll learn from

Jason Liu

Jason Liu

Consultant at the intersection of Information Retrieval and AI

Leonard Tang

Leonard Tang

Co-Founder & CEO @ Haize Labs

worked with

Haize Labs
Stitch Fix
Meta
University of Waterloo
New York University
See all products from Applied LLMs