Staging environment
Free Lesson

Evals Are Almost Never All You Need

30 min
Jun 26, 2025 12:00 PM
Virtual (Zoom)

In this video

What you'll learn

Common problems with benchmarks and evals

Understand how task contamination, conflicts of interest, and lack of scientific measurement can compromise evals.

What to do instead of benchmarks and evals

Learn about embedding-based approaches, red-teaming, and field testing for evaluating generative AI systems.

When to use benchmarks and evals

Developers need benchmarks. They're an important tool for development. They're the wrong tool for real-world assessment.

Why this topic matters

Generative AI is important. It has tangible real world impacts. Let's measure those real world impacts directly instead of guessing at them with ill-suited benchmarks. There are many technical and socio-technical measurement approaches for generative AI that are better for real world measurement than benchmarks. This lightening-session will provide an introduction to those approaches.

You'll learn from

Patrick Hall

Patrick Hall

Principal Scientist, HallResearch.ai

See all products from Hall Research AI