Staging environment
Free Lesson

Modern Information Retrieval Evaluation In The RAG Era

Part of Featured Lightning Lessons

45 min
Jul 2, 2025 3:30 PM
Virtual (Zoom)

In this video

What you'll learn

Traditional Retrieval Evaluations Are Stale

Why is a model topping every leader-board not the one you should use? Find out the pitfalls of stale benchmarks.

Rigorous Academic Evaluations Still Power Real-World Evals

Understand why, despite benchmarks being flawed, rigorous academic evaluations are more relevant than ever.

Evaluation Research Is Evolving To Meet New Needs

Hear about the new methods researchers are creating to construct evaluations that match real-world needs.

Why this topic matters

Traditional IR benchmarks fall short for real-world RAG applications due to stale data, incomplete labels, and unrealistic queries. This talk introduces FreshStack, a new benchmark built from recent StackOverflow and GitHub content, designed to reflect real programming queries.

You'll learn from

Nandan Thakur

Nandan Thakur

RAG researcher @ UWaterloo. Creator of BEIR and MIRACL benchmarks

Hamel Husain

Hamel Husain

ML Engineer with 20 years of experience

Shreya Shankar

Shreya Shankar

ML Systems Researcher Making AI Evaluation Work in Practice

See all products from Hamel Husain & Shreya Shankar