Free Lesson
Modern Information Retrieval Evaluation In The RAG Era
Part of Featured Lightning Lessons
45 min
Jul 2, 2025 3:30 PM
Virtual (Zoom)
In this video
What you'll learn
Traditional Retrieval Evaluations Are Stale
Why is a model topping every leader-board not the one you should use? Find out the pitfalls of stale benchmarks.
Rigorous Academic Evaluations Still Power Real-World Evals
Understand why, despite benchmarks being flawed, rigorous academic evaluations are more relevant than ever.
Evaluation Research Is Evolving To Meet New Needs
Hear about the new methods researchers are creating to construct evaluations that match real-world needs.
Why this topic matters
Traditional IR benchmarks fall short for real-world RAG applications due to stale data, incomplete labels, and unrealistic queries. This talk introduces FreshStack, a new benchmark built from recent StackOverflow and GitHub content, designed to reflect real programming queries.
You'll learn from

Nandan Thakur
RAG researcher @ UWaterloo. Creator of BEIR and MIRACL benchmarks

Hamel Husain
ML Engineer with 20 years of experience

Shreya Shankar
ML Systems Researcher Making AI Evaluation Work in Practice
.png&w=1536&q=75)