Free Lesson
Interleaving for better RAG evaluation
60 min
Nov 19, 2025 1:00 PM
Virtual (Zoom)
In this video
What you'll learn
Why summaries silently kill your retriever metrics
Why an LLM summary above search results reduces clicks, starves your A/B tests, and eliminates your ability to evaluate
Team-draft interleaving for RAG (not just for humans)
How team-draft interleaving blends results from two or more engines into a single ranked list. How to apply it to RAG
Using LLM citations as implicit relevance judgments
Treating summary citations to credit each retriever: a direct, comparable measure of the LLM's preference
Compare multiple retrievers and multiple LLMs at once
Three example search engines + Four LLM models to answer which retriever works best for my RAG stack?
Why this topic matters
Most teams still tune RAG retrievers like classic search—optimize NDCG, pick a config, move on. But summaries change everything: users read, don’t click, and metrics go blind. Interleaving for RAG blends multiple retrievers, lets the LLM summarize with citations, and shows which retriever actually powered the answer—giving real feedback in a post-click world.
You'll learn from

Max Irwin
Founder/CEO, Max.io

Doug Turnbull
Search consultant

Trey Grainger
Founder & CEO SearchKernel
Previously at
%2520(1).png&w=1536&q=75)