Staging environment
Free Lesson

Interleaving for better RAG evaluation

60 min
Nov 19, 2025 1:00 PM
Virtual (Zoom)

In this video

What you'll learn

Why summaries silently kill your retriever metrics

Why an LLM summary above search results reduces clicks, starves your A/B tests, and eliminates your ability to evaluate

Team-draft interleaving for RAG (not just for humans)

How team-draft interleaving blends results from two or more engines into a single ranked list. How to apply it to RAG

Using LLM citations as implicit relevance judgments

Treating summary citations to credit each retriever: a direct, comparable measure of the LLM's preference

Compare multiple retrievers and multiple LLMs at once

Three example search engines + Four LLM models to answer which retriever works best for my RAG stack?

Why this topic matters

Most teams still tune RAG retrievers like classic search—optimize NDCG, pick a config, move on. But summaries change everything: users read, don’t click, and metrics go blind. Interleaving for RAG blends multiple retrievers, lets the LLM summarize with citations, and shows which retriever actually powered the answer—giving real feedback in a post-click world.

You'll learn from

Max Irwin

Max Irwin

Founder/CEO, Max.io

Doug Turnbull

Doug Turnbull

Search consultant

Trey Grainger

Trey Grainger

Founder & CEO SearchKernel

Previously at

Max Irwin
Wolters Kluwer
OpenSource Connections
Reddit
Lucidworks
See all products from AI-Powered Search