Free Lesson
Beyond chunks - how to actually data model for RAG
60 min
Nov 21, 2025 11:00 AM
Virtual (Zoom)
In this video
What you'll learn
Why preferring large documents works best for RAG
Trade-offs between large documents with chunk arrays and chunk documents with denormalized metadata
The ideal data model for RAG retrieval
Combining chunk and document-level features into a single searchable entity
How to surface the most relevant chunks
Using rank profiles to compute chunk relevance within each document and return top N chunks for each
Why this topic matters
In enterprise and web search, many questions are answered by separate bits of documents, yet semantics and properties of the containing entity are also important. While there's no silver bullet - because data modeling is hard - we'll explore techniques to navigate the large vs small trade-off.
You'll learn from

Radu Gheorghe
Software Engineer, Vespa.ai

Doug Turnbull
Agentic Search Consultant
Previously at
.png&w=1536&q=75)