Staging environment
Free Lesson

Beyond chunks - how to actually data model for RAG

60 min
Nov 21, 2025 11:00 AM
Virtual (Zoom)

In this video

What you'll learn

Why preferring large documents works best for RAG

Trade-offs between large documents with chunk arrays and chunk documents with denormalized metadata

The ideal data model for RAG retrieval

Combining chunk and document-level features into a single searchable entity

How to surface the most relevant chunks

Using rank profiles to compute chunk relevance within each document and return top N chunks for each

Why this topic matters

In enterprise and web search, many questions are answered by separate bits of documents, yet semantics and properties of the containing entity are also important. While there's no silver bullet - because data modeling is hard - we'll explore techniques to navigate the large vs small trade-off.

You'll learn from

Radu Gheorghe

Radu Gheorghe

Software Engineer, Vespa.ai

Doug Turnbull

Doug Turnbull

Agentic Search Consultant

Previously at

Vespa.Ai
Reddit
Shopify.com
See all products from Doug