Free Lesson
Understanding Embedding Performance through Generative Evals
Part of The AI Evaluation Handbook
60 min
May 21, 2025 1:00 PM
Virtual (Zoom)
In this video
What you'll learn
AI Evaluation Challenges
Discover why AI systems need specialized benchmarking beyond traditional testing methods.
Benchmark Limitations
Identify the shortcomings of public benchmarks, including clean datasets and potential training data contamination.
Representativeness in Testing
Apply techniques to generate benchmark tests that accurately reflect real-world user queries and production conditions.
Why this topic matters
Effective AI evaluation is critical as systems move from labs to production. Understanding generative benchmarking helps you build AI that performs well on real-world tasks, not just academic tests. This knowledge bridges the gap between theoretical capabilities and practical performance, giving you a competitive edge in developing AI solutions that deliver genuine value to users.
You'll learn from

Jason Liu
Consultant at the intersection of Information Retrieval and AI

Kelly Hong
Researcher at Chroma
Previously at
.png&w=1536&q=75)