

Free Lesson
Understanding Innovations leading up to DeepSeek R1
45 min
Feb 7, 2025 12:00 PM
What you'll learn
What data and training innovations underly R1.
What architecture choices enable R1's performance.
What these innovations mean for the community.
Why this topic matters
We will be looking into some of the technical choices that empowered the impressive performance of the base model of R1 like Multi-head Latent Attention, Load balancing for MoE models, Fill-in-the-middle Learning Objective, FP8 training, and Multi-token Prediction.
You'll learn from

Amir Feizpour
Founder @ Aggregate Intellect

Suhas Pai
CTO @ Hudson Labs