Free Lesson
What happens when you make an LLM call?
45 min
Feb 4, 2026 11:00 AM
Virtual (Zoom)
What you'll learn
Run through the Full LLM Call Stack
Understand the end‑to‑end journey from a user’s HTTP/API request down through inference engines, runtimes, and hardware
Let's dissect the Inference Process Step‑by‑Step
We'll break down how prompts are tokenized, transformed, and processed, including core mechanics of prefill v/s decode
Systems Level Thinking in Action
We'll do a case study diagnosing bottlenecks, choosing inference engines and reasoning for latency, throughput & scaling
Why this topic matters
We will break down behind the scenes of an LLM call to understand exactly what happens from prompt to output. This will be done using an architecture diagram as well as the code. If you have been curious about how HF or PyTorch or vLLM or Ray and CUDA everything orchestrates-you'd love it.
We will look into the stack, the bottlenecks, and the system-level trade-offs everything ML engineers need.
You'll learn from

Abi Aryan
Lead Research Engineer @ Abide
