Staging environment
Free Lesson

What happens when you make an LLM call?

45 min
Feb 4, 2026 11:00 AM
Virtual (Zoom)

What you'll learn

Run through the Full LLM Call Stack

Understand the end‑to‑end journey from a user’s HTTP/API request down through inference engines, runtimes, and hardware

Let's dissect the Inference Process Step‑by‑Step

We'll break down how prompts are tokenized, transformed, and processed, including core mechanics of prefill v/s decode

Systems Level Thinking in Action

We'll do a case study diagnosing bottlenecks, choosing inference engines and reasoning for latency, throughput & scaling

Why this topic matters

We will break down behind the scenes of an LLM call to understand exactly what happens from prompt to output. This will be done using an architecture diagram as well as the code. If you have been curious about how HF or PyTorch or vLLM or Ray and CUDA everything orchestrates-you'd love it. We will look into the stack, the bottlenecks, and the system-level trade-offs everything ML engineers need.

You'll learn from

Abi Aryan

Abi Aryan

Lead Research Engineer @ Abide

Abide AI
O'Reilly Media
@Packtpub
See all products from goabiaryan