Course 5 · Unit 3
Distributed AI
Part of Advanced level (optional): specialisations
- 7 lessons
- ≈ 21 h of study
- Level: advanced
Topics covered
- Distributed systems
- Data parallelism
- Model parallelism
- Distributed training
- Distributed inference
- Fault tolerance
- Orchestration
Lessons in this unit
- Data parallelism and all-reduce 70 min
Replicate the model across N GPUs, split the batch and average gradients with a ring all-reduce to train as if you had a single giant GPU. - Model and optimizer sharding (ZeRO and FSDP) 70 min
Stop storing N copies of everything, split optimizer states, gradients and parameters across GPUs, and compute how much memory each one needs. - Fault tolerance, checkpoints and resumption 70 min
In a large cluster something fails every day. Learn to save a checkpoint that lets you resume exactly where you left off and to choose how often to save it. - Tensor parallelism in depth — the complete Transformer block 80 min
From splitting one matmul to splitting a whole Transformer — per-head attention, the f and g operators of the backward pass, sequence parallelism and cross-entropy with a sharded vocabulary. - Pipeline schedules — GPipe, 1F1B and interleaved 70 min
The pipeline with a real forward and backward pass — why GPipe and 1F1B take the same time but don't use the same memory, how the bubble is computed and what interleaved schedules gain. - Mixed precision, loss scaling and overlapping communication 75 min
Why you train in fp16/bf16 with an fp32 master copy, how dynamic loss scaling rescues tiny gradients, and how DDP groups gradients into buckets to communicate while the backward pass keeps computing. - Distributed inference — serving a model that doesn't fit on one GPU 70 min
How an LLM is served split across several GPUs - tensor and pipeline parallelism in the decode, sharding the KV cache by heads and by layers, replicas and routing, and how to choose the configuration that gives the most throughput without breaking the promised latency.
Prerequisites
Before this unit it helps to have done:
- LLM performance (Course 5 · Unit 2)
The full explanations, auto-graded exercises, exams and projects are inside the academy.
Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.
Shall we start?
Create your account and activate your subscription: you get the whole syllabus from day one.