Course 5 · Unit 2
LLM performance
Part of Advanced level (optional): specialisations
- 5 lessons
- ≈ 16 h of study
- Level: advanced
Topics covered
- Quantisation
- Pruning
- Distillation
- Batching
- KV cache
- Memory optimisation
- Inference optimisation
- Tensor parallelism
- Pipeline parallelism
Lessons in this unit
- Training and inference memory 65 min
Where every byte of the GPU goes — weights, gradients, Adam states, activations and the KV cache — and how to calculate it before you hit "run". - Quantization, pruning and distillation 85 min
Three ways to make a model smaller or cheaper — and how to measure what it costs in quality. - Batching, throughput and latency 75 min
How an LLM serves many users at once — prefill and decode, TTFT and TPOT, static, dynamic and continuous batching, queues, Little's law and percentiles. - Tensor and pipeline parallelism 85 min
How to split a model that doesn't fit on one GPU — data, ZeRO, column and row tensor parallelism, pipeline and its bubble — and how much communication costs. - Kernel fusion and FlashAttention — the online softmax 70 min
How to compute an exact softmax in a single pass over data that arrives in chunks, and how FlashAttention uses that trick to fuse attention into a kernel that never writes the N × N matrix to memory.
Prerequisites
Before this unit it helps to have done:
- GPUs and CUDA (Course 5 · Unit 1)
The full explanations, auto-graded exercises, exams and projects are inside the academy.
Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.
Shall we start?
Create your account and activate your subscription: you get the whole syllabus from day one.