Course 5 · Unit 1
GPUs and CUDA
Part of Advanced level (optional): specialisations
- 4 lessons
- ≈ 13 h of study
- Level: advanced
Topics covered
- GPU architecture
- CUDA
- Kernels
- Threads
- Blocks
- Memory hierarchy
- GPU profiling
- Optimisation
Lessons in this unit
- GPU architecture 75 min
SMs, 32-thread warps, SIMT divergence, the memory hierarchy and why bandwidth matters as much as FLOPs. - The CUDA programming model 80 min
Threads, blocks and grids; the indexing formula you'll use a thousand times; shared memory and synchronisation, simulated in Python. - Roofline and kernel optimisation 80 min
Arithmetic intensity, the roofline model, coalesced access, tiling in shared memory and kernel fusion, with simulations that count bytes. - The memory hierarchy in depth — caches, loop order and tile shape 70 min
Why the same computation can move 30 times more bytes depending on the loop order, how an LRU cache decides what stays, and how a tile's shape is chosen according to its arithmetic intensity and the shared memory available.
Prerequisites
Before this unit it helps to have done:
- LLM engineering (Course 3 · Unit 6)
The full explanations, auto-graded exercises, exams and projects are inside the academy.
Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.
Shall we start?
Create your account and activate your subscription: you get the whole syllabus from day one.