Marsof Academy
Course 5 · Unit 1

GPUs and CUDA

Part of Advanced level (optional): specialisations

  • 4 lessons
  • ≈ 13 h of study
  • Level: advanced

Topics covered

  • GPU architecture
  • CUDA
  • Kernels
  • Threads
  • Blocks
  • Memory hierarchy
  • GPU profiling
  • Optimisation

Lessons in this unit

  1. GPU architecture 75 min
    SMs, 32-thread warps, SIMT divergence, the memory hierarchy and why bandwidth matters as much as FLOPs.
  2. The CUDA programming model 80 min
    Threads, blocks and grids; the indexing formula you'll use a thousand times; shared memory and synchronisation, simulated in Python.
  3. Roofline and kernel optimisation 80 min
    Arithmetic intensity, the roofline model, coalesced access, tiling in shared memory and kernel fusion, with simulations that count bytes.
  4. The memory hierarchy in depth — caches, loop order and tile shape 70 min
    Why the same computation can move 30 times more bytes depending on the loop order, how an LRU cache decides what stays, and how a tile's shape is chosen according to its arithmetic intensity and the shared memory available.

Prerequisites

Before this unit it helps to have done:

The full explanations, auto-graded exercises, exams and projects are inside the academy.

Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.

Shall we start?

Create your account and activate your subscription: you get the whole syllabus from day one.