Marsof Academy
Course 5 · Unit 2

LLM performance

Part of Advanced level (optional): specialisations

  • 5 lessons
  • ≈ 16 h of study
  • Level: advanced

Topics covered

  • Quantisation
  • Pruning
  • Distillation
  • Batching
  • KV cache
  • Memory optimisation
  • Inference optimisation
  • Tensor parallelism
  • Pipeline parallelism

Lessons in this unit

  1. Training and inference memory 65 min
    Where every byte of the GPU goes — weights, gradients, Adam states, activations and the KV cache — and how to calculate it before you hit "run".
  2. Quantization, pruning and distillation 85 min
    Three ways to make a model smaller or cheaper — and how to measure what it costs in quality.
  3. Batching, throughput and latency 75 min
    How an LLM serves many users at once — prefill and decode, TTFT and TPOT, static, dynamic and continuous batching, queues, Little's law and percentiles.
  4. Tensor and pipeline parallelism 85 min
    How to split a model that doesn't fit on one GPU — data, ZeRO, column and row tensor parallelism, pipeline and its bubble — and how much communication costs.
  5. Kernel fusion and FlashAttention — the online softmax 70 min
    How to compute an exact softmax in a single pass over data that arrives in chunks, and how FlashAttention uses that trick to fuse attention into a kernel that never writes the N × N matrix to memory.

Prerequisites

Before this unit it helps to have done:

The full explanations, auto-graded exercises, exams and projects are inside the academy.

Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.

Shall we start?

Create your account and activate your subscription: you get the whole syllabus from day one.