Marsof Academy
Course 5 · Unit 3

Distributed AI

Part of Advanced level (optional): specialisations

  • 7 lessons
  • ≈ 21 h of study
  • Level: advanced

Topics covered

  • Distributed systems
  • Data parallelism
  • Model parallelism
  • Distributed training
  • Distributed inference
  • Fault tolerance
  • Orchestration

Lessons in this unit

  1. Data parallelism and all-reduce 70 min
    Replicate the model across N GPUs, split the batch and average gradients with a ring all-reduce to train as if you had a single giant GPU.
  2. Model and optimizer sharding (ZeRO and FSDP) 70 min
    Stop storing N copies of everything, split optimizer states, gradients and parameters across GPUs, and compute how much memory each one needs.
  3. Fault tolerance, checkpoints and resumption 70 min
    In a large cluster something fails every day. Learn to save a checkpoint that lets you resume exactly where you left off and to choose how often to save it.
  4. Tensor parallelism in depth — the complete Transformer block 80 min
    From splitting one matmul to splitting a whole Transformer — per-head attention, the f and g operators of the backward pass, sequence parallelism and cross-entropy with a sharded vocabulary.
  5. Pipeline schedules — GPipe, 1F1B and interleaved 70 min
    The pipeline with a real forward and backward pass — why GPipe and 1F1B take the same time but don't use the same memory, how the bubble is computed and what interleaved schedules gain.
  6. Mixed precision, loss scaling and overlapping communication 75 min
    Why you train in fp16/bf16 with an fp32 master copy, how dynamic loss scaling rescues tiny gradients, and how DDP groups gradients into buckets to communicate while the backward pass keeps computing.
  7. Distributed inference — serving a model that doesn't fit on one GPU 70 min
    How an LLM is served split across several GPUs - tensor and pipeline parallelism in the decode, sharding the KV cache by heads and by layers, replicas and routing, and how to choose the configuration that gives the most throughput without breaking the promised latency.

Prerequisites

Before this unit it helps to have done:

The full explanations, auto-graded exercises, exams and projects are inside the academy.

Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.

Shall we start?

Create your account and activate your subscription: you get the whole syllabus from day one.