Marsof Academy
Course 3 · Unit 5

Transformers

Part of Deep learning and Transformers

  • 5 lessons
  • ≈ 19 h of study
  • Level: intermediate to advanced

Building a small, working Transformer piece by piece.

Topics covered

  • Token embeddings
  • Positional encoding
  • Attention
  • Scaled dot-product attention
  • Multi-head attention
  • MLP
  • Residual connections
  • LayerNorm
  • Transformer block
  • Causal mask
  • Language head

Lessons in this unit

  1. Embeddings and positional encoding — from ids to vectors that know where they are 95 min
    Turn token ids into vectors and give them a notion of order with sinusoidal encoding, learned positions and RoPE.
  2. Multi-head attention — many views at once, without a single loop 100 min
    Go from the pure-Python attention head to vectorised causal multi-head attention, with the shapes (B, T, C) → (B, h, T, d) under control.
  3. The Transformer block — residual, LayerNorm and MLP 85 min
    Assemble attention, an MLP with GELU, LayerNorm and residual connections into the pre-norm block that is repeated N times throughout every GPT.
  4. The full decoder — from ids to logits, loss and generated text 100 min
    Stack N blocks, add the final LayerNorm and the language head, compute the cross-entropy with shifted targets, count the parameters and generate text token by token.
  5. Training a small language model — data, gradients and curves that talk 80 min
    Build batches of sequences from a token stream, train a tiny language model with hand-written gradients in NumPy and learn to read what its loss curves are telling you.

Prerequisites

Before this unit it helps to have done:

The full explanations, auto-graded exercises, exams and projects are inside the academy.

Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.

Shall we start?

Create your account and activate your subscription: you get the whole syllabus from day one.