Course 3 · Unit 5
Transformers
Part of Deep learning and Transformers
- 5 lessons
- ≈ 19 h of study
- Level: intermediate to advanced
Building a small, working Transformer piece by piece.
Topics covered
- Token embeddings
- Positional encoding
- Attention
- Scaled dot-product attention
- Multi-head attention
- MLP
- Residual connections
- LayerNorm
- Transformer block
- Causal mask
- Language head
Lessons in this unit
- Embeddings and positional encoding — from ids to vectors that know where they are 95 min
Turn token ids into vectors and give them a notion of order with sinusoidal encoding, learned positions and RoPE. - Multi-head attention — many views at once, without a single loop 100 min
Go from the pure-Python attention head to vectorised causal multi-head attention, with the shapes (B, T, C) → (B, h, T, d) under control. - The Transformer block — residual, LayerNorm and MLP 85 min
Assemble attention, an MLP with GELU, LayerNorm and residual connections into the pre-norm block that is repeated N times throughout every GPT. - The full decoder — from ids to logits, loss and generated text 100 min
Stack N blocks, add the final LayerNorm and the language head, compute the cross-entropy with shifted targets, count the parameters and generate text token by token. - Training a small language model — data, gradients and curves that talk 80 min
Build batches of sequences from a token stream, train a tiny language model with hand-written gradients in NumPy and learn to read what its loss curves are telling you.
Prerequisites
Before this unit it helps to have done:
- Language processing (Course 3 · Unit 4)
The full explanations, auto-graded exercises, exams and projects are inside the academy.
Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.
Shall we start?
Create your account and activate your subscription: you get the whole syllabus from day one.