Course 3 · Unit 6
LLM engineering
Part of Deep learning and Transformers
- 7 lessons
- ≈ 23 h of study
- Level: intermediate to advanced
Topics covered
- Tokenisers
- BPE
- Pretraining
- Datasets
- Training
- Evaluation
- Instruction tuning
- SFT
- RLHF
- DPO
- LoRA
- QLoRA
- Quantisation
- Inference
- KV cache
- Batching
- Context windows
- Hallucination
- Alignment
Lessons in this unit
- Pretraining and scaling laws 65 min
From the raw web to trillions of clean tokens, and how much compute it takes to turn them into a model. - Decoding — from logits to text 95 min
Greedy, temperature, top-k, top-p, repetition penalty and beam search, implemented and compared. - Inference and the KV cache 80 min
Why generating text is expensive, how the KV cache makes it cheaper and how much memory it costs you in return. - From base model to assistant — SFT and DPO 75 min
Chat templates, a loss masked to the answer, and preference optimisation with the DPO loss. - LoRA and QLoRA — fine-tuning without touching the weights 85 min
Low-rank adapters, how many parameters they really train, how they're merged and how their gradients are computed. - Quantization — fewer bits, almost the same model 80 min
Symmetric and asymmetric int8 and int4, per-tensor, per-channel and per-group scales, and the error each decision introduces. - Evaluating LLMs without fooling yourself 75 min
Perplexity and bits per byte, exact match and F1, unbiased pass@k, LLM-as-judge, contamination and hallucinations.
Prerequisites
Before this unit it helps to have done:
- Transformers (Course 3 · Unit 5)
The full explanations, auto-graded exercises, exams and projects are inside the academy.
Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.
Shall we start?
Create your account and activate your subscription: you get the whole syllabus from day one.