Course 5 · Unit 7
The AI frontier
Part of Advanced level (optional): specialisations
- 8 lessons
- ≈ 26 h of study
- Level: advanced
A living area that evolves with the field. Reasoning and inference-time compute, long context, Mixture of Experts, synthetic data and alignment with preferences. Model compression, agents and multimodality have their own units ("LLM performance", "AI agents" and "Multimodal AI").
Topics covered
- Reasoning systems
- Inference-time compute
- Long context
- Efficient architectures
- Mixture of Experts
- Synthetic data
- Alignment
- Reward models
- RLHF
- DPO
- Alignment evaluation
Lessons in this unit
- Reasoning and inference-time compute 85 min
Reasoning chains, self-consistency, best-of-n with verifiers and why spending more compute when answering helps… up to the point where it helps. - Long context and efficient attention 80 min
Why attention costs O(n²), how much the KV cache takes up, how RoPE is stretched with position interpolation and what sparse and linear attention promise (and don't). - Mixture of Experts 85 min
How a router sends each token to a few experts, why a load-balancing loss is needed and what happens when an expert fills up. - Alignment and synthetic data 90 min
RLHF, RLAIF and Constitutional AI at a conceptual level, the Bradley–Terry preference loss, and how to generate, filter and deduplicate synthetic data without falling into model collapse. - The reward model — from human rankings to a number 70 min
How RLHF's reward model is trained — from rankings to pairs, the Bradley–Terry gradient, why each prompt counts once, the accuracy ceiling imposed by human disagreement and the biases the model learns without anyone asking it to. - RLHF with PPO — reward with KL, advantages and clipping 75 min
The numerical pieces of RLHF's RL stage — the per-token reward with a KL penalty, the KL estimators, advantages with GAE, PPO's clipped objective and its critic-free variant, GRPO. - DPO — the full derivation and why it works 70 min
From the optimal policy of the KL objective to the Direct Preference Optimization loss, step by step, with its gradient, what the "implicit reward" means, and a numerical check that optimizing DPO recovers exactly the optimal policy. - Evaluating alignment — reward hacking, over-optimization and biased judges 70 min
How to know whether an aligned model is really better — Goodhart's law in numbers, the over-optimization curve of best-of-n, the best-of-n KL, and how to measure and correct a judge's position and length biases.
Prerequisites
Before this unit it helps to have done:
- Multimodal AI (Course 5 · Unit 6)
The full explanations, auto-graded exercises, exams and projects are inside the academy.
Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.
Shall we start?
Create your account and activate your subscription: you get the whole syllabus from day one.