Marsof Academy
Course 5 · Unit 7

The AI frontier

Part of Advanced level (optional): specialisations

  • 8 lessons
  • ≈ 26 h of study
  • Level: advanced

A living area that evolves with the field. Reasoning and inference-time compute, long context, Mixture of Experts, synthetic data and alignment with preferences. Model compression, agents and multimodality have their own units ("LLM performance", "AI agents" and "Multimodal AI").

Topics covered

  • Reasoning systems
  • Inference-time compute
  • Long context
  • Efficient architectures
  • Mixture of Experts
  • Synthetic data
  • Alignment
  • Reward models
  • RLHF
  • DPO
  • Alignment evaluation

Lessons in this unit

  1. Reasoning and inference-time compute 85 min
    Reasoning chains, self-consistency, best-of-n with verifiers and why spending more compute when answering helps… up to the point where it helps.
  2. Long context and efficient attention 80 min
    Why attention costs O(n²), how much the KV cache takes up, how RoPE is stretched with position interpolation and what sparse and linear attention promise (and don't).
  3. Mixture of Experts 85 min
    How a router sends each token to a few experts, why a load-balancing loss is needed and what happens when an expert fills up.
  4. Alignment and synthetic data 90 min
    RLHF, RLAIF and Constitutional AI at a conceptual level, the Bradley–Terry preference loss, and how to generate, filter and deduplicate synthetic data without falling into model collapse.
  5. The reward model — from human rankings to a number 70 min
    How RLHF's reward model is trained — from rankings to pairs, the Bradley–Terry gradient, why each prompt counts once, the accuracy ceiling imposed by human disagreement and the biases the model learns without anyone asking it to.
  6. RLHF with PPO — reward with KL, advantages and clipping 75 min
    The numerical pieces of RLHF's RL stage — the per-token reward with a KL penalty, the KL estimators, advantages with GAE, PPO's clipped objective and its critic-free variant, GRPO.
  7. DPO — the full derivation and why it works 70 min
    From the optimal policy of the KL objective to the Direct Preference Optimization loss, step by step, with its gradient, what the "implicit reward" means, and a numerical check that optimizing DPO recovers exactly the optimal policy.
  8. Evaluating alignment — reward hacking, over-optimization and biased judges 70 min
    How to know whether an aligned model is really better — Goodhart's law in numbers, the over-optimization curve of best-of-n, the best-of-n KL, and how to measure and correct a judge's position and length biases.

Prerequisites

Before this unit it helps to have done:

The full explanations, auto-graded exercises, exams and projects are inside the academy.

Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.

Shall we start?

Create your account and activate your subscription: you get the whole syllabus from day one.