Marsof Academy
Course 3 · Unit 4

Language processing

Part of Deep learning and Transformers

  • 5 lessons
  • ≈ 13 h of study
  • Level: intermediate to advanced

Topics covered

  • Text processing
  • Tokenisation
  • Embeddings
  • Sequence modelling
  • Attention
  • Language modelling
  • Transformers

Lessons in this unit

  1. Text as data 65 min
    A language model doesn't see letters or words, it sees integers. Here you build the bridge between the two.
  2. BPE tokenisation from scratch 55 min
    The 1994 compression algorithm that tokenises the text of GPT, Llama and almost every current LLM.
  3. N-gram language models 55 min
    The first real language model, built just by counting. It's the baseline your Transformer will have to beat.
  4. Embeddings — tokens as vectors 65 min
    An id says nothing about its meaning; a vector does. That's how text gets into a neural network.
  5. Attention — every token looks at the others 70 min
    The mechanism that makes Transformers work, computed first by hand and then in pure Python.

Prerequisites

Before this unit it helps to have done:

The full explanations, auto-graded exercises, exams and projects are inside the academy.

Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.

Shall we start?

Create your account and activate your subscription: you get the whole syllabus from day one.