Marsof Academy
Course 3

Deep learning and Transformers

From the neuron to your own language model

  • ≈ 289 h
  • Level: intermediate to advanced
  • 6 units
  • 36 lessons

Neural networks with backpropagation written by hand, professional PyTorch, computer vision, natural language processing and the Transformer piece by piece, up to training and serving a small LLM.

What you'll be able to do

  • A neural network from scratch, with no deep learning libraries
  • A CNN that recognises digits and an NLP system on real text
  • A complete Transformer and TuLLM, a language model you train yourself

Prerequisites

It helps to have done first (or to master what it teaches):

Units of the course

Unit 1 · Deep Learning

Neural networks from the neuron to attention, with backpropagation implemented by hand.

9 lessons · ≈ 24 h of study

  1. The artificial neuron
  2. Multilayer networks (MLP)
  3. Computational graphs and autograd
  4. Backpropagation
  5. Training a network
  6. Activations, initialisation and normalisation
  7. Dropout and regularisation in networks
  8. Recurrent networks and the gradient through time
  9. LSTM and GRU — memory with gates

See the unit: Deep Learning

Unit 3 · Computer vision

4 lessons · ≈ 13 h of study

  1. Convolution — a magnifying glass that sweeps across the image
  2. Convolutional networks — stacking ever wider views
  3. Detection and segmentation — what's there and where it is
  4. Vision Transformers — an image is worth 16 × 16 words

See the unit: Computer vision

Unit 5 · Transformers

Building a small, working Transformer piece by piece.

5 lessons · ≈ 19 h of study

  1. Embeddings and positional encoding — from ids to vectors that know where they are
  2. Multi-head attention — many views at once, without a single loop
  3. The Transformer block — residual, LayerNorm and MLP
  4. The full decoder — from ids to logits, loss and generated text
  5. Training a small language model — data, gradients and curves that talk

See the unit: Transformers

Unit 6 · LLM engineering

7 lessons · ≈ 23 h of study

  1. Pretraining and scaling laws
  2. Decoding — from logits to text
  3. Inference and the KV cache
  4. From base model to assistant — SFT and DPO
  5. LoRA and QLoRA — fine-tuning without touching the weights
  6. Quantization — fewer bits, almost the same model
  7. Evaluating LLMs without fooling yourself

See the unit: LLM engineering

Course projects

  • A neural network from scratch ≈ 35 h
    A multilayer neural network with hand-written backpropagation, trained on real data.
  • A Transformer from scratch ≈ 50 h
    Implement and train a complete Transformer, verifying each component against a reference.
  • TuLLM — a language model from scratch ≈ 80 h
    Your long-haul project. Build a complete language model, from raw text to optimized inference, as you progress through the academy.

The full explanations, auto-graded exercises, exams and projects are inside the academy. Estimated hours for someone starting from scratch: lessons with practice and review, challenges, exams and the recommended projects. It's the same estimate you'll see inside the academy.

Along the way you earn yang, the academy's currency: 10 for every lesson, 50 for every unit exam you pass and 150 for every verified project, plus 15 for today's challenge and 100 for the weekly one. You spend it in the shop on mascots and colours; it can't be bought with money.

Shall we start?

Create your account and activate your subscription: you get the whole syllabus from day one.