Marsof Academy
Course 2 · Unit 2

Maths for AI

Part of Machine learning

  • 12 lessons
  • ≈ 29 h of study
  • Level: intermediate

Linear algebra, probability, statistics, calculus and optimisation, implemented in code.

Topics covered

  • Vectors
  • Matrices
  • Matrix multiplication
  • Linear transformations
  • Eigenvalues
  • Probability
  • Conditional probability
  • Bayes
  • Statistics
  • Distributions
  • Derivatives
  • Partial derivatives
  • Gradients
  • Chain rule
  • Optimisation
  • Gradient descent
  • SGD
  • Adam
  • Regularisation

Lessons in this unit

  1. Vectors and the dot product 45 min
    AI's fundamental data structure, from Python lists to the similarity between embeddings.
  2. Matrices and matrix multiplication 55 min
    Shape, dimension compatibility and the row × column product, the operation every layer of a neural network runs.
  3. Derivatives: slope, limit and the power rule 35 min
    The derivative as a local slope, the limit as "bringing two points together", the power rule and the numerical approximation.
  4. Differentiation rules 35 min
    The product rule, the basic chain rule, the derivatives of eˣ and ln x, and the second derivative as curvature.
  5. Gradient descent 45 min
    The algorithm almost every AI model learns with, understood with a single variable.
  6. Probability and conditional probability 40 min
    The language a model uses to express its uncertainty, from counting cases to P(A|B).
  7. Bayes' theorem 55 min
    How to update a belief with evidence, and why a 99 % detector can be wrong almost every time.
  8. Statistics and distributions 65 min
    Mean, variance, the normal distribution and why averages over lots of data are reliable.
  9. Partial derivatives and the gradient 65 min
    From one variable to millions: how to differentiate a loss that depends on many parameters at once.
  10. The chain rule 75 min
    Differentiating composite functions, the one mathematical idea behind backpropagation.
  11. SGD, momentum and Adam 80 min
    The optimisers networks are really trained with, built step by step from gradient descent.
  12. Linear transformations and eigenvalues 90 min
    Seeing a matrix as a transformation of space and finding its special directions, the ones that don't rotate, to understand PCA and why a training run blows up.

Prerequisites

Before this unit it helps to have done:

The full explanations, auto-graded exercises, exams and projects are inside the academy.

Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.

Shall we start?

Create your account and activate your subscription: you get the whole syllabus from day one.