Marsof Academy
Course 3 · Unit 3

Computer vision

Part of Deep learning and Transformers

  • 4 lessons
  • ≈ 13 h of study
  • Level: intermediate to advanced

Topics covered

  • CNN
  • Classification
  • Object detection
  • Segmentation
  • Embeddings
  • Vision Transformers
  • Multimodal vision

Lessons in this unit

  1. Convolution — a magnifying glass that sweeps across the image 80 min
    The operation that lets a network "see" edges, textures and shapes, computed by hand and then with NumPy.
  2. Convolutional networks — stacking ever wider views 70 min
    Convolutions, activations and pooling chained into a complete CNN; receptive field, parameter counting and data augmentation.
  3. Detection and segmentation — what's there and where it is 80 min
    Boxes, IoU, non-maximum suppression and average precision; the pieces any detector is built and evaluated with.
  4. Vision Transformers — an image is worth 16 × 16 words 75 min
    Cutting the image into patches, turning them into tokens and letting attention do the rest; shapes, parameters and cost.

Prerequisites

Before this unit it helps to have done:

The full explanations, auto-graded exercises, exams and projects are inside the academy.

Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.

Shall we start?

Create your account and activate your subscription: you get the whole syllabus from day one.