Course 3 · Unit 3
Computer vision
Part of Deep learning and Transformers
- 4 lessons
- ≈ 13 h of study
- Level: intermediate to advanced
Topics covered
- CNN
- Classification
- Object detection
- Segmentation
- Embeddings
- Vision Transformers
- Multimodal vision
Lessons in this unit
- Convolution — a magnifying glass that sweeps across the image 80 min
The operation that lets a network "see" edges, textures and shapes, computed by hand and then with NumPy. - Convolutional networks — stacking ever wider views 70 min
Convolutions, activations and pooling chained into a complete CNN; receptive field, parameter counting and data augmentation. - Detection and segmentation — what's there and where it is 80 min
Boxes, IoU, non-maximum suppression and average precision; the pieces any detector is built and evaluated with. - Vision Transformers — an image is worth 16 × 16 words 75 min
Cutting the image into patches, turning them into tokens and letting attention do the rest; shapes, parameters and cost.
Prerequisites
Before this unit it helps to have done:
- Engineering with PyTorch (Course 3 · Unit 2)
The full explanations, auto-graded exercises, exams and projects are inside the academy.
Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.
Shall we start?
Create your account and activate your subscription: you get the whole syllabus from day one.