Course 5 · Unit 6
Multimodal AI
Part of Advanced level (optional): specialisations
- 5 lessons
- ≈ 18 h of study
- Level: advanced
Topics covered
- Vision-language
- Audio
- Voice
- Image generation
- Multimodal embeddings
- Multimodal agents
Lessons in this unit
- Contrastive learning and CLIP — images and texts on the same map 85 min
How to train two encoders so that a photo and its description land on the same point, and how to use that space to classify without training. - Audio and speech — from the microphone to the mel spectrogram 85 min
Sampling, Nyquist and aliasing, windows and the DFT, spectrograms and the mel scale; how audio turns into something a network can read. - Vision-language models and multimodal agents 85 min
How an image encoder is connected to an LLM (encoder → projector → tokens), how much an image costs in context and how to build an agent that sees the screen. - Generative models — autoencoders, VAEs and GANs 70 min
How to go from classifying images to making them — compressing into a latent space, the VAE and its ELBO with the reparameterisation trick, and the duel between a GAN's generator and discriminator, with its mode collapse. - Image generation — diffusion models 105 min
How an image is generated from text today — destroying data with Gaussian noise and learning to undo it step by step, the closed form of the forward process, the noise-prediction objective, sampling, classifier-free guidance, latent diffusion and how quality is measured with FID.
Prerequisites
Before this unit it helps to have done:
- Research Engineering (Course 5 · Unit 5)
- Computer vision (Course 3 · Unit 3)
The full explanations, auto-graded exercises, exams and projects are inside the academy.
Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.
Shall we start?
Create your account and activate your subscription: you get the whole syllabus from day one.