Marsof Academy
Course 5 · Unit 6

Multimodal AI

Part of Advanced level (optional): specialisations

  • 5 lessons
  • ≈ 18 h of study
  • Level: advanced

Topics covered

  • Vision-language
  • Audio
  • Voice
  • Image generation
  • Multimodal embeddings
  • Multimodal agents

Lessons in this unit

  1. Contrastive learning and CLIP — images and texts on the same map 85 min
    How to train two encoders so that a photo and its description land on the same point, and how to use that space to classify without training.
  2. Audio and speech — from the microphone to the mel spectrogram 85 min
    Sampling, Nyquist and aliasing, windows and the DFT, spectrograms and the mel scale; how audio turns into something a network can read.
  3. Vision-language models and multimodal agents 85 min
    How an image encoder is connected to an LLM (encoder → projector → tokens), how much an image costs in context and how to build an agent that sees the screen.
  4. Generative models — autoencoders, VAEs and GANs 70 min
    How to go from classifying images to making them — compressing into a latent space, the VAE and its ELBO with the reparameterisation trick, and the duel between a GAN's generator and discriminator, with its mode collapse.
  5. Image generation — diffusion models 105 min
    How an image is generated from text today — destroying data with Gaussian noise and learning to undo it step by step, the closed form of the forward process, the noise-prediction objective, sampling, classifier-free guidance, latent diffusion and how quality is measured with FID.

Prerequisites

Before this unit it helps to have done:

The full explanations, auto-graded exercises, exams and projects are inside the academy.

Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.

Shall we start?

Create your account and activate your subscription: you get the whole syllabus from day one.