8.14 Multimodal models
You can build features over images, audio and documents.
Before:07. Natural Language ProcessingUnlocks:09. Agentic AI12. Frontier Topics
Multimodal models take images, audio and documents as first-class input: vision-language understanding, speech recognition and synthesis, OCR-and-layout document AI — and real products are usually pipelines of these rather than one model doing everything. It closes the LLM module by widening it. The evaluation gap is modality-specific: a model that describes photographs charmingly can still misread tables, and each input type needs its own test set.
Work through these
Vision-language models and image understanding
Models that take images alongside text, which enables description, question answering and extraction from pictures. This is the most mature of the multimodal capabilities.
Speech: ASR, TTS, realtime voice
Turning speech into text, text into speech, and doing both fast enough for conversation. Latency is the hard constraint that shapes every design here.
Document AI: OCR, layout, extraction
Extracting structure from documents, which requires reading not only the characters but the layout that gives them meaning. It is one of the highest-value applications in business settings.
Video understanding and generation, briefly
Understanding and generating video, treated briefly because it is the least settled of these areas. What is worth learning now is the shape of the problem rather than any particular system.
Sign in to keep your progress.
Free resources
We haven't checked most of these for screen reader use yet.
Links last checked 29 Aug 2026.
Stuck here?
Ask a mentor. A real person answers, and they can see exactly which topic you're on. Usually within a couple of working days.
Checking your session…
Topics shown in module order.