All courses
AI path Β· course 23 of 54
Multimodal AI Systems
Advanced Β· 5 lessons Β· 0 complete
The capstone of the NLP & Language Systems track, this course explains how modern AI systems combine text, images, audio, and video into a single reasoning system by learning a shared embedding space. You'll see how CLIP-style contrastive training aligns modalities, what capabilities that alignment unlocks like visual question answering and image captioning, and where multimodal systems still genuinely struggle. Built for learners who've completed NLP Fundamentals, Speech Recognition & Synthesis, and Building Conversational AI & Chatbots.
