Emotion and multimodal foundation models

This course, delivered by the University of Trento in Italy, introduces multimodal foundation models and their application to emotion analysis. Learners cover both how these models are developed and what they can and cannot do when queried about emotions, gaining the tools to judge which foundation model suits a given application and how to use it.

About this course

Foundation models trained across text, images, audio and video have become the default starting point for building AI systems, and emotion analysis is one of the domains where their promise and their limits are both most visible. This course provides a structured introduction to how these models are built and to what happens when they are asked to interpret human affect.

The module opens with the background needed to read the field (i.e., a recap of deep neural networks, the Transformer architecture, and how text and images are encoded) before turning to contrastive learning and CLIP as the canonical vision–language model. From there it examines what CLIP can and cannot do, applies it directly to emotion recognition, and broadens out through large language models, multimodal fusion strategies, and the design of multimodal foundation models themselves. Later lectures cover working with more than two modalities, in-context learning and retrieval, post-training, adapters and agentic systems, giving participants a practical vocabulary for adapting existing models rather than training them from scratch.

The final part of the course is devoted to emotion specifically: how foundation models perform on emotion analysis, applications in emotion recognition and in tasks that go beyond standard recognition, and the ethical challenges of fairness and privacy that arise when systems make inferences about people’s emotional states. Practical topics are distributed through the course, and the material is self-paced throughout.

The course is delivered at an advanced level. It assumes working knowledge of Python and familiarity with deep learning fundamentals; the companion module Deep Network Development provides suitable preparation.

Course structure

  • Foundations: introduction to multimodal foundation models, recap of deep neural networks, Transformers.
  • Representing modalities: encoding text, encoding images, contrastive learning.
  • Vision–language models: CLIP, what CLIP can and cannot do, emotion recognition with CLIP (with practice).
  • Building multimodal systems: large language models, fusing multimodal information, multimodal foundation models in two parts, two things to keep in mind (with practice), beyond two modalities.
  • Adapting and deploying: in-context learning and retrieval, post-training, adapters, agentic systems.
  • Emotion and responsibility: foundation models for emotion analysis, applications in emotion recognition, applications beyond standard emotion recognition, and the ethical challenges of fairness and privacy.

The course comprises 24 lessons, 4 practical topics and 21 quizzes, and carries a course certificate on completion.

Learning outcomes

Upon successful completion of this course, students will be able to:

  • Explain what multimodal foundation models are and how they are developed, from encoding individual modalities to aligning them in a shared representation space.
  • Describe the role of Transformers, contrastive learning and vision–language models such as CLIP in current multimodal systems.
  • Assess the capabilities and limitations of foundation models when applied to emotion analysis.
  • Apply in-context learning, retrieval, post-training and adapter-based methods to adapt a pre-trained model to a specific task.
  • Compare strategies for fusing information across two or more modalities, and design agentic systems built on foundation models.
  • Select an appropriate foundation model for a given application need and use it effectively.
  • Identify the fairness and privacy risks raised by emotion-inference systems and account for them in system design.

Further details

This course is developed within the framework of the EMAI4EU project, with the support of the Digital Europe Programme of the European Union under grant agreement no. 101123289. More information on the course is available on the corresponding website.

Course Content

Lesson Content
Lesson Content
0% Complete 0/1 Steps
1 of 2

About Instructor

Not Enrolled

Course Includes

  • 24 Lessons
  • 4 Topics
  • 21 Quizzes
  • Course Certificate