Research area

Multimodal Learning and Reasoning

We investigate whether models can connect vision and language beyond surface correlations—grounding concepts in time, composing known ideas, and explaining what they observe.

ViLMA benchmark overview showing video frames of paper being folded, followed by a basic proficiency test and a harder temporal change-of-state test.
ViLMA evaluates video–language models in two stages: a proficiency test first confirms the prerequisite concept, then a controlled main test probes deeper temporal understanding. Figure from the ViLMA paper.

What we study

Multimodal systems should do more than associate captions with pictures. We study models that must identify entities and events, track their relationships, connect language to the right moment in a video, and recombine familiar concepts in unfamiliar sequences.

Our research spans vision–language learning, multilingual representation learning, video–language grounding, procedural understanding, and behavioral evaluation of foundation models.

Current questions

  • Can a model solve a new composition after learning its individual parts?
  • Does a video–language model distinguish actions, roles, and event order?
  • Which examples teach robust concepts, and which encourage shortcuts?
  • How should multimodal reasoning be evaluated across languages and cultures?
  • Can benchmarks reveal why a model succeeds or fails?

Our approach

We pair new models with diagnostic datasets and targeted evaluation. CompAct isolates sequential compositional generalization; ViLMA probes linguistic and temporal grounding with controlled contrasts; dataset cartography reveals how training examples shape generalization.

The goal is measurable reasoning: systems whose behavior can be tested with specific hypotheses rather than summarized by a single aggregate score.

From the lab

Selected publications

2024 · NAACL (2024, oral)

Sequential Compositional Generalization in Multimodal Models

We investigate sequential compositional generalization capabilities in multimodal models, introducing the CompAct benchmark.

2024 · ICLR (2024)

ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models

We introduce ViLMA, a zero-shot benchmark designed to evaluate linguistic and temporal grounding capabilities of video-language models.

2023 · Findings of EMNLP (2023)

Harnessing Dataset Cartography for Improved Compositional Generalization in Transformers

We leverage dataset cartography to improve compositional generalization in Transformer models.

2019 · CoNLL (2019)

Procedural Reasoning Networks for Understanding Multimodal Procedures

We introduce Procedural Reasoning Networks for understanding multimodal procedural content such as cooking recipes.

Browse all publications