Research area
Multimodal Learning and Reasoning
We investigate whether models can connect vision and language beyond surface correlations—grounding concepts in time, composing known ideas, and explaining what they observe.
What we study
Multimodal systems should do more than associate captions with pictures. We study models that must identify entities and events, track their relationships, connect language to the right moment in a video, and recombine familiar concepts in unfamiliar sequences.
Our research spans vision–language learning, multilingual representation learning, video–language grounding, procedural understanding, and behavioral evaluation of foundation models.
Current questions
- Can a model solve a new composition after learning its individual parts?
- Does a video–language model distinguish actions, roles, and event order?
- Which examples teach robust concepts, and which encourage shortcuts?
- How should multimodal reasoning be evaluated across languages and cultures?
- Can benchmarks reveal why a model succeeds or fails?
Our approach
We pair new models with diagnostic datasets and targeted evaluation. CompAct isolates sequential compositional generalization; ViLMA probes linguistic and temporal grounding with controlled contrasts; dataset cartography reveals how training examples shape generalization.
The goal is measurable reasoning: systems whose behavior can be tested with specific hypotheses rather than summarized by a single aggregate score.
From the lab
Selected publications