Research area
Structured Video Representations
We represent videos as evolving structure—not just stacks of pixels—to make motion coherent, compact, and directly editable.
What we study
Conventional video models repeatedly encode dense frames. We ask whether a scene can instead be described through persistent elements, smooth trajectories, and continuous functions of time. These representations can reduce redundancy while exposing useful controls for interpolation, resampling, editing, and style transfer.
Current questions
- Which scene elements should persist across time?
- How can continuous motion be learned without dense optical-flow supervision?
- How should appearance and motion be disentangled?
- Can one representation support reconstruction, interpolation, resampling, and editing?
- How can hierarchical learning capture both large motion and fine detail efficiently?
Our approach
GaussianVideo combines a Gaussian scene representation with continuous camera and object motion. Spatial and temporal hierarchies progressively refine the representation, while neural differential equations model smooth evolution. Earlier work such as VidStyleODE and SLAMP likewise separates the factors of a video so that they can be modeled and manipulated independently.
This direction connects computer vision, graphics, dynamical systems, and generative modeling.
From the lab
Selected publications