Research area

Structured Video Representations

We represent videos as evolving structure—not just stacks of pixels—to make motion coherent, compact, and directly editable.

GaussianVideo illustration showing coherent Gaussian trajectories across successive video frames.
GaussianVideo models coherent motion in the underlying Gaussians, enabling continuous and efficient video reconstruction. Figure adapted from the authors’ paper.

What we study

Conventional video models repeatedly encode dense frames. We ask whether a scene can instead be described through persistent elements, smooth trajectories, and continuous functions of time. These representations can reduce redundancy while exposing useful controls for interpolation, resampling, editing, and style transfer.

Current questions

  • Which scene elements should persist across time?
  • How can continuous motion be learned without dense optical-flow supervision?
  • How should appearance and motion be disentangled?
  • Can one representation support reconstruction, interpolation, resampling, and editing?
  • How can hierarchical learning capture both large motion and fine detail efficiently?

Our approach

GaussianVideo combines a Gaussian scene representation with continuous camera and object motion. Spatial and temporal hierarchies progressively refine the representation, while neural differential equations model smooth evolution. Earlier work such as VidStyleODE and SLAMP likewise separates the factors of a video so that they can be modeled and manipulated independently.

This direction connects computer vision, graphics, dynamical systems, and generative modeling.

From the lab

Selected publications

2025 · arXiv preprint (2025)

GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting

We combine 3D Gaussian splatting, Neural ODE camera modeling, and hierarchical spatiotemporal learning for fast, memory-efficient, and temporally consistent video representation.

2023 · ICCV (2023)

VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEs

We propose VidStyleODE, a method for disentangled video editing that combines StyleGAN with Neural ODEs for temporally consistent video manipulation.

2021 · ICCV (2021)

SLAMP: Stochastic Latent Appearance and Motion Prediction

We introduce SLAMP, a stochastic video prediction model that disentangles appearance and motion in a latent space.

Browse all publications