Research area
Generative Modeling for Controllable Visual Media
We build generative systems that turn human intent into editable images, coherent videos, and explicit motion—without sacrificing visual quality.
Forward process · perturbation
Reverse process · denoising
What we study
Generative models are most useful when people can direct them precisely. Our work connects high-level instructions—language, style, layout, camera intent, or object motion—to representations that a model can execute and revise.
We study diffusion models, generative adversarial networks, transformers, hypernetworks, neural ordinary differential equations, and Gaussian representations. Across these families, the central question stays the same: how can generation become controllable, compositional, and temporally consistent?
Current questions
- How can natural language specify camera movement and object trajectories?
- How can image edits preserve everything outside the intended change?
- How can video models separate appearance, motion, geometry, and style?
- How can constraints from physics and cinematography guide generation?
- How can high-quality generation become faster and more memory-efficient?
Our approach
We expose structure rather than hiding every decision inside a single latent vector. LAMP, for example, translates cinematic language into symbolic motion programs and explicit 3D trajectories. CLIPAway uses focused semantic embeddings to localize an edit. GaussianVideo represents dynamic content with moving Gaussians whose trajectories remain continuous through time.
This combination of learned generation and interpretable intermediate representations makes systems easier to steer, inspect, and reuse.
From the lab
Selected publications