Language-driven cinematographic camera control

Auteur: Language-Driven Cinematographic Framing for
Human-Centric Video Generation


M. Burak Kizil1 Enes Sanli1 Niloy J. Mitra2,4 Xuelin Chen4 Erkut Erdem3 Aykut Erdem1 Duygu Ceylan4

1Koç University 2University College London 3Hacettepe University 4Adobe Research
arXiv Dataset Demo

Auteur Teaser

Auteur turns language into actor-relative camera framing. Given a natural-language description, a fine-tuned LLM acts as a virtual director: it first predicts a sparse actor-motion program, then a camera-framing program expressed in our cinematographic DSL. The DSL is decoded into a 6-DoF trajectory that keeps the subject correctly framed and can drive a range of downstream video generators.

We replace world-space camera trajectories with actor-relative shot compositions: a cinematographic DSL and a fine-tuned LLM map natural language and human motion into camera keyframes that are geometrically consistent and narratively intentional.

The SOMA visualizations shown in this page are generated with Kimodo.

Abstract

Generative video models have achieved remarkable visual fidelity and temporal coherence, yet intentional camera control remains elusive. Existing frameworks treat camera motion as a byproduct of pixel synthesis, producing trajectories that are stochastic, spatially inconsistent, and indifferent to the human subject driving the scene. In this work, we present Auteur, a method for language-driven, human-centric camera framing in generative video. Our core insight is that professional filmmakers conceive shots not as world-space trajectories but as framings defined relative to the actor, encoding shot size, angle, and composition as functions of human pose and motion. We formalize this intuition as a human-centric camera parameterization and introduce a Domain-Specific Language (DSL) that is convertible to standard 6-DoF camera parameters. A fine-tuned multimodal large language model then acts as a virtual director, mapping natural language descriptions and coarse human motion to sparse DSL keyframes that are deterministically interpolated into continuous camera trajectories, which are then provided as input to video generators. We train and evaluate Auteur on a new dataset of 34K aligned text, human motion, and DSL-annotated camera trajectories drawn from procedural synthesis and real-world movie footage from the CondensedMovies dataset. Auteur enables cinematographic framing of human-centered scenes, a capability largely absent in prior generative models. To assess this behavior, we propose new framing-focused metrics, and our experiments show that Auteur consistently outperforms existing methods on actor-centric framing and subject visibility.

Contributions

🎥

Human-Centric Camera Parameterization & DSL

An actor-relative representation of camera state grounded in professional cinematographic conventions, with framing — rather than raw trajectory — as the primary compositional primitive. Its discrete, LLM-generatable DSL is deterministically convertible to standard 6-DoF parameters.

🗣️

Language-to-Human-to-Camera Pipeline

A two-stage method mapping natural language to coarse 3D human trajectories, then to sparse DSL keyframe programs that are deterministically interpolated into continuous, actor-aware 6-DoF camera paths.

🎞️

Auteur Dataset

A dataset of 34K aligned samples — caption, SOMA parameters over time, camera trajectory, and DSL program — from procedural synthesis and real movie footage, enabling learning of human-aware cinematic priors.

📊

Cross-Dataset Evaluation & Auteur Score

A new framing-centric benchmark suite measuring subject visibility, compositional adherence, temporal framing stability, and actor-camera coordination. Across our internal benchmark and the external PulpMotion benchmark, Auteur delivers stronger, more stable cinematic framing.

Motion and Framing Primitives

Each primitive isolates one axis of our DSL. The panels show the initial and final DSL states (changed fields highlighted), and the decoded camera motion rendered around a SOMA actor.

  • ← Jitter
    Arc
    Dolly in →
    The camera arcs from the left side to the front and then right, shifting the perspective smoothly across the space.
    Initial DSL
    Orientation: Left
    Depth: Medium
    Camera Level: Eye
    Lookat Level: Eye
    Dutch Angle: Normal
    Framing: Center
    Jitter: None
    Ease: None
    Final DSL
    Orientation: Right
    Depth: Medium
    Camera Level: Eye
    Lookat Level: Eye
    Dutch Angle: Normal
    Framing: Center

Model Comparisons

The same natural-language prompt is given to every method. The orange text describes the actor and the blue text the requested camera behavior. Auteur keeps the subject correctly framed while following the intended shot, where PulpMotion and LAMP drift or lose the actor.

  • ← Tail Track
    Side Track
    Dolly In →
    A woman runs forward through the scene. The camera tracks from the right side and switches to left, maintaining a consistent side profile of the subject.
    Pulp Motion
    LAMP
    Auteur (Ours)

VerseCrafter Pipeline

Inputs to VerseCrafter: object 3D Gaussians, motion trajectories, and the text prompt.

Auteur's predicted trajectory becomes the guidance signal. Toggle Scene View / Camera View on the left video to inspect the generated trajectory, then compare against the final rendered output on the right.

  • ← Tail Track
    Over-the-Shoulder
    Anchor Middle →
    A man and a woman are sitting faced to each other. Camera changes from over-the-shoulder perspective to a close-up to person sitting in front.
    Generated Trajectory / Guidance
    Final Output

VACE

Inputs to VACE: a Kimodo-generated control video and the text prompt.

Actor description and camera instruction are highlighted in each prompt. With Auteur's control video, the generated shot follows the requested framing; the text-only baseline (T2V) does not.

  • ← Over-the-Shoulder
    Arc
    Curved Path →
    Wild West scene set in a vast sunlit desert town, with a rugged cowboy standing still while adjusting his hat in the exact center of the frame, wearing .... The background .... Golden sand drifts lightly across the ground while the warm afternoon sun casts long dramatic shadows. The camera performs a smooth cinematic arc around the cowboy, beginning from his left side profile, slowly circling in front of him, then continuing toward his right side, revealing more of the town and desert landscape as it moves.
    Control Video
    Ours
    Without Guidance (T2V)

Cross-Renderer Generalization

Auteur outputs a standard 6-DoF trajectory, so it drives independently built generators without modification. Below, the same predicted trajectory is rendered through three different camera-conditioned generators and compared against a text-only baseline that receives the same prompt but no explicit trajectory. Every Auteur-guided renderer follows the requested framing more faithfully than the baseline.

  • ← Scenario 3
    Orbit
    Scenario 2 →
    The same Auteur trajectory, rendered by four different pipelines.
    Baseline (Wan2.2 TI2V)
    Ours + VerseCrafter Ours
    Ours + ReCamMaster Ours

Quantitative Results

🧭

Positioning relative to LAMP

Auteur adopts the established LLM → DSL → interpolation pipeline; our novelty is the actor-relative framing representation it carries. Where LAMP's DSL is a small handcrafted rule set, ours is a parametric, invertible parameterization of the camera–actor relationship — which is what unlocks full framing expressiveness and training on real movie footage. On the world-space PulpMotion benchmark, LAMP attains higher Camera F1 (both splits) and higher CLaTr (mixed split); Auteur is strongest on out-of-frame rate and, decisively, on actor-centric framing. The tables below isolate that gain.

Per-frame framing scores on our controlled benchmark, reported as min / mean / max over the full trajectory (replacing the earlier endpoint-only Auteur Score). Ranking is unchanged under mean and min — Auteur's worst frame beats the baselines' average frame on every axis.

MethodF-OriF-RoTF-ScaleF-TiltF-RollCam F1
LAMP0.183 / 0.451 / 0.7070.047 / 0.257 / 0.6890.484 / 0.729 / 0.9510.644 / 0.832 / 0.9650.308 / 0.745 / 0.9750.112
PulpMotion0.281 / 0.514 / 0.7480.086 / 0.351 / 0.7880.456 / 0.712 / 0.8920.805 / 0.891 / 0.9660.548 / 0.839 / 0.9740.091
Auteur (Ours)0.924 / 0.962 / 0.9990.368 / 0.745 / 0.9640.760 / 0.888 / 0.9880.943 / 0.968 / 0.9920.900 / 0.937 / 0.9730.623

Higher is better. Cam F1 aggregates adherence across all five framing axes.

Isolating the actor-relative representation with the pipeline, decoder, and data held fixed. The world-space DSL rows reconstruct from ground truth (an upper bound), and Direct-6DoF predicts raw poses with no parameterization. Neither recovers the actor-relative gain — the advantage is the representation itself.

RepresentationF-OriF-RoTF-ScaleF-TiltF-RollF1
World-space DSL, 4-seg (GT recon., upper bound)0.4990.3710.7300.6750.3110.262
World-space DSL, 8-seg (GT recon., upper bound)0.5000.3720.7330.6740.3130.280
Direct-6DoF (no DSL, no parameterization)0.7640.4760.7500.9040.8990.170
Actor-relative (Ours)0.9620.7450.8880.9680.9370.623

Direct-6DoF matches static framing targets per-axis but fails to produce coherent motion between keyframes, so its trajectory F1 collapses (0.170 vs. our 0.623).

Generation order. Camera-first, joint, and actor-first factorizations perform comparably (accuracy, %), so we adopt actor-first for its usefulness and because the camera is framed relative to the actor — generating camera tags first is not even well-defined.

Synth. Test SetCamera-FirstJointOurs (Actor-first)
Translation99.699.699.8
Human Yaw99.399.398.5
OA98.798.698.6
SS94.193.095.2
CL98.798.296.7
LL98.598.296.3
FO93.393.097.9
DA96.896.799.8

Training data composition. Adding real movie data on top of procedural synthesis substantially improves generalization on the challenging real test set.

60.1 → 85.8
Shot-scale accuracy on the real test set
(synth-only → synth + real)

Two-alternative forced-choice study on VerseCrafter-rendered videos. Across 10 scenarios, paired videos differ only in the predicted camera trajectory; method identities are hidden and left–right order is randomized. 24 participants, 225 valid pairwise judgments.

82.7%
preferred over LAMP
— more professionally filmed
80.4%
preferred over LAMP
— better instruction adherence
81.3%
preferred over PulpMotion
— more professionally filmed
83.1%
preferred over PulpMotion
— better instruction adherence

Automatic video-level quality on the same 10 videos with VBench (%). Auteur leads on most dimensions, including a large gap on Dynamic Degree.

VBench DimensionPulpMotionLAMPAuteur (Ours)
Subject Consistency87.7289.5091.23
Dynamic Degree81.8254.5590.91
Background Consistency90.8691.3793.01
Motion Smoothness98.3299.0898.99
Aesthetic Quality59.8859.3563.11
Image Quality69.3669.6568.71

Higher is better. Best value in each row is bold.

VQAScore (p(yes)) — an automatic prompt-adherence measure — quantifies how well a rendered video follows the requested camera instruction.

Method (same videos)VQAScore
LAMP0.36
PulpMotion0.31
Auteur (Ours)0.54

The gain also transfers across independently built renderers: every Auteur-guided generator beats the text-only baseline.

ConfigurationVQAScore
Baseline (TI2V, no trajectory)0.39
Ours + VerseCrafter0.50
Ours + ReCamMaster0.49
Ours + EPiC0.47

BibTeX

@article{kizil2026auteur,
  title   = {Auteur: Language-Driven Cinematographic Framing for
             Human-Centric Video Generation},
  author  = {Kizil, M. Burak and Sanli, Enes and Mitra, Niloy J. and
             Chen, Xuelin and Erdem, Erkut and Erdem, Aykut and
             Ceylan, Duygu},
  journal = {arXiv preprint},
  year    = {2026}
}