Camera trajectories control viewpoint changes in video generation, scene reconstruction, and robotic perception. Generating them from language requires both scene geometry and target-aware framing. We introduce OmniCam, an autoregressive model that generates camera pose sequences from a single panorama and textual trajectory descriptions. Its geometry-grounded pose token learning combines three components: a panoramic point-cloud encoder for omnidirectional geometric context; hybrid absolute-rotation and relative-translation tokenization with temporally consistent quaternion signs; and separate geometric and semantic conditioning streams with an explicit 3D target anchor. We also construct OmniCaT, containing 267,700 trajectories across four camera behaviors. On the reported OmniCaT evaluation, OmniCam reduces trajectory errors by 28--47% and collision rate by 65.8% relative to GenDoP retrained on OmniCaT. Against the best baseline for each metric, the ATE and collision reductions are 43.0% and 62.3%, respectively. Component ablations support the use of geometric and target-aware conditioning, while downstream experiments examine camera-controlled video generation and robotic active perception.
Figures & tables
Figure 2 : OmniCaT Dataset Overview. Left : (a) Construction pipeline comprising (1) Geometric and Semantic Grounding to extract nav-meshes and 3D targets, (2) Spatial Trajectory Synthesis to generate diverse camera behaviors, and (3) Rendering and Multimodal Annotation to produce text-trajectory pairs. Right : Dataset statistics showing (b) task distribution, (c) top object categories, (d) motion primitive diversity, and (e) target direction distribution.
Figure 3 : OmniCam Architecture. Three parallel branches encode the caption, panoramic image, and point cloud. A Language-Guided Resampler produces target-aware semantic tokens while a Geometry Resampler yields geometric descriptors; an attention-aggregated 3D target centroid provides an explicit spatial anchor for the decoder. The autoregressive decoder generates pose tokens from the concatenated condition, which are mapped to continuous SE(3) trajectories by a de-tokenizer. Right bottom: SE(3) tokenization with 9 tokens for each pose frame.
Method
Dataset
ATE ↓
FDE ↓
RPE-R ↓
RPE-T ↓
Coll. (%) ↓
Vis. (%) ↑
Angle ( ∘ ) ↓
Evaluation on the OmniCaT benchmark
CCD [ 18 ]
Pre-trained
2.483
3.396
3.112
0.082
36.9
12.3
32.5
E.T. [ 7 ]
Pre-trained
1.983
2.368
2.151
0.069
46.2
21.3
40.2
Director3D [ 23 ]
Pre-trained
1.828
2.942
2.309
0.052
28.4
28.1
27.2
Director3D
OmniCaT
1.522
2.648
2.187
0.049
27.6
29.3
26.5
GenDoP [ 34 ]
Pre-trained
1.865
2.951
1.681
0.044
32.6
31.5
29.1
Table 1 : Comparison with representative trajectory generation baselines on OmniCaT and the public DataDoP benchmark.
Figure 4 : Qualitative comparison on representative panoramic scenes. Each column shows the trajectory produced by a specific method. The target object is indicated by a red bounding box.
Figure 5 : Qualitative comparison on panoramic scenes. Each row displays an input panorama, generated video sequences, and 3D reconstruction results. The two rightmost columns compare reconstruction outputs without and with the OmniCam trajectory; no quantitative reconstruction metric is reported here.
Figure 6 : Real-robot demonstration of OmniCam-guided active perception. OmniCam plans a camera trajectory that actively repositions the observation camera to resolve target occlusion.
Variant
ATE ↓
FDE ↓
RPE-R ↓
RPE-T ↓
Coll. (%) ↓
Vis. (%) ↑
Angle ( ∘ ) ↓
Encoder
w/o semantic branch
2.186
3.247
2.694
0.072
13.6
11.5
64.7
w/o geometric branch
1.774
3.084
2.881
0.049
45.8
33.6
19.7
Geo. Input
Panorama RGB only
1.536
2.687
2.143
0.044
38.7
36.2
18.4
Panorama RGB-D
1.247
2.118
1.685
0.036
22.6
42.8
14.7
Grounding
w/o target cross-attention
1.673
2.259
2.774
0.042
15.2
25.8
35.6
w/o target localization loss
1.582
2.148
2.661
0.038
14.8
32.4
28.7
Table 2: Ablation studies on OmniCaT. The dual-branch panoramic encoder, target injection components, and pose tokenization schemes are investigated.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Training Time
Inference Latency
#Params
GenDoP
∼ 18h (8 × H20)
0.42s / trajectory
∼ 380M
OmniCam (Base)
∼ 24h (8 × H20)
0.84s / trajectory
∼ 500M
OmniCam (Large)
∼ 40h (8 × H20)
1.20s / trajectory
∼ 1B
Appendix
Table 3: Reported computational costs for GenDoP and OmniCam Base/Large.
Method
Dataset
Wander
Target
Surround
Reconstruct
ATE ↓
FDE ↓
ATE ↓
FDE ↓
ATE ↓
FDE ↓
ATE ↓
FDE ↓
CCD [ 18 ]
Pre-trained
3.718
4.907
1.485
2.285
2.524
3.419
2.206
2.972
E.T. [ 7 ]
Pre-trained
3.151
4.083
1.735
1.945
1.580
1.640
1.467
1.804
Director3D [ 23 ]
Pre-trained
2.969
4.779
1.444
2.207
1.620
2.688
1.279
2.094
Director3D
OmniCaT
2.562
4.250
1.130
2.089
1.351
2.304
1.045
1.950
GenDoP [ 34 ]
Pre-trained
2.667
4.408
1.362
2.131
2.176
3.248
1.255
2.016
Appendix
Table 4: Per-behavior decomposition of the main result on OmniCaT. Metrics are reported separately for Wander, Target, Surround, and Reconstruct trajectories.
#Trajectories
ATE ↓
FDE ↓
RPE-R ↓
RPE-T ↓
Coll. (%) ↓
Vis. (%) ↑
27K (10%)
1.483
2.586
1.836
0.054
19.8
34.6
53K (20%)
1.206
2.080
1.528
0.042
15.8
42.6
107K (40%)
1.003
1.709
1.297
0.033
12.7
48.2
267K (100%)
0.868
1.462
1.135
0.027
10.4
52.0
Appendix
Table 5: Reported effect of nominal data scale using the base architecture. Exact training-subset counts and the fixed evaluation split require verification.
Size
#Params
ATE ↓
FDE ↓
RPE-R ↓
RPE-T ↓
Coll. (%) ↓
Vis. (%) ↑
Small
∼ 120M
1.127
1.893
1.428
0.036
13.8
44.6
Base
∼ 500M
0.868
1.462
1.135
0.027
10.4
52.0
Large
∼ 1B
0.743
1.248
0.982
0.022
8.7
56.3
Appendix
Table 6: Reported effect of model scale. The common training subset and parameter-count scope require verification against training records.
Figure 7 : Attention visualization of the Language-Guided Semantic Resampler. Each group shows (from left to right) the input panorama, the reconstructed point cloud, and the 3D attention map. The maps visualize target-related attention on reconstructed geometry for the displayed examples; they are qualitative illustrations rather than a quantitative localization evaluation.
Dropout Strategy
ATE ↓
FDE ↓
RPE-R ↓
RPE-T ↓
Coll. (%) ↓
Vis. (%) ↑
Angle ( ∘ ) ↓
No dropout
0.856
1.438
1.118
0.026
10.1
38.7
18.4
Single-variant
0.862
1.451
1.126
0.027
10.3
45.8
14.2
Full (5-variant)
0.868
1.462
1.135
0.027
10.4
52.0
10.9
Appendix
Table 7: Caption-dropout ablation. The five-variant scheme favors visibility and angle, while no dropout has slightly lower geometric errors.
Method
NavMesh/annotation access
Runtime
Coll. (%) ↓
Vis. (%) ↑
Angle ( ∘ ) ↓
Heuristic Planner
Yes (privileged)
∼ 16.5s
5.2
58.7
9.4
GenDoP [ 34 ]
No
0.42s
30.4
33.6
27.6
OmniCam (Ours)
No
0.84s
10.4
52.0
10.9
Appendix
Table 8: Privileged-information planner comparison. The planner receives reconstructed geometry, navigation meshes, and semantic annotations. OmniCam estimates geometry from the input panorama and text.
Figure 8 : Additional qualitative comparison on panoramic scenes. Extending the comparison in Figure 4 , each row shows a different scene spanning indoor, outdoor, and stylized environments, and each column shows the trajectory produced by a specific method. The examples illustrate differences in target-directed trajectories across scene categories. The trajectory captions used for each scene are listed in Table 9 .
Row
Type
Target
Caption
1
Target
door
<TASK: TARGET> <TARGET_OBJ: door> <TARGET_DIR: Back> Approach the back door to bring it into focus. The sequence concludes by arriving at the final <END_OBJ: door> position within the <END_DIR: Back> sector of the panoramic scene.
2
Target
light fixture
<TASK: TARGET> <TARGET_OBJ: light fixture> <TARGET_DIR: Front> Approach the front light fixture to bring it into focus. The sequence concludes by arriving at the final <END_OBJ: light fixture> position within the <END_DIR: Front> sector of the panoramic scene.
3
Wander
free space
<TASK: WANDER> <TARGET_OBJ: free space> <TARGET_DIR: Front> Approach the front workstation to explore the layout of the industrial corridor. The sequence concludes by arriving at the final <END_OBJ: free space> position within the <END_DIR: Front> sector of the panoramic scene.
4
Target
door
<TASK: TARGET> <TARGET_OBJ: door> <TARGET_DIR: Back> Approach the back door to bring it into focus. The sequence concludes by arriving at the final <END_OBJ: door> position within the <END_DIR: Back> sector of the panoramic scene.
5
Target
ice tunnel
<TASK: TARGET> <TARGET_OBJ: ice tunnel> <TARGET_DIR: Front> Approach the ice tunnel to bring it into focus. The sequence concludes by arriving at the final <END_OBJ: ice tunnel> position within the <END_DIR: Front> sector of the panoramic scene.
6
Wander
free space
<TASK: WANDER> <TARGET_OBJ: free space> <TARGET_DIR: Front> Approach the distant peaks to explore the mountain valley. The sequence concludes by arriving at the final <END_OBJ: free space> position within the <END_DIR: Front> sector of the panoramic scene.
Appendix
Table 9: Trajectory captions for Figure 8 . Each row corresponds to a scene in the qualitative comparison figure, listing the scene type, target object, and the hierarchical textual instruction provided to OmniCam.
Method
CLIP-S ↑
Adj.-SSIM ↑
Aest. proxy ↑
Fidelity ↑
Coverage
GenDoP [ 34 ] + WorldStereo
0.2644
0.4248
7.647
0.6924
35.8%
OmniCam (Ours) + WorldStereo
0.3018
0.4904
9.021
0.7560
38.5%
Appendix
Table 10: Downstream evaluation: camera-controlled video generation with WorldStereo [ 36 ] . Different trajectory sources are compared using the same video generation model.
Trajectory Source
Grasp SR (%) ↑
Exec. Vis. (%) ↑
π0.5 [ 16 ] (fixed camera)
56.3
47.2
GenDoP [ 34 ]
48.0
52.6
OmniCam (Ours)
72.5
71.8
Appendix
Table 11: Downstream evaluation: robotic manipulation with π0.5 . Reported grasp success (SR) and execution visibility (Vis.). The stated 200-episode setup and fixed-camera SR aggregation require reconciliation with the episode records.
Figure 9 : Qualitative downstream results across diverse panoramic scenes. Each row shows the input panorama (left), four selected frames from camera-controlled video generation using an OmniCam trajectory (middle), and 3D reconstruction results without and with the OmniCam trajectory (two rightmost columns). The examples include indoor, outdoor, and stylized environments. These visual comparisons illustrate the pipeline outputs; they do not quantify reconstruction accuracy or guarantee collision-free motion.
Method
ATE ↓
FDE ↓
RPE-R ↓
RPE-T ↓
Coll. (%) ↓
Vis. (%) ↑
Angle ( ∘ ) ↓
GenDoP (OmniCaT)
1.624 ± 0.087
2.483 ± 0.124
1.577 ± 0.068
0.038 ± 0.003
30.4 ± 1.8
33.6 ± 2.1
27.6 ± 1.4
OmniCam (Ours)
0.868 ± 0.042
1.462 ± 0.071
1.135 ± 0.053
0.027 ± 0.002
10.4 ± 1.1
52.0 ± 1.9
10.9 ± 0.8
Appendix
Table 12: Bootstrap 95% confidence intervals on OmniCaT evaluation scenes (1,000 resamples). Intervals describe evaluation-scene variability for the two listed methods.
We present TCAM (Track and Caption Any Motion), a generative framework that watches a video and with no text query and no region prompt decides what is moving, describes each motion in open vocabulary, locates it in time, and points to the exact trajectories that carry it. Two mature lines of work make this possible yet leave it unsolved: dense point trackers follow pixels with sub-object precision but emit no language, while video-language models produce fluent descriptions only when handed a query and only from clip-level features that cannot resolve which pixels move. Object-level captioners narrow the gap but still reason over detector boxes or masks, never reaching individual trajectories. TCAM couples tracking and language at point granularity through a Caption-Aware Resampler, where a small set of learnable queries cross-attends to dense point trajectory tokens and distills them into a fixed-length motion context that conditions a language decoder. The decoder generates an entire video's events in a single pass, each with a free-form caption, a start and end time, and a pointer to the trajectories it refers to, for sequential events and several subjects active at once. Training uses only existing segmentation annotations, with no extra event labeling, to supervise caption quality, pointer-mask alignment, and pointer diversity. On over 50K clips, TCAM outperforms dense video captioning baselines and matches dedicated, query-based grounding and point-tracking methods despite using no query, showing that trajectory-conditioned generation is a direct route to motion-driven video understanding.
Automatically generating cinematically expressive camera trajectories through 3D scenes from natural language descriptions is a challenging task of high practical value, with applications ranging from real-estate advertising to virtual tour creation. Existing methods either lack true 3D spatial awareness by relying on 2D image priors, or treat trajectory generation as a geometric path planning problem divorced from cinematographic semantics. We present CinemaTraj, a framework that reframes camera trajectory planning as a language-grounded spatial reasoning problem. Given a set of RGB-D images and a user prompt, CinemaTraj equips an LLM agent with a structured 3D scene graph: the agent decomposes the prompt into a sequence of atomic cinematographic movements (dolly, orbit, crane, pan, tilt, zoom, arc). Each movement is instantiated via a novel parametric trajectory representation that is both cinematographically expressive and optimizable for collision avoidance. The scene graph acts as a structured spatial prior, grounding the agent's reasoning in accurate geometric and semantic knowledge of the environment. CinemaTraj further generates synchronized voiceover and subtitles aligned with camera motion, producing narrated cinematic video outputs. We evaluate CinemaTraj on real-world ScanNet++ environments, and show that it produces prompt-faithful, collision-free trajectories with high cinematographic quality, outperforming existing approaches on prompt alignment, trajectory quality, and safety metrics.
Qianru Li, Xuyang Chen, Erkin Türköz +5
Technical University of Munich Munich, Germany · Huawei Dresden Research Center Munich, Germany
For artistic applications, video generation requires fine-grained control over both performance and cinematography, i.e., the actor's motion and the camera trajectory. We present ActCam, a zero-shot method for video generation that jointly transfers character motion from a driving video into a new scene and enables per-frame control of intrinsic and extrinsic camera parameters. ActCam builds on any pretrained image-to-video diffusion model that accepts conditioning in terms of scene depth and character pose. Given a source video with a moving character and a target camera motion, ActCam generates pose and depth conditions that remain geometrically consistent across frames. We then run a single sampling process with a two-phase conditioning schedule: early denoising steps condition on both pose and sparse depth to enforce scene structure, after which depth is dropped and pose-only guidance refines high-frequency details without over-constraining the generation. We evaluate ActCam on multiple benchmarks spanning diverse character motions and challenging viewpoint changes. We find that, compared to pose-only control and other pose and camera methods, ActCam improves camera adherence and motion fidelity, and is preferred in human evaluations, especially under large viewpoint changes. Our results highlight that careful camera-consistent conditioning and staged guidance can enable strong joint camera and motion control without training. Project page: https://elkhomar.github.io/actcam/.
Omar El Khalifi, Thomas Rossi, Oscar Fossey +6
Kinetix, France · University of Oxford, United Kingdom · MBZUAI, United Arab Emirates