Camera trajectories control viewpoint changes in video generation, scene reconstruction, and robotic perception. Generating them from language requires both scene geometry and target-aware framing. We introduce OmniCam, an autoregressive model that generates camera pose sequences from a single panorama and textual trajectory descriptions. Its geometry-grounded pose token learning combines three components: a panoramic point-cloud encoder for omnidirectional geometric context; hybrid absolute-rotation and relative-translation tokenization with temporally consistent quaternion signs; and separate geometric and semantic conditioning streams with an explicit 3D target anchor. We also construct OmniCaT, containing 267,700 trajectories across four camera behaviors. On the reported OmniCaT evaluation, OmniCam reduces trajectory errors by 28--47% and collision rate by 65.8% relative to GenDoP retrained on OmniCaT. Against the best baseline for each metric, the ATE and collision reductions are 43.0% and 62.3%, respectively. Component ablations support the use of geometric and target-aware conditioning, while downstream experiments examine camera-controlled video generation and robotic active perception.
Figures & tables
Figure 2 : OmniCaT Dataset Overview. Left : (a) Construction pipeline comprising (1) Geometric and Semantic Grounding to extract nav-meshes and 3D targets, (2) Spatial Trajectory Synthesis to generate diverse camera behaviors, and (3) Rendering and Multimodal Annotation to produce text-trajectory pairs. Right : Dataset statistics showing (b) task distribution, (c) top object categories, (d) motion primitive diversity, and (e) target direction distribution.
Figure 3 : OmniCam Architecture. Three parallel branches encode the caption, panoramic image, and point cloud. A Language-Guided Resampler produces target-aware semantic tokens while a Geometry Resampler yields geometric descriptors; an attention-aggregated 3D target centroid provides an explicit spatial anchor for the decoder. The autoregressive decoder generates pose tokens from the concatenated condition, which are mapped to continuous SE(3) trajectories by a de-tokenizer. Right bottom: SE(3) tokenization with 9 tokens for each pose frame.
Method
Dataset
ATE ↓
FDE ↓
RPE-R ↓
RPE-T ↓
Coll. (%) ↓
Vis. (%) ↑
Angle ( ∘ ) ↓
Evaluation on the OmniCaT benchmark
CCD [ 18 ]
Pre-trained
2.483
3.396
3.112
0.082
36.9
12.3
32.5
E.T. [ 7 ]
Pre-trained
1.983
2.368
2.151
0.069
46.2
21.3
40.2
Director3D [ 23 ]
Pre-trained
1.828
2.942
2.309
0.052
28.4
28.1
27.2
Director3D
OmniCaT
1.522
2.648
2.187
0.049
27.6
29.3
26.5
GenDoP [ 34 ]
Pre-trained
1.865
2.951
1.681
0.044
32.6
31.5
29.1
Table 1 : Comparison with representative trajectory generation baselines on OmniCaT and the public DataDoP benchmark.
Figure 4 : Qualitative comparison on representative panoramic scenes. Each column shows the trajectory produced by a specific method. The target object is indicated by a red bounding box.
Figure 5 : Qualitative comparison on panoramic scenes. Each row displays an input panorama, generated video sequences, and 3D reconstruction results. The two rightmost columns compare reconstruction outputs without and with the OmniCam trajectory; no quantitative reconstruction metric is reported here.
Figure 6 : Real-robot demonstration of OmniCam-guided active perception. OmniCam plans a camera trajectory that actively repositions the observation camera to resolve target occlusion.
Variant
ATE ↓
FDE ↓
RPE-R ↓
RPE-T ↓
Coll. (%) ↓
Vis. (%) ↑
Angle ( ∘ ) ↓
Encoder
w/o semantic branch
2.186
3.247
2.694
0.072
13.6
11.5
64.7
w/o geometric branch
1.774
3.084
2.881
0.049
45.8
33.6
19.7
Geo. Input
Panorama RGB only
1.536
2.687
2.143
0.044
38.7
36.2
18.4
Panorama RGB-D
1.247
2.118
1.685
0.036
22.6
42.8
14.7
Grounding
w/o target cross-attention
1.673
2.259
2.774
0.042
15.2
25.8
35.6
w/o target localization loss
1.582
2.148
2.661
0.038
14.8
32.4
28.7
Table 2: Ablation studies on OmniCaT. The dual-branch panoramic encoder, target injection components, and pose tokenization schemes are investigated.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Training Time
Inference Latency
#Params
GenDoP
∼ 18h (8 × H20)
0.42s / trajectory
∼ 380M
OmniCam (Base)
∼ 24h (8 × H20)
0.84s / trajectory
∼ 500M
OmniCam (Large)
∼ 40h (8 × H20)
1.20s / trajectory
∼ 1B
Appendix
Table 3: Reported computational costs for GenDoP and OmniCam Base/Large.
Method
Dataset
Wander
Target
Surround
Reconstruct
ATE ↓
FDE ↓
ATE ↓
FDE ↓
ATE ↓
FDE ↓
ATE ↓
FDE ↓
CCD [ 18 ]
Pre-trained
3.718
4.907
1.485
2.285
2.524
3.419
2.206
2.972
E.T. [ 7 ]
Pre-trained
3.151
4.083
1.735
1.945
1.580
1.640
1.467
1.804
Director3D [ 23 ]
Pre-trained
2.969
4.779
1.444
2.207
1.620
2.688
1.279
2.094
Director3D
OmniCaT
2.562
4.250
1.130
2.089
1.351
2.304
1.045
1.950
GenDoP [ 34 ]
Pre-trained
2.667
4.408
1.362
2.131
2.176
3.248
1.255
2.016
Appendix
Table 4: Per-behavior decomposition of the main result on OmniCaT. Metrics are reported separately for Wander, Target, Surround, and Reconstruct trajectories.
#Trajectories
ATE ↓
FDE ↓
RPE-R ↓
RPE-T ↓
Coll. (%) ↓
Vis. (%) ↑
27K (10%)
1.483
2.586
1.836
0.054
19.8
34.6
53K (20%)
1.206
2.080
1.528
0.042
15.8
42.6
107K (40%)
1.003
1.709
1.297
0.033
12.7
48.2
267K (100%)
0.868
1.462
1.135
0.027
10.4
52.0
Appendix
Table 5: Reported effect of nominal data scale using the base architecture. Exact training-subset counts and the fixed evaluation split require verification.
Size
#Params
ATE ↓
FDE ↓
RPE-R ↓
RPE-T ↓
Coll. (%) ↓
Vis. (%) ↑
Small
∼ 120M
1.127
1.893
1.428
0.036
13.8
44.6
Base
∼ 500M
0.868
1.462
1.135
0.027
10.4
52.0
Large
∼ 1B
0.743
1.248
0.982
0.022
8.7
56.3
Appendix
Table 6: Reported effect of model scale. The common training subset and parameter-count scope require verification against training records.
Figure 7 : Attention visualization of the Language-Guided Semantic Resampler. Each group shows (from left to right) the input panorama, the reconstructed point cloud, and the 3D attention map. The maps visualize target-related attention on reconstructed geometry for the displayed examples; they are qualitative illustrations rather than a quantitative localization evaluation.
Dropout Strategy
ATE ↓
FDE ↓
RPE-R ↓
RPE-T ↓
Coll. (%) ↓
Vis. (%) ↑
Angle ( ∘ ) ↓
No dropout
0.856
1.438
1.118
0.026
10.1
38.7
18.4
Single-variant
0.862
1.451
1.126
0.027
10.3
45.8
14.2
Full (5-variant)
0.868
1.462
1.135
0.027
10.4
52.0
10.9
Appendix
Table 7: Caption-dropout ablation. The five-variant scheme favors visibility and angle, while no dropout has slightly lower geometric errors.
Method
NavMesh/annotation access
Runtime
Coll. (%) ↓
Vis. (%) ↑
Angle ( ∘ ) ↓
Heuristic Planner
Yes (privileged)
∼ 16.5s
5.2
58.7
9.4
GenDoP [ 34 ]
No
0.42s
30.4
33.6
27.6
OmniCam (Ours)
No
0.84s
10.4
52.0
10.9
Appendix
Table 8: Privileged-information planner comparison. The planner receives reconstructed geometry, navigation meshes, and semantic annotations. OmniCam estimates geometry from the input panorama and text.
Figure 8 : Additional qualitative comparison on panoramic scenes. Extending the comparison in Figure 4 , each row shows a different scene spanning indoor, outdoor, and stylized environments, and each column shows the trajectory produced by a specific method. The examples illustrate differences in target-directed trajectories across scene categories. The trajectory captions used for each scene are listed in Table 9 .
Row
Type
Target
Caption
1
Target
door
<TASK: TARGET> <TARGET_OBJ: door> <TARGET_DIR: Back> Approach the back door to bring it into focus. The sequence concludes by arriving at the final <END_OBJ: door> position within the <END_DIR: Back> sector of the panoramic scene.
2
Target
light fixture
<TASK: TARGET> <TARGET_OBJ: light fixture> <TARGET_DIR: Front> Approach the front light fixture to bring it into focus. The sequence concludes by arriving at the final <END_OBJ: light fixture> position within the <END_DIR: Front> sector of the panoramic scene.
3
Wander
free space
<TASK: WANDER> <TARGET_OBJ: free space> <TARGET_DIR: Front> Approach the front workstation to explore the layout of the industrial corridor. The sequence concludes by arriving at the final <END_OBJ: free space> position within the <END_DIR: Front> sector of the panoramic scene.
4
Target
door
<TASK: TARGET> <TARGET_OBJ: door> <TARGET_DIR: Back> Approach the back door to bring it into focus. The sequence concludes by arriving at the final <END_OBJ: door> position within the <END_DIR: Back> sector of the panoramic scene.
5
Target
ice tunnel
<TASK: TARGET> <TARGET_OBJ: ice tunnel> <TARGET_DIR: Front> Approach the ice tunnel to bring it into focus. The sequence concludes by arriving at the final <END_OBJ: ice tunnel> position within the <END_DIR: Front> sector of the panoramic scene.
6
Wander
free space
<TASK: WANDER> <TARGET_OBJ: free space> <TARGET_DIR: Front> Approach the distant peaks to explore the mountain valley. The sequence concludes by arriving at the final <END_OBJ: free space> position within the <END_DIR: Front> sector of the panoramic scene.
Appendix
Table 9: Trajectory captions for Figure 8 . Each row corresponds to a scene in the qualitative comparison figure, listing the scene type, target object, and the hierarchical textual instruction provided to OmniCam.
Method
CLIP-S ↑
Adj.-SSIM ↑
Aest. proxy ↑
Fidelity ↑
Coverage
GenDoP [ 34 ] + WorldStereo
0.2644
0.4248
7.647
0.6924
35.8%
OmniCam (Ours) + WorldStereo
0.3018
0.4904
9.021
0.7560
38.5%
Appendix
Table 10: Downstream evaluation: camera-controlled video generation with WorldStereo [ 36 ] . Different trajectory sources are compared using the same video generation model.
Trajectory Source
Grasp SR (%) ↑
Exec. Vis. (%) ↑
π0.5 [ 16 ] (fixed camera)
56.3
47.2
GenDoP [ 34 ]
48.0
52.6
OmniCam (Ours)
72.5
71.8
Appendix
Table 11: Downstream evaluation: robotic manipulation with π0.5 . Reported grasp success (SR) and execution visibility (Vis.). The stated 200-episode setup and fixed-camera SR aggregation require reconciliation with the episode records.
Figure 9 : Qualitative downstream results across diverse panoramic scenes. Each row shows the input panorama (left), four selected frames from camera-controlled video generation using an OmniCam trajectory (middle), and 3D reconstruction results without and with the OmniCam trajectory (two rightmost columns). The examples include indoor, outdoor, and stylized environments. These visual comparisons illustrate the pipeline outputs; they do not quantify reconstruction accuracy or guarantee collision-free motion.
Method
ATE ↓
FDE ↓
RPE-R ↓
RPE-T ↓
Coll. (%) ↓
Vis. (%) ↑
Angle ( ∘ ) ↓
GenDoP (OmniCaT)
1.624 ± 0.087
2.483 ± 0.124
1.577 ± 0.068
0.038 ± 0.003
30.4 ± 1.8
33.6 ± 2.1
27.6 ± 1.4
OmniCam (Ours)
0.868 ± 0.042
1.462 ± 0.071
1.135 ± 0.053
0.027 ± 0.002
10.4 ± 1.1
52.0 ± 1.9
10.9 ± 0.8
Appendix
Table 12: Bootstrap 95% confidence intervals on OmniCaT evaluation scenes (1,000 resamples). Intervals describe evaluation-scene variability for the two listed methods.