Real-world capture is heterogeneous: perspective, fisheye, and 360∘ panoramic images can coexist within a single reconstruction task, yet most feed-forward 3D reconstruction models assume perspective imagery and a uniform input representation. Recent models handling several camera types are either informed of the camera type for each view or reconstruct one image pair at a time. No single-pass method reconstructs mixed-camera tuples containing full panoramas from images alone. We present MEOW, a feed-forward system that jointly reconstructs metric pointmaps and camera poses from one N-view tuple mixing perspective, fisheye and full-panorama images, in a single forward pass from images alone: no calibration, distortion parameters, camera-type labels or poses are supplied for any view. Our guiding design philosophy is to treat heterogeneous-camera reconstruction as a data-adaptation problem rather than an architectural redesign. MEOW retains a perspective-pretrained backbone and learns heterogeneous cameras entirely from a procedural data engine, which renders each scene across a continuous manifold of camera models with exact rays and depth, and certifies covisibility for every camera-sampled training tuple. Trained on synthetic tuples only, MEOW transfers zero-shot to real captures: on heterogeneous 2D3DS tuples it achieves 79.9 mAA@30 against 53.8 for Wid3R given the camera type of every view; on our laser-scanned mixed-camera benchmark it registers every four-view mixed tuple with 79.4 AUC@30. The data engine, benchmark, and complete evaluation pipeline will be released.
Figures & tables
Figure 1: MEOW reconstructs a shared 3D scene from mixed-camera images in one forward pass without supplied camera calibration or camera-type labels. Left: Lenscope generates mixed-camera training views from procedural scenes, checks their covisibility, and provides exact rays, metric depth, and camera poses for supervision. Yellow markers indicate the selected panorama centres. Right: reconstructions of mixed-camera tuples constructed from Stanford 2D3DS and laser-scanned panoramic captures. MEOW receives images only, while Wid3R additionally receives camera-type labels. Each point cloud is independently similarity-aligned to the ground truth for visualisation; VGGT panels are magnified 2 × .
Method
Camera input at test time
Mixed tuple
Full 360 ∘
N -view pass
Poses
Metric
Wid3R
Class per view
✓
✓
✓
✓
–
CAM3R
None
(✓) pairs
✓
–
✓
–
Fisheye3R
None
(✓)
(✓)
✓
✓
–
X-Lens
Intrinsics or rays, and class
✓
–
✓
–
✓
G-ray
Ray map per view
✓
–
✓
✓
–
PanoVGGT
Panoramas only
–
✓
✓
✓
–
Table 1: Input requirements and outputs of the closest systems, read from their papers and code (all cited in Section 2 ). ✓: supported; (✓): partial or qualitative by the authors’ own account; –: not supported or not claimed. Only MEOW takes a mixed tuple with full panoramas from pixels alone and returns metric pointmaps and poses in one pass.
Figure 2: Offline, once per scene: a seed generates a procedural scene, panorama poses are placed in its free space, each pose is rendered with exact rays and depth; pairwise covisibility is computed by mutual visibility tests (lines: ckl≥0.40 ). Online, per tuple: a connected walk plans 2–8 poses and each is resampled through a sampled camera. Covisibility is recomputed (edge width; dashed below 0.25) and only connected tuples are kept.
Figure 3: One forward pass. Views from different cameras are resized to one tensor shape with their full field of view; the native aspect of each (fixed to 2:1 for detected full panoramas) enters through an MLP added to its patch tokens. Sixteen alternating global and frame attention layers, a dense head per view (circular padding for panoramas), a pose head and a scale head give rays, depth, camera-to-world poses (OpenCV axes) and metric scale in one shared world. Inputs: a real tuple, zero-shot.
Figure 4: Laser-scanned mixed tuples, zero-shot. Each row shows the four inputs, the laser ground truth and the fused pointmaps of MEOW (pixels only) and Wid3R (camera types given), placed by the scorer’s alignment. Top-down floor-to-wall slices coloured by height; triangles mark stations or predicted cameras.
Method
Camera input
RRA@30
RTA@30
mAA@30
ATE ↓
VGGT
None
32.6
47.3
10.5
1.65
π3
None
45.9
59.7
19.8
1.24
MapAnything
None
52.8
53.9
16.4
1.48
CAM3R †
None
–
–
–
–
MEOW (ours)
None
95.2
96.2
80.4
0.62
Wid3R
Class per view
96.9
86.5
54.3
0.83
Table 2: Heterogeneous 2D3DS tuples ( Armeni et al., 2017 ) : 88 tuples of 3–24 views, one forward pass. Pose metrics are computed per tuple and averaged over the 88 tuples; ATE is after Sim(3) alignment. † No multi-view weights released.
Mixed tuples (panorama + fisheye + two pinholes)
Single-camera tuples, AUC@30
Method
Camera input
RRA@30
RTA@30
AUC@30
Acc ↓
Comp ↓
NC
Panorama
Pinhole
Fisheye
DUSt3R
None
62.5
67.4
32.9
0.164
0.863
0.758
0.5
96.0
73.8
MASt3R
None
66.7
66.0
31.8
0.182
0.508
0.737
2.7
96.2
56.2
VGGT
None
45.8
63.9
19.0
0.279
0.591
0.665
1.9
97.1
55.0
π3
None
66.0
63.9
26.6
0.240
0.693
0.736
1.6
97.9
75.1
MapAnything
None
66.0
57.3
24.5
0.255
0.830
0.653
1.5
86.8
75.1
Table 3: Laser-scanned benchmark, 24 four-view tuples per track. AUC@30 is the mean over tuples; Acc and Comp are in metres after one alignment per tuple; NC is normal consistency.
2D3DS mAA@30
Laser mixed AUC@30
16:9 panoramas AUC@30
Final weights (full model)
80.4
79.4
78.6
Crop loader instead of full-field-of-view resizing
77.1
74.4
–
Aspect-ratio embedding withheld
57.2
68.1
42.4
Stage-2 run with the embedding (carried into Stage 3)
63.8
65.4
63.8
Stage-2 run without the embedding
60.6
61.7
53.9
Table 4: Ablations. Top: the final weights with one input change at a time; the crop-loader row routes panoramas by the dataset annotation, with which the full model scores as with the detector (Table 8 ). Bottom: the two Stage-2 runs, identical except for the aspect-ratio embedding (Table 10 ). Columns: per-tuple mAA@30 on the 88 heterogeneous 2D3DS tuples; laser mixed-track AUC@30; AUC@30 on 2D3DS panorama tuples squeezed to 16:9.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Camera model
Family
Sampling weight
Field-of-view parameter
OpenCV
Rectilinear with distortion
0.40
48 – 120∘
Pinhole
Rectilinear
0.15
48 – 95∘
Fisheye624
Fisheye
0.13
120 – 210∘
EUCM
Fisheye
0.07
90 – 180∘
Mei
Fisheye
0.05
140 – 200∘
Spherical crop
Spherical
0.05
110 – 300∘
Appendix
Table 5: Camera models of the manifold with their sampling weights and initial field-of-view ranges. Rectilinear draws are log-normal around 80∘ (log-space standard deviation 0.32) and clipped to the listed range; fisheye models and spherical crops are drawn log-uniformly. The parameter is diagonal for rectilinear models, nominal across the image circle for fisheye models (equidistant focal length, before distortion) and horizontal for spherical crops. Tuples in the aimed mode raise the lower bound to 90∘ , and the source-resolution floor and the covisibility repair can adjust a draw. Full panoramas span 360∘ and receive a random SO(3) content rotation.
Set
Views
Misses (panoramas not detected)
False alarms (other views flagged)
Heterogeneous 2D3DS tuples, with guard
503
0 / 189
2 / 314
Laser mixed track, with guard
96
0 / 24
0 / 72
Appendix
Table 6: Full-panorama detector audit on the two headline sets with our detector (two image cues plus the black-circle guard, thresholds fixed on training renders): panoramas missed and other views flagged as panoramas. The 2D3DS panorama tuples are also routed with the detector; single panoramas and Matterport3D are flagged as panoramas throughout and the perspective video sets not at all. Counts on the other laser tracks, the 2D3DS panorama tuples and the single panoramas will accompany the code release.
Drawn per view, 2–8 views, distinct optical centres
Covisibility control
Position prior, no overlap test
Offline baseline and angle thresholds
GT-depth overlap ≥25% before distortion
Room or trajectory grouping
Checked after the cameras are drawn (mutual projection, depth agreement); failures repaired or dropped
Appendix
Table 7: Training-data construction of the compared methods, from released code where available (X-Lens, a calibrated-rig depth method, omitted; test-time camera inputs are in Table 1 ). Of the compared methods, Lenscope draws each view’s camera from a continuous manifold, checks covisibility after the cameras are drawn, and uses no real image after the public perspective pretraining.
2D3DS tuples
Laser mixed track
Method
Input resizing and panorama flag
mAA@30
ATE ↓
AUC@30
Acc ↓
Comp ↓
MEOW
Full-FoV resizing, detector flag
80.4
0.62
79.4
0.142
0.367
MEOW
Full-FoV resizing, annotation flag
80.4
0.62
79.4
0.142
0.367
MEOW
Public crop loader, annotation flag
77.1
0.75
74.4
0.168
0.349
Stage-1 checkpoint (renders only)
Public crop loader
68.3
0.99
69.0
0.120
0.322
MapAnything
Public crop loader
16.4
1.48
24.5
0.255
0.830
Appendix
Table 8: Resizing and routing controls for the final checkpoint and the public MapAnything weights on the two mixed-camera benchmarks; 2D3DS pose metrics per tuple, averaged over tuples. The Stage-1 row is the checkpoint adapted on first-generation renders with the public loader and loss, before the full-field-of-view (full-FoV) resizing, the embedding and the camera manifold. The detector reproduces the dataset annotation on every tuple of the mixed laser track (189 of 189 panoramas found on 2D3DS, 2 of 314 other views flagged). On the single-camera laser tracks the annotation gives 71.0 / 86.6 / 73.2 AUC@30 (panorama / pinhole / fisheye) against the detector’s 71.0 / 85.7 / 73.2; on the 16 panorama tuples inside the covisibility envelope of Appendix F the final model reaches 92.6.
2D3DS tuples
2D3DS single panorama
MP3D tuples
Model
Checkpoint
AUC@30
Chamfer L1 ↓
Depth rel. ↓
Acc / Comp ↓
MEOW
Final
75.4
0.296
0.195
0.237 / 0.962
MEOW
90-epoch run
70.5
0.328
0.315
0.249 / 0.952
Wid3R (camera type given) ( Jung et al., 2026 )
Released
92.2
0.090
0.044
–
PanoVGGT † ( Guo et al., 2026 )
Released
99.9
0.063
0.028
–
MapAnything ( Keetha et al., 2026 )
Public
7.0
0.852
0.520
0.211 / 2.096
Appendix
Table 9: Real panoramas. 2D3DS tuples : same-room tuples of 2–13 panoramas, 19 cases, seven areas; Wid3R reaches RRA@30 94.0 and RTA@30 96.1, PanoVGGT 100 and 100. Single panorama : 40 panoramas of areas 5a and 5b, per-view Sim(3)-aligned Chamfer L1 in metres and median relative depth error. Matterport3D (MP3D) tuples : 18 scans, 8 panoramas per tuple, pointmap accuracy and completeness in metres. † PanoVGGT trains on 2D3DS and Matterport3D, so its 2D3DS numbers are in-domain; our model, Wid3R and the open baselines are zero-shot here. Wid3R’s published 79.9 on 2D3DS uses a different protocol and is not the number in this column.
Heterogeneous 2D3DS tuples
Laser tracks, AUC@30
Stage-2 run
RRA@30
RTA@30
mAA@30
Mixed
Panorama
With the embedding
91.9
93.3
63.8
65.4
45.0
Without the embedding
90.0
92.8
60.6
61.7
46.7
2D3DS panorama tuples, AUC@30: native 58.1 against 53.9; squeezed to 16:9, 63.8 against 53.9
Appendix
Table 10: Stage-2 runs with and without the aspect-ratio embedding. Both start from the Stage-1 checkpoint and train the Stage-2 configuration (first-generation scenes, full-field-of-view resizing, online camera sampling, 100 epochs, seed 0); the only configuration difference is the embedding. Same protocol as the main tables; 2D3DS pose metrics per tuple.
Input shape
Aspect input
90-epoch run
Earlier interpolation
Final
Native 2:1
Content aspect (2.0)
70.5
74.9
75.4
16:9 squeeze
Tensor aspect (1.78)
66.4
68.8
–
16:9 squeeze
Content aspect, from the detector (2.0)
75.9
78.8
78.6
16:9 squeeze
Embedding withheld
–
–
42.4
4:3 squeeze
Content aspect, from the detector (2.0)
71.8
–
76.2
1:1 squeeze
Content aspect, from the detector (2.0)
67.9
75.1
76.9
Appendix
Table 11: Squeezed input of real panoramas (2D3DS panorama tuples, AUC@30). The rows squeeze every panorama to the stated shape before it enters the pipeline. The panorama detector supplies the content aspect and panorama wrap in the final-model rows. At fixed 16:9 pixels, withholding the aspect-ratio embedding lowers the final model from 78.6 to 42.4. Dashes mark shapes that were not run for that checkpoint.
Set (tuples)
MEOW
MapAnything
VGGT
π3
DUSt3R
MASt3R
Replica (48)
69.0
78.4
93.5
98.5
87.1
98.0
ADT pinhole (80)
39.4
60.8
53.7
80.4
49.3
83.4
ADT fisheye (80)
38.4
41.0
42.2
46.1
28.4
34.0
Appendix
Table 12: Real perspective video with centimetre baselines, outside the room-scale tuples of the engine (AUC@30, same tuples, ground truth and scorer for every model). Rotation is solved by every model (RRA@30 of 99–100); the gap is translation direction under micro-baselines. Our row uses the final checkpoint with no view flagged as a panorama. VGGT trains on Replica and ADT.
Per-tuple wall clock, I/O included (s)
MEOW
MapAnything
DUSt3R
MASt3R
Replica, 8 views (RTX 5090)
0.66
0.66
5.54
8.36
Laser mixed track, 4 views (RTX 5090)
0.48
–
3.65
4.83
Network forward pass, N views (RTX 5090)
2
4
8
16
24
32
BF16 mean (ms)
67.5
126.9
266.8
583.0
965.8
1359.9
FP32 mean (ms)
129.2
267.7
610.9
1558.4
2867.0
4479.4
Peak memory BF16 (GB)
7.85
9.03
11.42
16.20
20.98
25.77
Appendix
Table 13: Efficiency. Top: full pipelines per tuple on one RTX 5090, including image loading, first tuple excluded as warm-up. Bottom: the network forward pass alone at 518 px (random inputs pre-built on the GPU, 30 iterations after 8 warm-up iterations, CUDA-event timing). A six-point log-log fit gives exponents 1.09 (bf16) and 1.28 (fp32); the local slope rises with N , so this is an empirical scaling and not a complexity statement.
MEOW (pixels only)
Wid3R (camera type given)
Views
Tuples
mAA@30
RRA@15
RTA@15
mAA@30
RRA@15
RTA@15
Ours ahead
3
36
87.5
99.1
100.0
53.6
91.7
64.8
35/36
4
13
85.0
98.7
100.0
56.3
87.2
72.4
13/13
5–8
28
79.7
92.9
93.0
54.0
88.8
67.3
25/28
9–14
7
67.9
83.9
81.2
56.9
83.6
70.1
6/7
15–24
4
29.3
46.7
49.2
52.1
71.1
71.2
0/4
Appendix
Table 14: Heterogeneous 2D3DS tuples by tuple size; the final model leads on 79 of the 88 tuples. Per-tuple mAA@30 and per-tuple accuracy at 15∘ , averaged over the tuples of each size range, for the final model and for Wid3R given the camera type of every view; the last column counts the tuples on which our per-tuple mAA@30 is higher. Training tuples have at most eight views.
Figure 5: Same-scene camera-model gallery from Lenscope : two poses of one procedural scene (rows) rendered through five camera models (columns) at fixed illustrative fields of view; panel titles use the descriptive model names (Brown–Conrady for OpenCV, extended unified for EUCM). The gallery illustrates the camera models; training tuples are assembled separately under the covisibility check of Section 3.2 .
Figure 6: As-fed distributions of the training stream, read from the per-tuple sampling log of one data-parallel rank of Stage 2 (438,001 tuples): (a) camera-model share, (b) the logged field-of-view parameter, one density strip per camera model (diagonal for rectilinear models, across the image circle for fisheye models, horizontal for spherical crops; full panoramas are 360∘ ), (c) tuple size, (d) source aspect ratio, (e) roll magnitude and tilt, and (f) the principal-point offset magnitude; (e) and (f) use a log density so the 10% uniform tails are visible. The fallback bar in (a) counts views delivered as unaugmented renders (Appendix B ).
Figure 7: What the network receives for one heterogeneous tuple. The public loader picks one bucket from the mean aspect of the tuple, cover-scales and centre-crops every view, so the panorama loses a third of its longitudes and the square views lose a quarter of their height. Our resizing maps every view anisotropically to the same bucket, keeps all content, and passes the aspect ratio and the panorama flag.
Figure 8: The laser-scanned benchmark. (a) The merged point cloud of the twelve registered scanner stations, numbered 0–11: a corridor chain and a meeting room. The highlighted stations and links are one mixed tuple of the benchmark (stations 5, 4, 7 and 6); the numbers on the links are the covisibility between the station panoramas along the chain (0.84, 0.68 and 0.87). (b) The four views of that tuple as delivered to the models: a full panorama, a fisheye and two pinhole views resampled from the station scans.
Figure 9: Cross-camera correspondences of the final model, obtained as mutual three-dimensional nearest neighbours between predicted world points. Rows one and two: real 2D3DS pairs with ground truth, panorama to fisheye and fisheye to perspective; yellow lines end within 3∘ of the true target, blue lines do not. Rows three and four: pairs without ground truth from a consumer panorama, a fisheye and a phone photograph of the same desk, coloured by the predicted three-dimensional agreement of the mutual match (brighter is closer).