Organizations: School of Computer Science, Wuhan University · Ecole Polytechnique F´ed´erale de Lausanne (EPFL), Switzerland · Department of Computer Science and Engineering, University of North Texas · Australian National University
Embodied agents benefit from 3D environments that combine visual fidelity to real-world observations with physical interactivity. Existing single-image tabletop reconstruction methods recover plausible scene geometry but typically represent objects as monolithic rigid bodies, limiting interaction to whole-object rigid motion and precluding executable part-level articulation. Meanwhile, recovering a scene layout consistent with the input view remains challenging because a single observation may admit multiple plausible pose-scale configurations. We present ArticuTable, a single-image 3D tabletop reconstruction framework that recovers both executable part-level articulation and an input-view-consistent scene layout. For object modeling, we introduce generation-robust articulation modeling (GRAM), which combines joint fitting guided by a multimodal large language model with semantic state reasoning to recover reliable joint parameters and valid motion ranges from imperfect monolithic proxy meshes, thereby converting them into executable articulated assets. For scene layout, we introduce progressive semantic-geometric scene registration (PSGSR), which progressively narrows the pose-scale search space under complementary metric, planar, and input-view constraints and resolves orientation ambiguity through structure-aware semantic correspondences, yielding a scene layout consistent with the input view. We further contribute ArticuTable-100, a curated collection of 100 simulation-ready tabletop scenes. Extensive evaluation, including a user study, demonstrates strong performance across visual fidelity, input-view consistency, articulation quality, physical plausibility, and simulation readiness.
Figures & tables
Figure 1: Single-image reconstruction of articulated tabletop scenes. Given one RGB image (left), ArticuTable reconstructs an input-view-consistent 3D scene containing rigid and articulated objects. Renderings of the same scene under different joint configurations (right) demonstrate part-level articulation while preserving the arrangement of object instances.
Figure 2: Overview of ArticuTable. From a single RGB tabletop image, ArticuTable reconstructs object-wise 3D proxies, converts operable proxies into executable URDF assets with GRAM, and registers them into a simulation-ready scene with PSGSR.
Figure 3: Qualitative comparison of scene-level reconstruction and registration on generated and real-world inputs.
Input-View Consistency
GPT-5 Evaluation
Physical Validity
Method
LPIPS ↓
DINOv2 ↑
CLIP ↑
VF ↑
IA ↑
PP ↑
Avg. ↑
OR ↓
ColO (%) ↓
ColS (%) ↓
ACDC ( Dai et al., 2025 )
0.4447
0.5135
0.8058
3.927
1.920
3.880
3.242
4.298
1.5302
32.00
Gen3DSR ( Ardelean et al., 2025 )
0.2716
0.5704
0.8386
2.220
4.680
4.013
3.638
3.349
11.3322
66.67
MIDI ( Huang et al., 2025 )
0.4302
0.6469
0.8668
2.787
2.933
2.360
2.693
4.142
37.9430
99.33
TabletopGen ( Wang et al., 2026 )
0.3935
0.7788
0.8908
5.380
4.860
5.893
5.378
2.162
1.4488
16.00
ArticuTable (Ours)
0.2398
0.8991
0.9265
5.867
6.253
6.600
6.240
1.049
0.3819
10.67
Table 1: Scene-level results. Avg. is the mean of VF, IA, and PP, and OR is the mean overall rank. Best and second-best values are bolded and underlined.
Method
Input
Task-Training-Free
Prec. ↑
Rec. ↑
Rest mIoU ↑
Art. gIoU ↑
OC ↓
AE ( ∘ ) ↓
LE ↓
SINGAPO ( Liu et al., 2025 )
Image
–
37.60
25.40
0.182
-0.263
0.030
30.30
0.077
PAct ( Liu et al., 2026 )
Image
–
24.60
16.80
0.092
-0.434
0.056
42.90
0.188
PhysX-Anything ( Cao et al., 2025 )
Image
–
20.30
19.20
0.093
-0.459
0.064
38.80
0.123
URDF-Anything+ ( Wu et al., 2026 )
Image
–
70.70
23.90
0.260
-0.138
0.033
50.40
0.128
Articulate AnyMesh ( Qiu et al., 2025 )
Mesh
✓
89.20
42.50
0.452
0.158
0.010
21.70
0.043
URDF-Anything+ ( Wu et al., 2026 )
Mesh
–
77.50
27.40
0.267
-0.110
0.026
49.90
0.111
Table 2: Results on the Lightwheel articulation benchmark ( Li et al., 2026b ) . Best and second-best results within each input modality are bolded and underlined, respectively.
Figure 4: Robustness on raw image-to-3D meshes. Particulate versus GRAM across motion states.
Component
Metric
Enabled
Disabled
Reduction
MLLM-Guided Joint Fitting
AE ↓
17.17∘
40.06∘
57.1%
LE ↓
0.0417
0.0527
20.9%
Semantic State Reasoning
Art. PC ↓
0.2502
0.2605
3.9%
OC ↓
0.0152
0.0176
13.6%
Table 3: GRAM component ablation on Lightwheel. Enabled and disabled values are averaged over both settings of the other component; lower is better.
Variant
Top-view Stage
Input-view Stage
SASPS
LPIPS ↓
DINOv2 ↑
CLIP ↑
Point-cloud alignment only
–
–
–
0.3439
0.8102
0.9077
w/o Top-view stage
×
✓
–
0.2564
0.8702
0.9210
Top-view stage only
✓
×
×
0.3235
0.8586
0.9156
w/o SASPS
✓
✓
×
0.2543
0.8762
0.9205
Full model
✓
✓
✓
0.2398
0.8991
0.9265
Table 4: PSGSR ablation. “–” denotes a component not applicable to the corresponding variant.
Table 5: User-study rankings from 59 participants. Each method receives 590 valid rankings; best and second-best results are bolded and underlined.
Figure 6: Scene editing and physical interaction in Isaac Sim. Top: object removal, insertion, and rearrangement. Bottom: gravity settling, contact-based transport, and articulated opening. Yellow dashed ellipses highlight the edited or manipulated regions.
Figure 7: Residual projection misalignment. Despite the overall scene alignment, the reconstructed cup exhibits less projected tilt than the input cup. Red boxes identify the cups; cyan dashed lines indicate their projected axes, and green dashed lines indicate the image vertical.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Evaluation level
N
Pass
Partial
Fail
Predicted joint
527
454 (86.1%)
57 (10.8%)
16 (3.0%)
Asset
418
346 (82.8%)
55 (13.2%)
17 (4.1%)
Appendix
Table 6: Manual audit of generated articulations over the complete 150-scene evaluation set. Joint-level results include predicted joints only. Asset-level percentages are computed over the 418 judgeable assets; seven unjudgeable assets are excluded. Clear missing articulations are counted as asset-level failures.
Measure
N
Exact agreement
Cohen’s κ
Asset-level result
85
83.5%
0.476
Structure score
101
93.1%
0.474 †
Motion score
101
89.1%
0.659 †
Appendix
Table 7: Inter-rater agreement on the random 20% re-evaluation sample. The dagger denotes quadratically weighted Cohen’s κ .
Configuration
AE ( ∘ ) ↓
LE ↓
Art. PC ↓
OC ↓
Neither component
40.06
0.0527
0.2556
0.0177
MLLM-Guided Joint Fitting only
17.17
0.0417
0.2653
0.0175
Semantic State Reasoning only
40.06
0.0527
0.2552
0.0149
Both components
17.17
0.0417
0.2452
0.0154
Appendix
Table 8: Full GRAM component ablation on Lightwheel. Lower is better for all metrics.
Figure 8: User-study interface. The website interface used to present method outputs and collect participant rankings.
Figure 9: Additional qualitative results of articulated tabletop-scene reconstruction. The first column shows the input images, while the remaining columns show the corresponding reconstructed scenes under different joint configurations.
Stage
Time (s) ↓
Table and camera initialization
8.3±2.4
Point-cloud alignment
14.1±2.0
Top-view registration
15.7±2.6
Input-view registration
103.3±11.3
Total PSGSR
141.4±14.0
Appendix
Table 9: PSGSR runtime. Per-scene wall-clock time over 150 scenes, reported as mean ± standard deviation.
Variant
Initialization
Candidate Rendering
Time (s) ↓
Relative
Full PSGSR
Top-view-guided
Batched
179.4
1.00×
Front-only SASPS
Input-view only
Batched
374.8
2.09×
Front-only SASPS
Input-view only
Sequential
595.3
3.32×
Appendix
Table 10: Pose-selection efficiency. Runtime averaged over the two scenes in Fig. 5 .
Setting
Configuration
Package format
Self-contained USDZ
Stage units
1 unit =1 m
Up axis
+Z
Gravity
9.81m/s2 along −Z
Visual representation
Original geometry with PBR materials and textures
Static table collision
Original unsimplified triangle mesh
Appendix
Table 11: Configuration used to export the simulation-ready USDZ packages.
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China · State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences, Beijing, China · D-Robotics, Beijing, China +2