ArticuTable: Generating Instance-Level Interactive Rigid-Articulated 3D Tabletop Scenes from a Single Image
Organizations: School of Computer Science, Wuhan University · Ecole Polytechnique F´ed´erale de Lausanne (EPFL), Switzerland · Department of Computer Science and Engineering, University of North Texas · Australian National University
Abstract
Embodied agents benefit from 3D environments that combine visual fidelity to real-world observations with physical interactivity. Existing single-image tabletop reconstruction methods recover plausible scene geometry but typically represent objects as monolithic rigid bodies, limiting interaction to whole-object rigid motion and precluding executable part-level articulation. Meanwhile, recovering a scene layout consistent with the input view remains challenging because a single observation may admit multiple plausible pose-scale configurations. We present ArticuTable, a single-image 3D tabletop reconstruction framework that recovers both executable part-level articulation and an input-view-consistent scene layout. For object modeling, we introduce generation-robust articulation modeling (GRAM), which combines joint fitting guided by a multimodal large language model with semantic state reasoning to recover reliable joint parameters and valid motion ranges from imperfect monolithic proxy meshes, thereby converting them into executable articulated assets. For scene layout, we introduce progressive semantic-geometric scene registration (PSGSR), which progressively narrows the pose-scale search space under complementary metric, planar, and input-view constraints and resolves orientation ambiguity through structure-aware semantic correspondences, yielding a scene layout consistent with the input view. We further contribute ArticuTable-100, a curated collection of 100 simulation-ready tabletop scenes. Extensive evaluation, including a user study, demonstrates strong performance across visual fidelity, input-view consistency, articulation quality, physical plausibility, and simulation readiness.
Figures & tables
| Input-View Consistency | GPT-5 Evaluation | Physical Validity | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Method | LPIPS | DINOv2 | CLIP | VF | IA | PP | Avg. | OR | (%) | (%) |
| ACDC ( Dai et al., 2025 ) | 0.4447 | 0.5135 | 0.8058 | 3.927 | 1.920 | 3.880 | 3.242 | 4.298 | 1.5302 | 32.00 |
| Gen3DSR ( Ardelean et al., 2025 ) | 0.2716 | 0.5704 | 0.8386 | 2.220 | 4.680 | 4.013 | 3.638 | 3.349 | 11.3322 | 66.67 |
| MIDI ( Huang et al., 2025 ) | 0.4302 | 0.6469 | 0.8668 | 2.787 | 2.933 | 2.360 | 2.693 | 4.142 | 37.9430 | 99.33 |
| TabletopGen ( Wang et al., 2026 ) | 0.3935 | 0.7788 | 0.8908 | 5.380 | 4.860 | 5.893 | 5.378 | 2.162 | 1.4488 | 16.00 |
| ArticuTable (Ours) | 0.2398 | 0.8991 | 0.9265 | 5.867 | 6.253 | 6.600 | 6.240 | 1.049 | 0.3819 | 10.67 |
| Method | Input | Task-Training-Free | Prec. | Rec. | Rest mIoU | Art. gIoU | OC | AE ( ∘ ) | LE |
|---|---|---|---|---|---|---|---|---|---|
| SINGAPO ( Liu et al., 2025 ) | Image | – | 37.60 | 25.40 | 0.182 | -0.263 | 0.030 | 30.30 | 0.077 |
| PAct ( Liu et al., 2026 ) | Image | – | 24.60 | 16.80 | 0.092 | -0.434 | 0.056 | 42.90 | 0.188 |
| PhysX-Anything ( Cao et al., 2025 ) | Image | – | 20.30 | 19.20 | 0.093 | -0.459 | 0.064 | 38.80 | 0.123 |
| URDF-Anything+ ( Wu et al., 2026 ) | Image | – | 70.70 | 23.90 | 0.260 | -0.138 | 0.033 | 50.40 | 0.128 |
| Articulate AnyMesh ( Qiu et al., 2025 ) | Mesh | 89.20 | 42.50 | 0.452 | 0.158 | 0.010 | 21.70 | 0.043 | |
| URDF-Anything+ ( Wu et al., 2026 ) | Mesh | – | 77.50 | 27.40 | 0.267 | -0.110 | 0.026 | 49.90 | 0.111 |
| Component | Metric | Enabled | Disabled | Reduction |
|---|---|---|---|---|
| MLLM-Guided Joint Fitting | AE | 57.1% | ||
| LE | 0.0417 | 0.0527 | 20.9% | |
| Semantic State Reasoning | Art. PC | 0.2502 | 0.2605 | 3.9% |
| OC | 0.0152 | 0.0176 | 13.6% |
| Variant | Top-view Stage | Input-view Stage | SASPS | LPIPS | DINOv2 | CLIP |
|---|---|---|---|---|---|---|
| Point-cloud alignment only | – | – | – | 0.3439 | 0.8102 | 0.9077 |
| w/o Top-view stage | – | 0.2564 | 0.8702 | 0.9210 | ||
| Top-view stage only | 0.3235 | 0.8586 | 0.9156 | |||
| w/o SASPS | 0.2543 | 0.8762 | 0.9205 | |||
| Full model | 0.2398 | 0.8991 | 0.9265 |
| Metric | ACDC | Gen3DSR | MIDI | TabletopGen | ArticuTable |
|---|---|---|---|---|---|
| Mean Rank | 4.08 | 3.75 | 3.75 | 2.24 | 1.17 |
| Median Rank | 4 | 4 | 4 | 2 | 1 |
| First-Place Rate | 0.17% | 1.02% | 1.02% | 12.20% | 85.59% |
| Top-2 Rate | 6.61% | 11.86% | 8.81% | 74.58% | 98.14% |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Evaluation level | Pass | Partial | Fail | |
|---|---|---|---|---|
| Predicted joint | 527 | 454 (86.1%) | 57 (10.8%) | 16 (3.0%) |
| Asset | 418 | 346 (82.8%) | 55 (13.2%) | 17 (4.1%) |
| Measure | Exact agreement | Cohen’s | |
|---|---|---|---|
| Asset-level result | 85 | 83.5% | 0.476 |
| Structure score | 101 | 93.1% | 0.474 † |
| Motion score | 101 | 89.1% | 0.659 † |
| Configuration | AE ( ∘ ) | LE | Art. PC | OC |
|---|---|---|---|---|
| Neither component | 40.06 | 0.0527 | 0.2556 | 0.0177 |
| MLLM-Guided Joint Fitting only | 17.17 | 0.0417 | 0.2653 | 0.0175 |
| Semantic State Reasoning only | 40.06 | 0.0527 | 0.2552 | 0.0149 |
| Both components | 17.17 | 0.0417 | 0.2452 | 0.0154 |
| Stage | Time (s) |
|---|---|
| Table and camera initialization | |
| Point-cloud alignment | |
| Top-view registration | |
| Input-view registration | |
| Total PSGSR |
| Variant | Initialization | Candidate Rendering | Time (s) | Relative |
|---|---|---|---|---|
| Full PSGSR | Top-view-guided | Batched | 179.4 | |
| Front-only SASPS | Input-view only | Batched | 374.8 | |
| Front-only SASPS | Input-view only | Sequential | 595.3 |
| Setting | Configuration |
|---|---|
| Package format | Self-contained USDZ |
| Stage units | unit m |
| Up axis | |
| Gravity | along |
| Visual representation | Original geometry with PBR materials and textures |
| Static table collision | Original unsimplified triangle mesh |