cs.CVOct 4, 2026

ArticuTable: Generating Instance-Level Interactive Rigid-Articulated 3D Tabletop Scenes from a Single Image

Authors: Kai Lv, Yibo Yin, Lijun Guo, Heng Fan, Kaihao Zhang, Xingping Dong

Organizations: School of Computer Science, Wuhan University · Ecole Polytechnique F´ed´erale de Lausanne (EPFL), Switzerland · Department of Computer Science and Engineering, University of North Texas · Australian National University

Abstract

Embodied agents benefit from 3D environments that combine visual fidelity to real-world observations with physical interactivity. Existing single-image tabletop reconstruction methods recover plausible scene geometry but typically represent objects as monolithic rigid bodies, limiting interaction to whole-object rigid motion and precluding executable part-level articulation. Meanwhile, recovering a scene layout consistent with the input view remains challenging because a single observation may admit multiple plausible pose-scale configurations. We present ArticuTable, a single-image 3D tabletop reconstruction framework that recovers both executable part-level articulation and an input-view-consistent scene layout. For object modeling, we introduce generation-robust articulation modeling (GRAM), which combines joint fitting guided by a multimodal large language model with semantic state reasoning to recover reliable joint parameters and valid motion ranges from imperfect monolithic proxy meshes, thereby converting them into executable articulated assets. For scene layout, we introduce progressive semantic-geometric scene registration (PSGSR), which progressively narrows the pose-scale search space under complementary metric, planar, and input-view constraints and resolves orientation ambiguity through structure-aware semantic correspondences, yielding a scene layout consistent with the input view. We further contribute ArticuTable-100, a curated collection of 100 simulation-ready tabletop scenes. Extensive evaluation, including a user study, demonstrates strong performance across visual fidelity, input-view consistency, articulation quality, physical plausibility, and simulation readiness.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PhyScene3D: Physically Consistent Interactive 3D Tabletop Scene Generation

    Jun 1, 2026Weixing Chen, Zhuoqian Feng, Yang Liu +63D Scene Generation3D Generation

  2. STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics-Physics Dual System

    May 15, 2026Zhen Luo, Yixuan Yang, Xudong Xu +5Physics SimulationEmbodied Artificial Intelligence

  3. TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation

    Dec 1, 2025Ziqian Wang, Yonghao He, Licheng Yang +6Indoor Scene GenerationRobotic Data Acquisition