Geometrically faithful and functional articulated 3D assets are essential for real-to-sim robot manipulation, where policies trained in simulation must transfer to physical objects. Recent mesh-based methods learn to infer articulation from annotated 3D assets, but deployment remains challenging when real-world objects fall outside the training distribution or their meshes are incomplete or corrupted. To address these limitations, we formulate articulated asset reconstruction as programmatic modeling grounded in partial geometric evidence and introduce USDCraft, a framework in which a pretrained LLM writes and revises executable programs for simulation-ready articulated assets without task-specific training. We propose source geometry analysis, which converts the source mesh into a metric textual description that distinguishes observed surface from unknown space, and iterative geometric rechecking, which re-encodes each candidate in the same representation so that discrepancies point to program edits while unobserved regions remain open to completion. Visual feedback and physical authoring guidance complete the modeling process, which produces articulated USD assets with explicit physical properties that load into Isaac Sim without manual adjustment. Experiments demonstrate leading articulation recovery on two benchmarks and validate USDCraft's effectiveness for real-to-sim-to-real robot manipulation.
Figures & tables
Figure 1: USDCraft enables generation, reconstruction, and real-to-sim-to-real manipulation.
Figure 2: USDCraft overview. The modeling agent refines an asset program, choosing which tools to use and in what order. Blue modules measure the source geometry (reconstruction only), purple modules are shared tools, and orange marks downstream use.
Table 3
Figure 3: USDCraft-bench reconstructions; each pair shows rest (left) and articulated (right) poses. Red rotation glyphs and yellow double arrows indicate revolute and prismatic joints, respectively.
Figure 4: Harness comparison on scanned toaster (a) and drawer (b), and generated chair (c) and game console (d). Pairs show first-run reconstructions at rest (left) and articulated (right) poses.
Table 6
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
USDCraft-bench
Chair
Laundry hamper
Stove
Clothes dryer
Manipulation assembly
Task lamp
Control knob
Microwave
Teapot
Dishwasher
Oven
Telescope
Door
Paper cutter
Toaster
Drawer
Recycling bin
Toaster oven
Appendix
Table 8: Object categories in USDCraft-bench and Lightwheel.
Metric
Mini Workflow–Astra
USDCraft–Astra
Part match
F1(%) ↑
74.48±1.00
84.83±0.88
Static geometry
gIoU ↑
0.512±0.023
0.694±0.010
PC ↓
0.070±0.005
0.052±0.002
mIoU ↑
0.646±0.013
0.784±0.009
Articulated geometry
gIoU ↑
0.495±0.022
0.670±0.009
PC ↓
0.119±0.008
0.093±0.001
Appendix
Table 9: Reconstruction consistency over three Astra runs on USDCraft-bench. Entries report mean ± standard deviation across runs.
Figure 5: Additional reconstruction comparisons. Each method shows rest and articulated poses from left to right; colors distinguish parts and joint markers follow Figure 3 .
Figure 6: USDCraft-bench reconstruction gallery. Each row shows the input image, source mesh, reconstructed appearance, and part colors with joint markers. Source and reconstruction share a camera and scale.
Figure 7: Lightwheel reconstruction gallery using USDCraft–Astra without +G. Columns follow Figure D.3 .
Figure 8: Text- and image-conditioned asset generation. Each example shows its input evidence, the generated mesh with authored appearance, and the same geometry colored by rigid part with joint markers.
Method
Input
Cache read
Output
Cost ($)
Mini Workflow–Sol
513.4
402.3
18.9
0.6878
Mini Workflow–Astra
128.3
101.0
5.7
0.4619
USDCraft–Sol
1,164.1
1,000.3
16.5
0.9704
USDCraft–Astra
560.0
508.5
3.8
0.8510
USDCraft–Claude
6,493.0
6,313.2
101.1
4.5115
USDCraft–Gemini
4,916.8
4,250.4
80.5
0.8643
Appendix
Table 12: Mean modeling-agent token usage (thousands) and estimated cost (USD) for the methods in Table 3 . Cache reads are included in input.
Figure 9: Asset preparation for robot manipulation: reference images, scanned geometry, and USDCraft reconstructions. Scan and reconstruction share the same camera and scale within each row.
Figure 10: Simulation and real-world execution of the three manipulation tasks. Each row shows four chronological video frames; timestamps refer to its own recording.
Figure 11: Fine-lattice reconstruction: the shopping cart retains its overall form but differs in local wire geometry. The overlay compares source and reconstruction in their shared coordinate frame.
Figure 12: Generated assets with flexible members, shown with native materials and part colors. The umbrella’s kinematic display uses annotated attachments.
Object
Object-specific request
Baby grand piano
Create a baby grand piano with an opening lid and keyboard cover. Include the curved case on three rigid legs, a large top lid hinged along the straight long side, a separate hinged keyboard fallboard and a hinged lid-prop stick mounted to the case. Model a complete recognizable keyboard as fixed geometry, a recessed soundboard and simplified strings. The lid, fallboard and prop are independent limited revolute joints, without a closed-loop support constraint. Start with both covers closed and prop stowed. Approximate length 1.5 m.
Compact excavator
Create a compact excavator with a three-stage digging arm. Include a fixed track-shaped undercarriage, a cab and upper platform on one vertical slew joint, a boom on a horizontal shoulder hinge, a stick on a parallel elbow hinge and a hollow toothed bucket on a wrist hinge. Add one side-hinged cab door. Author simplified hydraulic-cylinder appearances without redundant closed-loop constraints. Use limited boom, stick and bucket travel with meaningful folded and reaching poses. Tracks remain rigid decorative assemblies. Approximate chassis length 2.5 m. Start with door closed and arm in a compact collision-free resting pose.
Desktop 3D printer
Create an enclosed desktop Cartesian 3D printer. Include a rigid cubic frame, one left-hinged transparent front door, fixed transparent side panels, a print head sliding along an X carriage nested on a vertically sliding Z gantry, and a print bed sliding front to back on Y rails. Provide three independent prismatic axes plus the door hinge. Represent belts and wiring as fixed visual details rather than articulated links. Start with door closed, bed centered and nozzle safely above the bed. Approximate outer width 0.5 m.
Appendix
Table 13: Representative text prompts for the generation gallery; append the shared instructions to each row.
A bottleneck in learning to understand articulated 3D objects is the lack of large and diverse datasets. In this paper, we propose to leverage large language models (LLMs) to close this gap and generate articulated assets at scale. We reduce the problem of generating an articulated 3D asset to that of writing a program that builds it. We then introduce a new agentic system, Articraft, that writes such programs automatically. We design a programmatic interface and harness to help the LLM do so effectively. The LLM writes code against a domain-specific SDK for defining parts, composing geometry, specifying joints, and writing tests to validate the resulting assets. The harness exposes a restricted workspace and interface to the LLM, validates the resulting assets, and returns structured feedback. In this way, the LLM is not distracted by details such as authoring a URDF file or managing a complex software environment. We show that this produces higher-quality assets than both state-of-the-art articulated-asset generators and general-purpose coding agents. Using Articraft, we build Articraft-10K, a curated dataset of over 10K articulated assets spanning 245 categories, and show its utility both for training models of articulated assets and in downstream applications such as robotics simulation and virtual reality.
Matt Zhou, Ruining Li, Xiaoyang Lyu +6
University of Cambridge · University of Oxford · Nanyang Technological University
Replicating real-world environments into simulation by realistic visual representation like NeRF and 3D Gaussian Splatting (3DGS) has emerged as an effective strategy to reduce the sim-to-real gap in robot learning. However, implementing object articulation during the real-to-sim process is still a challenging task. Existing motion tracking or learning based articulation methods shows low success rates on complex kinematic structures having multiple joints. Furthermore, those methods require scan of dynamic motion of objects, which makes reconstruction process much complicated. In this work, we propose the first end-to-end pipeline that reconstructs simulation-ready assets with accurate articulation from a single static object video input through suggestion based human-in-the-loop process. Our approach exports a hybrid representation combining 3DGS for photorealistic rendering and mesh-based geometry for physical interaction. In the reconstruction process, our pipeline performs convex decomposition followed by user grouping for intuitive part segmentation, subsequently binding 3D Gaussians to the corresponding mesh parts. An Automatic Joint Suggestion Algorithm then calculates candidate joint axes from local boundary geometries and presents them to users for efficient articulated asset reconstruction. We have shown that our method achieves precise articulation results on partnet-mobility-v0 dataset and real objects. Additionally we presented a potential usage of our framework on robot learning, deploying the reconstructed assets in Unreal Engine and NVIDIA Isaac Sim, demonstrating real-time dexterous hand manipulation tasks.
Hyesung Lee, Youngseon Lee, Kyutae Lee +2
Department of Mechanical Engineering, Seoul National University, Seoul, Korea · Department of Robotics and Mechatronics Engineering, DGIST, Daegu, Korea
Physically grounded 3D assets are increasingly important for embodied AI and robotic simulation. However, most existing 3D assets lack unified physical semantics, including articulation semantics and intrinsic physical properties, required for realistic interaction. Current approaches either treat these semantics independently or rely on canonicalized object structures, limiting robustness across heterogeneous 3D assets. We present UniPhys, a scalable framework for automatically transforming raw 3D assets into simulation-ready assets with unified physical semantics. Based on UniPhys, we construct UniPhys-40K, a large-scale physically grounded dataset, together with UniPhys-Bench, a carefully verified benchmark for unified physical grounding evaluation. We further introduce UniPhysGen, a unified physical grounding model that jointly reasons over articulation semantics and intrinsic physical properties. UniPhysGen incorporates geometry-robust articulation grounding to mitigate geometric shortcut bias under heterogeneous part decompositions. Extensive experiments demonstrate state-of-the-art performance across articulation grounding and intrinsic physical property estimation tasks, while the resulting assets can be directly deployed in robotic simulation environments for realistic physical interaction. Our code and dataset will be available at https://github.com/breezexian/UniPhysGen.
Xian Li, Rong Wei, Lujie Yang +6
Zhejiang University · Manycore Tech Inc. · University of Electronic Science and Technology of China