Organizations: College of AI, Tsinghua University · School of Computer Science, Chongqing University · School of Cyber Science and Engineering, Huazhong University of Science and Technology · College of Design and Engineering, National University of Singapore
Robot learning in simulation depends on the objects the simulator offers. Many tasks need objects with separate parts, joints that allow the required motion, and physical properties that remain valid under contact. Existing methods recover this structure anew for every image: generative models predict parts and joints that mostly fail to settle or move in simulation, and general-purpose agents need a long session of model calls for each photograph. AffordCraft builds such an asset from a single RGB image and a task instruction by retrieval instead of generation: it locates the object and the part to operate, selects a matching entry from a library of articulated assets, and fits it to the image while keeping its parts and joints intact. Without any box or mask marking the object, AffordCraft produces a physically valid asset for 1,703 of 2,000 photographs from 31 categories. Five generative methods pass on at most 45% of the same photographs and, at the median, need 10 to 78 times our GPU time per valid asset. On 50 cluttered images, 162 of 237 annotated objects pass the same physical test after automatic detection. Growing the library from 141 to 11,372 entries needs no change to the method and raises category coverage from 46% to 100% and the share of selections with the requested label from 18% to 51%. We also build manipulation tasks from the constructed assets, both with single objects and in composed scenes; policies trained on scripted demonstrations complete both kinds of tasks from initial states unseen in training.
Figures & tables
Figure 1: AffordCraft overview. (a) A cluttered photograph with the objects to build; the task names the microwave (orange). (b) Retrieval ranks library entries by appearance, and a multimodal model picks the fourth by its parts and joints. (c) Built assets, operated part driven and tinted. (d) The assets composed into one scene, where a scripted teacher opens the microwave door (right: zoom). (e) Accepted share and GPU minutes per accepted asset, as in Table 1 , native where it applies and adapted otherwise (last two rows: 200-input subset).
Figure 2: Construction pipeline. The construction check (5) and the final validation (6) apply the same gate (Equation 5 ) in separate processes; “next candidate” draws from the stage-2 pool.
Export and physical gate (%)
Per input
Method
N
Export
Native
Adapted
Artic.
Median time (s)
Pass in 5 min (%)
GPU min per pass
Peak GiB
Generative reconstruction or mesh generation from one RGB image
PhysX-Anything
2,000
97.85
44.55
30.10
16.35
303.3
34.4
11.3
19.6
PhysX-Omni
2,000
98.05
36.50
26.60
16.35
464.6
15.6
21.2
60.5
PAct
2,000
94.60
23.70
16.80
14.65
108.8
15.6
7.7
21.8
PartCrafter
2,000
99.95
–
6.10
0.00
230.0 ‡
3.1 ‡
62.8 ‡
13.6
Table 1: Single-photograph construction from the full image. Native: passes the physical gate as delivered; Adapted: passes after our common physical adaptation; Artic.: accepted with a movable joint. Per-input columns: clean timing run on 32 subset inputs with one RTX 5090 per worker, AffordCraft with its library’s collision geometry built once (Appendix A.9 ). ‡ : adapted path (no physical parameters); \lx@sectionsign : timed in the method’s own API run; –: not applicable; blank: not measured.
Figure 3: Constructed assets in simulation. Rows 1–3: eighteen inputs with the grounded box, and each constructed asset after a joint drive moved its operated part (tinted). Row 4: six of these assets operated by the robot in recorded episodes (Section 4.5 ). Row 5: photograph, PhysX-Omni output, and AffordCraft asset for three inputs, at rest in the same simulator.
Figure 4: Outcomes of all 2,000 inputs. Inputs per category, stacked by physical pass and no structured export. Inset: the 297 inputs without an asset by terminal reason, and the 5,874 candidate attempts by outcome.
Figure 5: Accepted assets against the time budget per input. Share of the timed inputs whose asset passes the physical gate within a wall-clock budget: 32 inputs for the local methods, 200 for the two API routes. Dotted line: the five minutes of Table 1 (Appendix A.9 ).
Configuration
Pass
Correct
Rate (%)
Δ (pp)
Gain
Loss
p
AffordCraft (full method)
169
169
84.5
–
–
–
–
Encoder replacement (CLIP for DINOv2)
176
173
86.5
+ 2.0
7
3
0.34
Without task condition
173
125
62.5
− 22.0
3
47
3.7×10−11
Without multimodal selection
175
173
86.5
+ 2.0
4
0
0.12
Without scale adaptation
167
166
83.0
− 1.5
0
3
0.25
Without articulation adapter
93
89
44.5
− 40.0
0
80
1.7×10−24
Table 2: Paired one-factor study on the 200-input subset. Correct: passes whose asset has the requested category; Rate, Δ , Gain, and Loss refer to it (Gain/Loss: inputs correct only with the variant or only with the full method, shaded). p : exact two-sided McNemar test.
Figure 6: Library growth. Five nested libraries answer the same 2,000 queries. Left: category coverage and top-1 agreement (%). Right: median and P95 retrieval latency (ms).
Figure 7: Recorded episodes in the ten scenes , four frames of one successful episode each. Left column: the trained head on unseen initial states; right column: the scripted teacher (Appendix A.11 ).
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Window
Door
Storage
Total
Configuration
(263)
(145)
(140)
(548)
Rate (%)
AffordCraft (full method)
157
118
136
411
75.0
Without installation evidence
157
118
136
411
75.0
Without decomposition cascade
157
115
136
408
74.5
Appendix
Table 3: Installation evidence and the decomposition cascade on all 548 Window, Door, and StorageFurniture inputs, the categories these two parts of the method were written for. Cells are physical passes, and the three category columns sum to the total. Same study and paired design as Table 2 .
Figure 8: All 31 source categories under the full-RGB protocol. Each row gives passes over inputs, the physical pass rate as a bar with its Wilson 95% interval, and the rate in percent; categories are in descending input count, and every input without a structured export counts as a failure. Small categories carry wide intervals, and the 1,703 passes sum over all rows.
Figure 9: Thirty campaign assets. Each group: the input photograph with the grounding box, the delivered asset at rest, and the same asset with its driven joint moved, rendered from one camera in Isaac Sim; the link the drive moves is tinted in both states. Rows 1–6 are the assets of Figure 3 . Table 4 lists inputs, entries, scales, and joint travel; the bottle is rigid, so its two states coincide.
Asset
Category
Input
Entry
Role
Scale
Joint
Travel
Door
Door
OI 8d64efb8
9288
mounted
1.05
revolute
70 ∘
Window
Window
OI 5476ddec
103312
mounted
0.78
prismatic
-36 cm
Cabinet
Storage
OI 01c07a25
46277
free
0.91
revolute
85 ∘
Drawers
Storage
OI 26234158
46130
free
0.74
prismatic
26 cm
Fridge
Refrigerator
OI a10ea74b
10849
free
0.47
revolute
100 ∘
Microwave
Microwave
OI 26315460
7310
free
0.38
revolute
85 ∘
Appendix
Table 4: Assets of the gallery figure. Rows follow the figure. Input: Open Images id prefix (OI) or the COCO image and detection index of a cluttered-image input. Entry: the PartNet-Mobility id or the Objaverse id prefix of the selected library entry. Role: free-standing or mounted support. Scale: the uniform scale the adapter applied. Joint and travel: type of the driven joint and its displacement between the two rendered states, read back from the simulator and signed along the joint axis; the bottle is a rigid entry with no joint.
Configuration
Pass / 2,000
Rate (%)
Single-candidate construction
Single-candidate + adaptation
685
34.25
+ scale estimation
1,043
52.15
+ selected scale hypotheses
1,172
58.60
Candidate selection and repair
Multimodal candidate selection
887
44.35
Appendix
Table 5: Development configurations (target-assisted protocol). Each row received the registered target region and category label. All 2,000 inputs are retained. Rows are successive development versions, not one-factor ablations; the last row reuses the retrieval above it and adds the articulation adapter.
Configuration
Exported
Passed
Rate (%)
All-pass images
Crop-only control
34
29
12.2
0
Category-conditioned
237
236
99.6
49
Multimodal selection
192
191
80.6
19
Appendix
Table 6: Cluttered images under the target-assisted protocol. Three development configurations on the 237 annotated instances, each given the registered target crop. Object counts use N=237 ; image counts use N=50 . These rates measure construction from a located target and are not comparable with the detection-based result of Section 4.3 .
Library
Entries
Coverage (%)
Top-1 (%)
Median (ms)
P95 (ms)
Base
141
46.15
17.50
44.89
68.38
More instances
418
46.15
22.30
44.96
68.47
More mechanisms
472
46.15
24.75
44.94
68.52
More categories
614
100.00
30.30
45.00
68.41
Full
11,372
100.00
51.45
50.13
74.43
Appendix
Table 7: Library growth and retrieval diagnostics. Every setting answers the same 2,000 queries. Coverage: fraction whose category the library holds. Top-1: fraction whose selected asset carries the requested label. Latency covers retrieval only, excluding adaptation and validation.
Category
All 2,000
Subset
Category
All 2,000
Subset
In.
Pass
In.
Pass
In.
Pass
In.
Pass
Table
264
231
24
24
Toilet
25
20
3
2
Window
263
157
23
11
Keyboard
24
17
3
2
Bottle
159
127
15
10
Kettle
23
23
3
3
Chair
153
111
14
10
Coffee maker
19
19
3
3
Door
145
120
13
9
Printer
19
19
3
3
Appendix
Table 8: Composition of the 200-input subset. Per category: inputs and physical passes of the full 2,000-input campaign and of the registered 200-input subset shared with the external methods (seed 20260916, at least one input per category, the rest in proportion to category size). The subset passes are a recount of the campaign records on those inputs, not a new run.
Figure 10: Qualitative comparison of all methods. Columns: the input (cropped around our grounding box for display; every method saw the whole image), AffordCraft, PhysX-Anything, PhysX-Omni, PAct, PartCrafter, TRELLIS.2, Articulate-Anything, and the GPT-6 Astra agent; rows: six registered-subset inputs of articulated categories (refrigerator, microwave, laptop, storage furniture, door, washing machine). Each asset is shown as delivered, at rest, from one camera in Isaac Sim with a uniform clay material; links that a movable joint drives are tinted (for AffordCraft, the operated link). PhysX-Anything, PhysX-Omni, PAct, and the agent appear as our gate reads their native exports, and Articulate-Anything’s URDF in the same layout; PartCrafter and TRELLIS.2 deliver meshes only, and Articulate-Anything exports nothing for row 4. Every AffordCraft asset passes the physical gate. Of the external outputs, PhysX-Anything passes in row 5 as delivered and in row 4 after our adaptation, TRELLIS.2 in rows 3 and 4 after adaptation, Articulate-Anything in rows 2, 3, and 5 after adaptation, and the agent in rows 1, 2, and 6 as delivered and in rows 2 and 3 after adaptation; every other output fails both conditions.
Figure 11: Additional qualitative comparison. Columns as in Figure 10 . Rows: for each of six further articulated categories (oven, toilet, box, kettle, window, suitcase), the first registered-subset input, in the order of the subset, for which every method delivered an asset. Every AffordCraft asset passes the physical gate. Of the external outputs, PhysX-Anything passes in rows 5 and 6 as delivered, PhysX-Omni in rows 1, 4, and 6 as delivered and in rows 4 and 6 after our adaptation, PAct in row 1 as delivered, Articulate-Anything in rows 1, 3, 4, and 5 after adaptation, and the agent in rows 1, 3, and 4 as delivered; PartCrafter and TRELLIS.2 fail every row, and no external output passes for the toilet.
Metric
Afford- Craft
PhysX- Anything
PhysX- Omni
PAct
Part- Crafter
TRELLIS.2
Articulate- Anything
GPT agent
Perception / retrieval (s)
15.6
199.0
288.3
2.0
22.2
0.7
115.2
Generation (s)
37.9
104.3
125.0
94.0
46.0
142.7
58.5
1414.3
Adaptation to our contract (s)
237.2
287.8
332.2
300.4
253.1
168.8
64.7
Physical validation (s)
2.5
2.7
2.4
4.3
13.2
E2E median (s)
40.0
303.3
464.6
108.8
230.0 ‡
393.9 ‡
349.6 ‡§
1423.0 §
Throughput (passes / worker-h)
75.0
5.3
2.8
7.8
1.0 ‡
2.2 ‡
3.6 ‡§
0.9 §
Appendix
Table 9: Matched resource accounting on the full-RGB comparison. All times come from the clean timing run (one worker per exclusive RTX 5090; the PhysX-Omni geometry stage on an RTX PRO 6000) over a 32-input sample of the registered subset, AffordCraft with the geometry of its library entries built once; pass rates from the registered inputs of each Table 1 row. Per-input medians of the stage wall times each stage process recorded; adaptation is our common physical adaptation of the delivered asset (not applicable to AffordCraft, whose generation row is its own adaptation); E2E, minutes per accepted asset, throughput and GPU time follow the native path where it applies and the adapted path ( ‡ ) otherwise. API cost: official list prices of the model providers applied to the token counts each run recorded (gpt-6-astra: OpenAI Standard input, cached-input and cache-write rates; Gemini models: Google), 0 for methods that run only local models; \lx@sectionsign : times of the API run itself.
Stage
Group
n
Median (s)
Mean (s)
Grounding
perception
2,000
5.78
5.95
Retrieval
perception
1,767
0.09
0.10
Multimodal ranking
perception
2,347
25.00
25.64
Asset construction
adaptation
4,781
93.61
225.73
Repair
adaptation
3,138
0.00
10.03
Construction physics
validation
2,316
2.91
5.42
Appendix
Table 10: Recorded timing of the formal campaign. Wall-clock seconds over the 2,000 single-object inputs, from the verified per-input records of a throughput run (concurrent workers per GPU, pre-computed geometry cache), so not exclusive-device latencies. Phase rows: median and mean over the n recorded executions of the phase; repair runs at most once per attempt, after a failed export or a failed construction check, and returned a repaired asset in 179 attempts. Group rows: per input, over the n inputs that reached the group. End to end: per input; the P90 is 1,629 s.
Asset
Task
Untrained
Trained
Blocked
Rigid objects
Bottle
push into region
0
5
–
Box
push into region
0
6
–
Camera
push into region
0
19
–
Clock
push into region
0
14
–
Display
push into region
0
14
–
Appendix
Table 11: Ten single-asset tasks. Successes out of 20 held-out episodes per task for the same head before and after training; blocked: initial states without an episode record in either condition, counted as failures.
Figure 12: Single-asset tasks. Start, midpoint, and final frames of one recorded held-out episode of the trained head per task; the mark gives that episode’s outcome, and the numbers give untrained → trained successes out of 20 held-out episodes. Crops stay fixed within each sequence.
Scene and task
Assets
Teach. /32
Held-out task ⋅ sg1 /10
Seen /6
Untr. /4
Cabinet (2-door) + cup: open far door, put cup inside
46277, fae08c03
20
0 ⋅ 10
0
0
Cabinet (1-door) + box: open door, put box inside
45623, 160d40d9
24
4 ⋅ 10
3
0
Refrigerator + cup: open door, put cup inside
10849, fae08c03
32
0 ⋅ 10
0
0
Microwave + box: open door, put box inside
7310, 160d40d9
28
1 ⋅ 9
2
0
Drawer + cup: open top drawer, put cup inside
46130, fae08c03
29
3 ⋅ 8
2
0
Drawer + two cups: open top drawer, put both inside
46130, fae08c03 × 2
29
0 ⋅ 8
0
0
Appendix
Table 12: Per-scene tasks, assets, and results. Asset ids are PartNet-Mobility entries for articulated objects and 8-character Objaverse id prefixes for rigid ones. Teacher counts are out of 32 episodes. Trained-head columns show task successes and first-sub-goal counts on held-out states (10 episodes) and task successes on seen states (6); sg1: first sub-goal reached. Eight scenes include one distractor object.
Training set
Held-out
Seen
Config.
Teach.
Roll.
Frames
Task
Sg1
Task
Sg1
3rd-person, teacher only
272
0
152,469
6/40
28/40
8/40
29/40
+ wrist, crop; round 1
267
82
205,129
9/100
67/100
8/60
43/60
Round 2 (all rollouts)
267
397
373,031
11/100
66/100
8/60
42/60
Appendix
Table 13: Training configurations. Each row is one trained head. Episodes count teacher demonstrations plus mixed rollouts after a dynamics filter. The first row used 4 held-out episodes per scene; the others use 10 held-out and 6 seen. Sg1: first sub-goal reached. Untrained heads complete 0 held-out episodes in every configuration.
Scene and instruction
Teacher task ⋅ sg1 /32
Rollouts 1 task ⋅ sg1 /20
Rollouts 2 task ⋅ sg1 /24
Ticks
Train
Cabinet (2-door) + cup
20 ⋅ 24
8 ⋅ 20
10 ⋅ 24
528
54
open the cabinet door fully and put the cup inside the cabinet
Cabinet (1-door) + box
24 ⋅ 31
14 ⋅ 19
18 ⋅ 24
573
67
open the cabinet door fully and put the box inside the cabinet
Refrigerator + cup
32 ⋅ 32
10 ⋅ 20
10 ⋅ 24
716
73
open the refrigerator door fully and put the cup inside the refrigerator
Appendix
Table 14: Instructions and data-collection rounds per scene. Under each scene stands the instruction string every teacher and policy episode of that scene received. Each round lists task successes ⋅ first-sub-goal completions over its episodes: the teacher round (32 per scene) and the two rounds of mixed rollouts driven by the first and second trained heads (20 and 24 per scene). Ticks: median length of the passed teacher episodes at 10 Hz. Train: episodes of the scene in the final training set after the exclusions described under Data and training (Table 13 , last row).
Figure 13: Held-out episodes of the trained head, twelve frames each. The five scenes the head completed on held-out states (left column of Figure 7 ), each as its first passed held-out episode at twelve evenly spaced control ticks from the start to the final tick. Upper row: the third-person observation; lower row: the wrist camera at the same ticks. Both are the 320 × 240 images the policy received.
Condition
Episodes
Task success
Rate (%)
Sub-goal 1
Scripted teacher
320
273
85.3
298
Untrained head, held-out
40
0
0.0
1
Trained head, held-out
100
11
11.0
66
Trained head, seen
60
8
13.3
42
Appendix
Table 15: Manipulation in ten composed scenes. An independent evaluator reads success from simulator state. Held-out and seen refer to initial states unseen or seen during data collection; the teacher row scores the assets, not a policy. The untrained head uses random initial weights.
After first sub-goal
Mechanism short
Transfer, cup not out
Scene
Pass
Grasp
Lifted
Not grasped
Untouched
Partial
Grasp
Lifted
Cabinet (2-door) + cup
–
8
1
1
–
–
–
–
Cabinet (1-door) + box
4
4
2
–
–
–
–
–
Refrigerator + cup
–
5
1
4
–
–
–
–
Microwave + box
1
3
5
–
–
1
–
–
Drawer + cup
3
5
–
–
1
1
–
–
Appendix
Table 16: Outcome of the 100 held-out episodes of the trained head. Ten episodes per scene, classified from the recorded trace. After the first sub-goal: the gripper closed beside the object without securing it (grasp missed), lifted it and released it outside the target volume (lifted, not placed), or never closed on it (not grasped; one of the 13 approached without closing). Before the first sub-goal: the door, drawer, or lid stayed untouched or moved partly (mechanism columns), or, in the two transfer scenes, the cup was not carried out of the drawer (grasp missed, or lifted and dropped back).
Setting
Value
Backbone
OpenVLA-OFT, LIBERO-Spatial checkpoint, frozen
Per-image feature
mean of the 56 action-token states, 4,096-d
Images per step
2 (third-person crop, wrist), 320 × 240
Feature input to head
8,192 -d (two images concatenated)
Proprioception
9 joint positions (7 arm, 2 finger)
Head
GRU, 512 units
Appendix
Table 17: Action-head configuration of the scene study. Values of the registered round-2 training run. The backbone is never finetuned; the untrained condition uses the same head at initialization.
Figure 14: Scripted-teacher episodes, twelve frames each. The five scenes the trained head never completed on held-out states (right column of Figure 7 ), each as the first passed teacher episode of collection worker 0 at twelve evenly spaced ticks. Upper row: third-person observation; lower row: wrist camera. These episodes are demonstrations the head was trained on, not policy rollouts.
Training and evaluating robot policies in the real world is costly and difficult to scale. We introduce SimFoundry, a modular and automated system for zero-shot real-to-sim scene construction from a video. SimFoundry generates sim-ready digital twins and supports object, scene, and task editing, enabling the automated generation of diverse digital cousins: affordance-preserving variations of reconstructed real-world scenes. Policies trained on SimFoundry data transfer zero-shot to challenging real tasks involving multi-step manipulation, articulated object interaction, and bimanual interaction, and its digital cousins (variations of the original scene, objects, and tasks) facilitate generalization to new real-world conditions. Across 7 manipulation tasks and 5 policy architectures, SimFoundry simulation evaluations strongly predict real-world performance, with mean Pearson correlation 0.911 and mean maximum ranking violation 0.018. When evaluating sim-trained policies zero-shot in the real world, policies trained with object, scene, and task cousins in simulation show average task success rate improvements of 17%, 21%, and 40%, respectively. Additional details at https://research.nvidia.com/labs/gear/simfoundry/ .
Nadun Ranawaka, Josiah Wong, Wei-Lin Pai +15
NVIDIA · Stanford University · The University of Texas at Austin +2
Replicating real-world environments into simulation by realistic visual representation like NeRF and 3D Gaussian Splatting (3DGS) has emerged as an effective strategy to reduce the sim-to-real gap in robot learning. However, implementing object articulation during the real-to-sim process is still a challenging task. Existing motion tracking or learning based articulation methods shows low success rates on complex kinematic structures having multiple joints. Furthermore, those methods require scan of dynamic motion of objects, which makes reconstruction process much complicated. In this work, we propose the first end-to-end pipeline that reconstructs simulation-ready assets with accurate articulation from a single static object video input through suggestion based human-in-the-loop process. Our approach exports a hybrid representation combining 3DGS for photorealistic rendering and mesh-based geometry for physical interaction. In the reconstruction process, our pipeline performs convex decomposition followed by user grouping for intuitive part segmentation, subsequently binding 3D Gaussians to the corresponding mesh parts. An Automatic Joint Suggestion Algorithm then calculates candidate joint axes from local boundary geometries and presents them to users for efficient articulated asset reconstruction. We have shown that our method achieves precise articulation results on partnet-mobility-v0 dataset and real objects. Additionally we presented a potential usage of our framework on robot learning, deploying the reconstructed assets in Unreal Engine and NVIDIA Isaac Sim, demonstrating real-time dexterous hand manipulation tasks.
Hyesung Lee, Youngseon Lee, Kyutae Lee +2
Department of Mechanical Engineering, Seoul National University, Seoul, Korea · Department of Robotics and Mechatronics Engineering, DGIST, Daegu, Korea
RGB sim-to-real for deformable manipulation has remained largely unsolved without real-world fine-tuning. We present SimWeaver, which trains zero-shot RGB VLA policies on 200 simulated demonstrations per task, reaching above 80% per-task and 91% average real-world success across 5 diverse deformable tasks including plastic-bag manipulation, without teleoperation or per-task calibration. SimWeaver combines a reliable measurement-backed simulator (SimWeaver-Sim) with an extensible asset framework supporting single-image generation(SimWeaver-Asset), a deterministic topology-aware trajectory synthesizer (SimWeaver-Syn), and a sim-to-real protocol with ISP-aware photometric augmentation (SimWeaver-Real). On silk grasping, the sim-trained policy reaches 100% under visual distribution shifts where real-data baselines drop to 9-70%, at two orders of magnitude lower per-trajectory cost. We will release SimWeaver and a representative asset subset. Project page: https://simweaver.github.io/
Wenkang Hu, Haoran Wang, Yitong Li +10
1Shanghai Jiao Tong University · 2Horizon Robotics · 3Style3D Research