Understanding interaction in a 3D scene requires recovering movable parts, their motion, and where they can be operated. These quantities are related, and their predictions can inform one another. A closed cabinet door, for instance, reveals a movable surface but may leave the hinge side ambiguous; its handle helps resolve this ambiguity, while the part provides context for localizing and interpreting the small handle. Building on this observation, we present SEGMENT-SNAP, which combines geometric and semantic evidence through part-handle coupling. Three independently trained predictors recover movable parts, dense handles, and part-associated handle proposals. We couple their outputs in two directions. For part motion, a training-free geometric decoder fits predicted part surfaces under explicit physical priors and uses detected handles to select candidate hinge lines. For handle prediction, a part-conditioned branch proposes additional handles, while standalone part classes refine their rotation/translation labels, with dense-handle labels as a fallback. Each transfer is applied once, without iterative feedback. On the Articulate3D validation set, handle guidance raises motion-gated AP from 13.74 to 40.98 under fixed masks and axes. Additional handle proposals raise handle AP from 24.63 to 29.65, and full contextual class correction raises it to 30.99 in the reference configuration. Fixed-input controls, retraining ablations, learned-decoder comparisons, and paired visualizations together characterize the benefits and limits of this coupling. Our system also achieved first place in the Articulate3D Challenge.
Figures & tables
Figure 1 : Three predictors, two directed transfers. The same RGB point cloud and normals enter independently trained part, dense-handle, and joint part-handle predictors. Standalone parts supply all final part masks and classes. Dense handles provide locations to a training-free motion decoder, which fits part surfaces, applies explicit physical priors to the axes, and selects hinges opposite the handles. The joint predictor uses its own part features to propose additional handles. Standalone parts correct child classes, with dense labels as fallback, before the children are appended to unchanged dense detections. Only the added handles’ classes are corrected; their masks and scores remain fixed. The final handles do not feed back into motion decoding.
Method
Part AP 50
+Origin
+Axis
+Origin +Axis
Handle AP 50
SoftGroup
32.7
18.5
21.5
17.7
14.5
Mask3D
39.1
24.4
33.8
19.3
30.2
USDNet
41.8
31.4
34.6
25.0
31.1
\rowours Segment–Snap
47.93
43.11
43.75
40.98
30.99
Table 1 : Articulate3D validation results (%). Baseline values are reported in the USDNet benchmark paper ( Halacheva et al., 2025 ) . Origin and axis columns add the corresponding motion gates to part AP 50 .
+Origin +Axis
Method
Input
Part AP 50
All
Rotation
Translation
REACT3D
RGB-D, poses, mesh
14.46
8.09
12.29
3.89
\rowours Segment–Snap
RGB point cloud
47.93
40.98
56.37
25.59
Table 2 : Articulated-part results on all 42 validation scenes (AP, %). REACT3D receives RGB-D keyframes, camera poses, and a mesh, whereas Segment–Snap uses an RGB point cloud. Both systems are evaluated in the same point-cloud domain. The joint +Origin +Axis metric follows Tab. 1 ; class-specific results are also shown.
Figure 2 : Comparing part geometry and motion. Columns show the RGB scene, ground-truth parts and axes, Segment–Snap , and REACT3D. Purple/teal axes denote rotation/translation. Saturated ground-truth colors mark correct masks and motion; pale colors pass IoU >0.5 but fail a motion gate; gray predictions are unmatched. Missing colors reveal missed parts. These favorable scenes are selected by the displayed true-positive advantage; our display threshold is s≥0.50 , and all REACT3D predictions are shown.
Measurement
Segment–Snap
REACT3D
Median time (s/scene)
33.78
3,638.1
As-shipped peak host memory (GiB)
4.70
22.94
Logged GPU peak (MiB)
10,262
8,780
Learned models
2
5
Parameters (B)
0.201
5.398
Table 3 : Inference cost for articulated-part output. Each system is measured from its declared input through one complete pass per scene, including model loading. Host memory reports the as-shipped peak; GPU logging has unequal coverage and does not support a matched memory ranking. Appendix D specifies the measurement boundaries and REACT3D’s memory refit.
Controlled change
Before
After
Δ
Handles → parts: motion-gated AP
Centroid → handle-guided origin
13.74
40.98
+27.25
Parts → handles: handle AP 50
Dense field → append associated children
24.63
29.65
+5.01
Proposal union → parts-only class correction
29.65
30.63
+0.98
\rowours Proposal union → full contextual correction
29.65
30.99
+1.34
Table 4 : Two directed transfers on validation. Scores are AP (%); Δ is the change in percentage points, computed before rounding. Both class-correction rows start from the same proposal union, so their gains are not additive.
Motion decoder
Handle cue
AP50AO (three-seed range)
Reference training-free rule
✓
40.98
Box-axis rule with recomputed origin
✓
41.20
Continuous regression
–
17.24–18.24
Continuous regression
✓
29.61–31.30
Learned candidate selection
–
25.38–28.32
Learned candidate selection
✓
37.27–38.63
Table 5 : Matched learned motion decoders on frozen part inputs. Learned rows report minimum–maximum AP50AO scores (%) over three training seeds and receive the same predicted handle evidence when marked. All rows retain AP50=47.93 because masks are fixed. The bracket in the note gives lower and upper bounds of a paired 95% scene-jackknife confidence interval for the closest learned head minus the rule (pp); it is distinct from the three-seed score ranges.
Component
Add gain
Removal loss
Two-end difference
Per-query argmax selection
+0.13
+0.19
+0.06
Largest-CC support cleanup
−0.11
+2.95
+3.07
Dense-handle origins
+23.79
+27.25
+3.45
Connectivity rescoring
+1.49
+1.60
+0.11
Full versus naive pipeline
+28.74 pp; 95% CI [+22.12,+35.37]
Table 6 : Component changes measured from both ends of the part-motion pipeline. “Add gain” applies one component to the naive pipeline; “removal loss” is full decoder minus the result after removing that component. Their difference is a two-end non-additivity diagnostic, not a unique causal interaction. Values are changes in AP50AO (pp). The final row’s bracket gives the lower and upper bounds of the paired 95% scene-jackknife confidence interval for the full-minus-naive change, also in pp.
Proposal source
Handle AP50
Dense detections only
24.63
Dense + spatially permuted child masks
24.63
Dense + random size-matched child masks
24.63
\rowours Dense + query-associated child proposals
29.65
Table 7 : Complementary handle localization. Each configuration appends a child source below the same dense-handle detections without label correction. The corrupted-mask controls preserve child count, mask size and confidence while disrupting localization. Scores are validation AP50 percentages.
Label rule
Part cue
Dense cue
AP50
Δ
No correction
–
–
29.65
–
Parts only
✓
–
30.63
+0.98
Dense incumbents only
–
✓
30.66
+1.01
\rowours Full contextual rule
✓
✓
30.99
+1.34
Fallback only when no part speaks
✓
✓
30.99
+1.34
Part classes permuted within scene
corrupted
✓
28.62
−1.03
Table 8 : Contextual child-label controls. Child masks and scores are fixed; only labels of appended children can change. Scores are validation AP50 percentages.
Mechanism
Reference gain
Seed span
95% scene CI
Dense handles → hinge origin
+27.25
2.11 ( n=4 )
[+21.86,+32.63]
Child proposals appended to dense handles
+5.01
1.08 ( n=3 )
[+1.47,+8.56]
Contextual child-label vote
+1.34
1.15 ( n=3 )
[+0.37,+2.31]
Table 9 : Observed variation of the directed effects. All values are AP changes in pp. “Seed span” is the maximum minus minimum paired gain across n training draws, not a confidence interval. Brackets give the lower and upper bounds of the paired 95% scene-jackknife confidence interval (CI) for the reference gain. A CI containing zero leaves the direction of change unresolved.
Figure 3 : Paired hinge-origin cases on validation scenes. Each column compares centroid and handle-guided origins for the same predicted part. Black denotes the ground-truth hinge, and the colored segment depicts the perpendicular origin error. The first two columns are centroid-error ranks 61 and 109 among 121 rotations for which handle guidance passes but the centroid fails (median and 90th percentile). The remaining selections are the smallest centroid error among 37 both-pass cases, median axis error among 12 non-vertical hinges, median origin error among 19 missing-handle fallbacks, and the largest handle-guided error among 186 covered rotations. No case has a passing centroid but failing handle-guided origin in this population; the last column retains a large residual failure instead.
Figure 4 : Dense and child handle complementarity on validation scenes. The selected strata show dense successes, handles recovered only after appending child proposals, an over-extended dense prediction, and contextual child-label updates. The final panel is deliberately a relabeled child without an IoU >0.5 ground-truth match, making explicit that label accuracy is only directly assessable for the matched subset. Specimens are deterministic quantile selections within their stated strata.
Rank
Team
AP50
AP50A
AP50O
AP50AO
\rowours 1
TnG (ours)
55.38
52.34
49.76
48.28
2
coreMany
46.68
43.15
43.28
40.56
3
teamzb
47.20
44.99
38.55
37.11
4
core
44.44
41.05
39.55
36.89
5
ZZZ
47.72
44.81
37.77
35.72
6
linxii
43.52
41.35
35.47
33.60
Table 10 : Challenge leaderboard: movable-part motion result. All 11 public team entries and all four published metrics. Scores are percentages; ranking is by AP50AO .
Rank
Team
AP50
\rowours 1
TnG (ours)
34.46
2
coreMany
32.86
3
core
32.60
4
teamzb
31.10
5
ZZZ
30.69
6
coreN
28.86
Table 11 : Challenge leaderboard: handle-detection result. All eight public team entries. The board provides AP50 only; scores are percentages.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration or difference
Rotation
Translation
Macro
Part-motion centroid origin
1.88
25.59
13.74
Part-motion dense-handle origin
56.37
25.59
40.98
Dense-handle minus centroid
+54.49
+0.00
+27.25
95% paired CI
[+43.71,+65.27]
[0.00,0.00]
[+21.86,+32.63]
Handle child union
48.56
10.74
29.65
Handle contextual correction
48.54
13.45
30.99
Appendix
Table 12 : Per-class localization of the directed effects. Absolute rows are AP (%); difference rows and brackets are pp. Each bracket is the paired 95% scene-jackknife interval for the difference immediately above it. Macro AP averages the rotation and translation classes; differences use unrounded scores.
Prediction source
Without cue
With cue
Motion AP: fixed dense handles, varying part predictor
Table 13: Coupling under different perception configurations. Scores are validation AP (%); the complementary cue is fixed within each block. The random-initialization configurations also use longer training schedules, so these rows do not isolate pretraining.
Dense-model draw
Dense handles
+ Children
+ Context
Part motion
\rowours Reference
24.63
29.65
30.99
40.98
Repeat 1
23.76
27.68
28.81
40.66
Repeat 2
22.83
27.67
28.90
40.46
Repeat 3
25.16
29.70
30.82
39.51
Appendix
Table 14: Repeated dense-handle training with part and child predictions fixed. All scores are AP (%). The first three columns measure handles; the last measures part motion using that row’s dense field as the geometric cue.
Switched signal
Ranking metric
Configuration
Handle-guided origin
Context vote
Part motion AP50AO
Handle AP50
Neither signal
–
–
13.74
29.65
Handle-guided hinge only
✓
–
40.98
29.65
Contextual vote only
–
✓
13.74
30.99
\rowours Both signals (reference)
✓
✓
40.98
30.99
Appendix
Table 15 : Orthogonal one-pass coupling controls. The child-proposal union is fixed in every cell; the factorial switches only dense-handle guidance for part origins and contextual child-label voting for handles. Scores are validation percentages. Brackets in the note give lower and upper bounds of paired 95% scene-jackknife confidence intervals for each signal’s on-minus-off AP change, in pp.
Predictor / readout
Proposal union
After correction
(a) Parent conditioning with matched shared-field readout
Unconditioned child head
28.07
29.78
Conditioned child head
28.41
29.26
(b) Parent-child association on an alternate dense-model basis
Output-position association
27.72
29.18
Query-index association
28.96
30.37
Appendix
Table 16: Conditioning and association are separate comparisons. Scores are handle AP (%). Panel (a) holds the shared-field readout fixed when changing conditioning. Panel (b) holds a different dense prediction field fixed and changes child association; its absolute scores are not the reference chain.
Origin rule
Part AP 50
Axis-gated
Origin-gated
AP50AO
Centroid (no handle cue)
47.93
43.75
14.37
13.74
\rowours Handle-guided hinge
47.93
43.75
43.11
40.98
Appendix
Table 17: Fixed-input origin control (AP, %). Part masks, classes, scores, and axes are identical; only the representative-origin rule changes.
Decoder
Reference part model
Second part-model draw
Training-free rule
40.98
39.45
Learned candidate selection
38.63 [ −5.35,+0.63 ]
38.63 [ −4.54,+2.91 ]
Learned selection, box axes only
38.78 [ −5.33,+0.92 ]
39.39 [ −4.44,+4.31 ]
Rule axis + learned edge
39.36 [ −4.31,+1.05 ]
37.64 [ −4.66,+1.04 ]
Appendix
Table 18: Matched learned heads on two frozen part-model draws (joint line-distance AP, %). Brackets are paired 95% scene-jackknife intervals for learned-head minus training-free-rule AP (pp); each includes zero.
Population / diagnostic
n
Coverage
Selection given coverage
Movable parts matched by predicted mask
390
64.4%
–
Matched rotations: passing hinge candidate
1,423
94.4%
88.3%
GT rotations: passing hinge candidate
236
95.8%
88.1%
GT handles matched by child union
388
49.7%
–
Appendix
Table 19: Coverage and first-failure diagnostics on validation. Candidate coverage and selection are not end-to-end AP.
Figure 5 : Coverage and selection limits on the public validation split. Left: the movable-part census assigns each ground-truth instance to coverage, axis, origin, or fully scored outcomes. Centre: each metric’s macro AP is juxtaposed with a pooled instance-coverage fraction on a common 0–1 scale. These coverage fractions are diagnostic reachability statistics, not strict AP ceilings or quantities subtractable from AP. Right: handle coverage by ground-truth size quartile.
Figure 6 : Naive versus full geometric decoding. The top and bottom rows compare the naive and full decoders against the same ground-truth part; black denotes the annotated hinge. The first two columns are naive-error ranks 92 and 165 among 183 rotations covered by both configurations. The third is the median-size instance among three rotations covered only by the full configuration; its origin still fails. The last has the smallest naive-origin error among 38 both-pass cases, and the naive origin is closer. No translation is covered by only one configuration in the selection population. The gallery retains these exceptions rather than estimating their frequency.
Figure 7 : Dense-handle fields under source variation. The same validation crop and camera compare the reference and random-initialization configurations from Tab. 13 . Their training schedules also differ, so the comparison does not isolate pretraining. The scene has the median ground-truth-handle count (21st of 42 scenes). The random-initialized field emits more fragmented components (35 versus 14 in this scene; 1,195 versus 342 over validation). These are dense predictions before child augmentation.
Figure 8 : Hinge successes and failures. Within each pair, the top row uses a centroid origin and the bottom row uses handle guidance for the same predicted part. Upper group, left to right: centroid-error ranks 61 and 109 among 121 rotations that pass with guidance but fail with centroid origins, and the smallest centroid error among 37 both-pass cases. Lower group: median axis error among 12 non-vertical hinges, median origin error among 19 missing-handle fallbacks, and the largest guided-origin error among 186 covered rotations. Black lines denote annotated hinges. No case in this 186-instance population passes with the centroid but fails with guidance; the last column is a residual failure, not a reversal example.
Figure 9 : Where geometric components matter. Rows fix the part and camera. Dark points are ground truth, pale red the prediction, and black the annotated hinge; values are perpendicular origin errors. The upper case has the largest cleanup-induced change among 183 jointly covered rotations (its uncleaned origin is outside the view). The median-change case below is unchanged and has a horizontal hinge, violating the upright-axis prior.
Figure 10 : Distinct sources of error. Dark points are ground truth, pale red the prediction, and black the annotated hinge. Median-rank cases show uncovered rotations (25/50), uncovered translations (45/89), axis error (7/14), origin error (7/14), smallest-quartile handle misses (38/75), and over-extension (19/37). Frequencies come from the census, not this gallery.
Figure 11 : Handle guidance with different part predictors. Rows use reference, joint-model, and random-initialization parts. Each pair changes only the centroid origin (left) to dense-handle guidance (right); black marks the annotated hinge. Specimens have reference-origin-error ranks 61 and 110 among 122 rotations covered by all three sources. Guidance helps across sources. Random initialization also changes training duration and does not isolate pretraining.
Figure 12: Handle proposal and class behavior. Black points are annotated handles; colors show predictions. The top row includes two dense successes and the median-size child-only recovery. Below: the smallest recovery, largest dense over-extension ratio, matched class correction, and an unmatched corrected child. Correction cases have median annotated size in their respective strata. Selections are deterministic, not random.
Outcome
Rotation
Translation
Ground-truth parts
236
154
Predicted parts
155
44
Same-class mask match (IoU >0.5 )
61
7
No overlapping prediction
158
98
Overlap below the mask gate
16
48
Above-gate overlap, wrong class
1
1
Appendix
Table 20: REACT3D coverage and motion diagnostics on all 42 validation scenes. The four mask-outcome rows partition the ground-truth parts in each motion class. Counts describe coverage, not AP contributions.
Figure 13 : Additional qualitative comparison with REACT3D. The next two scenes under the ranking used in the main report: at least five annotated parts, ordered by displayed true-positive advantage and then motion-AP margin. Our threshold is s≥0.50 ; all REACT3D predictions are shown. Saturated colors indicate correct masks and motion, pale colors indicate mask matches that fail a motion gate, and gray indicates unmatched detections. Purple and teal axes denote rotation and translation. Input and training conditions are given in Appendix D .
Figure 14 : Motion accuracy and inference time. Full-validation motion AP versus median complete-pass time per scene (log scale), including model loading on the same RTX 5070 Ti. Segment–Snap ’s two-model part-motion path and REACT3D produce common outputs; input, training, and timing scopes are specified in Appendix D .
Figure 15 : Per-scene motion AP on all 42 validation scenes. The 41 scenes with defined AP are sorted by Segment–Snap ’s score; ∅ marks three verified empty REACT3D outputs. The scene without an annotated part (a980334473) is appended at rank 42: its AP is undefined, not zero. Segment–Snap has higher AP on 40 scored scenes and ties on one; evaluation conditions are specified in Appendix D .
Hand-object interaction (HOI) is a fundamental human behavior with broad applications in AR/VR, digital humans, and embodied interaction. Existing methods typically require predefined object geometry, object trajectories, or task-specific conditions, limiting their use with natural real-world inputs. To address this, we study a more practical problem of synthesizing 3D hand-object interaction sequences from a single RGB photograph and an open-vocabulary language instruction, and introduce PhotoHOI. PhotoHOI first uses a vision-language model to parse the input image and instruction into a structured task specification, including the interaction object, target region, and spatial relation. It then recovers a compact task-relevant 3D scene and plans a smooth collision-aware object trajectory based on the recovered object states, support relations, and surrounding scene geometry. To synthesize hand motion that generalizes to real-world photographs and unseen objects, it learns transferable task-conditioned contact and contact-conditioned grasp priors from large-scale affordance and HOI data. The grasp is further refined in a learned latent space, constraining the optimization to a plausible hand-pose manifold. Experiments on GRAB and H2O demonstrate improved contact quality and reduced penetration over representative baselines. Results on real-world photographs further demonstrate higher task success and scene consistency, together with generalization to unseen objects and open-vocabulary instructions.
Reconstructing physically plausible 3D human-scene interactions (HSI) from a single image currently presents a trade-off: optimization based methods offer accurate contact but are slow (~20s), while feed-forward approaches are fast yet lack explicit interaction reasoning, producing floating and interpenetration artifacts. Our key insight is that geometry-based human--scene fitting can be amortized into fast feed-forward inference. We present GRAFT (Geometric Refinement And Fitting Transformer), a learned HSI prior that predicts Interaction Gradients: corrective parameter updates that iteratively refine human meshes by reasoning about their 3D relationship to the surrounding scene. GRAFT encodes the interaction state into compact body-anchored tokens, each grounded in the scene geometry via Geometric Probes that capture spatial relationships with nearby surfaces. A lightweight transformer recurrently updates human meshes and re-probes the scene, ensuring the final pose aligns with both learned priors and observed geometry. GRAFT operates either as an end-to-end reconstructor using image features, or with geometry alone as a transferable plug-and-play HSI prior that improves feed-forward methods without retraining. Experiments show GRAFT improves interaction quality by up to 122% over state-of-the-art feed-forward methods and matches optimization-based interaction quality at ∼100× lower runtime, while generalizing seamlessly to in-the-wild multi-person scenes and being preferred in 64.8% of three-way user study. Project page: https://pradyumnaym.github.io/graft .
Pradyumna YM, Yuxuan Xue, Yue Chen +3
University of Tübingen · Tübingen AI Center · Westlake University +1
Synthesizing physically plausible human-scene interactions (HSI) remains a critical challenge in computer vision and the development of human avatars. Although recent generative models enable diverse motion synthesis, they suffer from an inductive bias referred to as semantic-geometric entanglement. Because spatial constraints often strongly correlate with specific actions in training data, monolithic models will learn the shortcut bias, aggressively overriding the semantic intent when faced with strict geometric cues. Furthermore, this entanglement exacerbates physical hallucinations, such as body-scene penetrations. To address these limitations, we propose DeSeG, a hierarchical framework that explicitly decouples semantic intent from geometric constraints. First, we introduce a Residual Semantic Planner that encodes textual instructions and canonicalized goal voxels into a compact latent space, enabling fine-grained semantic control independent of spatial trajectories. Second, we propose a physics regularized diffusion executor that incorporates differentiable repulsive potential fields directly into the diffusion objective, enforcing collision-aware motion generation. Extensive experiments on the Lingo dataset demonstrate that DeSeG achieves state-of-the-art performance, reducing mean scene penetration by 47% and improving semantic alignment by 29% over the SOTA baselines.
Jiakun Li, Zhe Li, Wenqiang Wu +4
1Southern University of Science and Technology · 1Wenqiang Wu1 · 1Zheng Chang1 +3