Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes
Organizations: National University of Singapore · The Hong Kong Polytechnic University
Abstract
Understanding interaction in a 3D scene requires recovering movable parts, their motion, and where they can be operated. These quantities are related, and their predictions can inform one another. A closed cabinet door, for instance, reveals a movable surface but may leave the hinge side ambiguous; its handle helps resolve this ambiguity, while the part provides context for localizing and interpreting the small handle. Building on this observation, we present SEGMENT-SNAP, which combines geometric and semantic evidence through part-handle coupling. Three independently trained predictors recover movable parts, dense handles, and part-associated handle proposals. We couple their outputs in two directions. For part motion, a training-free geometric decoder fits predicted part surfaces under explicit physical priors and uses detected handles to select candidate hinge lines. For handle prediction, a part-conditioned branch proposes additional handles, while standalone part classes refine their rotation/translation labels, with dense-handle labels as a fallback. Each transfer is applied once, without iterative feedback. On the Articulate3D validation set, handle guidance raises motion-gated AP from 13.74 to 40.98 under fixed masks and axes. Additional handle proposals raise handle AP from 24.63 to 29.65, and full contextual class correction raises it to 30.99 in the reference configuration. Fixed-input controls, retraining ablations, learned-decoder comparisons, and paired visualizations together characterize the benefits and limits of this coupling. Our system also achieved first place in the Articulate3D Challenge.
Figures & tables
| Method | Part AP 50 | +Origin | +Axis | +Origin +Axis | Handle AP 50 |
| SoftGroup | 32.7 | 18.5 | 21.5 | 17.7 | 14.5 |
| Mask3D | 39.1 | 24.4 | 33.8 | 19.3 | 30.2 |
| USDNet | 41.8 | 31.4 | 34.6 | 25.0 | 31.1 |
| \rowours Segment–Snap | 47.93 | 43.11 | 43.75 | 40.98 | 30.99 |
| +Origin +Axis | |||||
| Method | Input | Part AP 50 | All | Rotation | Translation |
| REACT3D | RGB-D, poses, mesh | 14.46 | 8.09 | 12.29 | 3.89 |
| \rowours Segment–Snap | RGB point cloud | 47.93 | 40.98 | 56.37 | 25.59 |
| Measurement | Segment–Snap | REACT3D |
| Median time (s/scene) | 33.78 | 3,638.1 |
| As-shipped peak host memory (GiB) | 4.70 | 22.94 |
| Logged GPU peak (MiB) | 10,262 | 8,780 |
| Learned models | 2 | 5 |
| Parameters (B) | 0.201 | 5.398 |
| Controlled change | Before | After | |
| Handles parts: motion-gated AP | |||
| Centroid handle-guided origin | 13.74 | 40.98 | +27.25 |
| Parts handles: handle AP 50 | |||
| Dense field append associated children | 24.63 | 29.65 | +5.01 |
| Proposal union parts-only class correction | 29.65 | 30.63 | +0.98 |
| \rowours Proposal union full contextual correction | 29.65 | 30.99 | +1.34 |
| Motion decoder | Handle cue | (three-seed range) |
| Reference training-free rule | ✓ | 40.98 |
| Box-axis rule with recomputed origin | ✓ | 41.20 |
| Continuous regression | – | 17.24–18.24 |
| Continuous regression | ✓ | 29.61–31.30 |
| Learned candidate selection | – | 25.38–28.32 |
| Learned candidate selection | ✓ | 37.27–38.63 |
| Component | Add gain | Removal loss | Two-end difference |
| Per-query argmax selection | |||
| Largest-CC support cleanup | |||
| Dense-handle origins | |||
| Connectivity rescoring | |||
| Full versus naive pipeline | pp; 95% CI | ||
| Proposal source | Handle |
| Dense detections only | 24.63 |
| Dense + spatially permuted child masks | 24.63 |
| Dense + random size-matched child masks | 24.63 |
| \rowours Dense + query-associated child proposals | 29.65 |
| Label rule | Part cue | Dense cue | ||
| No correction | – | – | 29.65 | – |
| Parts only | ✓ | – | 30.63 | |
| Dense incumbents only | – | ✓ | 30.66 | |
| \rowours Full contextual rule | ✓ | ✓ | 30.99 | |
| Fallback only when no part speaks | ✓ | ✓ | 30.99 | |
| Part classes permuted within scene | corrupted | ✓ | 28.62 |
| Mechanism | Reference gain | Seed span | 95% scene CI |
| Dense handles hinge origin | 2.11 ( ) | ||
| Child proposals appended to dense handles | 1.08 ( ) | ||
| Contextual child-label vote | 1.15 ( ) |
| Rank | Team | ||||
| \rowours 1 | TnG (ours) | 55.38 | 52.34 | 49.76 | 48.28 |
| 2 | coreMany | 46.68 | 43.15 | 43.28 | 40.56 |
| 3 | teamzb | 47.20 | 44.99 | 38.55 | 37.11 |
| 4 | core | 44.44 | 41.05 | 39.55 | 36.89 |
| 5 | ZZZ | 47.72 | 44.81 | 37.77 | 35.72 |
| 6 | linxii | 43.52 | 41.35 | 35.47 | 33.60 |
| Rank | Team | |
| \rowours 1 | TnG (ours) | 34.46 |
| 2 | coreMany | 32.86 |
| 3 | core | 32.60 |
| 4 | teamzb | 31.10 |
| 5 | ZZZ | 30.69 |
| 6 | coreN | 28.86 |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Configuration or difference | Rotation | Translation | Macro |
| Part-motion centroid origin | 1.88 | 25.59 | 13.74 |
| Part-motion dense-handle origin | 56.37 | 25.59 | 40.98 |
| Dense-handle minus centroid | |||
| 95% paired CI | |||
| Handle child union | 48.56 | 10.74 | 29.65 |
| Handle contextual correction | 48.54 | 13.45 | 30.99 |
| Prediction source | Without cue | With cue |
| Motion AP: fixed dense handles, varying part predictor | ||
| Standalone reference parts | 13.74 | 40.98 |
| Joint-model part head | 12.26 | 37.33 |
| Random-initialization part configuration | 5.86 | 25.42 |
| Handle AP: fixed child proposals, varying dense predictor | ||
| Reference dense model | 24.63 | 29.65 |
| Dense-model draw | Dense handles | + Children | + Context | Part motion |
| \rowours Reference | 24.63 | 29.65 | 30.99 | 40.98 |
| Repeat 1 | 23.76 | 27.68 | 28.81 | 40.66 |
| Repeat 2 | 22.83 | 27.67 | 28.90 | 40.46 |
| Repeat 3 | 25.16 | 29.70 | 30.82 | 39.51 |
| Switched signal | Ranking metric | |||
| Configuration | Handle-guided origin | Context vote | Part motion | Handle |
| Neither signal | – | – | 13.74 | 29.65 |
| Handle-guided hinge only | ✓ | – | 40.98 | 29.65 |
| Contextual vote only | – | ✓ | 13.74 | 30.99 |
| \rowours Both signals (reference) | ✓ | ✓ | 40.98 | 30.99 |
| Predictor / readout | Proposal union | After correction |
| (a) Parent conditioning with matched shared-field readout | ||
| Unconditioned child head | 28.07 | 29.78 |
| Conditioned child head | 28.41 | 29.26 |
| (b) Parent-child association on an alternate dense-model basis | ||
| Output-position association | 27.72 | 29.18 |
| Query-index association | 28.96 | 30.37 |
| Origin rule | Part AP 50 | Axis-gated | Origin-gated | |
| Centroid (no handle cue) | 47.93 | 43.75 | 14.37 | 13.74 |
| \rowours Handle-guided hinge | 47.93 | 43.75 | 43.11 | 40.98 |
| Decoder | Reference part model | Second part-model draw |
| Training-free rule | 40.98 | 39.45 |
| Learned candidate selection | 38.63 [ ] | 38.63 [ ] |
| Learned selection, box axes only | 38.78 [ ] | 39.39 [ ] |
| Rule axis + learned edge | 39.36 [ ] | 37.64 [ ] |
| Population / diagnostic | Coverage | Selection given coverage | |
| Movable parts matched by predicted mask | 390 | 64.4% | – |
| Matched rotations: passing hinge candidate | 1,423 | 94.4% | 88.3% |
| GT rotations: passing hinge candidate | 236 | 95.8% | 88.1% |
| GT handles matched by child union | 388 | 49.7% | – |
| Outcome | Rotation | Translation |
| Ground-truth parts | 236 | 154 |
| Predicted parts | 155 | 44 |
| Same-class mask match (IoU ) | 61 | 7 |
| No overlapping prediction | 158 | 98 |
| Overlap below the mask gate | 16 | 48 |
| Above-gate overlap, wrong class | 1 | 1 |