Accurate contact modeling is fundamental to understanding hand-object interaction, yet existing contact representations are typically restricted to object surfaces and rely on hand-crafted rules to recover contact details, leading to severe penetrations and implausible results. To better exploit the rich detail in motion-capture data, we introduce Volumetric Contact (VolCo), a representation that expands surface points to a set of 3D volumetric grids. VolCo encodes 3D contact that allows precise hand part recovery, and is organized in an inherent hierarchy: local contact details within each volume and global hand geometry across all volumes. Our framework, VolCoDiff, employs two modules to capture local and global features following this hierarchy. For local contact details, we use a 3D variational autoencoder to model the possible hand configurations conditioned on the local object signed distance field (SDF). For global hand geometry, we design a prior-guided diffusion model that learns the distribution of compressed latent features aggregated from the volumetric grids. We evaluate our method on two benchmark datasets and demonstrate state-of-the-art performance in penetration and stability, indicating the capability to generate tight grasps with much less severe penetrations. Our code is available at https://github.com/chzh9311/volco.
Figures & tables
Figure 1 : 3D contact volumes vs. 2D contact maps. Previous contact maps query points only on object surfaces, which can retain only a small fraction of hand surface. Thus it requires nearest-neighbor (NN) queries to avoid penetrations, which can easily stick in local minima. Contact volumes, however, sample points around the object surfaces and is able to retain the whole local hand configuration in the volume. Consequently, contact volumes can deterministically and precisely reconstruct local hand part and guide the pose fitting without instable penetration losses.
Figure 2 : VolCo and VolCoDiff framework . Within the local volumes sampled from the object surface, the hand mesh and VolCo representation are bidirectionally transferable. In each volume, a uniform grid is sampled, and each grid point encodes the contact likelihood and CSE, forming VolCo. Conversely, local hand configurations can be reconstructed from each volume using weighted average and concatenated to recover the hand geometry. VolCoDiff operates on the latent codes extracted from VolCo via VolumeVAE, with an auxiliary hand latent vector for supervision. The generated VolCo restores high-fidelity hand meshes around the object, with the hand decoder output subsequently fitted to VolCo to produce the final grasp.
Figure 3 : VolCo hierarchy and calculation. Around each of N object surface points, a k×k×k volume grid is sampled, where each grid point carries the contact likelihood as a function of d and the CSE of its nearest hand surface point y .
Method
Corr
Contact
EPE( mm ) ↓
F@5 mm↑
F@15 mm↑
AUC ↑
ContactOpt [ 11 ]
–
Map
89.42
0.125
0.365
0.067
ContactGen [ 24 ]
Part
Map + D
59.38
0.405
0.773
0.619
ManiDext [ 44 ]
CSE
Map
11.19
0.616
0.923
0.785
Sparse ( 64×43 )
CSE
Volume
6.70 ↓ 40.1%
0.739 ↑ 20.0%
0.970 ↑ 5.1%
0.868 ↑ 10.6%
Normal ( 128×83 )
CSE
Volume
6.33 ↓ 43.4%
0.752 ↑ 22.1%
0.971 ↑ 5.2%
0.875 ↑ 11.4%
Table 1 : Quantitative results of hand reconstruction from ground truth contact on GRAB test set. ‘ ↑ ’ after a criterion means the higher the better, while ‘ ↓ ’ means the opposite. 2D contact maps use the same number of sampled points (4096) as our sparse setting, yet our contact still demonstrates much better performance across all criteria.
Figure 4 : Reconstruction from previous contact maps vs. VolCo. The rightmost column represents VolCo using the reconstructed local vertices. VolCo demonstrates the best capability to restore details.
Table 2 : Quantitative results of grasp synthesis on GRAB Dataset. ‘ ↑ ’ after a criterion means the higher the better, while ‘ ↓ ’ means the opposite. Best result in each column is marked in bold and the second best is underlined . Our method achieves the best stability, smallest PD and largest CA.
Table 3 : Quantitative results of grasp synthesis on HO3D Dataset. ‘ ↑ ’ after a criterion means the higher the better, while ‘ ↓ ’ means the opposite. Best result in each column is marked in bold. Our method shows strong adaptivity to out-of-domain objects.
Figure 5 : Qualitative comparison with previous methods [ 24 , 38 , 6 ] . Hovering fingers are marked with purple boxes while penetrations are marked in red. While previous methods tend to produce excessively loose or penetrating grasps, our method achieves a physically plausible balance.
No.
VolCo
Prior Guidance
Stability Loss
Initia- lization
SD ( cm ) ↓
PD ( cm ) ↓
IV ( cm ) ↓
CR ↑
CA ↑
1
✗
✗
✗
✗
1.72
0.42
4.64
0.92
24.7
2
✓
✗
✗
✗
1.28
0.42
2.34
1.00
19.5
3
✓
✗
✓
✗
1.24
0.39
2.11
1.00
19.1
4
✓
✓
✗
✗
1.63
0.23
4.16
1.00
26.5
5
✓
✓
✗
✓
0.78
0.26
3.61
1.00
28.5
6
✓
✓
✓
✗
1.10
0.21
3.53
1.00
30.8
Table 4 : Ablation study on key design choices. The best are marked in bold and the second best are underlined . VolCo brings obvious improvement, and the stability loss is more effective in reducing SD with geometric plausibility ensured by prior guidance and initialization.
Generator Framework
Inference time (s)
Optim. Iterations
Optim. time (s)
Total time (s)
SD (cm)
PD (cm)
GrabNet [ 34 ]
VAE
0.14
-
-
0.14
1.02
0.40
GraspTTA † [ 17 ]
VAE
0.17
1000
6.83
7.00
3.35
0.64
ContactGen [ 24 ]
VAE
0.34
1200
43.6
44.0
2.32
0.52
FastGrasp [ 38 ]
Diffusion
8.60
-
-
8.60
2.55
0.53
FAGrasp [ 6 ]
VAE
0.415
1200
39.1
39.5
0.61
0.73
Point-Contact Diff [ 23 ]
Diffusion
2.90
1000
4.69
7.59
1.72
0.42
Table 5 : Per sample runtime breakdown and performance, tested on GRAB dataset. We provide a spars version (VolCoDiff-S) to study the influence of volume resolutions. † GraspTTA is not trained on GRAB, so the SD and PD results are not strictly comparable.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : VolumeVAE architecture
Figure 7 : The UNet architecture of VolCoDiff.
Figure 8 : The initial guess, predicted VolCo, and the final grasps of several objects, with each sample shown by two viewpoints.
Figure 9 : Qualitative results on GRAB dataset (2/3)
Figure 10 : Qualitative results on GRAB dataset (3/3).
Figure 11 : Qualitative results on out-of-domain HO3D dataset (1/5).
Figure 12 : Qualitative results on out-of-domain HO3D dataset (3/5).
Figure 13 : Qualitative results on out-of-domain HO3D dataset (5/5).
Generating 3D hand-object interactions is essential for applications in robotics, XR, and synthetic data generation, where flexible controllability and strong generalization to diverse object geometries are required. However, existing methods rarely satisfy these requirements, limiting their practical applicability. We present DCGrasp, a distance-aware controllable grasp generation system built on a novel grasp energy term. This term computes Distance Profile, a signed distance from each hand vertex to the nearest object point, coupled with distance-aware weighting, effectively capturing the semantically similar hand-object interaction in near-contact regions while remaining invariant to object and hand identity. Given various controllable signals, DCGrasp first generates a Distance Profile based on a Diffusion Transformer, together with a corresponding candidate hand pose. We then refine the candidate pose through optimization, enforcing consistency between the optimized hand pose and the generated Distance Profile in near-contact regions. Our experiments show that DCGrasp produces high-quality, physically plausible grasps with flexible user control, generalizing to diverse object and hand shapes and scales. Our work establishes a robust and versatile pipeline for the synthesis of controllable 3D hand-object interactions.
Hiroyasu Akada, Jesús Pérez, Emre Aksan +6
Google · Max Planck Institute for Informatics, SIC
Dense hand contact estimation requires both high-level semantic understanding and fine-grained geometric reasoning of human interaction to accurately localize contact regions. Recently, multi-modal large language models (MLLMs) have demonstrated strong capabilities in understanding visual semantics, enabled by vision-language priors learned from large-scale data. However, leveraging MLLMs for dense hand contact estimation remains underexplored. There are two major challenges in applying MLLMs to dense hand contact estimation. First, encoding explicit 3D hand geometry is difficult, as MLLMs primarily operate on vision and language modalities. Second, capturing fine-grained vertex-level contact remains challenging, as MLLMs tend to focus on high-level semantics rather than detailed geometric reasoning. To address these challenges, we propose ContactPrompt, a training-free and zero-shot approach for dense hand contact estimation using MLLMs. To effectively encode 3D hand geometry, we introduce a detailed hand-part segmentation and a part-wise vertex-grid representation that provides structured, localized geometric information. To enable accurate and efficient dense contact prediction, we develop a multi-stage structured contact reasoning with part conditioning, progressively bridging global semantics and fine-grained geometry. Therefore, our method effectively leverages the reasoning capabilities of MLLMs while enabling precise dense hand contact estimation. Surprisingly, the proposed approach outperforms previous supervised methods trained on large-scale dense contact datasets without requiring any training. The codes will be released.
Daniel Sungho Jung, Kyoung Mu Lee
1IPAI, 2Dept. of ECE & ASRI, Seoul National University, Korea
Text driven hand object interaction (HOI) generation is gaining attention for immersive applications and robotics, yet producing physically plausible interactions remains challenging. Even when individual motions appear natural, small contact errors can cause conspicuous artifacts such as floating and interpenetration. Prior methods mitigate these issues using explicit contact cues or implicit grasp priors, but typically rely on multi stage pipelines and fail to model temporally evolving contact. We present JointHOI, a single stage diffusion framework that jointly generates 3D hand object motion and dynamic, distance based contact maps from text. By treating contact as an auxiliary inner modality, joint generation enables the model to learn contact motion coupling during training. At inference, contact guided sampling enforces consistency between generated contact maps and motion implied geometry, improving temporal stability and reducing penetration and floating. Experiments on GRAB and ARCTIC demonstrate consistent improvements in text adherence and physical plausibility over prior methods.
Mingyeong Song, Jungbin Cho, Jisoo Kim +5
Ewha Womans University, Seoul, Korea · Yonsei University, Seoul, Korea · Carnegie Mellon University, Pittsburgh PA, USA +1