Hand anthropometry supports protective-glove design, but existing measurement methods often require trained operators, specialized hardware, or manual landmarking. We present HandAnthro, which estimates 44 projected hand dimensions from a smartphone photograph of a palm-up hand on US letter-size paper. The pipeline reconstructs wrist-occluded paper boundaries for rectification, whitens non-hand pixels, and refines 41 anthropometry-specific landmarks from a fine-tuned You Only Look Once (YOLO) pose model using image-specific geometry and contours. Controlled evaluation comprised 720 captures from 45 held-out participants, each contributing 16 images across two smartphones, two backgrounds, two angles, and two nominal illumination settings. HandAnthro produced complete outputs for 704 captures (97.8%); among these, mean absolute error (MAE) was 3.80 mm per dimension against two trained operators' caliper measurements. Regional MAEs were 2.48 mm for non-thumb fingers, 6.04 mm for thumbs, and 6.17 mm for palm and wrist. In a researcher-assisted mobile-app pilot, automated batch processing returned all 44 dimensions for 260 of 268 retained, researcher-screened firefighter images (97.0%). A descriptive, unpaired comparison with an independent national firefighter reference yielded a mean absolute difference of 2.40 mm across 28 sex-by-dimension group-mean contrasts. These results characterize controlled measurement performance and researcher-assisted field feasibility for future distributed hand-anthropometry studies.
Figures & tables
Method
Automation and equipment requirements
Study sample; dimensions
Reference
Hand-dimension MAE (mm)
Han and Park ( 2016 )
Automated ; required equipment: flatbed scanner and enclosure
11 people; 17 tabulated dimensions
Live-hand calipers
—
Kaashki et al. ( 2022 )
Automated ; required equipment: depth sensor and tablet
20 people; 11 real-hand dimensions
Anthropometrist; instrument unspecified
4.5
Nguyen, Le, and La ( 2025 )
Automated ; required equipment: industrial camera, light box, ring light, and calibration target
539 people; 24 dimensions
Gauge-block calibration
—
HandAnthro
Automated ; required equipment: smartphone and known-size paper
Figure 2: Occlusion-aware PR workflow: (a) input image; (b,c) initial paper and hand masks; (d,e) construction of the inpainting region; (f) completed paper mask; (g) detected paper quadrilateral (green frame); and (h) perspective-rectified RGB image.
Figure 3: Geometry-constrained correction of groupwise rotation-and-scale misalignment.
Figure 4: Perspective-rectification comparison.
Figure 5: Capture-condition accuracy and stability in 704 completed captures from 45 participants. (a) Caliper-referenced MAE averaged across 44 dimensions for each condition. (b) Absolute differences between the two condition-specific mean predictions, computed for each participant and dimension and then averaged across participants and dimensions. Both panels summarize completed outputs; seven participants had incomplete capture-condition grids.
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S1: One author’s hand under the eight test-set conditions (two devices, two backgrounds, and two capture angles; illumination pooled). Fingerprints are removed.
Figure S2: Controlled capture rig: the left clamp provides the top-down view and the right clamp the oblique view. Distances are in centimeters; the measured oblique configuration is approximately 25∘ from vertical.
Figure S3: The 41 anthropometry-specific landmarks on a palmar hand, labeled by abbreviation (Table S1 ).
Region
Abbreviation range
Count
Role
Thumb
T1 to T6
6
Thumb-specific landmarks
Index finger
I1 to I7
7
Fingertip to finger root
Middle finger
M1 to M7
7
Fingertip to finger root
Ring finger
R1 to R7
7
Fingertip to finger root
Little finger
L1 to L7
7
Fingertip to finger root
Inter-finger webbing
RO1 to RO3
3
Root-of-finger webbing points
Appendix
Table S1: The 41 anthropometry-specific landmarks predicted directly by the YOLO model, grouped by anatomical region.
Abbrev.
Expansion
Anatomical joint (palmar crease)
Formula
MFKI
Mid First Knuckle (Index)
Index-finger DIP crease
midpoint( I2 , I3 )
MFKM
Mid First Knuckle (Middle)
Middle-finger DIP crease
midpoint( M2 , M3 )
MFKR
Mid First Knuckle (Ring)
Ring-finger DIP crease
midpoint( R2 , R3 )
MFKL
Mid First Knuckle (Little)
Little-finger DIP crease
midpoint( L2 , L3 )
MFKT
Mid First Knuckle (Thumb)
Thumb IP crease
midpoint( T2 , T3 )
MSKI
Mid Second Knuckle (Index)
Index-finger PIP crease
midpoint( I4 , I5 )
Appendix
Table S2: The 14 midpoint landmarks derived from the 41 directly predicted anthropometry-specific landmarks.
Figure S4: Illustrative SAM-HQ prompts for (a) paper-mask and (b) hand-mask generation. Red points are positive prompts; blue points are negative prompts.
Figure S5: Background whitening using the warped hand mask: (a) raw hand mask; (b) detected paper quadrilateral (green); (c) rectified hand mask; (d) background-whitened image.
Item
Value
Epochs / batch size
8000 / 40
Early stopping patience
10000
Compute
CUDA GPU; dataloader workers: 8
Optimizer
auto (Ultralytics default selection)
Learning rate schedule
lr0=0.001 , lrf=0.01
Momentum / weight decay
0.937 / 5×10−4
Appendix
Table S3: Key hyperparameters and settings for YOLO fine-tuning.
Figure S6: Finger-axis estimation and fingertip correction for finger f .
Variant
Mean (mm)
Median (mm)
Δ MAE (mm / %)
Runtime (ms)
Ours
3.80
3.45
—
17
REMBG
4.26
3.79
+0.46 / +12.0%
93
Background Remover
4.26
3.82
+0.46 / +12.1%
188
CarveKit
4.71
4.48
+0.91 / +23.9%
379
Appendix
Table S4: Background-whitening alternatives under identical PR, landmark-prediction, and PP stages. MAE is measured against calipers on the 704 PR-complete captures, all of which yielded complete outputs for every variant; runtime is BW-only mean latency over the same 704 images. Positive Δ MAE indicates degradation relative to mask reuse.
Figure S7: BW-stage runtime versus caliper-referenced dimension MAE on the 704 complete captures. Runtime is measured on an NVIDIA RTX 4000 Ada Generation GPU; the horizontal axis is logarithmic. Values match Table S4 .
Figure S8: Our BW versus the three off-the-shelf baselines on representative images. REMBG, BackgroundRemover, and CarveKit often retain inter-finger shadows that blur the hand boundary, whereas our BW gives a clean separation.
Figure S9: Per-dimension MAE for E16 across 44 dimensions (704 images, 45 participants). (a) Distribution (median 3.45 mm, mean 3.80 mm). (b) Spatial map in which each line connects a dimension’s two endpoints and is colored by its MAE magnitude.
ID
Measurement
Endpoints
Mean ± SD
Median
P95
Range
Thumb (5)
D1
Thumb IP breadth
T2→T3
3.84±1.46
3.77
6.57
0.98–10.01
D2
Thumb MCP breadth
T4→T5
6.51±3.71
6.45
12.63
0.06–18.17
D3
Thumb distal-segment length
T1→MFKT
5.69±3.09
5.36
10.87
0.00–14.03
D4
Thumb proximal-segment length
MFKT→T6
7.82±4.43
7.57
15.65
0.05–23.46
D5
Thumb total length
T1→T6
6.32±4.08
6.01
13.49
0.04–20.18
Appendix
Table S5: Definitions and per-dimension absolute-error statistics for E16 on the 704 complete outputs. Endpoints use the notation in Tables S1 – S2 ; all statistics are in millimeters, and the Range column reports minimum and maximum errors.
Metric
E15 (A2)
E16 (A3)
%Δ
KP Mean (px)
11.18
7.83
−29.94%
KP Std (px)
8.73
6.74
−22.77%
Median error (px)
8.87
6.12
−30.98%
IQR (px)
9.13
6.26
−31.47%
P(error≤5px)
21.52%
39.02%
+81.3%
P(error≤10px)
57.56%
75.77%
+31.6%
Appendix
Table S6: Pixel-level comparison of E15 (A2, no PP) and E16 (A3, with PP) on 45 annotated test images (1,845 keypoint pairs). %Δ=(E16−E15)/E15 .
Group
E15 (A2)
E16 (A3)
%Δ
Index
8.57
5.59
−34.8%
Middle
8.70
6.05
−30.5%
Ring
10.81
6.90
−36.1%
Little
13.13
7.35
−44.0%
Thumb
17.35
14.54
−16.2%
Root
5.91
5.85
−1.0%
Appendix
Table S7: Mean keypoint pixel error (px) by anatomical group for E15 (A2) and E16 (A3) on 45 real test images.
Figure S10: Representative outputs across the progressive ablation configurations.
Group ( n )
A2
A3
Δ [95% CI]
Overall (44)
3.707
3.804
+0.097[0.002,0.193]
Finger length (19)
3.196
3.097
−0.100[−0.235,0.038]
Finger width (14)
2.722
2.910
+0.188[0.035,0.340]
Palm/wrist (11)
5.842
6.165
+0.323[0.167,0.494]
Appendix
Table S8: Paired PP ablation on 704 common completed captures from 45 participants. A2 and A3 columns report MAE; n counts dimensions. All values are in millimeters, and negative Δ indicates improvement. Confidence intervals use participant-cluster resampling.
Mean absolute error
Signed bias
P95 absolute error
ID
A2
A3
Δ
95% CI for Δ
A2
A3
A2
A3
D1
2.709
3.843
+1.134
[0.775,1.504]
+2.700
+3.809
4.830
6.565
D2
8.450
6.509
−1.941
[−2.389,−1.501]
+8.426
+6.324
15.552
12.629
D3
3.280
5.691
+2.410
[1.967,2.878]
−2.600
−5.540
7.641
10.866
D4
11.083
7.819
−3.264
[−4.055,−2.474]
+10.884
+7.318
18.193
15.655
D5
9.927
6.317
−3.610
[−4.720,−2.540]
+9.794
+3.336
17.745
13.486
Appendix
Table S9: Complete dimension-level PP comparison on the same 704 captures. IDs follow Table S5 , and all values are in millimeters. Δ=MAEA3−MAEA2 ; negative values indicate improvement. The paired participant-bootstrap intervals are exploratory, without multiplicity adjustment. Signed bias is the mean signed error (prediction minus caliper reference); P95 is the per-dimension 95th percentile of per-capture absolute errors, using linear interpolation.
ID
YOLO
Train aug.
Case
Pose (Train)
Pose (Val)
Without training data augmentation
E1
v8
no
A0
7.29355
3.42462
E2
v8
no
A1
0.94881
0.78883
E3
v8
no
A2
0.93568
0.83675
E4
v8
no
A3
–
–
E5
v11
no
A0
9.02002
8.49234
Appendix
Table S10: Training and validation pose loss for all 16 configurations.
Factor
A / B
NA / NB
MAE A / MAE B
Δ
Device
Android / iPhone
350 / 354
3.83 / 3.78
+0.05
Background
Complex / Simple
354 / 350
3.80 / 3.80
−0.00
Angle
Top-down / Oblique
346 / 358
4.18 / 3.44
+0.74
Appendix
Table S11: Descriptive accuracy by capture factor among complete outputs. NA and NB count images, not independent participants; Δ MAE = MAE A− MAE B , and all MAE values are in millimeters.
Within-person CV (%)
Mean absolute condition-mean difference (mm)
ID
Measurement
Median
P95
Angle
Smartphone
Background
D1
Thumb IP breadth
3.16
7.15
1.02
0.37
0.62
D2
Thumb MCP breadth
4.39
8.42
1.88
0.84
1.23
D3
Thumb distal-segment length
9.43
12.36
3.87
0.76
1.42
D4
Thumb proximal-segment length
7.14
9.80
3.77
0.85
1.52
D5
Thumb total length
7.93
10.17
7.73
1.28
2.24
Appendix
Table S12: Within-person capture variability across 44 dimensions. For each dimension, median and P95 summarize the CVs of 45 participants; each condition-mean difference is the average across participants of the absolute difference between their two condition-specific mean predictions. Summaries use the 704 completed captures. Dimension IDs and endpoints follow Table S5 .
Table S13: Largest and recurrent dimension-specific capture-angle contrasts. Positive Δd indicates higher top-down MAE; all contrasts are descriptive. Dimension definitions appear in Table S5 .
Stage
CPU (M3)
GPU (RTX 4000 Ada)
Total execution time
26.36±5.59
3.86±0.61
Perspective rectification
24.98±5.25
2.30±0.38
Background whitening
00.02±0.01
0.02±0.01
Inference and PP
01.16±0.37
1.51±0.31
Appendix
Table S14: Per-image stage runtime (mean ± SD, seconds). CPU: Apple M3, 16 successful images from a 20-image batch; GPU: RTX 4000 Ada, 704 complete outputs. BW reports mask-to-white composition only; cached-mask I/O is excluded.
Figure S11: Bland–Altman plot of signed relative differences between the two trained operators ( N=1,980 across 45 hands and 44 dimensions). Horizontal lines indicate the pooled mean bias and 95% limits of agreement.
Reference dimension
HandAnthro
Note
Hand length
D40 + D19
palm length + middle-finger length
Hand breadth
D34
across the metacarpals
Palm length
D40
wrist-midpoint to middle MCP
Palm breadth
D34
proxy (no palm-crease breadth)
Thumb length
D5
tip to thumb root
Thumb breadth
D1
IP-joint breadth
Appendix
Table S15: Mapping from Hsiao (2015) Table 2 dimensions to HandAnthro outputs.
Dimension
Ours
Ref
Δ
Δ%
g
FDR
Hand length
204.9 (10.6)
197.6 (9.3)
+7.33
+3.7
+0.77
yes
Hand breadth
100.8 (6.0)
97.2 (4.6)
+3.58
+3.7
+0.73
yes
Palm length
121.5 (6.0)
113.8 (5.8)
+7.75
+6.8
+1.33
yes
Palm breadth
100.8 (6.0)
96.0 (4.6)
+4.78
+5.0
+0.98
yes
Thumb length
76.0 (5.8)
70.8 (4.3)
+5.18
+7.3
+1.12
yes
Thumb breadth
27.7 (7.4)
24.4 (1.6)
+3.31
+13.6
+0.94
yes
Appendix
Table S16: Sex-specific comparison of 260 successful dominant-hand palmar cases (197 male and 63 female) with Hsiao et al. Ours and Ref are mean (SD) in millimeters; Δ= Ours − Ref, g is Hedges’ g , and FDR indicates Benjamini–Hochberg significance at q<0.05 .
Accurate metric-space hand pose estimation (HPE) is essential for immersive human-computer interaction and robotics. However, most existing methods predict poses in a root-relative coordinate system and cannot estimate the hand in absolute metric scale. In this work, we observe that the intrinsic proportional relationships among human hand bones encode stable anthropometric priors that implicitly correlate with the overall metric size of the hand. Leveraging this insight, we present ScaleHP, an end-to-end one-stage hand pose estimation framework that bypasses fragile extrinsic depth modules to recover the hand in metric space. ScaleHP employs a transformer-based decoder with a novel scale token to fuse multi-scale morphological and appearance features. By solving for metric coordinates through a perspective-constrained least-squares approach, we achieve high-precision pose estimation in the camera coordinate system. ScaleHP delivers state-of-the-art performance, including 35.8 CS-MPJPE on FreiHand and 4.6/5.9 PA-MPJPE on DexYCB and HO3Dv3. These results demonstrate that internal biological constraints significantly reduce relative geometry and absolute metric errors, offering a robust solution for generalized, real-world hand tracking.
Ruitao Jing, Xingyu Chen, Hongyang Li +3
Tsinghua University · Visincept · International Digital Economy Academy (IDEA Research) +2
Robust, high-fidelity 3D hand capture, while fundamental to digital human creation, remains challenging with practical multi-view systems that balance rich photometry with the geometric ambiguities of reconstruction arising from limited viewpoint density. This paper presents an end-to-end pipeline for dynamic hand performance capture and registration, specifically designed for view-efficient setups (∼20 views). We address key challenges with two primary innovations. First, to overcome reconstruction difficulties like limited view overlap and background clutter, our mask-free neural method robustly extracts detailed hand geometry and appearance from unmasked images using scene parameterization and scenario-specific density regularization. Second, addressing registration challenges such as accurately capturing non-linear skin deformations and ensuring plausible results during severe self-contact, we propose a physics-inspired framework. It aligns reconstructions to a personalized hand model by optimizing intrinsic volumetric offsets within its canonical tetrahedral mesh, alongside pose parameters. This approach, supported by robust losses and optimization, captures fine surface deformations, ensures plausible results under severe articulation and self-contact, and demonstrates strong tolerance to input noise. We demonstrate the scalability and robustness of our automated pipeline on an extensive dataset of over 12,000 sequences, from which we also derive a large-scale, high-quality synthetic 2D/3D hand dataset for training downstream tasks. This showcases its effectiveness for single hands, intricate two-hand interactions, and natural hand-object manipulations. Our method achieves state-of-the-art reconstruction fidelity in view-efficient, unmasked scenarios and highly accurate registration. Our project page are available at https://vephand.github.io/.
We introduce a novel 3D hand pose estimator that can accurately recover the shape and pose of people's hands in a room from afar, typically from fixed cameras at room corners, in extremely low-resolution and frequently occluded views. Our key idea is to fully leverage hand-body coordination, its temporal progression, and multiview observations. We achieve this with a novel Transformer-based model, in which hand and body configurations are modeled through correlations between their visual features expressed as per-view tokens, and their temporal coordination is exploited in an autoregressive manner. We introduce a novel dataset, which we refer to as REACH, Room-Environment dataset Annotated with Chest cameras for Hand pose estimation, to train and test our method. REACH is a first-of-its-kind large-scale hand pose dataset that captures accurate hand movements of 50 participants across a wide variety of daily activities. In order to avoid interfering with natural movements while annotating the hands with accurate shape and pose, we leverage concealed chest cameras. Through extensive experiments, including comparative studies with existing methods, we show that our model, REACH-Net, achieves highly accurate 3D hand pose estimation from afar. These results broaden the horizon of 3D hand pose estimation, especially towards "in-the-wild" continuous human behavior analysis.
Shu Nakamura, Ryo Kawahara, Genki Kinoshita +4
Graduate School of Informatics, Kyoto University · RIKEN · Kyoto Institute of Technology