As smart glasses and lightweight MR devices become increasingly practical, input remains a key challenge. The bare palm is an always-available, tactile, and proprioceptively accessible surface, but it has neither an explicit coordinate system nor embedded touch sensing. Prior on-palm systems typically expose isolated touch events, discrete regions, continuous trajectories, or task-specific gestures, limiting the palm's ability to support precise selection and gesture manipulation through a common input representation. We present PalmSpace, a wrist-worn infrared system that exposes mode-aware, body-referenced absolute input on the bare palm without per-user sensing calibration. At the interaction level, PalmSpace jointly represents contact occurrence, interaction mode, and palm-referenced absolute location; at the model level, it learns these coupled outputs through a shared real-time representation. In leave-one-participant-out evaluation with 17 participants, PalmSpace achieved 6.7 mm mean localization error, 98.9% contact detection accuracy, and 96.7% F1 for four-class interaction-state recognition. User studies further demonstrated absolute pointing and dragging, eyes-free digit input, and representative multi-finger controls including scrolling and pinch-based map manipulation. These results show that a morphologically variable bare palm can function as a transferable, mode-aware interaction surface.
Figures & tables
Figure 1 . Diagram of the core concept of PalmSpace. PalmSpace turns the palm into a structured and eyes-free interaction space. The four panels demonstrate: (1) a flip-up camera design, (2) a structured interaction space that supports complex map manipulation requiring both single-finger localization and multi-finger scroll/pinch input; (3) absolute positioning across the entire palm area including fingers, and (4) support for precision-demanding interactions like handwriting. Four panels show the PalmSpace wristband with its camera flipped up, map manipulation using single- and multi-finger palm input, tracked touch locations across the palm and fingers, and a handwriting example drawn on the palm.
Table 1 . Comparison of sensing setup, input representation, and evaluation scope across existing on-palm techniques. PalmSpace combines contact-aware continuous absolute positioning with multi-finger mode recognition without per-user sensing calibration. (‘C’: continuous, ‘D’: discrete)
Figure 2 . PalmSpace’s hardware prototype overview. Two views of the PalmSpace prototype show the infrared camera and LED hardware, followed by the wristband worn on the inner wrist with the camera module raised above the palm.
Figure 3 . Illustration of finger-palm interaction modes. (a)-(c) Single-finger modes: users operating in LY, VT, and RY modes respectively. (d) The pitch angle between the finger and palm surface can be freely adjusted. (e)-(f) Multi-finger gestures: Scroll and Pinch are performed on the palm surface. Six panels illustrate three single-finger yaw orientations, free adjustment of finger pitch, a two-finger scrolling gesture, and a pinch gesture performed on the palm.
Figure 4 . PalmSpace network architecture. A multi-task Transformer pipeline processes wrist-camera palm images, extracts shared features, and jointly predicts contact and gesture class together with a normalized two-dimensional interaction location.
Figure 5 . Error analysis and ablation results. Error bars in (c) denote SE; (d) reports the three users with the smallest palms. Four plots summarize localization performance: a two-dimensional error scatter, an error distribution, mean absolute error by touch mode and hand region with standard-error bars, and a comparison showing the effect of palm-size normalization for three small-palm participants.
Metric
Non-Touch
Touch Gestures *
Overall
Single
Scroll
Pinch
F1 Score (%)
97.9
97.2
88.1
97.4
96.7
MAE (mm)
–
6.1
6.3
10.6
6.7
SD (mm)
–
4.3
3.9
11.4
5.7
* The binary Touch Detection Accuracy (Touch vs. Non-Touch) is 98.9%.
Table 2 . Performance overview across different interaction modalities. ‘Single’ refers to single-finger touch.
Figure 6 . Setup and results for the indoor single-finger user studies. A two-by-two layout shows the Fitts' law target-selection interface and palm keypad mapping, selection time by condition, mean movement time by index of difficulty, and eyes-free digit-entry accuracy for PalmSpace and touchscreen input.
Figure 7 . Interaction logic and recognition results for multi-finger input. Two panels show PalmSpace's finite state machine and the confusion matrix for four-direction list navigation.
Figure 8 . Applications in the PalmSpace design space: (a) T9 keyboard, (b) handwriting, (c) AI-assisted drawing, (d–e) MR controllers, (f) content editing, (g) map manipulation, and (h) page control. Eight panels show example applications including keypad entry, handwriting, AI-assisted drawing, mixed-reality control, content editing, map manipulation, and page control.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9 . System workflow: The user interacts by touching their non-dominant palm with their dominant hand’s finger. A deep learning model processes this input in real-time, predicting both touch events and precise touch positions. Upon detecting valid touch interactions, the system maps these absolute coordinates to specific applications in the real world. A left-to-right workflow shows infrared palm capture, neural-network inference of touch state, gesture class and location, and mapping of those predictions to single-finger drawing or multi-finger application commands.
Figure 10 . Blender render of the wristband casing. A three-dimensional rendering shows the compact rectangular wristband casing and its hinged camera mount.
Figure 11 . Data collection setup: (a) Overall architecture of the capture system; (b) 3D-printed calibration device and corresponding calibration system; (c) Tracking of the dominant hand’s index fingertip pixel coordinates via MediaPipe ( Zhang et al., 2020 ) and their transformation into the calibrated palm plane coordinate system. Three panels show the complete data-capture setup, the three-dimensional printed calibration fixture and calibration interface, and fingertip tracking transformed from camera pixels into coordinates on the palm plane.
Palm-Area
Finger-Area
Overall
Mode
Metric
xerror
yerror
lerror
xerror
yerror
lerror
xerror
yerror
lerror
LY
MAE
2.8
3.9
5.4
3.5
5.0
6.6
3.1
4.4
5.8
SD
2.6
3.1
3.3
3.2
4.3
4.6
2.9
3.6
3.9
VT
MAE
3.1
4.1
5.5
4.5
5.8
8.1
3.6
4.7
6.4
SD
2.9
3.0
3.4
4.6
5.6
6.5
3.6
4.2
4.9
RY
MAE
2.8
3.9
5.4
3.8
5.4
7.2
3.2
4.5
6.1
Appendix
Table 4 . The average MAE and SD in different interaction modes and touch regions. Errors are reported in mm.
Figure 12 . Confusion matrices for digit input: (a) PalmSpace, (b) Touchscreen. Side-by-side ten-class confusion matrices compare intended and recognized digits for PalmSpace and the touchscreen baseline. PalmSpace errors are more evenly distributed, whereas touchscreen errors are concentrated among several neighboring digits.
Digit Recognition Accuracy (%)
Transition Accuracy (%)
Method
0
1
2
3
4
5
6
7
8
9
Identical
Adjacent
Non-Adj.
PalmSpace
98.55
100.00
90.74
89.47
95.16
91.94
97.06
100.00
92.31
95.00
95.16
94.90
95.91
Touchscreen
95.65
46.77
85.19
70.18
54.84
74.19
75.00
48.15
90.38
88.33
97.87
79.60
69.29
Appendix
Table 5 . Comparison of Digit Recognition and Transition Accuracy (%) between PalmSpace and Touchscreen.
Figure 13 . Handwriting recognition and subjective evaluation results. The top row contains a 36-class handwritten-character confusion matrix and representative character traces from multiple participants. The bottom panel compares participant ratings of PalmSpace and touchscreen input.
Figure 14 . Sample images from various outdoor data collection scenarios. A grid of infrared wrist-camera frames shows palm interactions collected outdoors across varied locations, daylight levels, and weather conditions.
Capturing hand motion and interaction forces is critical for interactive computing, VR, and high-fidelity tactile demonstrations for robot learning. We introduce a wrist-worn pressure-sensing wristband that recovers continuous full-hand pose and distributed contact force on a single wearable. The system consists of flexible capacitive sensor arrays around the wrist, which require no electrical skin contact, and a recurrent network that maps the resulting pressure signal to hand state. Our key insight is that muscle contraction and tendon displacement produce pressure patterns, which correlate strongly with hand pose and interaction force. To validate this, we collect synchronized recordings of wrist pressure, optical motion-capture hand pose, and tactile-glove interaction force, covering isolated finger motion, fingertip-force stress tests, and natural hand-object manipulation. On isolated single-user motion the wristband attains 4.6∘ mean finger-joint MAE, and across four users manipulating everyday objects it estimates per-finger contact force at R2=0.57, which an external pose signal brings up to 0.75. We see the wristband as one node in a constellation of everyday wearables -- e.g. paired with an egocentric camera -- adding the contact force that vision cannot observe and taking over when the hand is occluded.
We present ART-Glove, an articulated tactile glove designed to capture contact-grounded dexterous demonstrations while preserving human dexterity. ART-Glove makes hand-side contact geometry explicit with 16 rigid functional surfaces covering the fingers, thumb, and palm. Twenty-two anatomically aligned joints connect these surfaces and allow them to follow human hand motion during dexterous manipulation. Encoder-based sensing tracks surface motion, while dense piezoresistive tactile sensing records contact over the same surfaces. The complete system captures synchronized 22-DoF joint measurements and 2048-taxel tactile measurements at 120 Hz. We evaluate ART-Glove across experiments on motion freedom, joint sensing, tactile sensing, and contact-rich interaction capture, demonstrating its ability to preserve human dexterity while recording contact-grounded information that can support downstream dexterous robot learning.
Tactile gloves digitize contact and force during hand-object interactions, enabling robotics applications in dexterous manipulation, teleoperation, and learning from demonstration. To preserve hand dexterity and capture the nuances of natural interactions, these gloves and the integrated tactile sensors are designed to be soft, flexible, and comfortable. However, such flexible sensors are sensitive not only to contact forces but also unavoidably to hand pose changes, resulting in pose-related artifacts (PRAs). PRAs are especially problematic in the low-force range, resulting in misdetections or late-onset detections of contact, which raises the minimum detectable force (MDF) of the glove. In this work, we characterize the PRAs in relation to pose and force. Building on these insights, we introduce a glove-agnostic algorithmic framework that leverages hand pose information, which is increasingly available, to mitigate PRAs without glove modifications. Our pose-aware force estimation model augments tactile-to-force pipelines with a residual prediction branch that explicitly accounts for pose-induced sensor deformations. We validate our approach across 3 glove designs and 15 users, reducing MDF by 10.4%, 12.2%, and 18.3%, with consistent improvements across all evaluated metrics. This method provides a practical path to improving the usability of tactile gloves in data collection and diverse robotic applications.
Tianhong Catherine Yu, Ziyi Kou, Mia Huang +4
Cornell Univeristy · 2Meta Reality Labs · University of Washington