World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video world model for future-frame prediction of egocentric manipulation from both observed video and tactile signals from a glove worn on each hand. We condition a pretrained video diffusion transformer on each hand's tactile signal through a zero-initialized residual at the corresponding hand location in the video tokens, while a causal mask prevents predicted frames from accessing future information. Compared with a vision-only model matched in architecture, parameters, and training, DTWM reduces the underestimation of hand motion from 23% to 9%, while reducing the perceptual error in the hand region by 7.4% across three training runs per model. The benefit also increases over the prediction horizon, with the improvement in the later predicted chunks being about 4.1x larger than in the first. DTWM also outperforms other visual-tactile world models under the same setting, and training with touch improves future-frame prediction even when no touch is available at inference. Ablations show that the model benefits from both the magnitude and spatial location of force: replacing the tactile signal with binary contact states, either per hand or per location, increases prediction error. The observed course of the force indicates whether the interaction will persist or change.
Figures & tables
Figure 1: DTWM shows better future-frame prediction of human manipulation with whole-hand tactile sensing. Left: a held-out clip in which both hands open a spray can; dashed boxes are the ground-truth hand regions. Right: we achieve better visual quality of the entire frame and the hand crop, and more accurate hand location and motion, than existing visual-tactile world models.
Figure 2: Pipeline of DTWM. During both training and inference, the model observes RGB frames, hand poses, and tactile signals from both hands for time steps t≤K . We sample one reading from both gloves ( 255 pressure taxels and 30 finger-flexion channels in total) per video frame. In the touch panel, color shows pressure and blue bars show flexion. The task instruction is encoded by the text encoder, the observed frames are represented as clean latents, and the hand skeleton is encoded using a widened patch embedding. Each hand’s tactile signal is embedded as e and added to the video tokens at the hand’s projected location using a normalized Gaussian footprint g : u←u+proj(LN(e))g . The model then predicts the future RGB frames at time steps t>K . During training, K is sampled from {4,…,44} , and the model is optimized with a flow-matching loss against the ground-truth future frames. At inference, the model observes 13 frames ( K=12 ) and predicts all 36 future frames in a single pass.
Figure 3: Comparison with the vision-only baseline. The hands of the vision-only model lose their definition or stop moving, while ours keep their shape and their motion. Warm colors: pressure, blue bars: finger flexion. Dashed boxes: ground-truth hand regions.
Figure 4: Comparison with other visual-tactile world models. In these examples, DTWM preserves hand shape and follows the ground-truth hand positions more closely than the other models.
Figure 5: Examples on OOD tasks. On tasks with no close counterpart in training, DTWM preserves more hand detail and follows the ground-truth hand positions more closely than the vision-only baseline in the examples shown.
ID (2,728 clips)
OOD (272 clips)
Hand, all 3,000 clips
Model
LPIPS ↓
LPIPS hand↓
LPIPS ↓
LPIPS hand↓
IoU ↑
flow dir. ↑
Vision-only baseline
0.4243
0.3273
0.4882
0.3765
0.475
0.039
TouchWorld ( Zhou et al., 2026b )
0.4899
0.3979
0.5341
0.4204
0.376
0.031
VT-WM ( Higuera et al., 2026 )
0.4237
0.3324
0.4854
0.3797
0.462
0.032
FeelWorld ( Ma et al., 2026 )
0.4488
0.3465
0.5071
0.3964
0.430
0.027
DTWM (ours)
0.4157
0.3227
0.4774
0.3668
0.476
0.041
Table 1: Prediction quality on the held-out set. DTWM achieves the lowest full-frame and hand-crop LPIPS on both task splits. Hand overlap (IoU) and flow direction agreement are measured over all clips. All rows report one training run.
ID (2,728 clips)
OOD (272 clips)
Tactile input
LPIPS
LPIPS hand
LPIPS
LPIPS hand
zeroed reading (vision-only baseline)
0.4243
0.3273
0.4882
0.3765
pathway disabled at inference
0.4208
0.3297
0.4838
0.3754
one contact bit per hand
0.4307
0.3318
0.5014
0.3864
per-taxel contact, no force
0.4200
0.3246
0.4823
0.3745
DTWM (ours, full reading)
0.4157
0.3227
0.4774
0.3668
Table 2: Ablation of tactile input. The full reading gives the lowest perceptual error. Per-taxel contact retains contact locations and finger flexion but removes pressure magnitude; one bit per hand retains only contact state. The pathway-disabled row evaluates the full-reading checkpoint with its tactile residual switched off. Lower is better for all metrics.
Figure 6: Touch history correlates with future manipulation. The observed frames look alike, but the observed touch evolves differently, and so do the manipulations that follow.
Figure 7: The LPIPS improvement grows over the prediction horizon. (a) Cumulative scoring windows: DTWM has lower LPIPS at every window, with a larger gap beyond the first chunk. (b) Disjoint temporal segments: on held-out episodes, the average per-frame gap after the first chunk is about 4.1 times that within it. Scores are averaged over three runs per model.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Group
Examples (held-out task: training task)
Tasks
Clips
Same task
pick up toothpaste; organize a suitcase; use a microwave
6
479
Near-duplicate
pick up earphones: over-ear headphones; grasp sunscreen: grasp body lotion; fold shorts: fold clothes; organize medicine: sort medicine
18
1,251
Related object or action
squeeze a duck toy: pick up a toy racket; use a G-clamp: rotate a table clamp; pick up a thermos cup: wash a cup
18
998
Nothing similar
push a cart; drag a chair; use a thermometer
3
272
Appendix
Table 3: Only three held-out tasks have no counterpart in training. Held-out tasks grouped by their closest counterpart in training, by task name.
Backbone
Pretrained model
Wan2.2-Fun-5B-InP (flow-matching video transformer)
Transformer parameters
5,001,016,512 (after the channel extension)
Hidden dimension, blocks, heads
3072, 30, 24 (head dimension 128)
Feed-forward dimension
14,336
Patchifier
3D convolution, kernel and stride (1,2,2)
Video autoencoder
48 latent channels, 4× temporal and 16× spatial compression
Appendix
Table 4: One recipe for every model in the paper. Complete training recipe.
TouchWorld ( Zhou et al., 2026b )
VT-WM ( Higuera et al., 2026 )
FeelWorld ( Ma et al., 2026 )
Original sensor
the same pressure gloves (our corpus), and tactile gloves on two robot hands
four Digit 360 optical fingertip sensors on an Allegro hand
DM tactile sensors, 3D contact point cloud
Original predictor
video model finetuned from Wan2.2-TI2V-5B
action-conditioned latent transformer over Cosmos tokens
latent dynamics predictor over V-JEPA 2 features
Original mechanism
readings rendered as tactile images beside the RGB views
Sparsh-X tactile tokens concatenated with the visual tokens
one global tactile token per fingertip sensor, attended by the visual tokens through a contact-gated attention
Our implementation
grids drawn as image panels in the hand-skeleton conditioning stream
one token per latent frame and hand appended to the video tokens
one tactile feature per hand added to all video tokens, scaled by a learned gate
Tactile parameters
51,581,725 (zeroed residual pathway kept)
23,236,381
51,597,090
Appendix
Table 5: Each existing visual-tactile world model keeps its original mechanism and receives the same glove reading as DTWM. The original system and our implementation.
Held-out episodes
Training episodes
Model
LPIPS
LPIPS hand
LPIPS
LPIPS hand
Vision-only baseline
0.4905 ± 0.0328
0.3817 ± 0.0439
0.5024 ± 0.0417
0.3992 ± 0.0488
DTWM (ours)
0.4654 ± 0.0109
0.3536 ± 0.0119
0.4777 ± 0.0170
0.3733 ± 0.0149
Difference
−0.0252
−0.0281
−0.0247
−0.0259
Appendix
Table 6: Three independent training runs per model. Mean ± standard deviation over runs. The difference is clip-paired.
Held-out episodes
Training episodes
control
DTWM
control
DTWM
run 1
0.5281
0.4777
0.5505
0.4963
run 2
0.4677
0.4571
0.4771
0.4630
run 3
0.4758
0.4612
0.4796
0.4737
mean over runs
0.4905
0.4654
0.5024
0.4777
run comparisons won by DTWM
7 of 9
7 of 9
Appendix
Table 7: DTWM wins 7 of the 9 run pairings on each split, and its worst run is below the baseline’s mean. Mean LPIPS of each independently trained run in the main comparison, and the number of run comparisons in which DTWM obtains the lower score.
Held-out episodes
Training episodes
Predicted frames scored
control
DTWM
difference
control
DTWM
difference
Cumulative window, mean over the frames scored
frames 13 to 20
0.2727
0.2644
−0.0083
0.2812
0.2770
−0.0042
frames 13 to 28
0.3593
0.3501
−0.0092
0.3624
0.3559
−0.0065
frames 13 to 36
0.4265
0.4087
−0.0178
0.4329
0.4168
−0.0161
frames 13 to 48
0.4905
0.4654
−0.0252
0.5024
0.4777
−0.0247
Appendix
Table 8: The per-frame improvement is smallest in the first block and several times larger in the later blocks. LPIPS as a function of the scored window, and per frame within each block of predicted frames, averaged over the runs of Table 6 .
Model
LPIPS ↓
LPIPS hand↓
IoU ↑
centroid (px) ↓
DTW (px) ↓
flow dir. ↑
TouchWorld
0.494
0.400
0.376
28.0
14.2
0.031
FeelWorld
0.454
0.351
0.430
25.1
11.9
0.027
VT-WM
0.429
0.337
0.462
23.3
11.2
0.032
DTWM (ours)
0.421
0.327
0.476
21.9
10.4
0.041
Appendix
Table 9: DTWM is ahead of every existing visual-tactile world model on all six metrics. The metrics of Figure 1 on the 3,000-clip held-out set, predicted frames 13 to 48.
DTWM against the vision-only baseline
Metric
vision-only
DTWM
difference
silhouette IoU ↑
0.475
0.476
+0.001
silhouette area ratio (ideal 1)
0.920
1.032
+0.111
trajectory DTW (px) ↓
10.41
10.37
−0.04
flow direction agreement ↑
0.039
0.041
+0.003
flow magnitude ratio (ideal 1)
0.771
0.912
+0.141
Appendix
Table 10: Touch changes how much of the hands is drawn and how much they move: the silhouette area and flow magnitude ratios move toward one. Hand-behavior metrics over frames 13 to 48 on the held-out set, clip-paired.
Prediction
LPIPS ↓
PSNR ↑
SSIM ↑
Vision-only baseline
0.414
15.24
0.547
DTWM
0.412
14.94
0.535
DTWM, Gaussian blur σ=1
0.476
15.08
0.551
DTWM, Gaussian blur σ=2
0.565
15.21
0.564
DTWM, Gaussian blur σ=3
0.625
15.30
0.571
reconstruction of the observed frames
0.031
35.50
0.960
Appendix
Table 11: PSNR and SSIM improve when the prediction is blurred; LPIPS penalizes blur. LPIPS, PSNR and SSIM of the two models and of degraded versions of the DTWM prediction on 150 held-out clips, frames 13 to 48; the last row is the models’ reconstruction of the observed frames, which no prediction can pass. Best value of each column in bold.
ID (2,728 clips)
OOD (272 clips)
Model
Params
LPIPS ↓
LPIPS hand↓
PSNR ↑
SSIM ↑
LPIPS ↓
LPIPS hand↓
PSNR ↑
SSIM ↑
Vision-only baseline
51.6M
0.4243
0.3273
15.28
0.5095
0.4882
0.3765
13.99
0.5144
TouchWorld ( Zhou et al., 2026b )
51.6M
0.4899
0.3979
14.80
0.5088
0.5341
0.4204
13.88
0.5218
VT-WM ( Higuera et al., 2026 )
23.2M
0.4237
0.3324
15.11
0.5102
0.4854
0.3797
13.91
0.5203
FeelWorld ( Ma et al., 2026 )
51.6M
0.4488
0.3465
14.40
0.4896
0.5071
0.3964
13.36
0.5012
DTWM (ours)
51.6M
0.4157
0.3227
15.06
0.4994
0.4774
0.3668
13.87
0.5068
Appendix
Table 12: PSNR and SSIM order the models differently from LPIPS. Table 1 with PSNR and SSIM added, on the same clips.
Held-out episodes
Training episodes
Model
LPIPS ↓
LPIPS hand↓
PSNR ↑
SSIM ↑
LPIPS ↓
LPIPS hand↓
PSNR ↑
SSIM ↑
Vision-only baseline
0.4905
0.3817
14.98
0.5271
0.5024
0.3992
13.83
0.4751
DTWM (ours)
0.4654
0.3536
14.89
0.5180
0.4777
0.3733
13.95
0.4705
Difference
−0.0252
−0.0281
−0.09
−0.0091
−0.0247
−0.0259
+0.12
−0.0045
Appendix
Table 13: Over three runs, PSNR and SSIM favor the baseline where LPIPS favors DTWM. Table 6 with PSNR and SSIM added, three runs per model.
Figure 8: Further in-distribution examples show the same behavior as Figure 3 : the hands stay in place with the reading and drift or dissolve without it. As Figure 3 .
Figure 9: The same behavior on further in-distribution examples , continued from Figure 8 .
Figure 10: On further clips, the hands of DTWM remain closest to the ground truth among the tactile models. As Figure 4 .
First chunk
Second chunk
(frames 13 to 28)
(frames 29 to 44)
Share of predicted chunks in which, relative to the last observed frame, …
the contact state of the two hands taken together changes
5.0%
6.2%
the contact state of a hand changes
16.8%
21.0%
the force of a hand in contact changes by more than 20%
46.1%
60.6%
the force of a hand in contact changes by more than 50%
22.4%
36.4%
Appendix
Table 14: The force and its location change far more often than the contact state, and the observed force trend signals an imminent release. Training clips in the evaluation layout; the second and third blocks use the hands in contact at the last observed frame.
Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largely vision-centric and therefore cannot directly model these contact dynamics. We present DexTacWAM, a visuo-tactile WAM that encodes each fingertip independently, aggregates the resulting features through a finger- and pose-aware tactile compressor, and injects the tactile latent into a video diffusion world model for joint visuo-tactile world modeling. Across six contact-rich dexterous manipulation tasks on a 22-DoF bimanual platform, DexTacWAM achieves the highest score on every task, averaging 70.6 versus 38.0 for the strongest baseline. Ablations attribute the gain to modeling contact evolution as part of the predicted world state rather than tactile conditioning alone: removing tactile world modeling reduces the four-task mean from 74.7 to 26.6 while keeping the same tactile features and action expert. After four hours of tactile-encoder adaptation with a frozen pretrained vision VAE, our continual vision-to-touch learning extends the pretrained video model to touch using roughly 100 demonstrations per task without tactile midtraining, while retaining visual prediction quality within 0.5 dB of vision-only counterparts. The compressor retains 89.4% of pre-fusion contact recall while enabling 2.26x faster training and 1.29x faster inference. Together, these results show that pretrained video priors can be extended to distributed multi-finger contact dynamics in a data- and compute-efficient manner.
Haoran Yuan, Zekai Wang, Boning Shao +4
University of Illinois Urbana-Champaign · University of California, Berkeley · Northwestern University
Learning predictive models of contact-rich dexterous manipulation requires dense tactile interaction, but such data are costly to scale on real robots and remain tied to embodiment-specific sensors. We introduce DexTouch-WM, an action-conditioned world model that learns from scalable human touch to jointly predict future RGB observations and bilateral tactile dynamics. Our insight is that human and robot manipulation share transferable contact dynamics when their tactile observations and action spaces are made compatible. We deploy flexible piezoresistive arrays with a shared sensing layout on both human and dexterous robot hands, and retarget human motion into the robot action space so that human interaction can supervise the same dynamics model used for real-robot prediction. DexTouch-WM couples a pretrained video expert with a lightweight tactile expert using anatomy-aware tactile tokens and aligned action conditioning. In human-to-robot scaling experiments, we keep five hours of real-robot supervision fixed while increasing human interaction from 0 to 100 hours, and observe substantial improvements in held-out robot-domain visual, geometric, and contact prediction despite disjoint human and robot task sets. Beyond prediction, we evaluate the world models as surrogate environments for policy evaluation and as generators of synthetic trajectories for real-robot policy learning, showing that scalable human interaction provides a complementary data axis for learning dexterous robot world models.
World Action Models bring the predictive capabilities of video models into robot action generation, providing a rich foundation for modeling future visual states. Tactile sensing complements this foundation with direct measurements of physical interaction. Some existing methods use tactile features as conditioning inputs without jointly predicting future tactile states, visual observations, and actions. Our key insight is that tactile signals, like video, provide observations of the evolving world state and should be modeled as future observations alongside video. We present ME-Dex-1.0 (MachEmbodied-Dex-1.0), a unified World Action Tactile Model for joint visual, tactile, and action learning. ME-Dex-1.0 adopts a Mixture-of-Transformers architecture comprising a Video Expert, a Tactile Expert, and an Action Expert, all trained with flow matching. We use shared attention connects the experts in intermediate layers, allowing action generation to draw on learned representations of visual and tactile dynamics during joint denoising. To support multi-source heterogeneous tactile inputs, a Canonical Hand Model and a Unified Tactile Autoencoder map tactile observations from different embodiments and sensing layouts into shared spatial and latent spaces. To address the limited availability of paired visual, tactile, and action data, we develop the Agentic Tactile Data Engine, an agent-based data production platform. It supplements RoboTwin and DexJoCo with tactile data recorded directly from force sensors during trajectory replay in simulation. Experiments on the RoboTwin, DexJoCo, and ManiFeel simulation platforms, together with real robot evaluations, demonstrate improved manipulation performance using both grippers and dexterous hands equipped with tactile sensing.