TacDyn-WAM: Learning Implicit Tactile Dynamics in a Heterogeneous Visuo-Tactile World Action Model
Authors: Enyi Wang, Mingxin Wang, Quan Shi, Hetian Guo, Hongyu Wang, Xi Wang, Bin Qian, Yupeng Zheng, +7 more
Organizations: Institute for AI Industry Research (AIR), Tsinghua University · Tsinghua University · Institute of Automation, Chinese Academy of Sciences · The Hong Kong University of Science and Technology (Guangzhou) · School of Information, Renmin University of China · Fudan University · TARS Robotics
World action models improve robotic manipulation by conditioning actions on predicted futures, yet existing tactile variants largely inherit video-generation pipelines that reconstruct future tactile observations through iterative denoising. Such prediction can become unreliable under deployment drift: small changes in contact position or force may substantially alter tactile pixels even when the underlying contact evolution remains predictable. We introduce TacDyn-WAM, a heterogeneous visuo-tactile world action model that predicts implicit tactile dynamics rather than reconstructing future tactile observations. It learns TacRep, a dynamics-aware tactile target space trained through masked spatio-temporal prediction on tactile clips and regularized by relational structure distillation. A visual expert and an Implicit Tactile Dynamics Expert predict future visual and tactile representations in separate target spaces while interacting through joint attention; the tactile expert predicts future representations and their changes at multiple horizons in a single forward pass, and a read-only tactile memory supplies the current tactile state. On UniVTAC, TacDyn-WAM achieves an average success rate of 81.5% using only the provided demonstrations, reaching state-of-the-art-level performance and remaining competitive with models pretrained on large-scale visuo-tactile trajectories. Ablations confirm the benefits of both tactile pathways and TacRep over pixel-reconstruction and static alternatives. On five real-robot tasks, TacDyn-WAM reaches 71.0% average success, and modest-scale pretraining raises it to 85.0%, further validating our method.
Figures & tables
Figure 1: Overview of TacDyn-WAM. The visual and tactile experts form a heterogeneous world model, while the Tactile Understanding Memory provides the current tactile state as read-only keys and values.
Figure 2: Stage 1: TacRep training. Tactile Dynamics Prediction trains the encoder against an EMA target on masked four-frame clips. Relational Structure Distillation keeps local patch relations close to a frozen DINOv2. Only the EMA encoder is kept.
Figure 3: Implicit Tactile Dynamics Expert. A 0.4B dynamics predictor maps context tokens and future queries to the Future and Delta representations at three horizons.
Figure 4: Attention visibility and trainable modules in Stages 2 to 4. Rows are queries and columns are keys; U, M, V, T, and A denote the prefix, memory, visual, tactile, and action blocks. Bold marks connections unlocked in each stage and a flame marks trained modules.
Method
Insert Hole
Insert Tube
Lift Can
Pull-out Key
Put Bottle
Lift Bottle
Grasp Classify
Insert HDMI
Avg.
Vision-only Policies
π0.5 ( Physical Intelligence et al., 2025 )
25
74
6
35
34
100
49
8
41.4
StarVLA- α ( Ye et al., 2026b )
52
69
65
51
88
32
68
24
56.1
InternVLA-A1 ( Cai et al., 2026a )
70
56
58
67
37
58
88
12
55.8
Xiaomi-Robotics-0 ( Cai et al., 2026b )
96
98
13
80
12
21
45
69
54.3
GigaWorld-Policy ( Ye et al., 2026a )
12
9
0
32
21
38
20
0
16.5
LingBot-VA ( Li et al., 2026 )
42
96
0
58
0
0
17
38
31.4
Table 1: Success rates (%) on the eight UniVTAC tasks. Baselines are grouped into vision-only policies, visuo-tactile policies, and policies with large-scale visuo-tactile trajectory pretraining. Gray rows are excluded from comparison, “–” marks results not reported, and the best participating result per column is bolded.
Method
Insert Hole
Insert Tube
Lift Can
Pull-out Key
Put Bottle
Lift Bottle
Grasp Classify
Insert HDMI
Avg.
w/o Tactile World Model
79
72
70
88
52
76
95
11
67.9
w/o Tactile Understanding Memory
84
83
77
82
58
86
91
9
71.3
TacRep → Cosmos VAE
74
69
63
61
47
70
97
15
62.0
TacRep → DINOv2
85
83
80
89
61
84
98
12
74.0
TacDyn-WAM (full)
91
97
87
97
72
97
99
12
81.5
Table 2: Ablation study on UniVTAC (success rate, %). All variants share the same base model, data, and per-stage training hyperparameters. More details are given in Appendix E .
Method
Stack Cups
Remove Plug
Insert Plug
Unscrew Cup Lid
Wipe Whiteboard
Avg.
LingBot-VA
40
10
45
15
50
32.0
InternVLA-A1
55
20
45
10
70
40.0
FTP-1
65
45
65
20
80
55.0
TacDyn-WAM
80
60
80
40
95
71.0
TacDyn-WAM (pretrained)
95
85
90
55
100
85.0
Table 3: Real-world success rates (%) over 20 trials per task. All baselines are fine-tuned per task on the same demonstrations.
Method
Actions per chunk
Latency per chunk (ms)
Latency per action (ms)
LingBot-VA (naive)
40
2,211
55.3
LingBot-VA (optimized)
40
1,338
33.5
TacDyn-WAM (ours)
50
484
9.7
Table 4: Inference latency on one A100 GPU with batch size 1. LingBot-VA (naive) recomputes the history at every step; LingBot-VA (optimized) uses the deployment settings of its paper.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Examples of tactile distribution shift between expert demonstrations (top) and policy execution (bottom) on Pull-out Key (left) and Put Bottle (right). In each panel, rows show the left and right sensors without and with markers, and columns are six frames over time. The imprint follows a similar trend in both cases, but its position, extent, and timing differ during execution, and the marker view makes the difference more visible.
Figure 6: Tactile input preprocessing. Left: raw RGB frames from the two sensors of a gripper. Middle: the static no-contact background frame. Right: the signed residual after subtracting the background and adding a mid-gray offset. Static gel texture and illumination are removed, and only the current deformation remains.
Stage 1
Stage 2
Stage 3
Stage 4
Batch size
64
128
128
64
Peak learning rate
10−4 / 10−5 / 5×10−6
10−4
10−4 / 5×10−5
10−4 / 5×10−5
Weight decay
0.04
0.05
0.01
0.01
Warmup
5%
5%
5%
2,000 steps
Length
20 epochs
15 epochs
5 epochs
12,000 steps
Appendix
Table 5: Optimization settings of the four stages on UniVTAC.
Data
Length
Global batch
Stage 1: TacRep Training
OmniViTac + real
2 epochs
256
Stage 2: Tactile World Grounding
OmniViTac + real
2 epochs
128
Stage 3: Tactile–Action Alignment
OmniViTac
0.5 epochs
128
Stage 4: Joint Training
OmniViTac
3 epochs
128
Real-robot fine-tuning (Stage 4, per task)
60 demonstrations
8,000 steps
64
Appendix
Table 6: Data and schedule of modest-scale pretraining and real-robot fine-tuning.
Figure 7: The five real-world tasks. Each row shows one task from left to right over time.
Figure 8: Two failed Wipe Whiteboard trials. Left: the letters are not wiped clean. Right: most of the letters are removed, but part of a letter remains. Both are counted as failures.
Task
Category
Description
Lift Bottle
Pose reasoning
Grasp a bottle standing against a wall and lift it vertically without hitting the wall.
Lift Can
Pose reasoning
Grasp a horizontally placed can of one of three diameters and lift it without slipping.
Put Bottle
Pose reasoning
Grasp a standing bottle and place it into the cavity of a shelf.
Grasp Classify
Shape perception
Touch two visually similar cylinders with different surface textures, then place each on the target pad of its class.
Insert Hole
Contact-rich
Insert a tube into a tilted hole whose orientation must be found by contact.
Insert Tube
Contact-rich
Insert a tube into a narrow hole on a tilted surface with a small clearance.
SKL-MAIS, Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · TARS Robotics +2