TacDyn-WAM: Learning Implicit Tactile Dynamics in a Heterogeneous Visuo-Tactile World Action Model
Organizations: Institute for AI Industry Research (AIR), Tsinghua University · Tsinghua University · Institute of Automation, Chinese Academy of Sciences · The Hong Kong University of Science and Technology (Guangzhou) · School of Information, Renmin University of China · Fudan University · TARS Robotics
Abstract
World action models improve robotic manipulation by conditioning actions on predicted futures, yet existing tactile variants largely inherit video-generation pipelines that reconstruct future tactile observations through iterative denoising. Such prediction can become unreliable under deployment drift: small changes in contact position or force may substantially alter tactile pixels even when the underlying contact evolution remains predictable. We introduce TacDyn-WAM, a heterogeneous visuo-tactile world action model that predicts implicit tactile dynamics rather than reconstructing future tactile observations. It learns TacRep, a dynamics-aware tactile target space trained through masked spatio-temporal prediction on tactile clips and regularized by relational structure distillation. A visual expert and an Implicit Tactile Dynamics Expert predict future visual and tactile representations in separate target spaces while interacting through joint attention; the tactile expert predicts future representations and their changes at multiple horizons in a single forward pass, and a read-only tactile memory supplies the current tactile state. On UniVTAC, TacDyn-WAM achieves an average success rate of 81.5% using only the provided demonstrations, reaching state-of-the-art-level performance and remaining competitive with models pretrained on large-scale visuo-tactile trajectories. Ablations confirm the benefits of both tactile pathways and TacRep over pixel-reconstruction and static alternatives. On five real-robot tasks, TacDyn-WAM reaches 71.0% average success, and modest-scale pretraining raises it to 85.0%, further validating our method.
Figures & tables
| Method | Insert Hole | Insert Tube | Lift Can | Pull-out Key | Put Bottle | Lift Bottle | Grasp Classify | Insert HDMI | Avg. | |
| Vision-only Policies | ( Physical Intelligence et al., 2025 ) | 25 | 74 | 6 | 35 | 34 | 100 | 49 | 8 | 41.4 |
| StarVLA- ( Ye et al., 2026b ) | 52 | 69 | 65 | 51 | 88 | 32 | 68 | 24 | 56.1 | |
| InternVLA-A1 ( Cai et al., 2026a ) | 70 | 56 | 58 | 67 | 37 | 58 | 88 | 12 | 55.8 | |
| Xiaomi-Robotics-0 ( Cai et al., 2026b ) | 96 | 98 | 13 | 80 | 12 | 21 | 45 | 69 | 54.3 | |
| GigaWorld-Policy ( Ye et al., 2026a ) | 12 | 9 | 0 | 32 | 21 | 38 | 20 | 0 | 16.5 | |
| LingBot-VA ( Li et al., 2026 ) | 42 | 96 | 0 | 58 | 0 | 0 | 17 | 38 | 31.4 |
| Method | Insert Hole | Insert Tube | Lift Can | Pull-out Key | Put Bottle | Lift Bottle | Grasp Classify | Insert HDMI | Avg. |
|---|---|---|---|---|---|---|---|---|---|
| w/o Tactile World Model | 79 | 72 | 70 | 88 | 52 | 76 | 95 | 11 | 67.9 |
| w/o Tactile Understanding Memory | 84 | 83 | 77 | 82 | 58 | 86 | 91 | 9 | 71.3 |
| TacRep Cosmos VAE | 74 | 69 | 63 | 61 | 47 | 70 | 97 | 15 | 62.0 |
| TacRep DINOv2 | 85 | 83 | 80 | 89 | 61 | 84 | 98 | 12 | 74.0 |
| TacDyn-WAM (full) | 91 | 97 | 87 | 97 | 72 | 97 | 99 | 12 | 81.5 |
| Method | Stack Cups | Remove Plug | Insert Plug | Unscrew Cup Lid | Wipe Whiteboard | Avg. |
|---|---|---|---|---|---|---|
| LingBot-VA | 40 | 10 | 45 | 15 | 50 | 32.0 |
| InternVLA-A1 | 55 | 20 | 45 | 10 | 70 | 40.0 |
| FTP-1 | 65 | 45 | 65 | 20 | 80 | 55.0 |
| TacDyn-WAM | 80 | 60 | 80 | 40 | 95 | 71.0 |
| TacDyn-WAM (pretrained) | 95 | 85 | 90 | 55 | 100 | 85.0 |
| Method | Actions per chunk | Latency per chunk (ms) | Latency per action (ms) |
|---|---|---|---|
| LingBot-VA (naive) | 40 | 2,211 | 55.3 |
| LingBot-VA (optimized) | 40 | 1,338 | 33.5 |
| TacDyn-WAM (ours) | 50 | 484 | 9.7 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Stage 1 | Stage 2 | Stage 3 | Stage 4 | |
|---|---|---|---|---|
| Batch size | 64 | 128 | 128 | 64 |
| Peak learning rate | / / | / | / | |
| Weight decay | 0.04 | 0.05 | 0.01 | 0.01 |
| Warmup | 5% | 5% | 5% | 2,000 steps |
| Length | 20 epochs | 15 epochs | 5 epochs | 12,000 steps |
| Data | Length | Global batch | |
|---|---|---|---|
| Stage 1: TacRep Training | OmniViTac + real | 2 epochs | 256 |
| Stage 2: Tactile World Grounding | OmniViTac + real | 2 epochs | 128 |
| Stage 3: Tactile–Action Alignment | OmniViTac | 0.5 epochs | 128 |
| Stage 4: Joint Training | OmniViTac | 3 epochs | 128 |
| Real-robot fine-tuning (Stage 4, per task) | 60 demonstrations | 8,000 steps | 64 |
| Task | Category | Description |
|---|---|---|
| Lift Bottle | Pose reasoning | Grasp a bottle standing against a wall and lift it vertically without hitting the wall. |
| Lift Can | Pose reasoning | Grasp a horizontally placed can of one of three diameters and lift it without slipping. |
| Put Bottle | Pose reasoning | Grasp a standing bottle and place it into the cavity of a shelf. |
| Grasp Classify | Shape perception | Touch two visually similar cylinders with different surface textures, then place each on the target pad of its class. |
| Insert Hole | Contact-rich | Insert a tube into a tilted hole whose orientation must be found by contact. |
| Insert Tube | Contact-rich | Insert a tube into a narrow hole on a tilted surface with a small clearance. |