Contact-rich manipulation requires adapting to contact states that can evolve substantially within an action horizon. However, chunk-based vision-language-action models predict complete action chunks from observations collected before execution, leaving tactile conditioning stale during execution. Existing tactile-reactive approaches typically rely on separate high-frequency controllers, which increase both architectural and training complexity. In this paper, we introduce TacForcing, a streaming action-generation framework incorporating execution-time tactile feedback. TacForcing replaces the standard action expert with a streaming expert that generates action blocks sequentially while preserving intermediate states of unfinished blocks. After each block is executed, the expert resumes generation from these states using newly acquired tactile feedback. To better align tactile conditioning with action execution, we further introduce Execution-Aware Tactile Attention (EATA), which restricts direct access to each tactile update to the next block scheduled for execution. Across six UniVTAC simulation tasks and six real-world contact-rich manipulation tasks on two robot platforms, TacForcing achieves average success rates of 65% in simulation and 66% in the real world, outperforming the strongest baselines by 6 and 15 percentage points, respectively.
Figures & tables
Figure 1
Figure 2 : Visual and tactile dynamics during dropper squeezing. (a) Visual observations at the beginning and end of the action horizon. (b) Deformation maps from the thumb and index finger, sampled every five actions. (c) Cosine distances from the initial visual and tactile representations; the inset provides an enlarged view of the visual cosine distances.
Figure 3 : Overview of TacForcing . The VLM encodes the visual observation and language instruction once, and the resulting task context is reused throughout streaming generation. The Streaming Action Expert progressively refines action blocks according to block-specific flow times, allowing successive blocks to become ready for sequential execution. After each ready block is executed, the resulting tactile feedback is encoded before generation resumes from the retained intermediate states of unfinished blocks. Execution-Aware Tactile Attention restricts direct access to the latest tactile feedback to the block scheduled for execution next, thereby reducing the temporal mismatch between tactile acquisition and action execution.
Method
Lift Bottle
Pull-out Key
Lift Can
Put Bottle in Shelf
Insert Hole
Insert Tube
Avg.
π0.5
88
43
46
43
39
48
51
UniVTAC-ACT
59
41
24
6
36
58
37
RDP
84
18
12
41
23
75
42
FTP-1
89
35
66
23
62
76
59
Table 1 : Results on the UniVTAC simulation benchmark. Success rates (%). The highest and second-highest distinct values in each column are shown in bold and underlined , respectively.
Figure 4 : Real-world platforms and tasks. Top: Close Cap , Insert Battery , and Close Lid on Sharpa North. Bottom: Transfer Liquid , Wipe Board , and Stand Bottle on Sharpa & RealMan. Platform views are shown on the left; task scenes are shown on the right.
Sharpa North
Sharpa & RealMan
Method
Close Cap
Insert Battery
Close Lid
Transfer Liquid
Wipe Board
Stand Bottle
Avg.
π0.5
25
38
13
6
31
44
26
GR00T N1.7
44
75
25
19
50
56
45
FTP-1
50
63
38
6
75
75
51
T-Rex
44
50
44
31
63
75
51
TacForcing (Ours)
56
81
50
50
75
81
66
Table 2 : Results on six real-world tasks. All values are success rates (%). The highest and second-highest distinct rates in each column are shown in bold and underlined , respectively.
Figure 5 : Ablation results. All variants retain streaming generation. Fixed tactile holds the initial tactile observation constant within each action chunk, while tactile updates refresh the observation during execution; both disable EATA. TacForcing enables both updates and EATA. Average denotes the mean success rate over the three tasks in each panel.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Real-world platform. The setup comprises (a) Sharpa Wave dexterous hands, (b) Manus Pro data gloves, (c) two RealMan RM75 robot arms, and (d) top- and wrist-mounted RealSense cameras.
Figure 7 : Sharpa North tasks. Representative execution sequences in the task order of Table 2 . Each row progresses from left to right.
Figure 8 : Sharpa & RealMan tasks. Representative execution sequences in the task order of Table 2 . Each row progresses from left to right.
Figure 9 : UniVTAC simulation tasks. Representative execution sequences for the six tasks used in our simulation experiments, reproduced from UniVTAC [ 4 ] .
World action models inherit the predictive capability of world models, enabling action generation to be guided by anticipated future observations. However, they rely primarily on vision and often fail in contact-rich manipulation, where critical cues arise from physical interaction. In this paper, we propose Dream-Tac, a unified Tactile-World Action Model that jointly models actions, future visual observations, and tactile dynamics. Specifically, Dream-Tac introduces (i) contact-gated visuotactile fusion to selectively integrate tactile signals and (ii) a contact-aware attention bias to better regulate cross-modal interactions during manipulation. To support real-time deployment, we further design a dual-level acceleration strategy, reformulating the contact-aware bias to preserve the fused attention path during training and introducing cache-based diffusion acceleration at inference, achieving up to 2.9× faster training and 1.8× faster inference. Across six contact-rich manipulation tasks, Dream-Tac improves action accuracy by 31.7% on average, demonstrating the effectiveness of unified visuotactile world modeling.Code is available at https://github.com/LYFCLOUDFAN/Dream-Tac.
Yunfan Lou, Yifan Ye, Yankai Fu +7
Peking University Beijing, China · Nanjing University Nanjing, China · State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University Beijing, China
Contact-rich manipulation requires robots to continuously perceive and regulate evolving physical interactions under dynamic contact transitions or complex surface geometries. Recent imitation learning methods improve contact-aware control by incorporating tactile or force feedback, but they rarely model the asymmetric spatiotemporal roles of global force and local tactile sensing. To address this, we propose TacForeSight, a lightweight force-conditioned tactile foresight framework for real-time manipulation. The core component is TacForceWM, a tactile world model that predicts short-horizon tactile latent dynamics from dual-finger tactile observations conditioned on high-frequency wrist force and torque signals. Another key component, the Predictive Tactile-Conditioned Policy, leverages the predicted latents as anticipatory contact priors, models the current-to-future tactile evolution via cross-attention, and adaptively fuses visuo-tactile features through a tactile-guided gating module. By forecasting purely within a compact latent space, TacForeSight enables proactive contact reasoning with efficient real-time inference suitable for high-frequency manipulation control. Real-robot experiments on five representative tasks and three in-process perturbation settings show that TacForeSight consistently outperforms existing baselines, particularly under dynamic contact disturbances. All models and datasets will be made publicly available on the project website at https://tacforesight.github.io/ProjectPage.
Yujie Zang, Yuhang Zheng, Xian Nie +7
TARS Robotics · National University of Singapore · Shanghai Jiao Tong University +2
Fusing tactile signals has proven effective for contact-rich manipulation, enabling robots to perceive contact states and adapt to rapidly changing physical interactions. Yet effectively integrating tactile feedback into dexterous manipulation remains underexplored. In this work, we introduce ReTouch, a vision-language-action model (VLA) that supports contact-rich dexterous manipulation through tactile predictions continually refined online using execution-time feedback. ReTouch builds on two main innovations for tactile representation and closed-loop action generation. First, its Tactile-Patch Encoder represents tactile observations as structured tactile patch features that preserve finger identity and local contact structure, providing contact cues for fine-grained dexterous control. Second, its high-frequency action module jointly predicts future tactile states and action chunks and refines both using incoming tactile feedback during execution. This closed-loop refinement keeps tactile predictions aligned with evolving physical interactions, enabling responsive action correction and improving robustness to contact changes and execution errors. We further introduce XHT-Dataset, comprising 900 real-world demonstrations across seven contact-rich tasks collected on an XHand--UR7e platform, and evaluate ReTouch through closed-loop real-robot experiments. ReTouch surpasses the strongest baseline by 18.4 and 23.8 percentage points in average success rate under standard and challenging conditions, respectively, demonstrating its effectiveness and robustness.
Shiqi Zhang, Xin Zhang, Yedong Shen +9
University of Science and Technology of China · The Chinese University of Hong Kong · iFLYTEK