Force and torque (F/T) sensors enable contact-aware control by providing reactive feedback, but they are often fragile and expensive. To overcome these limitations, sensorless methods estimate F/T or wrench solely from robot proprioception, and have shown success in slow interaction tasks such as grasping. However, their low-pass characteristics limit the estimation of high-frequency signals, which are critical in rapid-contact tasks such as grinding. Communication delays can also make their estimates outdated during deployment, but few methods address this directly. To bridge these gaps, we propose a Frequency-aware Decomposition Network (FDN) to estimate vibration-rich wrench in a sensorless, multi-step-ahead manner. Considering higher-frequency stochasticity, FDN spectrally decomposes the wrench horizon into a low-frequency trend and a high-frequency residual, and estimates each by pointwise regression and a learned conditional distribution, respectively. The frequency-aware layers impose band decomposition priors on the outputs and adaptively enhance frequency amplitudes of the inputs. FDN requires neither an identified robot model nor an F/T sensor during estimation. On real-world grinding data from our 6-DoF hydraulic manipulator, FDN reduces high-frequency amplitude error by up to 47% over the baselines under assumed time delays and maintains competitive low-frequency pointwise accuracy, while the baselines fail to balance these two. We also find multi-step-ahead estimation feasible, with FDN estimating a 1,000 ms horizon within 11 ms on a single CPU thread. Ablation studies further support our design choices. In an exploratory study, transferring wrench dynamics learned from an open-source everyday manipulation dataset reduces low-frequency error by 8%, while high-frequency dynamics appear domain-specific.
Figures & tables
Figure 1: Illustration of our grinding task and the Frequency-aware Decomposition Network (FDN). J1 to J6 represent the joint numbers, and ‘EE’ denotes the end-effector. Cylinder drive types are either revolute (‘R’) or prismatic (‘P’), indicated above the joint numbers. Fixed and controlled components are colored in grey and yellow, respectively. The robot grinds the block by moving in −x and −z directions. FDN produces decomposed wrench estimates over a prediction horizon from a proprioception history, trained using ground truth from a wrist F/T sensor.
Figure 2: Illustration of the proposed Frequency-aware Decomposition Network (FDN).
Figure 3: Frequency-aware layers. FFT and IFFT denote the Fast Fourier Transform and its inverse. (a) A learnable frequency enhancement filter ( FEF ) layer. (b) Denoising high-pass filter ( FPFhigh ) and low-pass filter ( FPFlow ).
Figure 4: Visualization of the ‘Stiff-2’ episode in the collected hydraulic dataset. (a) Trajectories of Δq4 , q˙4 , q¨4 and Δp4 . (b) Denoised and decomposed trajectories of Fx and My with cutoffs fc=1 Hz and fcdn=15 Hz.
Episode
Dataset
Duration
In-contact
Nq,Δp
NW
max∣F∣
max∣M∣
dx
dz
E[vx]
Unit
-
[s]
[s]
-
-
[N]
[Nm]
[mm]
[mm]
[mm/s]
Soft-1
Test
289.18
214.11
28,861
28,924
∣Fx∣=122
∣Mx∣=35
158.31
11.75
-0.74
Soft-2
Test
149.34
114.15
14,909
14,937
∣Fx∣=96
∣My∣=27
154.70
12.58
-1.36
Soft-3
Training
230.33
141.70
22,933
23,035
∣Fx∣=155
∣My∣=47
213.84
13.79
-1.51
Soft-4
Training
152.04
108.32
15,121
15,205
∣Fz∣=116
∣Mx∣=42
160.00
12.57
-1.48
Soft-5
Training
176.63
130.39
17,618
17,663
∣Fx∣=134
∣My∣=41
206.74
13.23
-1.59
Table 1: Overview of the collected hydraulic grinding dataset. For max∣F∣ and max∣M∣ , the channel with the maximum magnitude is indicated.
Figure 5: Power spectrum of raw training wrench windows. Each channel is normalized to zero mean and unit variance before the Fourier transform. Power spectrum before denoising (a) and after denoising (b). Yellow vertical lines indicate the decomposition cutoff fc=1 Hz and the denoising cutoff fcdn=15 Hz applied during training.
Multi-step-ahead
wRMSE , High-frequency Windowed RMS Error
Time Delay
100 ms
1,000 ms
Force/Torque Unit
[N]
[Nm]
[N]
[Nm]
MINN
24.359 ± 0.035
4.909 ± 0.003
24.431 ± 0.032
4.925 ± 0.003
RBF
24.130 ± 0.110
4.753 ± 0.018
24.164 ± 0.109
4.759 ± 0.017
GPR
24.974 ± 0.004
4.958 ± 0.001
25.013 ± 0.004
4.970 ± 0.001
LSTM
18.592 ± 0.237
3.867 ± 0.046
19.110 ± 0.211
3.938 ± 0.043
Table 2: High-frequency windowed RMS errors of models.
Multi-step-ahead
pRMSE , Low-frequency Pointwise RMSE
Time Delay
100 ms
1,000 ms
Force/Torque Unit
[N]
[Nm]
[N]
[Nm]
MINN
12.230 ± 0.494
3.612 ± 0.095
12.274 ± 0.489
3.623 ± 0.093
RBF
10.455 ± 0.256
3.076 ± 0.109
10.432 ± 0.265
3.067 ± 0.109
GPR
9.362 ± 0.016
2.285 ± 0.019
9.362 ± 0.018
2.281 ± 0.019
LSTM
13.040 ± 1.102
3.289 ± 0.192
13.189 ± 1.029
3.300 ± 0.189
Table 3: Low-frequency pointwise RMSE of models.
Multi-step-ahead
CRPS
Time Delay
100 ms
1,000 ms
Force/Torque Unit
[N]
[Nm]
[N]
[Nm]
MINN
13.006 ± 0.100
3.900 ± 0.038
13.029 ± 0.101
3.905 ± 0.038
RBF
12.561 ± 0.134
3.855 ± 0.063
12.574 ± 0.142
3.856 ± 0.064
GPR
12.095 ± 0.005
3.520 ± 0.007
12.124 ± 0.005
3.528 ± 0.007
LSTM
14.622 ± 0.518
4.332 ± 0.029
14.847 ± 0.581
4.368 ± 0.058
Table 4: Continuous ranked probability scores of models.
Figure 6: Test episode reconstructions with tdelay=100 ms. We visualize the Fx and My , which are representative channels in our experimental setting described in Section 5.1 . The upper two rows correspond to the ‘Soft-1’ episode, and the lower rows correspond to the ‘Stiff-1’ episode. Colored areas illustrate the prediction interval defined by μ±3σ .
Figure 7: Band-specific errors evaluated at every horizon step from t+1 to t+100 , computed on a normalized scale and aggregated over channels and test episodes. (a) and (b) report pRMSE and wRMSE , respectively. The upper axis gives the corresponding time delay at the 100 Hz sampling rate, where t+10 and t+100 are the delays reported in Tables 2 and 3 .
HF wRMSE
LF pRMSE
CRPS
w/o FEF
0.518 ± 0.002 (+3.4%)
0.569 ± 0.006 (+6.8%)
0.534 ± 0.003 (+3.1%)
w/o FEF-W
0.518 ± 0.003 (+3.4%)
0.572 ± 0.009 (+7.3%)
0.534 ± 0.005 (+3.1%)
w/o FEF-MoE
0.515 ± 0.002 (+2.8%)
0.565 ± 0.000 (+6.0%)
0.531 ± 0.001 (+2.5%)
w/o FPF
0.500 ± 0.001 (-0.2%)
0.537 ± 0.012 (+0.8%)
0.520 ± 0.005 (+0.4%)
w/o ModSpec
0.553 ± 0.002 (+10.4%)
0.606 ± 0.001 ( +13.7% )
0.559 ± 0.001 ( +7.9% )
w/o TrdHead
0.615 ± 0.049 ( +22.8% )
0.631 ± 0.063 ( +18.4% )
0.513 ± 0.013 (-1.0%)
Table 5: Architectural ablation results. Metrics are aggregated on the normalized scale.
Ablated Inputs
HF wRMSE
LF pRMSE
CRPS
q˙
0.500 ± 0.004 (+0.2%)
0.583 ± 0.014 (+9.6%)
0.542 ± 0.007 (+4.6%)
q¨
0.499 ± 0.002 (+0.0%)
0.574 ± 0.026 (+7.9%)
0.538 ± 0.013 (+3.9%)
q˙,q¨
0.531 ± 0.002 (+6.4%)
0.587 ± 0.041 (+10.3%)
0.539 ± 0.024 (+4.1%)
u(=Δp)
0.572 ± 0.003 (+14.6%)
0.604 ± 0.003 (+13.5%)
0.554 ± 0.002 (+6.9%)
q˙,q¨,u
0.701 ± 0.050 (+40.5%)
0.666 ± 0.014 (+25.2%)
0.583 ± 0.007 (+12.5%)
∅
0.499 ± 0.003
0.532 ± 0.007
0.518 ± 0.002
Table 6: Effect of input variables on estimation performance. Metrics are aggregated on the normalized scale.
Figure 8: Channel-wise marginal distributions of the normalized training residuals. Each panel overlays histograms of the three force residual channels (Left) and three torque residual channels (Right). Solid lines indicate fitted Gaussian probability density functions.
Figure 9: Correlation analysis of the normalized training residuals. (a) Channel-wise temporal autocorrelation functions of the six wrench residual channels over the T=100 prediction horizon. (b) Pearson cross-channel correlation matrix averaged over the horizons.
HF wRMSE
LF pRMSE
Correlation DoF
IT⊗RC (Channel)
0.533 ± 0.004 (+6.8%)
0.555 ± 0.020 (+4.3%)
15
RT⊗IC (Time)
0.602 ± 0.003 (+20.6%)
0.551 ± 0.018 (+3.6%)
4,950
RT⊗RC (Time-Channel)
0.638 ± 0.005 (+27.9%)
0.554 ± 0.018 (+4.1%)
4,965
I (Independent)
0.499 ± 0.003
0.532 ± 0.007
0
Table 7: Effect of residual correlation structure on band-specific estimation performance. Metrics are aggregated on the normalized scale.
Wall Time [ms]
Relative
Optimized Parameters
MINN
0.038 ± 0.003
0.004
3,974
RBF
0.062 ± 0.004
0.006
868
GPR
16.663 ± 0.737
1.582
3,302,563
LSTM
6.229 ± 0.423
0.591
80,902
CNN
5.936 ± 0.226
0.564
81,766
LSTM-ED
13.186 ± 0.520
1.252
151,814
Table 8: CPU inference time and model sizes.
Datasets
Pretraining (RH20T)
Downstream
HF W Energy
9.839%
85.188%
LF W Energy
90.161%
14.812%
Task
Everyday Manipulation
Block Grinding
Actuation
Electric
Hydraulic
Actuation Signal
Motor Torque
Differential Pressure
Drive Type
Revolute
Prismatic and Revolute
Table 10: Properties of pretraining and downstream datasets, including band energy ratios of the wrench. ‘LF’ and ‘HF’ indicate bands in f≤1 Hz and f>1 Hz, respectively.
Figure 10: Channel-averaged power spectrum of W in the pretraining (RH20T) and downstream (Hydraulic) datasets, computed from the training splits only. Each channel is normalized to zero mean and unit variance before the Fourier transform, and the spectra are shown without denoising to display the full frequency content. Both axes are logarithmic. Vertical lines mark the cutoffs fc=1 Hz and fcdn=15 Hz applied during training.
Humanoid robots are entering our physical world at scale, yet as oversized toys--good at singing and dancing, but short on force-interaction capabilities for practical tasks. Bridging this gap necessitates prioritizing reliable contact perception as a fundamental requirement. Estimating external wrenches in humanoids is complicated by floating-base dynamics and indeterminate contact locations. Existing analytical frameworks require idealistic assumptions and hard-to-obtain measurements, which are often unavailable in practice. To bridge this gap, we propose SixthSense, a task-agnostic approach that infers whole-body contact timing, location, and wrenches from proprioception and IMU data alone. To capture the multi-modal dynamics between unstructured contact inputs and the uncertain motion outputs, we employ conditional flow matching to tokenize proprioceptive histories and estimate a spatiotemporally sparse contact-event flow. SixthSense serves as a plug-and-play perception module for applications including collision detection, physical human-robot interaction, and force-feedback teleoperation. Experiments across standing, walking, and whole-body motion-tracking policies showcased unprecedented performance in diverse behaviors.
Xingzhou Chen, Xiayan Xu, Yan Ning +8
The Hong Kong University of Science and Technology, Hong Kong SAR, China · Zhejiang University, Hangzhou, China · Tencent Robotics X, Shenzhen, China
This paper presents a scalable, open-source visuotactile sensing system for tensegrity robots that enables six-axis wrench estimation and contact detection. The proposed endcap sensor integrates an elastomeric shell, a 3D-printed thermoplastic polyurethane (TPU) interface, and a rigid base housing an embedded camera and LED illumination ring. A novel gyroid-infill bonding technique is introduced to form a durable elastomer-TPU interface without adhesives, yielding a lightweight and modular design compatible with large-scale tensegrity structures. A tactile-to-wrench neural network maps shear vector fields to six-dimensional force and torque measurements. Experimental results demonstrate accurate and stable wrench estimation with a mean squared error (MSE) of 0.1531 on static validation data and out-of-domain generalization under dynamic motion. Furthermore, full-system integration on a 12 kg tensegrity robot confirms the sensor's ability to reliably identify ground contacts. The system substantially improves the practicality of tactile feedback for tensegrity robots, offering a low-cost, reproducible, and physically interpretable pathway toward contact-aware proprioception and state estimation. Open source files are available at \href{https://github.com/Jonathan-Twz/tensegrity-gelfoot}{github.com/Jonathan-Twz/tensegrity-gelfoot}
Wenzhe Tong, Jonathan Mi, Xili Yi +2
Robotics Department, University of Michigan, Ann Arbor, MI, 48109, USA
Robotic surface swabbing requires sustained interaction between a compliant tool and heterogeneous environments, where accurate estimation of tip-level contact force is critical for consistent sampling performance. However, deformable tool dynamics introduce nonlinear viscoelastic hysteresis that decouples wrist-mounted force measurements from true contact forces, while tool-integrated sensors are impractical for deployment due to sterility and disposability constraints. This paper presents a data-driven framework for contact force estimation in Deformable Tool Manipulation (DTM) that leverages proprioceptive sensing without requiring explicit physical models or permanent embedded sensing hardware at the tool tip. A recurrent architecture is first identified through a comparative evaluation of temporal models, where a compact LSTM achieves the lowest estimation error and sub-millisecond inference latency. To address generalization across unseen surfaces and tool compliance conditions, we introduce a parameter-isolated few-shot adaptation strategy that augments a frozen recurrent backbone with low-dimensional context embeddings using feature-wise linear modulation (FiLM). Experiments on a UR5e platform across nine tool-surface interaction regimes demonstrate that the proposed approach significantly improves robustness under domain shift, reducing zero-shot estimation error by up to 63% while preserving baseline performance without catastrophic forgetting. These results show that separating shared deformation-history dynamics from domain-specific conditioning enables reliable force estimation for DTM in non-stationary environments.
Siavash Mahmoudi, Chaitainya Kuppar Reddy, Yang Tian +1
Department of Biological and Agricultural Engineering, University of Arkansas, Fayetteville, AR 72701, USA · Department of Food Science, University of Arkansas, Fayetteville, AR 72701, USA