SurgGMF: Fully Causal Gaussian Motion Forecasting for Anticipatory Surgical Scene Rendering
Authors: Jingqian Sun, Yichao Tang
Organizations: Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University, Shanghai, China · School of Mechanical Engineering, Tongji University, Shanghai, China · Shanghai Innovation Institute, Shanghai, China
Dynamic surgical scene modeling is essential for robotic perception, simulation, and decision support. Although existing neural rendering methods enable efficient reconstruction and rendering of deformable surgical scenes, they remain primarily focused on observed-frame reconstruction rather than forecasting future scene states. To this end, we present SurgGMF, a fully causal Gaussian motion forecasting framework for anticipatory surgical scene rendering. Rather than predicting future RGB images directly, SurgGMF forecasts future Gaussian motion states represented by position, scale, and rotation residuals (X/S/R) from historical Gaussian motion fields. To prevent target leakage, we introduce a full-causal-last rendering protocol, where future Gaussian states are rendered without accessing target-frame Gaussian attributes while preserving causal appearance propagation. We evaluate SurgGMF on 12 EndoNeRF and StereoMIS video slices using neural temporal learners and classical dynamics baselines under a unified forecasting protocol. Learned Gaussian motion forecasting consistently outperforms classical dynamics baselines in render space, demonstrating gains beyond hand-crafted state extrapolation. Latency analysis further reveals an accuracy--efficiency trade-off: under the current implementations, TKAN achieves the highest accuracy, whereas GRU and LSTM provide more favorable module-level latency profiles. These results establish SurgGMF as a reproducible framework for causal Gaussian motion forecasting and advance surgical Gaussian representations from retrospective reconstruction toward predictive scene modeling.
Figures & tables
Figure 1: Overview of SurgGMF. A trained Deform3DGS teacher exports temporally aligned Gaussian states. SurgGMF constructs X/S/R motion histories, forecasts future Gaussian motion states with a temporal learner or a dynamics baseline, converts predicted residuals back into future Gaussian states, and evaluates rendered future frames under the full-causal-last protocol.
Figure 2: Qualitative future-frame rendering comparison under the full-causal-last protocol at an intermediate horizon ( k=3 ). The layout compares the target, persistence, constant velocity, GRU, LSTM, and TKAN forecasts, together with selected error maps. Error maps show absolute RGB differences to the target under the same valid mask; darker colors indicate smaller errors. Red regions denote surgical instrument masks, which are excluded from Deform3DGS reconstruction and rendering and are shown only for visualization.
Slice
Source
Pairs / method
Rendered instances
01
EndoNeRF cutting
75
675
02
EndoNeRF pulling
25
225
03
StereoMIS P1A
105
945
04
StereoMIS P1B
115
1,035
05
StereoMIS P2-0
90
810
06
StereoMIS P2-2
90
810
Table 1: Evaluation slices and valid render-space pairs. Pair counts denote valid frame–horizon pairs per forecasting method after boundary filtering; rendered instances multiply each count by the nine evaluated methods.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
MAE ↓
Persistence
29.480
0.8133
0.1818
0.0230
Const. Vel.
30.691
0.8343
0.1888
0.0203
Linear fit
28.717
0.7892
0.2004
0.0255
Const. Accel.
29.918
0.8145
0.2185
0.0234
Kalman-CV
26.638
0.7259
0.2820
0.0331
GRU
31.801
0.8611
0.1768
0.0175
Table 2: Average render-space performance over 12 surgical video slices. Higher PSNR and SSIM are better; lower LPIPS and MAE are better. Const. Vel. denotes constant velocity; Const. Accel. denotes constant acceleration.
Model
Target
PSNR ↑
SSIM ↑
LPIPS ↓
MAE ↓
LSTM
X
30.570
0.8401
0.1798
0.02039
LSTM
X+S
30.900
0.8438
0.1791
0.01952
LSTM
X+R
31.542
0.8589
0.1774
0.01809
LSTM
X+S+R
31.810
0.8614
0.1768
0.01746
TKAN
X
30.570
0.8404
0.1798
0.02037
TKAN
X+S
30.917
0.8442
0.1791
0.01946
Table 3: Forecasting-target ablation with independently trained variants. X predicts position only; X+S adds raw scale residuals; X+R adds raw rotation residuals; X+S+R predicts all three components.
Model
Precision
Latency (ms) ↓
FPS ↑
GRU
FP32
35.08
28.5
GRU
AMP
18.14
55.1
LSTM
FP32
46.67
21.4
LSTM
AMP
24.05
41.6
Transformer
FP32
151.99
6.6
Transformer
AMP
71.66
14.0
Table 4: Minimal forecast–render latency for the PyTorch backbones. FPS is derived from the corresponding module latency.
School of Engineering, Newcastle University, Newcastle upon Tyne NE1 7RU, UK · School of Computing, Newcastle University, Newcastle upon Tyne NE4 5TG, UK