SurgGMF: Fully Causal Gaussian Motion Forecasting for Anticipatory Surgical Scene Rendering
Authors: Jingqian Sun, Yichao Tang
Organizations: Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University, Shanghai, China · School of Mechanical Engineering, Tongji University, Shanghai, China · Shanghai Innovation Institute, Shanghai, China
Dynamic surgical scene modeling is essential for robotic perception, simulation, and decision support. Although existing neural rendering methods enable efficient reconstruction and rendering of deformable surgical scenes, they remain primarily focused on observed-frame reconstruction rather than forecasting future scene states. To this end, we present SurgGMF, a fully causal Gaussian motion forecasting framework for anticipatory surgical scene rendering. Rather than predicting future RGB images directly, SurgGMF forecasts future Gaussian motion states represented by position, scale, and rotation residuals (X/S/R) from historical Gaussian motion fields. To prevent target leakage, we introduce a full-causal-last rendering protocol, where future Gaussian states are rendered without accessing target-frame Gaussian attributes while preserving causal appearance propagation. We evaluate SurgGMF on 12 EndoNeRF and StereoMIS video slices using neural temporal learners and classical dynamics baselines under a unified forecasting protocol. Learned Gaussian motion forecasting consistently outperforms classical dynamics baselines in render space, demonstrating gains beyond hand-crafted state extrapolation. Latency analysis further reveals an accuracy--efficiency trade-off: under the current implementations, TKAN achieves the highest accuracy, whereas GRU and LSTM provide more favorable module-level latency profiles. These results establish SurgGMF as a reproducible framework for causal Gaussian motion forecasting and advance surgical Gaussian representations from retrospective reconstruction toward predictive scene modeling.
Figures & tables
Figure 1: Overview of SurgGMF. A trained Deform3DGS teacher exports temporally aligned Gaussian states. SurgGMF constructs X/S/R motion histories, forecasts future Gaussian motion states with a temporal learner or a dynamics baseline, converts predicted residuals back into future Gaussian states, and evaluates rendered future frames under the full-causal-last protocol.
Figure 2: Qualitative future-frame rendering comparison under the full-causal-last protocol at an intermediate horizon ( k=3 ). The layout compares the target, persistence, constant velocity, GRU, LSTM, and TKAN forecasts, together with selected error maps. Error maps show absolute RGB differences to the target under the same valid mask; darker colors indicate smaller errors. Red regions denote surgical instrument masks, which are excluded from Deform3DGS reconstruction and rendering and are shown only for visualization.
Slice
Source
Pairs / method
Rendered instances
01
EndoNeRF cutting
75
675
02
EndoNeRF pulling
25
225
03
StereoMIS P1A
105
945
04
StereoMIS P1B
115
1,035
05
StereoMIS P2-0
90
810
06
StereoMIS P2-2
90
810
Table 1: Evaluation slices and valid render-space pairs. Pair counts denote valid frame–horizon pairs per forecasting method after boundary filtering; rendered instances multiply each count by the nine evaluated methods.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
MAE ↓
Persistence
29.480
0.8133
0.1818
0.0230
Const. Vel.
30.691
0.8343
0.1888
0.0203
Linear fit
28.717
0.7892
0.2004
0.0255
Const. Accel.
29.918
0.8145
0.2185
0.0234
Kalman-CV
26.638
0.7259
0.2820
0.0331
GRU
31.801
0.8611
0.1768
0.0175
Table 2: Average render-space performance over 12 surgical video slices. Higher PSNR and SSIM are better; lower LPIPS and MAE are better. Const. Vel. denotes constant velocity; Const. Accel. denotes constant acceleration.
Model
Target
PSNR ↑
SSIM ↑
LPIPS ↓
MAE ↓
LSTM
X
30.570
0.8401
0.1798
0.02039
LSTM
X+S
30.900
0.8438
0.1791
0.01952
LSTM
X+R
31.542
0.8589
0.1774
0.01809
LSTM
X+S+R
31.810
0.8614
0.1768
0.01746
TKAN
X
30.570
0.8404
0.1798
0.02037
TKAN
X+S
30.917
0.8442
0.1791
0.01946
Table 3: Forecasting-target ablation with independently trained variants. X predicts position only; X+S adds raw scale residuals; X+R adds raw rotation residuals; X+S+R predicts all three components.
Model
Precision
Latency (ms) ↓
FPS ↑
GRU
FP32
35.08
28.5
GRU
AMP
18.14
55.1
LSTM
FP32
46.67
21.4
LSTM
AMP
24.05
41.6
Transformer
FP32
151.99
6.6
Transformer
AMP
71.66
14.0
Table 4: Minimal forecast–render latency for the PyTorch backbones. FPS is derived from the corresponding module latency.
Reconstructing dynamic surgical scenes is crucial for robot-assisted minimally invasive surgery; however, it continues to be difficult because of tissue deformation, occlusions, specular reflections, and restricted viewpoints. In this study, we introduce Endo-NeRF++, a neural rendering framework that accounts for uncertainty in the reconstruction of dynamic surgical scenes. Expanding on EndoNeRF, the suggested approach incorporates multi-resolution hash-grid encoding, temporal feature merging, and uncertainty-informed adaptive sampling to enhance reconstruction accuracy and temporal coherence in deformable endoscopic scenes.The multi-resolution hash-grid representation within the framework effectively captures both coarse and fine anatomical details, while temporal feature blending ensures stable reconstruction during tissue deformation and surgical tool occlusions. Additionally, uncertainty-driven adaptive sampling assigns more samples to uncertain areas to enhance rendering quality and geometric coherence. Experiments on robotic surgical video sequences demonstrate that the proposed uncertainty-guided adaptive sampling improves PSNR by up to 1.22dB (4.3%), increases SSIM by up to 5.3%, and reduces LPIPS by up to 55.1% compared with the EndoNeRF baseline.
Gousia Habib, Laura Ruotsalainen
Department of Computer Science, University of Helsinki, Finland
Dynamic endoscopic reconstruction is fundamental to robotic surgery and computer-assisted interventions. While 3D Gaussian Splatting (3DGS) realises real-time rendering, its application to deformable intraoperative environments remains constrained by spurious geometry and varying illuminations. To address these limitations, we introduce EndoPrior-GS, a novel pipeline that explicitly couples frame-extracted vision heuristics and estimated depth maps. EndoPrior-GS derives a joint texture prior from a tool-filtered valid tissue mask, a non-specular photometric filter, and anatomical structural salience, yielding a probability map that guides primitive initialisation and subsequent density control. The prior is further extended to the temporal domain through a texture-aware term that dynamically weighs pairwise primitive contributions during training. We conduct extensive experiments on benchmark datasets EndoNeRF and SCARED, and the obtained results show that our method EndoPrior-GS reduces Flow Error by 27.7% and 25.8% over the representative approaches while preserving competitive rendering quality and real-time rendering speed. Our project website is available at https://jiaqi-huang-77.github.io/EndoPrior-GS/.
Jiaqi Huang, Shidong Wang, Tong Xin +1
School of Engineering, Newcastle University, Newcastle upon Tyne NE1 7RU, UK · School of Computing, Newcastle University, Newcastle upon Tyne NE4 5TG, UK
Reliable surgical planning requires models that move beyond recognizing the current surgical step or imitating expert demonstrations, and instead anticipate how instrument motion reshapes subsequent operative states. Most surgical video understanding methods focus on recognizing phases, actions, or workflow states, while providing limited support for explicitly modeling instrument motion. Conversely, existing tool motion prediction methods can forecast instrument trajectories, but they generally do not capture the coupled evolution of future surgical video states. World models offer a natural framework for jointly modeling visual state transitions and instrument motion dynamics. However, existing surgical world model studies remain largely centered on visual generation quality, relying on generation-oriented metrics such as FVD and CD-FVD. These metrics are poorly aligned with instrument motion planning, as they do not directly measure whether predicted trajectories are geometrically accurate, temporally coherent, or actionable for downstream planning. This limitation is partly structural, since the field lacks public datasets and standardized evaluation protocols that provide the benchmarking infrastructure needed to assess motion-centric capabilities in surgical world models. In this paper, we introduce SurgWMBench, a vision-based benchmark for short-horizon surgical motion planning and dynamics prediction. Given intraoperative image sequences and historical instrument trajectory, SurgWMBench evaluates both near-future instrument motion prediction and stability under continuous rollout or input perturbations.
Huanrong Liu, Weiliang Huang, Bob Zhang +3
University of Macau · Xiamen University · National University of Singapore