Beyond Masks and Trajectories: Flow-Guided Latent Action Injection for Stable Surgical Video Generation
Organizations: The Hong Kong University of Science and Technology
Abstract
Surgical video generation holds substantial potential for surgical education, simulation, and data augmentation, yet generating surgical videos with realistic and clinically plausible motion remains challenging. Most existing methods rely on auxiliary conditions, such as masks, trajectories, depth, or reference videos, to achieve visually plausible synthesis. Yet, these auxiliary conditions typically require additional manual annotation or specialized acquisition, making it difficult to scale such methods beyond small, curated datasets. This motivates the need for a reference-free architecture capable of generating high-quality surgical video without requiring auxiliary visual conditions at inference time. We propose FLAIR, a Flow-guided LatentAction Injection framework for Reference-free surgical video generation. FLAIR learns action priors from optical flow of real surgical videos, dynamically predicts corresponding latent action representation from an input prompt, and injects it into a frozen base model to generate surgical videos with improved action consistency. We further construct SurgActionClip-30K, the first large-scale surgical vision dataset comprising action-centric segmented clips and structured caption labels, addressing the persistent lack of fine-grained, action-centric surgical datasets. Lastly, we introduce SurgMetrics, the first surgical domain-specific evaluation metrics for quantifying the quality of generated surgical videos, addressing the persistent absence of clinically grounded evaluation standards in this domain. Extensive experiments demonstrate that FLAIR enables generating high-quality surgical videos using text-only inference without auxiliary conditions, and validation in SurgMetrics demonstrates its strength in alignment with human perception compared to traditional metrics.
Figures & tables
| Method | General Video Metrics | SurgMetrics | ||||||||||
| V.CLIP | FID | FVD | TI | AF | Sem | App | Drift | Inst | Tissue | Track | Dom | |
| Cosmos-H-Surgical | – | – | – | 0.780 | 8.52 | – | – | – | – | – | – | – |
| KVLR | – | – | – | 0.893 | 4.48 | – | – | – | – | – | – | – |
| Wan2.1 Base | 8.80 | 246.8 | 3466 | 0.814 | 1.50 | – | – | – | – | – | – | 4.228 |
| Wan2.1 LoRA | 8.82 | 126.3 | 1340 | 0.877 | 0.11 | 0.097 | 0.132 | 0.027 | 0.059 | 0.085 | 0.100 | 0.931 |
| Wan2.1 LoRA-FLAIR (ours) | 9.08 | 117.2 | 1148 | 0.913 | 0.03 | 0.029 | 0.175 | -0.005 | 0.027 | -0.023 | -0.001 | 1.165 |
| Latent Condition | Next-Frame MSE |
|---|---|
| No latent | 0.002521 (reference) |
| Zero latent | 0.002525 |
| Shuffled FAE latent | 0.002601 |
| FLAIR FAE latent | 0.002356 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Source | Procedure or domain | Boundary supervision | Retained clips |
|---|---|---|---|
| CholecT50 | Laparoscopic cholecystectomy | Instrument–verb–target triplets | 3,773 |
| M2CAI16 | Laparoscopic cholecystectomy | Phase labels; detector refinement | 3,086 |
| AutoLaparo | Laparoscopic hysterectomy | Phase labels; detector refinement | 3,416 |
| SLAM | Laparoscopic cholecystectomy | Native action groups | 221 |
| PhaKIR | Laparoscopic surgery | Joint phase and instrument labels | 1,268 |
| SurgicalActions160 | Laparoscopic action demonstrations | Native action classes | 75 |
| ID | Procedure | Count | Design |
| 000–029 | Cholecystectomy | 30 | 10 actions 3 variants |
| 030–059 | Gastrectomy | 30 | 10 actions 3 variants |
| 060–089 | Hysterectomy | 30 | 10 actions 3 variants |
| 090–119 | Intestinal Resection | 30 | 10 actions 3 variants |
| 120–149 | Nephrectomy | 30 | 10 actions 3 variants |
| 150–179 | Prostatectomy | 30 | 10 actions 3 variants |
| Evaluation setting | AUC | |
|---|---|---|
| Held-out cases (multi) | ||
| External (single) | ||
| External (multi) |
| Dataset | Mean robust | AUROC vs. in-domain |
|---|---|---|
| In-domain unseen test | 0.5614 | 0.5000 |
| PhaKIR | 1.8538 | 0.7646 |
| SLAM | 1.9552 | 0.7933 |
| OphNet | 3.6019 | 0.9547 |
| SurgVU24 | 4.9917 | 0.9923 |
| UCF101 | 5.8199 | 0.9957 |
| Held-out | External | ||||||
|---|---|---|---|---|---|---|---|
| Configuration | Params | Test Acc | Internal | multi | AUC | single | single AUC |
| Full | 354,502 | 0.8250 | 0.5436 | 0.5661 | 0.9144 | 0.7307 | 0.9243 |
| SurgViSTA-only | 345,478 | 0.7917 | 0.5144 | 0.5278 | 0.9029 | 0.5934 | 0.8556 |
| Optical-only | 92,102 | 0.6569 | 0.2779 | 0.2452 | 0.6376 | 0.5752 | 0.7729 |
| Configuration | Adapter test | Internal |
|---|---|---|
| Drift Acc | Drift | |
| Full | 0.9500 | 0.6106 |
| SurgViSTA-only | 0.7250 | 0.4375 |
| Optical-only | 0.5333 | 0.5235 |
| Backbone | Params | Training budget | Resolution | Selected step | Adapter scale |
|---|---|---|---|---|---|
| Wan2.1-T2V-1.3B | 1,904,896 | 1k / 1 GPU | 480 832 | 500 | 0.30 |
| HunyuanVideo 1.5 | 2,430,208 | 1k / 4 GPU (DDP) | 320 576 | 1,000 | 0.50 |
| CogVideoX-2B | 1,576,576 | 1k / 8 GPU (DDP) | 480 832 | 400 | 0.50 (LoRA 1/32) |
| Parameter | Wan2.1-1.3B | HunyuanVideo 1.5 | CogVideoX-2B |
|---|---|---|---|
| Resolution (W H) | 832 480 | 576 320 | 832 480 |
| Frames | 81 | 81 | 81 |
| FPS | 16 | 24 | 16 |
| Denoising steps | 20 | 30 | 50 |
| Guidance scale | 5.0 | 6.0 | 6.0 |
| Flow shift | 5.0 | 2.0 | – |
| Degradation | Description | Head |
| A. Drift: temporal and global-motion degradations | ||
| local temporal rewind | Reverses or loops a local temporal segment (8%–38% of the clip), producing action rewind or repetition. | Drift |
| temporal shuffle | Swaps adjacent short temporal blocks while preserving intra-block motion order, disrupting local event ordering. | Drift |
| frame freeze drop | Replaces random frames with the preceding frame, producing freezing, duplication, and discontinuous jumps after drops. | Drift |
| motion jerk | Applies stride-and-hold sampling within a short segment, producing sudden acceleration, pauses, and jerks. | Drift |
| speed distortion | Applies non-linear temporal resampling across the clip, causing gradual speed-up or slow-down instead of uniform playback. | Drift |
| Family | Sem. | App. | Drift | Instr. | Tissue | Track. |
|---|---|---|---|---|---|---|
| local temporal rewind | 0.35 | – | 1.00 | – | – | – |
| temporal shuffle | – | – | 0.75 | – | – | – |
| frame freeze drop | – | – | 0.80 | – | – | – |
| motion jerk | – | – | 0.80 | – | – | – |
| speed distortion | – | – | 0.80 | – | – | – |
| global jitter | – | – | 0.90 | – | – | – |