Video2SwimFish: An Automated Pipeline for Reconstructing Controllable Fish Models and Biological Locomotion from Real Fish Videos
Authors: Hangong Chen, Linfeng Cheng, Tahsin Zaman Jilan, Ian Fuller, Lee Caesar, Jiaye Wu, Yantian Zha
Organizations: Department of Computer Science, North Carolina A&T State University · Department of Computer Science, University of Maryland, College Park
We present Video2SwimFish, an automated pipeline and benchmark for building controllable fish assets from real-fish videos for underwater embodied AI. Given synchronized multi-view videos of an individual fish, the pipeline reconstructs a metrically scaled deformable mesh from a VLM-selected canonical frame, generates internal articulation adapted to that individual's morphology through a VLM actor-critic loop, and extracts a Biological Locomotion Manifold (BLM) from the fish's observed midline curvature. The BLM provides a low-dimensional action space bounded by real-fish motion, enabling an individual swimming policy to be learned for each reconstructed fish. We release two paired datasets: synchronized top- and front-view recordings of 120 individual fish across 6 species, and the controllable assets and individual swimming policies derived from them. Because every asset is tied to the animal it came from, the dataset supports a benchmark that evaluates locomotion learning not only on task success but on fidelity to that individual in trajectory shape, body curvature, and tail-beat frequency, across trajectory following, reward-free swimming behavior transfer from video, and a downstream case study in which a simulated BlueROV underwater robot captures one of the assets. We find that task success and locomotion fidelity do not necessarily improve together: the method achieving the highest task completion is not the method achieving the highest locomotion fidelity, and we identify faithful reproduction of individual animal locomotion as an open challenge for the community. Project website: https://hangongchen.github.io/video2swimfish-web/.
Figures & tables
Figure 2: The Video2SwimFish articulation generation pipeline. A VLM selects the most canonical frame from the multi-view video, from which Meshy reconstructs a 3D fish mesh. An actor VLM then proposes a skeleton as a list of bone templates with positions and sizes, and a critic VLM scores it and returns feedback; the two iterate until the critic accepts.
Method
Completion ↑
Fréchet ↓
W1(κ)↓
Δf↓
Δv↓
Joint RL
0.70
0.48
1.43
1.61
0.52
Joint RL + AMP
0.51
0.34
0.78
2.11
0.64
CPG + RL
0.67
0.58
1.16
0.70
0.38
BCO + RL
0.47
0.56
1.89
0.81
0.44
BLM + RL (Ours)
0.89
0.65
1.01
0.56
0.33
Table 1: Task 1 trajectory following on twelve fish (two per species): completion rate, Fréchet distance (BL), curvature Wasserstein distance W1(κ) , dominant frequency error Δf (Hz) and speed error Δv (BL/s); mean over fish. Best per column in bold.
Figure 3: Qualitative comparison of five methods following the same reference trajectory in Task 1. Snapshots are shown at 0.4-s intervals. The dashed black curve denotes the reference trajectory, the orange trail shows the simulated fish’s executed path, and the star marks the real fish’s position on the reference trajectory at each time step. Completion status is indicated on the right.
Method
Fréchet ↓
Speed err. ↓
Heading stab. ↑
W1(κ)↓
Δf↓
BCO
2.79
0.39
0.56
1.63
0.012
BLM + IL
2.08
0.11
0.56
0.74
0.038
Table 2: Task 2 free swimming from video on six fish (one per species): 20 real initial states, 5 s roll-outs, no reward. Best per column in bold.
Prey
Capture success rate ↑
Epochs to convergence
Channel catfish 002
0.988
44
Lake sturgeon 016
0.992
49
Table 3: Case Study: ROV capture of video-driven fish.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Species
n
length [cm]
bones (med.)
coverage
alignment
own PCA basis
Channel catfish
20
12.6–31.1 (med. 20.4)
14.5
74%
10.9∘
20/20
Lake sturgeon
20
4.6–18.9 (med. 10.0)
15
72%
8.7∘
18/20
Bluegill
20
1.6–9.5 (med. 6.8)
15
74%
11.1∘
18/20
White bass
20
4.3–12.9 (med. 6.3)
15
74%
10.6∘
17/20
Brook Trout
20
5.3–16.7 (med. 10.2)
14
73%
10.0∘
18/20
Brown Trout
20
4.7–12.1 (med. 9.3)
15
75%
11.3∘
17/20
Appendix
Table 4: Dataset summary. Coverage is the fraction of body length spanned by the bone chain; alignment is the mean angle between each bone axis and the local body centerline.
Species
Fish
IoU (similarity)
IoU (affine)
IoU (body only)
bluegill
013
0.85
0.88
0.88
bluegill
019
0.80
0.82
0.81
catfish
002
0.86
0.90
0.89
catfish
013
0.77
0.84
0.78
lake sturgeon
002
0.88
0.91
0.93
lake sturgeon
006
0.78
0.85
0.84
Appendix
Table 5: Silhouette IoU between the generated mesh (orthographic lateral render) and its canonical input frame for the two most cleanly segmented fish per species. Similarity: 4-DOF alignment with the head end matched; affine: 6-DOF refinement; body only: both silhouettes with fins morphologically stripped before the affine IoU.
Figure 4: Mesh-versus-photo silhouettes for the eight fish of Table 5 . Each panel shows the canonical crop with its U 2 -Net segmentation (left), the lateral render of the generated mesh (middle), and the two silhouettes after head-matched affine alignment (right; white: both, red: photo only, blue: mesh only). The remaining disagreement sits in the fin pose and a slight body bend of the photographed animal.
We propose a new method for synthesizing physically-based swimming motions. Physically-based character animation aims to generate physically valid, controllable, and natural-looking motions which can respond to unexpected disturbances, where one dictating factor of difficulty is the complexity of the task, especially the level of sophistication of the required interactions with the environment. Existing research has succeeded in various tasks in static and dynamic environments. We push the difficulty further to swimming, which requires full-body coordination and continuous interactions with fluids, a new level of complexity when it comes to interacting with the environment. This complexity imposes challenges in learning control under volatile environmental forces, generalizing control to different environments and swimming styles, lack of data references, and prohibitively slow physical simulation which is inevitable during control learning. To this end, we propose SWIM, a new imitation method for swimming motions, which can learn from a single swimming motion and generalize to unseen environments, body conditions, and swimming styles. Extensive evaluation and comparison demonstrate that SWIM is data-efficient, stable, robust, and generalizable, outperforming alternative methods across multiple classes of tasks and metrics.
Complex tasks for underwater robots remain limited by the capabilities of their controllers. Learning a better one for a soft, underactuated robotic fish trades simulator cost against fidelity. We show that an intentionally low-fidelity simulator is enough: a stateless, quasi-steady fluid model with no wake and no added-mass history suffices to learn a \emph{general}, closed-loop controller that transfers to hardware without tuning. Our platform is a soft, single-motor, tendon-driven fish whose policy observes only what the hardware can measure. A staged pipeline grounds the simulator in two independent identifications, fixing the tail dynamics and a stateless fluid model; the policy then acts through a band-limited rhythmic trajectory generator rather than commanding the tail directly. Deployed unchanged in an outdoor pool, a single policy performs closed-loop target reaching, disturbance rejection, and out-of-distribution target acquisition and tracking. The transfer rests on the constraint rather than the fidelity: the generator cannot leave the band over which the fluid was identified. This raises the question of how much of the physics can reside in the controller rather than in the simulator.
Liam Maloney, Simon Ramchandani, Mike Y. Michelis +2
Soft Robotics Lab, ETH Zurich, Switzerland · ETH AI Center, ETH Zurich, Switzerland
Fish-like swimming has inspired the design of several dozens if not hundreds of bioinspired robots in the last few decades. But the control and motion planning of such robots has been challenging due to the poorly modeled fluid-structure interaction and the nonlinear underactuated dynamics of such robots. While reinforcement learning has allowed significant advances in the context of ground and aerial robots, the lack of a suitable simulation environment with appropriate computational speed and accuracy have prevented similar progress for fish-like robots. We address this two-fold problem by developing a simulation platform that approximates the motion of our fish-like robot with computational efficiency. Then the motion control and path tracking by the robot is performed using PID control where the (variable) gains are learned using back propagation through time and training on a curriculum. The policy learned in the simulation is then applied on the physical platform, demonstrating an excellent match.