FiVOS: A Fish Segmentation Algorithm Based on Interactive Video Object Segmentation and Filter Enhancement
Authors: Yuqing Duan, Song Zhang, Shili Zhao, Daoliang Li, Ran Zhao
Organizations: National Innovation Center for Digital Fishery, China Agricultural University, Beijing 10083, China · Key Laboratory of Smart Farming Technologies for Aquatic Animal and Livestock, Ministry of Agriculture and Rural Affairs, China Agricultural University, Beijing 10083, China · Beijing Engineering and Technology Research Center for Internet of Things in Agriculture, China Agricultural University, Beijing 10083, China · College of Information and Electrical Engineering, China Agricultural University, China Agricultural University, Beijing 10083, China
With the continuous expansion of aquaculture, precise and efficient monitoring of fish behavior has become increasingly critical for improving farming efficiency and reducing economic losses. In particular, with the ongoing enhancement of computational capabilities in deep learning models, vision-based fish segmentation methods are garnering growing attention. By analyzing video segmentation results, fish behavior can be effectively tracked, thereby providing reliable data support for the precise regulation of aquaculture environments. However, existing deep learning-based video segmentation methods for aquaculture scenarios often overlook the dynamic correlations between video frames. In contrast, Interactive Video Object Segmentation (IVOS) employs an interaction-propagation scheme to achieve high-precision segmentation while minimizing user effort, thereby enhancing monitoring efficiency. Yet, IVOS applications in aquaculture remain limited due to data scarcity, and are susceptible to error accumulation and mask loss over long sequence propagation due to high intra-class similarity. In response, this paper proposes an improved interactive video object segmentation method (FiVOS) and constructs two fish-specific datasets. FiVOS utilizes a mask block filter to enable early detection and correction of erroneous propagated mask blocks, enhancing filtering accuracy through a rule-based thresholding approach. Additionally, it serializes noise filters to further eliminate erroneous mask noise, thereby improving model robustness. Experimental results demonstrate that FiVOS achieves state-of-the-art (SOTA) performance in fish video segmentation tasks, providing robust technical support for fish behavior research.
Figures & tables
Fig. 1 : Users annotate a single frame with simple clicks or scribbles (red for positive, yellow for negative), generating a high-precision mask via Scribble-to-Mask (S2M). The propagation module thenpropagates the mask throughout the video sequence. The two rows on the right compare the propagation results of the baseline model and FiVOS. Unlike the baseline, which suffers from irreversible error accumulation, FiVOS effectively mitigates this issue.
Fig. 2 : Data Collection Schematic. Videos of fish swimming in a tank are captured using a high-resolution camera, covering the entire water surface. The videos are then processed on a computer to generate the Fish-static and Fish-DAVIS datasets.
Parameter
Value
Water level (m)
0.327 ± 0.059
Water temperature (°C)
21.85 ± 1.04
pH
7.44 ± 0.18
Dissolved oxygen (mg/L)
9.13 ± 1.21
Tab. 1 : Experimental environment water quality parameters.
Dataset Name
Configuration
value
Fish−static
Total Samples
1350
Image : Mask set
1: no *
Fish−DAVIS
Total Samples
23
Data sample frame number
10
Train : Val set
18:5
Tab. 2 : Configuration details of the dataset.
Fig. 3 : FiVOS Framework Diagram. In the initial round, all masks are initialized to zero. The user selects frame t and interactively refines the object mask using the S2M module, after which the propagation module bidirectionally propagates it throughout the video sequence to produce an initial mask. If the initial mask contains multiple mask blocks, the filtering module removes erroneous blocks, further refining the accuracy of the predicted mask.
Fig. 4 : Propagation model architecture. (a) The Space-Time Memory Reader generates an initial prediction mask for the query frame using memory frames. (b) The Mask Block Filter identifies and removes misidentified regions from the initial prediction mask. (c) The Noise Filter eliminates residual noise to further refine the mask.
Fig. 5 : Mask block filter. The centroid of each mask block is calculated, and the distance dn (with n as the mask block index) between the centroids of query and memory frame mask blocks is measured. Mask blocks with dn>DT are marked as ecognition errors and set to background; others are preserved.
Window size
AUC-J&F
J&F
1
91.72
91.66
3
93.84
93.89
5
92.54
92.49
11
91.32
91.25
21
90.91
90.83
Tab. 3 : Performance comparison of different median filter window sizes.
Parameter/Configuration
Result value
Operating System
Ubantu 20.04.2
CPU
Core i5-13600KF
GPU
NVIDIA RTX 4070
Video memory
16GB
CUDA
11.8
PyTorch
2.2.2+cu118
Tab. 4 : Experimental system and hardware configuration.
Model
Parameter/Configuration
Result value
Scribble To Mask
Epoch
80000
Batch size
2
The initial learning rate
1e-4
Gamma
0.1
Dataset
Fish-static
Mask Propagation-step1
Epoch
30000
Tab. 5 : Model parameters configuration.
AUC-J &F
J &F
MANet [ 17 ]
67.84
67.90
RGMP [ 19 ]
-
69.31
XMen [ 3 ]
-
89.14
STCN [ 5 ]
93.45
93.48
MiVOS
93.20
93.24
FIVOS
93.80
93.85
Tab. 6 : Performance comparison on the Fish-DAVIS validation set (10-frame short video).
AUC-J &F
J &F
MANet [ 17 ]
78.63
79.18
RGMP [ 19 ]
-
21.19
XMen [ 3 ]
-
29.94
STCN [ 5 ]
74.65
74.65
MiVOS
65.50
65.40
FIVOS
92.50
92.40
Tab. 7 : Performance comparison on the Fish-DAVIS validation set (80-frame long video).
Fig. 6 : Multi-object video segmentation example on the Fish-DAVIS validation set (10-frame short video). For a short video sequence, only the MANet method exhibited target confusion, while other methods produce comparable results with minor differences.
Fig. 7 : Single-object interaction performance on a 20-second long video sequence. MiVOS and STCN experienced target loss and recognition errors at frames 99 and 117, respectively. In contrast, our method consistently delivers accurate segmentation results throughout the sequence.
AUC-J &F
J &F
Baseline
93.198 -
93.242 -
(+) mask block filter
92.881 ↓ 0.317
92.930 ↓ 0.312
(+) mask block filter + noise filter
93.804 ↑ 0.612
93.852 ↑ 0.61
(+) mask block filter + noise filter +
93.844 ↑ 0.646
93.894 ↑ 0.652
Tab. 8 : Ablation study on the Fish-DAVIS short video.
AUC-J &F
J &F
Baseline
65.5 -
65.4 -
(+) mask block filter
70.1 ↑4.6
69.5 ↑4.1
(+) mask block filter + noise filter
83.1 ↑17.6
82.9 ↑17.5
(+) mask block filter + noise filter +
92.5 ↑27.0
92.4 ↑27.0
Tab. 9 : Ablation study on the Fish-DAVIS long video.
Fig. 8 : Ablation study of the mask block filter. S(a) displays the segmentation result of the baseline model, while S(b) shows the result after applying the mask block filter. The purple transparent box indicates the target recognition errors in the baseline model during propagation.
Fig. 9 : Qualitative ablation study: Mask Block Filter and Noise Filter. M(o) and M(o)-L display the segmentation results of the baseline model and an enlarged view of local errors. M(a) and M(a)-L show the interactive results and local magnification after adding the mask block filter. M(b) and M(b)-L illustrate the segmentation results after adding both the mask block filter and noise filter.
Fig. 10 : Ablation study of rule-based thresholding. Compared to manually set empirical thresholds (in the second row), the rule-based thresholding method (in the third row) significantly enhances filter performance.
AUC-J &F
J &F
MiVOS
81.450
81.833
MIVOS+mask block filter
80.921
81.392
MiVOS+mask+noise
80.658
81.127
Tab. 10 : Results of validation on the DAVIS dataset.
Fig. 11 : Experimental comparison in real farming scenarios
Fig. 13 : Visualization of FIVOS limitations: suboptimal performance of the fusion module during multiple rounds of interaction.
Fig. 14 : Limitations of small target segmentation. (a) Segmentation result when the target occupies 0.8% of the image. (b) Segmentation result when the target occupies 0.2% of the image.
Fig. 15 : Visualization of fish movement trajectory. On the left is a sample of mask frames, where the centroid of each frame’s mask is extracted and sequentially plotted on a coordinate axis, then connected in temporal order.The right side displays the final trajectory plot .
The aquaculture industry needs to address several challenges to secure sustainable seafood production that can serve an increasing global demand. One major challenge is to ensure good fish health and acceptable welfare during production since the improvement of fish welfare is of vital importance in current and future production systems. In this study, this is addressed by developing and implementing methods to identify fish behaviors in response to intrusive objects both on individual and on a group basis. A novel approach for detecting, tracking, and estimating the 3D position of individual fish has thus been developed, and specifically designed to track the caudal fins of farmed fish in industrial sea cages. The tracking data was subjected to a novel stereo-vision method adapted to estimate fish positions, velocities, accelerations, and turning and pitch angles. Datasets obtained from industrial-scale fish farms were then analyzed to identify the impact of structures of varying shapes, sizes, and colors on fish behavior. The method was trained using manually labeled caudal fins, and used YOLOv8 with ByteTrack as an object detector and tracker, SuperGlue for matching detections in the left and right frames, and triangulation to reconstruct the 3D positions of the fish. Different image pre-processing and augmentation methods for enhancing object detection accuracy were tested and their performance compared, while RAFT-Stereo was tested for depth estimation purposes. The obtained results both validate the method's performance against previous research efforts, and demonstrate the novelty and potential of this method in providing more insight into behavioral dynamics in sea-cages.
Hanne-Grete Alvheim, Stian Mjelde Jakobsen, Martin Føre +1
Department of Engineering Cybernetics, NTNU · Department of Aquaculture, SINTEF Ocean AS
Recent advances in field technology have led to a massive influx of in-the-wild video data for ecological science. The primary bottleneck in leveraging this data is the high cost of expert annotation. While computer vision offers a potential solution, current models frequently fail when deployed in complex marine environments. To characterize these failures, we introduce WildFin, a novel benchmark for fish behavior recognition collected and annotated by ecologists. WildFin spans two critical real-world paradigms: stationary cameras monitoring groups of fish and dynamic divers following individual subjects. The dataset represents a massive curation effort, involving 1,350 hours of fieldwork and 600 hours of expert annotation to produce 9 hours of behavioral data with over 2 million frame-by-frame labels. We benchmark modern vision foundation models and quantify tradeoffs between static and spatiotemporal architectures, revealing the substantial gap that remains between current model capabilities and the demands of real-world underwater behavioral analysis. Project website: https://team-wildfin.github.io/.
Abigail G. Grassick, Jerome Tze-Hou Hsu, Ethan Lin +10
Cornell University, 14850 Ithaca, NY · University of Colorado Boulder, 80309 Boulder, CO · HHMI Janelia Research Campus, 20147, Ashburn VA
Object-centric models inspired by DETR have become the dominant paradigm for open-vocabulary video instance segmentation (OV-VIS). While recent efforts have reduced the computational cost of pixel decoding, textual modality fusion, and object decoding to make these architectures more suitable for mobile devices, real-time on-device inference at high frame rates remains an open challenge. In this paper, we introduce SegFS, a dual-stream fast-slow framework that significantly improves efficiency without sacrificing accuracy. On sparse keyframes, an open-vocabulary object-based model predicts instance-level representations. These representations are then projected back into the backbone feature space to condition a lightweight fast network, which efficiently relocalizes and segments the instances in subsequent frames. By shifting instance propagation from object decoding to feature-space conditioning, our approach decouples multimodal semantic understanding from dense mask prediction and enables efficient temporal propagation. The proposed fast branch achieves up to 14x lower latency than the mobile-oriented MOBIUS model, while maintaining competitive segmentation performance on standard OV-VIS benchmarks.
Luca Barsellotti, Martin Sundermeyer, Mattia Segu +5
⋆Work done during an internship at Google. · Google · TU Munich; Munich Center for Machine Learning +1