FiVOS: A Fish Segmentation Algorithm Based on Interactive Video Object Segmentation and Filter Enhancement
Authors: Yuqing Duan, Song Zhang, Shili Zhao, Daoliang Li, Ran Zhao
Organizations: National Innovation Center for Digital Fishery, China Agricultural University, Beijing 10083, China · Key Laboratory of Smart Farming Technologies for Aquatic Animal and Livestock, Ministry of Agriculture and Rural Affairs, China Agricultural University, Beijing 10083, China · Beijing Engineering and Technology Research Center for Internet of Things in Agriculture, China Agricultural University, Beijing 10083, China · College of Information and Electrical Engineering, China Agricultural University, China Agricultural University, Beijing 10083, China
With the continuous expansion of aquaculture, precise and efficient monitoring of fish behavior has become increasingly critical for improving farming efficiency and reducing economic losses. In particular, with the ongoing enhancement of computational capabilities in deep learning models, vision-based fish segmentation methods are garnering growing attention. By analyzing video segmentation results, fish behavior can be effectively tracked, thereby providing reliable data support for the precise regulation of aquaculture environments. However, existing deep learning-based video segmentation methods for aquaculture scenarios often overlook the dynamic correlations between video frames. In contrast, Interactive Video Object Segmentation (IVOS) employs an interaction-propagation scheme to achieve high-precision segmentation while minimizing user effort, thereby enhancing monitoring efficiency. Yet, IVOS applications in aquaculture remain limited due to data scarcity, and are susceptible to error accumulation and mask loss over long sequence propagation due to high intra-class similarity. In response, this paper proposes an improved interactive video object segmentation method (FiVOS) and constructs two fish-specific datasets. FiVOS utilizes a mask block filter to enable early detection and correction of erroneous propagated mask blocks, enhancing filtering accuracy through a rule-based thresholding approach. Additionally, it serializes noise filters to further eliminate erroneous mask noise, thereby improving model robustness. Experimental results demonstrate that FiVOS achieves state-of-the-art (SOTA) performance in fish video segmentation tasks, providing robust technical support for fish behavior research.
Figures & tables
Fig. 1 : Users annotate a single frame with simple clicks or scribbles (red for positive, yellow for negative), generating a high-precision mask via Scribble-to-Mask (S2M). The propagation module thenpropagates the mask throughout the video sequence. The two rows on the right compare the propagation results of the baseline model and FiVOS. Unlike the baseline, which suffers from irreversible error accumulation, FiVOS effectively mitigates this issue.
Fig. 2 : Data Collection Schematic. Videos of fish swimming in a tank are captured using a high-resolution camera, covering the entire water surface. The videos are then processed on a computer to generate the Fish-static and Fish-DAVIS datasets.
Parameter
Value
Water level (m)
0.327 ± 0.059
Water temperature (°C)
21.85 ± 1.04
pH
7.44 ± 0.18
Dissolved oxygen (mg/L)
9.13 ± 1.21
Tab. 1 : Experimental environment water quality parameters.
Dataset Name
Configuration
value
Fish−static
Total Samples
1350
Image : Mask set
1: no *
Fish−DAVIS
Total Samples
23
Data sample frame number
10
Train : Val set
18:5
Tab. 2 : Configuration details of the dataset.
Fig. 3 : FiVOS Framework Diagram. In the initial round, all masks are initialized to zero. The user selects frame t and interactively refines the object mask using the S2M module, after which the propagation module bidirectionally propagates it throughout the video sequence to produce an initial mask. If the initial mask contains multiple mask blocks, the filtering module removes erroneous blocks, further refining the accuracy of the predicted mask.
Fig. 4 : Propagation model architecture. (a) The Space-Time Memory Reader generates an initial prediction mask for the query frame using memory frames. (b) The Mask Block Filter identifies and removes misidentified regions from the initial prediction mask. (c) The Noise Filter eliminates residual noise to further refine the mask.
Fig. 5 : Mask block filter. The centroid of each mask block is calculated, and the distance dn (with n as the mask block index) between the centroids of query and memory frame mask blocks is measured. Mask blocks with dn>DT are marked as ecognition errors and set to background; others are preserved.
Window size
AUC-J&F
J&F
1
91.72
91.66
3
93.84
93.89
5
92.54
92.49
11
91.32
91.25
21
90.91
90.83
Tab. 3 : Performance comparison of different median filter window sizes.
Parameter/Configuration
Result value
Operating System
Ubantu 20.04.2
CPU
Core i5-13600KF
GPU
NVIDIA RTX 4070
Video memory
16GB
CUDA
11.8
PyTorch
2.2.2+cu118
Tab. 4 : Experimental system and hardware configuration.
Model
Parameter/Configuration
Result value
Scribble To Mask
Epoch
80000
Batch size
2
The initial learning rate
1e-4
Gamma
0.1
Dataset
Fish-static
Mask Propagation-step1
Epoch
30000
Tab. 5 : Model parameters configuration.
AUC-J &F
J &F
MANet [ 17 ]
67.84
67.90
RGMP [ 19 ]
-
69.31
XMen [ 3 ]
-
89.14
STCN [ 5 ]
93.45
93.48
MiVOS
93.20
93.24
FIVOS
93.80
93.85
Tab. 6 : Performance comparison on the Fish-DAVIS validation set (10-frame short video).
AUC-J &F
J &F
MANet [ 17 ]
78.63
79.18
RGMP [ 19 ]
-
21.19
XMen [ 3 ]
-
29.94
STCN [ 5 ]
74.65
74.65
MiVOS
65.50
65.40
FIVOS
92.50
92.40
Tab. 7 : Performance comparison on the Fish-DAVIS validation set (80-frame long video).
Fig. 6 : Multi-object video segmentation example on the Fish-DAVIS validation set (10-frame short video). For a short video sequence, only the MANet method exhibited target confusion, while other methods produce comparable results with minor differences.
Fig. 7 : Single-object interaction performance on a 20-second long video sequence. MiVOS and STCN experienced target loss and recognition errors at frames 99 and 117, respectively. In contrast, our method consistently delivers accurate segmentation results throughout the sequence.
AUC-J &F
J &F
Baseline
93.198 -
93.242 -
(+) mask block filter
92.881 ↓ 0.317
92.930 ↓ 0.312
(+) mask block filter + noise filter
93.804 ↑ 0.612
93.852 ↑ 0.61
(+) mask block filter + noise filter +
93.844 ↑ 0.646
93.894 ↑ 0.652
Tab. 8 : Ablation study on the Fish-DAVIS short video.
AUC-J &F
J &F
Baseline
65.5 -
65.4 -
(+) mask block filter
70.1 ↑4.6
69.5 ↑4.1
(+) mask block filter + noise filter
83.1 ↑17.6
82.9 ↑17.5
(+) mask block filter + noise filter +
92.5 ↑27.0
92.4 ↑27.0
Tab. 9 : Ablation study on the Fish-DAVIS long video.
Fig. 8 : Ablation study of the mask block filter. S(a) displays the segmentation result of the baseline model, while S(b) shows the result after applying the mask block filter. The purple transparent box indicates the target recognition errors in the baseline model during propagation.
Fig. 9 : Qualitative ablation study: Mask Block Filter and Noise Filter. M(o) and M(o)-L display the segmentation results of the baseline model and an enlarged view of local errors. M(a) and M(a)-L show the interactive results and local magnification after adding the mask block filter. M(b) and M(b)-L illustrate the segmentation results after adding both the mask block filter and noise filter.
Fig. 10 : Ablation study of rule-based thresholding. Compared to manually set empirical thresholds (in the second row), the rule-based thresholding method (in the third row) significantly enhances filter performance.
AUC-J &F
J &F
MiVOS
81.450
81.833
MIVOS+mask block filter
80.921
81.392
MiVOS+mask+noise
80.658
81.127
Tab. 10 : Results of validation on the DAVIS dataset.
Fig. 11 : Experimental comparison in real farming scenarios
Fig. 13 : Visualization of FIVOS limitations: suboptimal performance of the fusion module during multiple rounds of interaction.
Fig. 14 : Limitations of small target segmentation. (a) Segmentation result when the target occupies 0.8% of the image. (b) Segmentation result when the target occupies 0.2% of the image.
Fig. 15 : Visualization of fish movement trajectory. On the left is a sample of mask frames, where the centroid of each frame’s mask is extracted and sequentially plotted on a coordinate axis, then connected in temporal order.The right side displays the final trajectory plot .