Reference-guided camouflaged object detection aims to segment a target whose visual appearance closely resembles its surroundings by exploiting auxiliary reference samples. The task remains difficult because reference samples contain inconsistent target cues, while generic visual representations are not inherently aligned with the target specified by the references. To handle these problems, we present a consensus-aware multi-source fusion framework. Reference-Conditioned Dual-Backbone Fusion (RCDF) couples trainable PVTv2 query features with frozen DINOv3 representations and uses reference-conditioned correlation to select foundation-model evidence before multi-scale fusion. The framework also aggregates multiple references through cross-reference consensus aggregation and injects reference information at semantic depths matched to the query features. Extensive experiments demonstrate the effectiveness of the proposed method. The results further show that reference consensus, target-conditioned foundation features, and hierarchical decoding provide complementary improvements under the evaluation protocol. The source code will be made publicly available upon acceptance.
Figures & tables
Figure 1: Representative results for reference-guided camouflaged object detection. Columns show the input, ground truth (GT), predictions from UAT, R2CNet, Qwen+SAM, and our method. The displayed Sm values accompany examples with low contrast, thin structures, and multiple targets.
Figure 2: Overview of the consensus-aware multi-source fusion framework. CRCA aggregates heterogeneous references into consensus priors. RCDF couples the trainable PVTv2 and frozen DINOv3 backbones by using reference-conditioned correlation to gate foundation features before multi-scale fusion. The HRFD stage performs spatial cross-attention at deep stages, vector-based modulation at shallow stages, and HMU decoding to produce the final mask and auxiliary predictions.
Figure 3: Internal structure of cross-reference consensus aggregation. The Reference Reliability Estimator (RRE) applies self-attention across the reference dimension and predicts reliability scores after normalization and projection. Reference-Guided Fusion (RGF) then uses the normalized weights {αk}k=1K to aggregate the reference features {Rk}k=1K into the consensus representation Rˉ .
Method
Backbone
Overall
Single-object
Multi-object
Sm↑
αE↑
wF↑
MAE ↓
Sm↑
αE↑
wF↑
MAE ↓
Sm↑
αE↑
wF↑
MAE ↓
R2CNet
ResNet-50
0.805
0.879
0.669
0.036
0.810
0.880
0.674
0.035
0.747
0.872
0.602
0.046
UAT
PVTv2-B2
0.855
0.912
0.757
0.026
0.859
0.913
0.761
0.025
0.805
0.900
0.701
0.033
Qwen+SAM
–
0.852
0.917
0.793
0.029
0.861
0.926
0.805
0.027
0.778
0.844
0.683
0.047
RefOnce
PVTv2-B2
0.890
0.937
0.819
0.019
0.894
0.937
0.825
0.018
0.853
0.930
0.767
0.024
Ours / All components
PVTv2-B2
0.895
0.941
0.825
0.017
0.899
0.943
0.830
0.016
0.859
0.930
0.780
0.021
Table 1: Quantitative comparison on the overall, single-object, and multi-object subsets. Higher values are better for Sm , αE , and wF ; lower values are better for MAE.
Figure 4: Qualitative comparison on representative query–reference pairs. From top to bottom, the rows show the query image, reference image, ground truth, and predictions from R2CNet, UAT, Qwen+SAM, and our method.
CRCA
RCDF
HRFD
Overall
Single-object
Multi-object
Sm↑
αE↑
wF↑
MAE ↓
Sm↑
αE↑
wF↑
MAE ↓
Sm↑
αE↑
wF↑
MAE ↓
−
−
−
0.8550
0.9120
0.7570
0.0260
0.8590
0.9130
0.7610
0.0250
0.8050
0.9000
0.7010
0.0330
✓
−
−
0.8610
0.9160
0.7660
0.0247
0.8654
0.9172
0.7722
0.0239
0.8210
0.9050
0.7100
0.0315
−
✓
−
0.8720
0.9230
0.7850
0.0229
0.8762
0.9242
0.7905
0.0222
0.8340
0.9120
0.7350
0.0288
−
−
✓
0.8660
0.9190
0.7740
0.0238
0.8703
0.9202
0.7800
0.0231
0.8270
0.9080
0.7200
0.0300
✓
✓
−
0.8810
0.9320
0.8060
0.0208
0.8847
0.9331
0.8111
0.0202
0.8480
0.9220
0.7600
0.0258
Table 2: Full-factorial ablation of CRCA, RCDF, and HRFD on the Overall, Single-object, and Multi-object subsets. Higher values are better for Sm , αE , and wF ; lower values are better for MAE.
Figure 5: Qualitative visualization of progressive component variants. The labels embedded in the figure denote the progressive variants and are distinct from the factorial combinations evaluated in Table 2 .
Study
Setting
Sm↑
αE↑
wF↑
MAE ↓
Edge-loss weight
0.05
0.8884
0.9402
0.8162
0.0174
0.10
0.8904
0.9334
0.8108
0.0177
0.50
0.8885
0.9318
0.8192
0.0170
0.20 (selected)
0.8946
0.9414
0.8253
0.0168
Correlation-scale initialization
1
0.8874
0.9368
0.8172
0.0172
3
0.8873
0.9394
0.8170
0.0172
Table 3: Three representative one-factor parameter studies. Each block changes only the named setting, while selected rows repeat the shared complete-model result. Higher values are better for Sm , αE , and wF ; lower values are better for MAE.
Figure 6: Qualitative comparison of correlation-scale initialization. Columns show the input, ground truth, scales of 1, 3, and 5, and the selected scale of 20 (Ours). Red boxes enlarge challenging regions and boundaries.
Figure 7: Qualitative comparison of ASPP dilation settings. Columns show the input, ground truth, groups (1,3,5) and (2,4,6) , and the selected group (3,6,9) (Ours). Red boxes enlarge challenging regions and boundaries.
Item
Measurement
Training iteration time (batch size 8)
140.68 ms/step
Inference latency (FP32, 5-shot)
24.31 ms/image
Inference throughput
41.13 FPS
Computational cost
143.54 GFLOPs / 71.77 GMACs
Peak inference memory
548.6 MiB
Peak training memory
7008.8 MiB
Table 4: Computational efficiency on one NVIDIA GeForce RTX 4090 D. Inference uses FP32, a 384×384 query, five references, and cached ICON-R features; latency excludes reference-feature extraction.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: Qualitative visualization of reference-conditioned DINOv3 feature selection. From left to right, the columns show the input, ground truth (GT), baseline prediction, output with the DINOv3 prior, output after the Semantic Gate, and gate map. In the present formulation, the Semantic Gate denotes RCDF’s reference-conditioned correlation map.
Figure A2: Qualitative visualization of progressive reference-guided refinement. From left to right, the columns show the input, GT, outputs after the Semantic Gate, Cross-Attention, and Ref. Interaction, followed by residual error. The latter two stages are the reference-guided interactions in HRFD.
Unsupervised Camouflaged Object Detection (UCOD) remains a challenging task due to the high intrinsic similarity between target objects and their surroundings, as well as the reliance on noisy pseudo-labels that hinder fine-grained texture learning. While existing refinement strategies aim to alleviate label noise, they often overlook intrinsic perceptual cues, leading to boundary overflow and structural ambiguity. In contrast, learning without pseudo-label guidance yields coarse features with significant detail loss. To address these issues, we propose a unified UCOD framework that enhances both the reliability of pseudo-labels and the fidelity of features. Our approach introduces the Multi-Cue Native Perception module, which extracts intrinsic visual priors by integrating low-level texture cues with mid-level semantics, enabling precise alignment between masks and native object information. Additionally, Pseudo-Label Evolution Fusion intelligently refines labels through teacher-student interaction and utilizes depthwise separable convolution for efficient semantic denoising. It also incorporates Spectral Tensor Attention Fusion to effectively balance semantic and structural information through compact spectral aggregation across multi-layer attention maps. Finally, Local Pseudo-Label Refinement plays a pivotal role in local detail optimization by leveraging attention diversity to restore fine textures and enhance boundary fidelity. Extensive experiments on multiple UCOD datasets demonstrate that our method achieves state-of-the-art performance, characterized by superior detail perception, robust boundary alignment, and strong generalization under complex camouflage scenarios. Code is available at https://github.com/JSLiam94/EReCu.
Shuo Jiang, Gaojia Zhang, Min Tan +2
Zhejiang Key Laboratory of Space Information Sensing and Transmission, Hangzhou Dianzi University · Laboratory of Complex Systems Modeling and Simulation, School of Computer Science and Technology, Hangzhou Dianzi University · College of Computer Science and Technology, Zhejiang University
Video camouflaged object detection (VCOD) is challenging due to dynamic environments. Existing methods face two main issues: (1) SAM-based methods struggle to separate camouflaged object edges due to model freezing, and (2) MLLM-based methods suffer from poor object separability as large language models merge foreground and background. To address these issues, we propose a novel VCOD method based on SAM and MLLM, called Phantom-Insight. To enhance the separability of object edge details, we represent video sequences with temporal and spatial clues and perform feature fusion via LLM to increase information density. Next, multiple cues are generated through the dynamic foreground visual token scoring module and the prompt network to adaptively guide and fine-tune the SAM model, enabling it to adapt to subtle textures. To enhance the separability of objects and background, we propose a decoupled foreground-background learning strategy. By generating foreground and background cues separately and performing decoupled training, the visual token can effectively integrate foreground and background information independently, enabling SAM to more accurately segment camouflaged objects in the video. Experiments on the MoCA-Mask dataset show that Phantom-Insight achieves state-of-the-art performance across various metrics. Additionally, its ability to detect unseen camouflaged objects on the CAD2016 dataset highlights its strong generalization ability.
Hua Zhang, Changjiang Luo
Institute of Information Engineering, Chinese Academy of Sciences · School of Cyber Security, University of Chinese Academy of Sciences
Camouflaged object detection is an emerging and challenging computer vision task that requires identifying and segmenting objects that blend seamlessly into their environments due to high similarity in color, texture, and size. This task is further complicated by low-light conditions, partial occlusion, small object size, intricate background patterns, and multiple objects. While many sophisticated methods have been proposed for this task, current methods still struggle to precisely detect camouflaged objects in complex scenarios, especially with small and multiple objects, indicating room for improvement. We propose a Multi-Scale Recursive Network that extracts multi-scale features using a Pyramid Vision Transformer backbone and combines them with specialized Attention-Based Scale Integration Units, thereby enabling selective feature merging. For more precise object detection, our decoder recursively refines features by incorporating Multi-Granularity Fusion Units. A novel recursive-feedback decoding strategy is developed to enhance the model's understanding of global context, thereby helping it overcome the challenges of this task. By jointly leveraging multi-scale learning and recursive feature optimization, our proposed method achieves performance gains, successfully detecting small and multiple camouflaged objects. Our model achieves state-of-the-art results on two benchmark datasets for camouflaged object detection and ranks second on the remaining two. Our code, model weights, and results are available at https://github.com/linaagh98/MSRNet.
Leena Alghamdi, Muhammad Usman, Hafeez Anwar +2
King Fahd University of Petroleum and Minerals, KSA · Ontario Tech University, Canada · National University of Computer and Emerging Sciences, Pakistan +2