Authors: Zengyi Yang, Shuai Yuan, Zhong-Cheng Wu, Juan Cheng, Huafeng Li, Yu Liu
Organizations: School of Instrument Science and Opto-electronics Engineering, Hefei University of Technology, Hefei 230009, China · Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Kunming 650500, China
Infrared and visible (IR-VIS) image fusion integrates complementary multimodal information into a single fused image to support downstream vision tasks. However, existing methods are typically tailored to seen tasks within a fixed task set and struggle to generalize to unseen tasks, which restricts their applicability in real-world open-task scenarios. To address this issue, this paper proposes CRT-HMAR, a Causal Requirement Tracing-Guided Hierarchical Multi-Agent Regulation Framework for open-task-aware IR-VIS image fusion. CRT-HMAR introduces a Causal Requirement Tracing Task Localization mechanism, which actively intervenes in key image information and observes task-network response variations to map task-specific semantic preferences into image-level causal requirement maps. Based on these maps, a requirement analysis agent aggregates task-specific requirement knowledge to adaptively guide requirement-customized image fusion. Moreover, CRT-HMAR incorporates History-Analysis Multi-Objective Balancing and Task-Level-Correction Conflict Mitigation mechanisms, jointly constructing a hierarchical regulation chain of "requirement interpretation - task balancing - conflict mitigation". Through multiple collaborative agents, CRT-HMAR dynamically regulates key processes including open-task requirement modeling, multi-task balanced optimization, and gradient conflict mitigation. Extensive experiments on open-task scenarios involving five downstream tasks demonstrate that CRT-HMAR significantly improves generalization to unseen tasks while maintaining the performance and balance of seen tasks. Overall, CRT-HMAR shifts IR-VIS image fusion from task-oriented modeling toward requirement-oriented modeling, promoting its extension from closed-task settings to real-world open-task scenarios.
Figures & tables
Fig. 1: Comparison between the proposed method and existing methods. (a) Performance comparison with existing methods in open-task scenarios. (b) Existing task-oriented modeling paradigm. (c) The proposed method adopts a Causal Requirement Tracing-guided Hierarchical multi-Agent Regulation framework and shifts toward a requirement-oriented modeling paradigm.
Fig. 2: Overview of the proposed CRT-HMAR. CRT-HMAR constructs causal requirement maps through the Causal Requirement Tracing Task Localization mechanism. Agentv is introduced to interpret these maps and aggregate requirement knowledge vectors from the RKB. Guided by these vectors, the TAFN generates requirement-customized fused images. In addition, the History-Analysis Multi-Objective Balancing and Task-Level-Correction Conflict Mitigation mechanisms employ Agentl/g , respectively, to dynamically regulate multi-task optimization, thereby alleviating multi-task imbalance and gradient conflicts.
Fig. 3: Network architecture of TAFN. TAFN performs semantic restoration on IR-VIS features according to task-specific requirement knowledge vectors, thereby reconstructing requirement-customized fused images.
Fig. 4: Architecture of the CRT module. The CRT module actively intervenes in key information within the fused image and observes the response variations of downstream task networks, thereby constructing causal requirement maps.
Fig. 5: Asynchronous evaluation strategy of Agentl .
Fig. 6: Quantitative visualization results of the proposed method compared with task-network retraining methods and multi-task-aware methods in diverse open-task scenarios.
Methods
Seen Tasks
Unseen Tasks
OD
Seg
SOD
PD
DE
mAP50→95
mIoU
mFβ
Em
mAP50→95
Δ<1.252
Δ<1.253
MLFusion
0.6172
57.47
0.7900
0.8950
0.4703
0.8681
0.9423
SuperFusion
0.6089
58.11
0.7977
0.9013
0.6019
0.8777
0.9530
TarDAL
0.6157
59.80
0.7767
0.8864
0.4424
0.8651
0.9455
MRFS
0.6189
58.28
0.7865
0.8939
0.5985
0.8826
0.9533
TABLE I: Quantitative comparison of the proposed method with task-network retraining methods and multi-task-aware methods in the OD-Seg open-task scenario. The best, second-best, and third-best metric values are highlighted with Red , Blue , and Green backgrounds, respectively.
Fig. 7: Qualitative comparison of the proposed method with task-network retraining methods and multi-task-aware methods in the OD-Seg open-task scenario. The first and second columns show the visible and infrared source images, respectively. The third to eleventh columns present the downstream task results of the compared methods, and the twelfth column shows the Ground Truth (GT). Every two rows correspond to one task, showing the results of OD, Seg, SOD, PD, and DE from top to bottom, respectively.
Fig. 8: Qualitative comparison of the proposed method with task-network retraining methods and multi-task-aware methods in the OD-SOD open-task scenario. The first and second columns show the visible and infrared source images, respectively. The third to eleventh columns present the downstream task results of the compared methods, and the twelfth column shows the Ground Truth (GT). Every two rows correspond to one task, showing the results of OD, Seg, SOD, PD, and DE from top to bottom, respectively.
Fig. 9: Qualitative comparison of the proposed method with task-network retraining methods and multi-task-aware methods in the Seg-SOD open-task scenario. The first and second columns show the visible and infrared source images, respectively. The third to eleventh columns present the downstream task results of the compared methods, and the twelfth column shows the Ground Truth (GT). Every two rows correspond to one task, showing the results of OD, Seg, SOD, PD, and DE from top to bottom, respectively.
Methods
Seen Tasks
Unseen Tasks
OD
SOD
Seg
PD
DE
mAP50→95
mFβ
Em
mIoU
mAP50→95
Δ<1.252
Δ<1.253
MLFusion
0.6172
0.7920
0.8970
55.89
0.4703
0.8681
0.9423
SuperFusion
0.6089
0.7963
0.8984
56.86
0.6019
0.8777
0.9530
TarDAL
0.6157
0.7977
0.8993
54.48
0.4424
0.8651
0.9455
MRFS
0.6189
0.8032
0.9026
57.96
0.5985
0.8826
0.9533
TABLE II: Quantitative comparison of the proposed method with task-network retraining methods and multi-task-aware methods in the OD-SOD open-task scenario. The best, second-best, and third-best metric values are highlighted with Red , Blue , and Green backgrounds, respectively.
Methods
Seen Tasks
Unseen Tasks
Seg
SOD
OD
PD
DE
mIoU
mFβ
Em
mAP50→95
mAP50→95
Δ<1.252
Δ<1.253
MLFusion
57.47
0.7920
0.8970
0.5467
0.4703
0.8681
0.9423
SuperFusion
58.11
0.7963
0.8984
0.5093
0.6019
0.8777
0.9530
TarDAL
59.80
0.7977
0.8993
0.4812
0.4424
0.8651
0.9455
MRFS
58.28
0.8032
0.9026
0.5479
0.5985
0.8826
0.9533
TABLE III: Quantitative comparison of the proposed method with task-network retraining methods and multi-task-aware methods in the Seg-SOD open-task scenario. The best, second-best, and third-best metric values are highlighted with Red , Blue , and Green backgrounds, respectively.
Fig. 10: Qualitative comparison of the proposed method with representative fusion methods on the M 3 FD, FMB, VT5000, LLVIP, and VTD datasets. The first and second columns show the visible and infrared source images, respectively, and the third to eleventh columns present the fused images of the compared methods. Every two rows correspond to one dataset, showing results from M 3 FD, FMB, VT5000, LLVIP, and VTD from top to bottom, respectively.
Fig. 11: Qualitative comparison between the complete model and ablation models in the OD-Seg open-task scenario. The first and second columns show the visible and infrared source images, respectively. The third to eighth columns present the downstream task results, and the ninth column shows the Ground Truth (GT). The first to fifth rows correspond to the results of OD, Seg, SOD, PD, and DE, respectively.
Methods
M 3 FD
FMB
VT5000
LLVIP
VTD
QCB
SSIM
CC
NAB/F
QCB
SSIM
CC
NAB/F
QCB
SSIM
CC
NAB/F
QCB
SSIM
CC
NAB/F
QCB
SSIM
CC
NAB/F
MLFusion
0.4272
1.3874
0.6251
0.01913
0.4313
1.4688
0.6638
0.02264
0.4356
1.5006
0.7014
0.02632
0.3643
1.2497
0.8428
0.01702
0.3578
1.2877
0.7398
0.04402
SuperFusion
0.3979
1.3696
0.6331
0.01251
0.4148
1.4663
0.6745
0.00942
0.4662
1.4610
0.7018
0.01710
0.3463
1.2498
0.8435
0.02642
0.3188
1.2226
0.7671
0.03353
TarDAL
0.3605
0.7739
0.5077
0.16685
0.3367
0.6790
0.5818
0.14852
0.4193
1.0466
0.6396
0.20763
0.3682
0.5994
0.7612
0.09780
0.3631
0.7658
0.6838
0.29679
MRFS
0.4218
1.4152
0.6370
0.01186
0.4329
1.4968
0.6701
0.01047
0.4507
1.5232
0.6957
0.01480
0.3695
1.3099
0.8364
0.00249
0.3707
1.3267
0.7570
0.01539
TIMFusion
0.3405
1.3523
0.6325
0.00753
0.3180
1.4126
0.6532
0.00749
0.5139
1.4101
0.6947
0.04872
0.3150
1.1370
0.8936
0.00522
0.3910
1.2601
0.7262
0.02043
TABLE IV: Quantitative comparison of the proposed method with representative fusion methods on the M 3 FD, FMB, VT5000, LLVIP, and VTD datasets. The best, second-best, and third-best metric values are highlighted with Red , Blue , and Green backgrounds, respectively.
Methods
Seen Tasks
Unseen Tasks
OD
Seg
SOD
PD
DE
mAP50→95
mIoU
mFβ
Em
mAP50→95
Δ<1.252
Δ<1.253
w/o CRT
0.6104
59.38
0.8076
0.9063
0.6032
0.8798
0.9484
w/o Agentv
0.6093
59.86
0.8078
0.9058
0.6072
0.8832
0.9522
w/o RKB
0.6139
59.84
0.8093
0.9069
0.6164
0.8852
0.9532
w/o Agentl
0.6113
59.23
0.8096
0.9075
0.6179
0.8827
0.9503
TABLE V: Quantitative comparison between the complete model and ablation models in the OD-Seg open-task scenario. The best metric values are highlighted with Red backgrounds.
School of Computer Science and Technology, Dalian University of Technology · Key Laboratory of Social Computing and Cognitive Intelligence (Dalian University of Technology), Ministry of Education · Institute of Image Processing and Understanding, North Minzu University
MoE Key Lab of Artificial Intelligence, Institute of AI, School of Computer Science, Shanghai Jiao Tong University, Shanghai, China · Tsinghua University, Beijing, China · Institute of Modern Optics, Nankai University, Tianjin, China