Publication Type
Conference Proceeding Article
Version
publishedVersion
Publication Date
7-2025
Abstract
Multimodal image fusion and object detection are critical tasks in computer vision, particularly in scenarios requiring robust perception under low illumination conditions. Existing approaches that attempt to combine these tasks often rely on cascaded or loosely coupled designs, which can result in suboptimal performance due to gradient conflicts and task imbalance. In this paper, we propose WA-FDNet, a novel Weight Adaptation Fusion Detection Network that unifies multimodal image fusion and object detection into a single end-to-end framework. WA-FDNet adopts a shared encoder–private decoder architecture, enabling efficient feature sharing while preserving task-specific characteristics. The image fusion branch employs a spatial attention-based feature reconstruction module to generate high-quality fused images by emphasizing semantically important regions. Meanwhile, the detection branch introduces a dual-cross attention feature interaction module that enhances inter-modal representation learning for accurate object detection. To address training instability caused by conflicting objectives, we propose a Dynamic Task Weight Adaptation (DTWA) strategy that dynamically balances gradient contributions across tasks based on optimization feedback. Extensive experiments on public benchmarks demonstrate that WA-FDNet achieves state-of-the-art performance in both fusion quality and detection accuracy, validating the effectiveness of our unified multitask learning approach.
Keywords
Image fusion, Object detection, Multimodal, Attention
Discipline
Artificial Intelligence and Robotics | Databases and Information Systems
Research Areas
Intelligent Systems and Optimization
Areas of Excellence
Digital transformation
Publication
Proceedings of the 42nd Computer Graphics International Conference, CGI 2025, Hong Kong, China, July 14-18
First Page
211
Last Page
223
ISBN
9783032222640
Identifier
10.1007/978-3-032-22264-0_17
Publisher
Springer
City or Country
Cham
Citation
GUO, Yanyin; LUO, Ying; LI, Junwei; and ZHANG, Zhiyuan.
WA-FDNet: A unified weight adaptation network for multimodal image fusion and object detection. (2025). Proceedings of the 42nd Computer Graphics International Conference, CGI 2025, Hong Kong, China, July 14-18. 211-223.
Available at: https://ink.library.smu.edu.sg/sis_research/11234
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-No Derivative Works 4.0 International License.
Additional URL
https://doi.org/10.1007/978-3-032-22264-0_17