A Frequency-guided Sparse Fusion Algorithm for Multispectral Object Detection in Complex Environments
-
摘要: 随着多光谱感知与智能视觉技术的发展, 如何在复杂环境中实现稳定而精确的目标检测已成为自动化视觉检测领域的重要研究方向. 针对传统单模态可见光目标检测在夜间、大雾及低照度等复杂环境中性能下降的问题, 提出一种基于频域特征细化导向的多光谱稀疏融合目标检测方法. 该方法利用共享权重的双分支编码器分别提取可见光与红外光特征, 并通过组稀疏自注意力模块实现跨模态长距离特征筛选, 以抑制冗余信息、增强显著特征表达. 同时, 设计频域自适应加权模块, 在频域空间中进行多光谱特征解耦与自适应融合, 实现不同光谱模态间的高效语义交互与动态权重分配. 该方法可在端到端框架下实现跨模态特征的高精度对齐与融合, 有效提升模型的检测精度与鲁棒性. 在M3FD和FLIR数据集上取得83.5%和81.6%的mAP50结果, 在 KAIST数据集上取得76.2%的AP50结果, 显著优于现有多光谱目标检测算法, 验证了所提方法在复杂场景下的优越性能和泛化能力.Abstract: With the development of multispectral perception and intelligent vision technologies, achieving reliable and accurate object detection in complex environments has become a critical research focus in automatic vision detection. To address the performance degradation problem of traditional unimodal visible-light object detection in complex environments such as darkness, dense fog, and low illumination, this paper proposes a multispectral sparse fusion object detection method guided by frequency-domain feature refinement. The proposed method employs a dual-branch encoder with shared weights to extract features from visible and infrared modalities. A group sparse self-attention module is designed to perform cross-modal long-range feature selection, suppress redundant information, and enhance salient feature representation. Furthermore, a frequency-domain adaptive weighting module is designed to decouple and adaptively fuse multispectral features in the frequency-domain space, achieving efficient semantic interaction and dynamic weighting allocation across different spectral modalities. Under an end-to-end architecture, this method realizes high-precision cross-modal feature alignment and fusion, significantly improving detection accuracy and robustness. The proposed method achieved mAP50 results of 83.5% and 81.6% on the M3FD and FLIR datasets, and AP50 results of 76.2% on the KAIST dataset. These results significantly outperform those of existing multispectral object detection algorithms, and validate the superior performance and generalization capability of the proposed method in complex scenes.
-
Key words:
- object detection /
- cross-spectral /
- feature selection /
- cross-modal feature alignment
-
表 1 不同方法在M3FD数据集上的实验评价结果比较
Table 1 Comparative experimental evaluation results of different methods on the M3FD dataset
架构 方法 $ {\rm{mAP}}_{50} $ (%) 精度(%) 召回率(%) $ {\rm{mAP}}_{50} $ (%) $ {\rm{mAP}}_{50:95} $ (%) Params (M) FPS 巴士 轿车 灯 摩托车 行人 卡车 单光谱模态 可见光 90.9 88.5 79.6 75.2 67.7 85.3 89.4 73.4 81.2 50.5 1.77 175.4 红外光 90.7 86.5 57.8 69.5 79.0 83.9 85.1 71.5 77.9 48.2 1.77 175.4 融合检测架构 U2Fusion[16] 80.8 86.3 73.8 56.4 73.2 78.2 84.1 70.1 74.8 45.4 — — RFN[27] 79.2 85.7 69.9 53.1 71.4 76.4 88.5 65.1 72.6 44.9 — — Tardal[17] 77.7 85.0 67.1 56.7 73.1 74.4 84.4 65.7 72.3 43.6 — — MFEIF[18] 79.7 84.9 68.2 56.9 72.9 73.5 85.2 65.9 72.7 44.6 — — HALDeR[28] 79.8 86.0 70.8 52.7 71.2 75.4 86.7 66.8 72.7 44.4 — — 端到端架构 Twin-YOLOv5-n 89.1 87.8 77.1 66.2 77.0 83.8 85.0 75.1 80.2 47.8 2.84 192.3 Twin-YOLOv8-n 89.8 88.6 78.2 67.5 78.3 85.1 85.2 73.0 81.2 49.6 4.28 128.2 DEYOLO[29] 89.3 87.5 77.7 67.8 77.4 85.2 84.0 75.9 80.8 49.3 4.57 188.5 ICAFusion[20] 90.4 87.6 78.5 65.2 77.2 86.0 83.6 76.1 80.8 48.5 5.94 144.9 CFT[19] 90.4 87.9 77.0 70.2 76.6 85.1 87.7 74.3 81.2 48.1 11.20 154.0 DAMSDet[21] 83.1 92.7 71.1 73.5 73.6 78.6 78.7 83.1 78.8 51.0 78.90 88.5 SuperYOLO[30] 89.1 87.6 74.7 68.9 76.6 83.2 88.2 70.1 80.0 48.0 1.80 212.7 FGSFNet 92.1 89.0 80.5 74.3 79.0 86.1 86.7 77.3 83.5 51.2 7.44 208.3 表 2 不同方法在KAIST和FLIR数据集上的实验评价结果比较(%)
Table 2 Comparative experimental evaluation results of different methods on the KAIST and FLIR datasets (%)
KAIST (单类别: 行人) FLIR (多类别) 架构 方法 精度 召回率 $ {\rm{AP}}_{50} $ $ {\rm{AP}}_{50:95} $ $ {\rm{mAP}}_{50} $ 精度 召回率 $ {\rm{mAP}}_{50} $ $ {\rm{mAP}}_{50:95} $ 自行车 汽车 行人 单光谱模态 可见光 71.6 54.4 62.0 24.6 68.9 83.5 70.0 80.2 67.6 74.1 34.8 红外光 75.8 65.4 73.5 32.3 75.8 87.5 81.2 84.5 72.0 81.5 40.6 融合检测架构 U2Fusion[16] 65.4 45.9 52.2 21.1 63.1 83.0 71.6 79.4 65.9 72.6 33.8 RFN[27] 59.6 43.7 48.9 18.9 63.1 82.0 71.4 78.7 65.2 72.1 33.4 Tardal[17] 62.4 45.0 50.6 20.6 60.1 81.5 68.5 77.1 63.3 70.0 32.3 MFEIF[18] 62.1 43.8 50.0 20.6 62.9 80.9 69.7 76.0 66.2 71.2 32.8 HALDeR[28] 63.3 45.7 51.6 21.1 62.9 82.1 70.1 79.6 64.9 71.7 32.9 端到端架构 Twin-YOLOv5-n 76.7 68.7 74.8 32.8 74.9 87.2 82.0 84.0 74.6 81.4 39.4 Twin-YOLOv8-n 78.2 68.0 74.2 32.9 73.8 87.5 81.2 83.0 72.6 80.8 39.7 DEYOLO[29] 74.9 68.2 74.4 31.9 74.5 86.9 82.0 84.7 72.7 81.1 40.0 ICAFusion[20] 75.5 69.1 75.7 33.2 74.8 87.4 81.6 82.2 75.9 81.3 39.6 CFT[19] 77.6 67.2 75.7 32.9 72.9 87.4 82.1 82.9 73.4 80.8 39.2 SuperYOLO[30] 74.3 65.5 71.8 31.5 70.6 86.3 78.3 82.1 69.1 78.4 37.7 FGSFNet 77.6 71.1 76.2 33.8 75.0 87.7 82.1 83.4 75.0 81.6 40.2 表 3 稀疏率组合选取
Table 3 Sparsity rate combination selection
实验 稀疏率参数组合 实验结果(%) 50% 67% 75% 80% $ {\rm{mAP}}_{50} $ $ {\rm{mAP}}_{50:95} $ I 0 0 0 1 78.4 46.2 II 0 1 0 1 82.3 49.7 III 1 1 0 1 83.5 51.2 IV 1 1 1 0 83.5 50.2 V 1 1 1 1 81.4 49.1 表 4 不同融合阶段性能研究
Table 4 Performance study at different fusion stages
融合策略 精度(%) 召回率(%) $ {\rm{mAP}}_{50} $ (%) $ {\rm{mAP}}_{50:95} $ (%) Params (M) 基线 85.0 75.1 81.2 47.7 7.12 早期融合 86.3 71.8 80.1 48.1 7.12 中期融合 86.7 77.3 83.5 51.2 7.44 晚期融合 86.8 75.2 82.5 49.1 7.39 表 5 消融实验
Table 5 Ablation experiments
实验 组稀疏
自注意力模块频域自适应
加权模块$ {\rm{mAP}}_{50} $ (%) $ {\rm{mAP}}_{50:95} $ (%) Params (M) I $ \times $ $ \times $ 80.2 47.8 2.84 II √ $ \times $ 81.2 47.7 7.12 III $ \times $ √ 82.8 49.5 3.20 IV √ √ 83.5 51.2 7.44 -
[1] Mao J G, Shi S S, Wang X G, Li H S. 3D object detection for autonomous driving: A comprehensive survey. International Journal of Computer Vision, 2023, 131(8): 1909−1963 doi: 10.1007/s11263-023-01790-1 [2] Shi X L, Song A J. Defog YOLO for road object detection in foggy weather. The Computer Journal, 2024, 67(11): 3115−3127 doi: 10.1093/comjnl/bxae074 [3] Wang J, Yang P, Liu Y S, Shang D, Hui X, Song J H, et al. Research on improved YOLOv5 for low-light environment object detection. Electronics, 2023, 12(14): Article No. 3089 doi: 10.3390/electronics12143089 [4] 闫梦凯, 钱建军, 杨健. 弱对齐的跨光谱人脸检测. 自动化学报, 2023, 49(1): 135−147 doi: 10.16383/j.aas.c210058Yan Meng-Kai, Qian Jian-Jun, Yang Jian. Weakly aligned cross-spectral face detection. Acta Automatica Sinica, 2023, 49(1): 135−147 doi: 10.16383/j.aas.c210058 [5] 赵兴科, 李明磊, 张弓, 黎宁, 李家松. 基于显著图融合的无人机载热红外图像目标检测方法. 自动化学报, 2021, 47(9): 2120−2131 doi: 10.16383/j.aas.c200021Zhao Xing-Ke, Li Ming-Lei, Zhang Gong, Li Ning, Li Jia-Song. Object detection method based on saliency map fusion for UAV-borne thermal images. Acta Automatica Sinica, 2021, 47(9): 2120−2131 doi: 10.16383/j.aas.c200021 [6] Hu S M, Zhao F, Lu H Z, Deng Y J, Du J M, Shen X L. Improving YOLOv7-tiny for infrared and visible light image object detection on drones. Remote Sensing, 2023, 15(13): Article No. 3214 doi: 10.3390/rs15133214 [7] Wang P, Wu J S, Fang A Q, Zhu Z X, Wang C W. Multi-spectral image fusion for moving object detection. Infrared Physics & Technology, 2024, 141: Article No. 105489 doi: 10.1016/j.infrared.2024.105489 [8] He X, Tang C, Zou X, Zhang W. Multispectral object detection via cross-modal conflict-aware learning. In: Proceedings of the 31st ACM International Conference on Multimedia. Ottawa, Canada: ACM, 2023. 1465−1474 [9] Girshick R, Donahue J, Darrell T, Malik J. Rich feature hierarchies for accurate object detection and semantic segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Columbus, USA: IEEE, 2014. 580−587 [10] 张鹏, 雷为民, 赵新蕾, 董力嘉, 林兆楠, 景庆阳. 跨摄像头多目标跟踪方法综述. 计算机学报, 2024, 47(2): 287−309 doi: 10.11897/SP.J.1016.2024.00287Zhang Peng, Lei Wei-Min, Zhao Xin-Lei, Dong Li-Jia, Lin Zhao-Nan, Jing Qing-Yang. A survey on multi-target multi-camera tracking methods. Chinese Journal of Computers, 2024, 47(2): 287−309 doi: 10.11897/SP.J.1016.2024.00287 [11] Redmon J, Divvala S, Girshick R, Farhadi A. You only look once: Unified, real-time object detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas, USA: IEEE, 2016. 779−788 [12] Redmon J, Farhadi A. YOLOv3: An incremental improvement. arXiv preprint arXiv: 1804.02767, 2018. [13] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, et al. Attention is all you need. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. Long Beach, USA: Curran Associates Inc., 2017. 6000−6010 [14] Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X H, Unterthiner T, et al. An image is worth 16 × 16 words: Transformers for image recognition at scale. In: Proceedings of the International Conference on Learning Representations (ICLR). Vienna, Austria: OpenReview.net, 2020. [15] Chen Y T, Shi J H, Ye Z L, Mertz C, Ramanan D, Kong S. Multimodal object detection via probabilistic ensembling. In: Proceedings of the 17th European Conference on Computer Vision. Tel Aviv, Israel: Springer, 2022. 139−158 [16] Xu H, Ma J Y, Jiang J J, Guo X J, Ling H B. U2Fusion: A unified unsupervised image fusion network. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44(1): 502−518 doi: 10.1109/TPAMI.2020.3012548 [17] Liu J Y, Fan X, Huang Z B, Wu G Y, Liu R S, Zhong W, et al. Target-aware dual adversarial learning and a multi-scenario multi-modality benchmark to fuse infrared and visible for object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans, USA: IEEE, 2022. 5792−5801 [18] Liu J Y, Fan X, Jiang J, Liu R S, Luo Z X. Learning a deep multi-scale feature ensemble and an edge-attention guidance for image fusion. IEEE Transactions on Circuits and Systems for Video Technology, 2022, 32(1): 105−119 doi: 10.1109/TCSVT.2021.3056725 [19] Fang Q Y, Han D P, Wang Z K. Cross-modality fusion Transformer for multispectral object detection. arXiv preprint arXiv: 2111.00273, 2021. [20] Shen J F, Chen Y F, Liu Y, Zuo X, Fan H, Yang W K. ICAFusion: Iterative cross-attention guided feature fusion for multispectral object detection. Pattern Recognition, 2024, 145: Article No. 109913 doi: 10.1016/j.patcog.2023.109913 [21] Guo J J, Gao C Q, Liu F C, Meng D Y, Gao X B. DAMSDet: Dynamic adaptive multispectral detection Transformer with competitive query selection and adaptive feature fusion. In: Proceedings of the 18th European Conference on Computer Vision. Milan, Italy: Springer, 2024. 464−481 [22] Zhu X Z, Su W J, Lu L W, Li B, Wang X G, Dai J F. Deformable DETR: Deformable Transformers for end-to-end object detection. In: Proceedings of the 9th International Conference on Learning Representations (ICLR). Vienna, Austria: OpenReview.net, 2021. [23] Xiao G B, Tang Z M, Guo H L, Yu J, Shen H T. FAFusion: Learning for infrared and visible image fusion via frequency awareness. IEEE Transactions on Instrumentation and Measurement, 2024, 73: Article No. 5015011 doi: 10.1109/tim.2024.3374294 [24] Ma J Y, Ma Y, Li C. Infrared and visible image fusion methods and applications: A survey. Information Fusion, 2019, 45: 153−178 doi: 10.1016/j.inffus.2018.02.004 [25] Hwang S, Park J, Kim N, Choi Y, Kweon I S. Multispectral pedestrian detection: Benchmark dataset and baseline. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Boston, USA: IEEE, 2015. 1037−1045 [26] Zhang H, Fromont E, Lefevre S, Avignon B. Multispectral fusion for object detection with cyclic fuse-and-refine blocks. In: Proceedings of the IEEE International Conference on Image Processing (ICIP). Abu Dhabi, United Arab Emirates: IEEE, 2020. 276−280 [27] Li H, Wu X J, Kittler J. RFN-Nest: An end-to-end residual fusion network for infrared and visible images. Information Fusion, 2021, 73: 72−86 doi: 10.1016/j.inffus.2021.02.023 [28] Liu J Y, Shang J J, Liu R S, Fan X. Halder: Hierarchical attention-guided learning with detail-refinement for multi-exposure image fusion. In: Proceedings of the IEEE International Conference on Multimedia and Expo (ICME). Shenzhen, China: IEEE, 2021. 1−6 [29] Chen Y S, Wang B R, Guo X Y, Zhu W B, He J S, Liu X B, et al. DEYOLO: Dual-feature-enhancement YOLO for cross-modality object detection. In: Proceedings of the 27th International Conference on Pattern Recognition. Kolkata, India: Springer, 2024. 236−252 [30] Zhang J Q, Lei J, Xie W Y, Fang Z M, Li Y S, Du Q. SuperYOLO: Super resolution assisted object detection in multimodal remote sensing imagery. IEEE Transactions on Geoscience and Remote Sensing, 2023, 61: Article No. 5605415 doi: 10.52783/jes.3664 -
下载: