Feature Alignment and Fusion in Infrared-visible Collaborative Object Detection: Challenges and Strategies
-
摘要: 红外−可见光协同目标检测是构建全天候、高可靠智能感知系统的核心技术, 在无人机巡航、自动驾驶及安防监控等领域具有极高的应用价值. 然而, 受成像机理异构、环境复杂及传感器失配等因素影响, 如何实现高效、鲁棒的跨模态信息融合仍是当前研究的难点. 首先系统分析红外−可见光协同目标检测在几何对齐、语义一致性及模态完备性等维度的技术挑战. 随后, 面向这些挑战, 综述了基于深度学习的红外−可见光协同目标检测算法, 重点梳理针对几何弱对齐的特征对齐与免配准融合方法、模态不平衡的自适应权重分配与训练优化策略、模态冲突的语义校正与特征解耦方法以及应对模态不完整的伪标注生成与跨域适应方法. 进一步地, 本文采用mAP、MR等精度指标及参数量、FPS、FLOPs等效率指标, 对主流基准数据集上的代表性方法进行性能对比, 展示不同类型方法的优势与局限. 最后, 立足于大模型时代背景, 对统一感知范式、跨模态语义对齐及多模态融合范式拓展等前沿方向进行展望, 旨在为该领域的后续研究与工程实践提供较为系统的参考依据.Abstract: Infrared-visible collaborative object detection is a pivotal technology for constructing all-weather, high-reliability intelligent perception systems, holding significant application value in fields such as UAV navigation, autonomous driving, and security surveillance. However, achieving efficient and robust cross-modal information fusion remains a formidable challenge due to heterogeneous imaging mechanisms, environmental complexity, and sensor misalignment. This paper first systematically deconstructs the technical challenges of infrared-visible collaborative object detection across the dimensions of geometric alignment, semantic consistency, and modal completeness. Subsequently, addressing these challenges, it provides a comprehensive review of deep learning-based infrared-visible collaborative object detection algorithms, with particular focus on feature alignment and alignment-free fusion methods for geometric weak alignment, adaptive weight allocation and training optimization strategies for modality imbalance, semantic correction and feature decoupling methods for modality conflict, and pseudo-label generation and cross-domain adaptation methods for modality incompleteness. Furthermore, representative methods on mainstream benchmark datasets are systematically compared by using accuracy metrics including mAP and MR, as well as efficiency metrics including parameters, FPS and FLOPs, demonstrating the strengths and limitations of different approaches. Finally, in the context of foundation models era, this paper offers a forward-looking perspective on frontier directions such as unified perception paradigms, cross-modal semantic alignment, and the extension of multimodal fusion paradigms, aiming to provide a systematic reference for future research and engineering practices in this field.
-
Key words:
- infrared-visible /
- object detection /
- multimodal fusion /
- deep learning /
- modality imbalance
-
表 1 红外−可见光协同目标检测四类挑战与方法思路归纳
Table 1 Summary of four challenges and methodological ideas for infrared-visible collaborative object detection
挑战类别 核心问题 主要解决思路 几何弱对齐 传感器视场差异、热形变等导致跨模态空间错位, 特征融合受限 1)几何驱动特征对齐(局部区域对齐、几何参数估计); 2)先验引导特征对齐; 3)免配准融合 模态不平衡 光照敏感度与成像噪声差异导致双模态贡献度动态失衡, 静态融合策略引发模态依赖与优化偏置 1)光照感知模态选择与权重分配; 2)基于模态质量的自适应融合; 3)训练优化层面均衡; 4)噪声抑制 模态冲突 成像机理本质差异导致跨模态语义不一致与特征冗余, 协同学习失效并产生负融合效应 1)语义校正(显式冲突感知); 2)动态交互(相关语义选择与无关信息过滤并行); 3)特征解耦 模态不完整 标注缺失或传感器失效导致训练数据不完备, 检测性能退化 1)标注缺失: 跨模态伪标注生成、主动学习样本筛选; 2)数据缺失: 自适应完整性感知融合、跨模态知识蒸馏与迁移、跨域适应 表 2 基于深度学习的代表性红外−可见光目标检测方法概览
Table 2 Overview of representative infrared-visible object detection methods based on deep learning
算法 年份 发表期刊/会议 改进模块 几何弱对齐 模态不平衡 模态冲突 模态不完整 TS-RPN[65] 2019 Information Fusion √ AR-CNN[9] 2019 International Conference on Computer Vision √ IAF R-CNN[41] 2019 Pattern Recognition √ CIAN[80] 2019 Information Fusion √ IATDNN+IASS[40] 2019 Information Fusion √ UMAD[66] 2019 Conference on Computer Vision and Pattern Recognition Workshops √ MBNet[42] 2020 European Conference on Computer Vision √ TCDet[74] 2020 European Conference on Computer Vision √ CFL[44] 2021 Transactions on Circuits and Systems for Video Technology √ ASPFF-Net[45] 2021 Infrared Physics & Technology √ GAFF[46] 2021 Winter Conference on Applications of Computer Vision √ 文献[47]中方法 2021 Transactions on Circuits and Systems for Video Technology √ 文献[68]中方法 2021 International Conference on Image Processing √ TSFADet[30] 2022 European Conference on Computer Vision √ LG-FAPF[81] 2022 Information Fusion √ CMPD[82] 2022 Transactions on Multimedia √ AMSF-net[12] 2022 Chinese Conference on Pattern Recognition and Computer Vision √ CMAFF[83] 2022 Pattern Recognition √ RISNet[84] 2022 Remote Sensing √ AANet[28] 2023 ACM International Conference on Multimedia √ √ CLGNet[27] 2023 Journal of Selected Topics in Applied Earth Observations and Remote Sensing √ √ √ CMLNet[11] 2023 International Conference on Real-Time Computing and Robotics √ MBM+ MHSM[49] 2023 Chinese Conference on Pattern Recognition and Computer Vision √ TINet[5] 2023 Transactions on Instrumentation and Measurement √ IGT[17] 2023 Knowledge-Based Systems √ MoE-Fusion[48] 2023 International Conference on Computer Vision √ HAFNet[16] 2023 Remote Sensing √ √ CALNet[57] 2023 ACM International Conference on Multimedia √ DATFF[58] 2023 International Conference on Artificial Neural Networks √ LRAF-Net[14] 2023 Transactions on Neural Networks and Learning Systems √ MCHE-CF[85] 2023 Transactions on Multimedia √ CMA-Det[32] 2024 Transactions on Intelligent Vehicles √ CF-Deformable DETR[35] 2024 International Joint Conference on Artificial Intelligence √ CPFM[36] 2024 Transactions on Multimedia √ CAGTDet[31] 2024 Information Fusion √ OAFA[37] 2024 Conference on Computer Vision and Pattern Recognition √ ICAFusion[86] 2024 Pattern Recognition √ CMRM[50] 2024 Sensors Journal √ MS-DETR[51] 2024 Transactions on Intelligent Transportation Systems √ √ DCSANet[87] 2024 International Conference on Intelligent Robots and Systems √ FMCFNet[54] 2024 Geoscience and Remote Sensing Letters √ GLFNet[88] 2024 Geoscience and Remote Sensing Letters √ TFDet[7] 2024 Transactions on Neural Networks and Learning Systems √ LF-MDet[53] 2024 Transactions on Geoscience and Remote Sensing √ CPCF[89] 2024 Transactions on Intelligent Transportation Systems √ M2FNet[63] 2024 Transactions on Multimedia √ D3T[67] 2024 Conference on Computer Vision and Pattern Recognition √ CBAM[90] 2024 Winter Conference on Applications of Computer Vision √ SCFR[91] 2024 Transactions on Intelligent Transportation Systems √ √ HalluciDet[75] 2024 Winter Conference on Applications of Computer Vision √ ModTr[76] 2024 European Conference on Computer Vision √ M2FP[77] 2024 Journal of Selected Topics in Applied Earth Observations and Remote Sensing √ UMFT[78] 2024 Information Fusion √ MMPedestron[79] 2024 European Conference on Computer Vision √ RGFNet[38] 2025 Transactions on Intelligent Transportation Systems √ √ CSSFDet[39] 2025 ACM International Conference on Multimedia √ EMOD[92] 2025 Information Fusion √ √ EI2Det[6] 2025 Transactions on Circuits and Systems for Video Technology √ WaveMamba[55] 2025 International Conference on Computer Vision √ DDFD[52] 2025 ACM International Conference on Multimedia √ DecomKD[73] 2025 ACM International Conference on Multimedia √ √ MCAFNet[93] 2025 Digital Signal Processing √ DMM[59] 2025 Transactions on Geoscience and Remote Sensing √ DDCINet[60] 2025 Transactions on Geoscience and Remote Sensing √ SemFusion[61] 2025 ACM International Conference on Multimedia √ RSDet[62] 2025 Transactions on Intelligent Transportation Systems √ JTMDet[94] 2025 Image and Vision Computing √ DANet[64] 2025 Transactions on Instrumentation and Measurement √ DACFusion[95] 2025 Neurocomputing √ PFGF[96] 2025 Conference on Computer Vision and Pattern Recognition √ 表 3 红外−可见光目标检测数据集
Table 3 Infrared-visible object detection datasets
数据集 年份 采集平台 图像对 标注数量 类别 适用场景 KAIST[23] 2015 车载 95328 103128 4 自动驾驶 CVC-14[97] 2016 车载 7085 7499 3 自动驾驶 VEDAI[106] 2016 卫星拍摄 1210 3700 9 无人机检测 Utokyo[100] 2017 车载 7512 5 自动驾驶 FLIR[98] 2018 车载 TIR: 9711 /RGB:9233 26442 15 自动驾驶 FLIR-aligned[99] 2020 车载 5142 3 自动驾驶 LLVIP[101] 2021 固定监控 16836 1 夜间监控 M3FD[102] 2022 固定监控 4200 33603 6 监控、自动驾驶 DroneVehicle[107] 2022 无人机 28439 953087 5 智能交通、自动驾驶 Multi-Spectral Stereo[103] 2023 车载 195000 4 自动驾驶 MMPD[79] 2024 混合/手持 260000 1 自动驾驶、机器人 InfraPairs[104] 2024 车载 7301 20 自动驾驶、多任务学习 DVTOD[32] 2024 无人机 2179 6142 3 无人机检测 MFAD[6] 2025 车载 12370 100000 6 自动驾驶 SMOD[105] 2025 车载 8676 31443 4 自动驾驶 表 4 FLIR数据集上的性能对比
Table 4 Performance comparison on the FLIR dataset
检测器 骨干网络 算法 发表期刊/会议 Bicycle Car Person mAP@0.5 FSSD VGG16 CFR[99] International Conference on Image Processing 57.77 84.91 74.49 72.39 SSD ResNet50 CMPD[82] Transactions on Multimedia 59.87 78.11 69.64 69.35 SSD VGG16 D3T[67] Conference on Computer Vision and Pattern Recognition 57.44 79.68 70.77 69.30 Faster R-CNN ResNet-FPN DetFusion[108] ACM International Conference on Multimedia 39.40 73.70 55.40 56.15 Faster R-CNN ResNet50 MFPT[109] Transactions on Intelligent Transportation Systems 67.70 89.00 83.20 80.00 Faster R-CNN VGG16 SMPD[110] Transactions on Circuits and Systems for Video Technology 56.20 85.80 78.74 73.58 Faster R-CNN VGG16 CMRM[50] Sensors Journal 68.90 89.10 85.10 81.10 Faster R-CNN ResNet50 SCFR[91] Transactions on Intelligent Transportation Systems 73.70 89.40 83.70 82.30 RoITransformer ResNet-FPN UA-CMDet[107] Transactions on Circuits and Systems for Video Technology 64.30 88.40 83.20 78.60 RetinaNet ResNet50 MFPT[109] Transactions on Intelligent Transportation Systems 65.00 87.30 78.10 76.80 Sparse R-CNN ResNet50 SCFR[91] Transactions on Intelligent Transportation Systems 60.60 87.30 79.00 75.60 FCOS ResNet50 ICAFusion[86] Pattern Recognition - - - 71.70 YOLOv5 CSPDarkNet53 ICAFusion[86] Pattern Recognition 66.90 89.00 81.60 79.20 YOLOv5 ResNet50 ICAFusion[86] Pattern Recognition - - - 72.00 YOLOv5 VGG16 ICAFusion[86] Pattern Recognition - - - 69.80 YOLOv5 CSPDarkNet53 EI2Det[6] Transactions on Circuits and Systems for Video Technology 66.30 89.40 84.90 80.20 YOLOv8 CSPDarkNet53 JTMDet[94] Image and Vision Computing - - - 81.20 YOLOv10 CSPDarkNet53 JTMDet[94] Image and Vision Computing - - - 80.90 YOLOXs CSPDarkNet53 EMOD[92] Information Fusion 67.71 89.86 87.23 81.60 表 5 LLVIP数据集上的性能对比
Table 5 Performance comparison on the LLVIP dataset
检测器 骨干网络 算法 发表期刊/会议 mAP@0.5↑ mAP@0.75↑ mAP@0.5:0.95↑ YOLOX CSPDarknet-53 AMSF-net[12] Chinese Conference on Pattern Recognition and Computer Vision 97.0 74.0 64.5 YOLOX CSPDarknet53 CPCF[89] Transactions on Intelligent Transportation Systems 96.4 75.4 65.0 Faster R-CNN ResNet-FPN DetFusion[108] ACM International Conference on Multimedia 80.7 - - Faster R-CNN ResNet50 CSSA[15] Conference on Computer Vision and Pattern Recognition Workshops 94.3 66.6 59.2 Faster R-CNN ResNet50 SCFR[91] Transactions on Intelligent Transportation Systems 97.5 - - Faster R-CNN ResNet50 TFDet[7] Transactions on Neural Networks and Learning Systems 96.0 - 59.4 Faster R-CNN ResNet50 HalluciDet[75] Winter Conference on Applications of Computer Vision $ 90.92\pm0.20 $ - - DETR ResNet50 DAMSDet[111] European Conference on Computer Vision 97.9 79.1 69.6 DETR ResNet50 MS-DETR[51] Transactions on Intelligent Transportation Systems 97.9 76.3 66.1 FCOS ResNet50 CPCF[89] Transactions on Intelligent Transportation Systems 96.0 69.5 60.6 FCOS ResNet50 HalluciDet[75] Winter Conference on Applications of Computer Vision $ 64.85\pm1.46 $ - - YOLOv5 CSPDarknet53 CPCF[89] Transactions on Intelligent Transportation Systems 96.1 70.1 62.0 YOLOv5 CSPDarknet53 CRSIOD[112] Transactions on Geoscience and Remote Sensing 98.10 - - YOLOv5 CSPDarknet53 TFDet[7] Transactions on Neural Networks and Learning Systems 97.9 - 71.1 YOLOv5 CSPDarknet-s CMA-Det[35] Transactions on Intelligent Vehicles 97.1 - - Sparse R-CNN ResNet50 SCFR[91] Transactions on Intelligent Transportation Systems 96.4 - - RetinaNet ResNet50 HalluciDet[75] Winter Conference on Applications of Computer Vision $ 56.78\pm3.85 $ - - 表 6 DroneVehicle数据集上的性能对比
Table 6 Performance comparison on the DroneVehicle dataset
检测器 骨干网络 算法 发表期刊/会议 Car↑ Freight car↑ Truck↑ Bus↑ Van↑ mAP@0.5↑ DroneVehicle测试集 RoITransformer ResNet-FPN UA-CMDet[107] Transactions on Circuits and
Systems for Video Technology87.51 46.80 60.70 87.08 37.95 64.01 RoITransformer ResNet-FPN CCLDet[113] Transactions on Intelligent Transportation Systems 97.70 68.80 75.40 95.70 59.50 79.40 Oriented R-CNN ResNet50 AMSF-net[12] Chinese Conference on Pattern Recognition and Computer Vision — — — — — 83.90 Oriented R-CNN ResNet50 文献[49]中算法 Chinese Conference on Pattern Recognition and Computer Vision 90.38 66.46 69.53 90.17 63.04 75.92 Oriented R-CNN ResNet50 DDCINet[60] Transactions on Geoscience and
Remote Sensing91.00 66.10 78.90 90.70 65.50 78.40 CALNet DarkNet53 CALNet[57] ACM International Conference
on Multimedia90.30 62.97 76.15 89.11 58.46 75.39 CALNet ResNet50 CALNet[57] ACM International Conference on Multimedia 90.32 73.83 60.92 88.74 51.68 73.08 DINO ResNet50 LF-MDet[53] Transactions on Geoscience and
Remote Sensing82.20 59.60 73.60 86.60 57.00 71.80 S2ANet ResNet50 CPCF[89] Transactions on Intelligent Transportation Systems — — — — — 79.20 S2ANet ResNet-FPN C2Former[114] Transactions on Geoscience and
Remote Sensing90.20 64.40 68.30 89.80 58.50 74.20 S2ANet ResNet50 DDCINet[60] Transactions on Geoscience and
Remote Sensing90.40 62.00 75.50 89.70 60.90 75.70 S2ANet VMamba DMM[59] Transactions on Geoscience and
Remote Sensing90.40 68.20 79.80 89.90 68.60 79.40 Faster R-CNN ResNet50 CPCF[89] Transactions on Intelligent Transportation Systems — — — — — 76.10 Faster R-CNN VMamba DMM[59] Transactions on Geoscience and
Remote Sensing90.40 63.00 77.80 88.70 66.00 77.20 RetinaNet ResNet50 CPCF[89] Transactions on Intelligent Transportation Systems — — — — — 72.90 PSC ResNet50 CPCF[89] Transactions on Intelligent Transportation Systems — — — — — 77.80 YOLOv8 CSPDarknet-s RGFNet[38] Transactions on Geoscience and
Remote Sensing98.40 68.70 81.10 95.80 63.00 81.40 DroneVehicle验证集 Cascade ResNet50 TSFADet[30] European Conference
on Computer Vision90.01 65.45 69.15 89.70 55.19 73.90 Oriented R-CNN ResNet50 TSFADet[30] European Conference
on Computer Vision89.88 63.74 67.87 89.81 53.99 73.06 Oriented R-CNN ResNet50 CAGTDet[31] Information Fusion 90.82 66.28 69.65 90.46 55.62 74.57 YOLOv5s CSPDarkNet SLBAF-Net[115] Multimedia Tools and Applications 90.20 68.60 72.00 89.90 59.90 76.10 YOLOv5s CSPDarkNet OAFA[37] Conference on Computer Vision and Pattern Recognition 90.30 73.30 76.80 90.30 66.00 79.40 CALNet DarkNet53 CALNet[57] ACM International Conference
on Multimedia90.27 68.67 73.69 89.70 59.74 76.41 CALNet ResNet50 CALNet[57] ACM International Conference
on Multimedia90.11 70.66 67.48 89.74 55.01 74.60 YOLOv8s YOLOv8s backbone ADMPF[116] Transactions on Geoscience and
Remote Sensing97.64 69.34 82.28 95.93 64.77 82.05 表 7 CVC-14数据集上的性能对比
Table 7 Performance comparison on the CVC-14 dataset
检测器 骨干网络 算法 发表期刊/会议 All↓ Day↓ Night↓ AR-CNN VGG16 AR-CNN[9] International Conference on Computer Vision 22.10 24.70 18.10 MBNet ResNet50 MBNet[42] European Conference on Computer Vision 21.10 24.70 13.50 SSD-Like VGG16 MLPD[69] Robotics and Automation Letters 21.33 24.18 17.97 Faster RCNN VGG16 文献[117]中算法 AAAI Conference on Artificial Intelligence 19.88 23.69 12.35 Faster RCNN VGG-16 SMPD[110] Transactions on Circuits and Systems for Video Technology 19.85 14.98 17.76 Faster RCNN VGG16 MCHE-CF[85] Transactions on Multimedia 21.32 24.01 17.53 Faster RCNN VGG16 CMRM[50] Sensors Journal 19.48 23.09 12.01 EAST-style decoder ResNeXt-50+DCN 文献[118]中算法 Transactions on Intelligent Transportation Systems 19.04 20.32 12.86 LG-FAPF VGG-16 LG-FAPF[81] Information Fusion 18.20 22.50 12.20 RetinaNet ResNet HAFNet[16] Remote Sensing 21.10 23.90 14.30 DETR ResNet50 MS-DETR[51] Transactions on Intelligent Transportation Systems 16.90 24.10 8.80 YOLOXs CSPDarkNet53 EMOD[92] Information Fusion 14.43 19.07 9.46 表 8 KAIST数据集上的性能对比
Table 8 Performance comparison on the KAIST dataset
检测器 骨干网络 算法 发表期刊/会议 All↓ Day↓ Night↓ MSDS-RCNN VGG16 MSDS-RCNN[119] British Machine Vision Conference 11.63 10.60 13.73 CIAN VGG16 CIAN[80] Information Fusion 14.12 14.77 11.13 AR-CNN VGG16 AR-CNN[9] International Conference on Computer Vision 9.34 9.94 8.38 MBNet ResNet50 MBNet[42] European Conference on Computer Vision 8.13 8.28 7.86 SSD VGG16 MLPD[69] Robotics and Automation Letters 7.58 7.95 6.95 SSD ResNet50 CMPD[82] Transactions on Multimedia 8.16 8.77 7.31 RISNet ResNet50 RISNet[84] Remote Sensing 7.89 7.61 7.08 BAANet ResNet50 BAANet[13] International Conference on Robotics and Automation 7.92 8.37 6.98 RetinaNet VGG CLGNet[27] Journal of Selected Topics in Applied Earth Observations and Remote Sensing 6.67 7.48 4.80 RetinaNet ResNet HAFNet[16] Remote Sensing 6.93 7.68 5.66 RetinaNet ResNet18 AMFD[105] Transactions on Multimedia 8.82 10.98 4.86 CMM VGG16 CMM[120] Conference on Computer Vision and Pattern Recognition 8.54 9.60 5.93 YOLOv5 CSPDarkNet53 ICAFusion[86] Pattern Recognition 7.17 6.82 7.85 M2FNet VGG16 M2FNet[63] Transactions on Multimedia 4.25 5.52 2.19 M2FNet ResNet50 M2FNet[63] Transactions on Multimedia 5.85 7.20 3.40 DETR ResNet50 MS-DETR[51] Transactions on Intelligent Transportation Systems 6.13 7.78 3.18 Faster RCNN ResNet50+FPN TINet[5] Transactions on Instrumentation and Measurement 9.15 10.25 7.48 Faster RCNN VGG16 CMRM[50] Sensors Journal 6.34 7.01 6.04 Faster RCNN VGG16 MCHE-CF[85] Transactions on Multimedia 6.71 7.58 5.52 Faster RCNN VGG16 TFDet[7] Transactions on Neural Networks and Learning Systems 4.47 5.22 3.36 Faster RCNN ResNet18 AMFD[105] Transactions on Multimedia 7.23 9.83 2.18 Faster RCNN ResNet50 SCFR[91] Transactions on Intelligent Transportation Systems 7.06 7.82 4.83 YOLOv5 CSPDarknet53 MCOR[121] Winter Conference on Applications of Computer Vision 6.76 7.41 7.83 Sparse R-CNN ResNet50 SCFR[91] Transactions on Intelligent Transportation Systems 13.42 14.20 10.46 表 9 不同方法的计算复杂度对比
Table 9 Comparison of computational complexity for different methods
算法 数据集 输入尺寸 检测器 骨干网络 参数量(M) FLOPs(G) FPS(Hz)/
时间(ms)平台 ICAFusion[86] KAIST $ 640\times512 $ YOLOv5 CSPDarknet53 120.21 — 38.46/− RTX 3090 JTMDet[94] FLIR $ 640\times640 $ YOLOv5 CSPDarknet53 237.40 88.2 −/29.70 RTX 3090 AMSF-net[12] FLIR $ 512\times640 $ YOLOX CSPDarknet53 68.24 147.38 −/− — DAMSDet[111] M3FD $ 640\times640 $ DETR-based ResNet50 — — −/117.00 RTX 3090 CPCF[89] FLIR $ 640\times512 $ YOLOX CSPDarknet53 14.61 15.13 −/26.70 RTX 2080 CPCF[89] FLIR $ 640\times512 $ YOLOv5 CSPDarknet53 12.66 10.61 −/35.10 RTX 2080 MS-DETR[51] — $ 640\times512 $ DETR ResNet50 76.38 258.84 — — CRSIOD[112] — $ 640\times640 $ YOLOv5 CSPDarknet 18.26 — — RTX 2080Ti TFDet[7] KAIST $ 512\times640 $ Faster R-CNN VGG16 — — −/130.00 RTX 3060 CMA-Det[32] DVTOD VIS: $ 1920\times1080 $
IR: $ 640\times512 $YOLOv5 CSPDarknet-s 33.00 — 123.00/− RTX 3060 LF-MDet[53] VEDAI $ 1024\times1024 $ DINO ResNet50 38.70 145.0 — — LF-MDet[53] DroneVehicle $ 840\times712 $ DINO ResNet50 38.70 77.7 — — C2Former[114] DroneVehicle $ 640\times512 $ S2Anet ResNet50 100.80 89.9 — NVIDIA TITAN V CCLDet[113] DroneVehicle $ 640\times512 $ RoITransformer ResNet-FPN 82.28 293.25 — — DDCINet[60] DroneVehicle $ 640\times512 $ S2ANet ResNet50 121.70 124.86 17.20/− NVIDIA V100 DMM[59] DroneVehicle — S2ANet VMamba 87.97 — — NVIDIA 4090 RGFNet[38] DroneVehicle $ 640\times640 $ YOLOv8 CSPDarknet-s 79.80 — 46.20/− NVIDIA 3090 TSFADet[30] DroneVehicle $ 512\times512 $ Oriented R-CNN ResNet50 — — 18.60/− NVIDIA GV100 OAFA[37] DroneVehicle $ 640\times640 $ YOLOv5s CSPDarkNet — — 33.10/− NVIDIA A6000 CAGTDet[31] DroneVehicle $ 512\times512 $ Oriented R-CNN ResNet50 — 120.6 17.80/− NVIDIA V100 ADMPF[116] DroneVehicle $ 640\times640 $ YOLOv8s YOLOv8s backbone 31.80 68.8 39.80/− NVIDIA A6000 文献[117]中算法 — $ 471\times640 $ Faster RCNN VGG16 139.00 — −/40.00 1080 TiCMRM[50] KAIST $ 640\times512 $ Faster R-CNN VGG16 — — −/120.00 1080Ti CLGNet[27] DroneVehicle $ 640\times512 $ Rotated RetinaNet ResNet50 — — 15.50/− RTX 3090 TINet[5] FLIR $ 640\times512 $ Faster RCNN ResNet50 97.00 — 26.80/− RTX 3090 MCOR[121] KAIST $ 640\times640 $ YOLOv5 CSPDarknet53 — — 35.40/− RTX 3090 表 10 红外−可见光协同目标检测挑战类别与代表方法总结
Table 10 Summary of challenge categories and representative methods for infrared-visible collaborative object detection
挑战类别 代表方法 典型适用数据集 主要优势 几何弱对齐 AR-CNN[9], CMLNet[11], CLGNet[27], AANet[28], TSFADet[30], CAGTDet[31], RGFNet[38], CSSFDet[39], BANet[29], CF-Deformable DETR[35], CPFM[36], OAFA[37], ICAFusion[86] DroneVehicle, DVTOD, CVC-14 对空间错位鲁棒, 无人机场景适应性强 模态不平衡 TINet[5], TFDet[7], IGT[17], IAF R-CNN[41], MBNet[42], IAF-RTDETR[43], MoE-Fusion[48], MBM[49], BAA-Gate[13], WaveMamba[55], GEM-YOLO[56] KAIST, LLVIP, FLIR 夜间及低光照场景检测稳定性高 模态冲突 LRAF-Net[14], CALNet[57], DMM[59], DDCINet[60], SemFusion[61], M2FNet[63], DANet[64], JTMDet[94] M3FD, FLIR, DroneVehicle 有效抑制负融合, 复杂背景判别性强 模态不完整 TS-RPN[65], D3T[67], MLPD[69], CIRDet[70], DecomKD[73], HalluciDet[75], ModTr[76], MMPedestron[79] KAIST, LLVIP 单模态条件下性能退化幅度小 -
[1] 李升波, 关阳, 侯廉, 高洪波, 段京良, 梁爽, 等. 深度神经网络的关键技术及其在自动驾驶领域的应用. 汽车安全与节能学报, 2019, 10(02): 119−145Li Sheng-Bo, Guan Yang, Hou Lian, Gao Hong-Bo, Duan Jing-Liang, Liang Shuang, et al. Key technique of deep neural network and its applications in autonomous driving. Journal of Automotive Safety and Energy, 2019, 10(02): 119−145 [2] 林露. 智能安防的感知和识别关键技术研究. 浙江大学, 2019Lin Lu. Research on key technology of perception and recognition for intelligent security[Ph. D. dissertation], Zhejiang University, China, 2019 [3] 车彦卓, 刘寿宝. 探析无人机低空遥感技术与人工智能技术融合发展. 中国安防, 2021(04): 34−38 doi: 10.3969/j.issn.1673-7873.2021.04.007Che Yan-Zhuo, Liu Shou-Bao. Research on integrated development of uav low-altitude remote sensing technology and artificial intelligence technology. China Security & Protection, 2021(04): 34−38 doi: 10.3969/j.issn.1673-7873.2021.04.007 [4] Ma J, Ma Y, Li C. Infrared and visible image fusion methods and applications: A survey. Information fusion, 2019, 45: 153−178 doi: 10.1016/j.inffus.2018.02.004 [5] Zhang Y, Yu H, He Y, et al. Illumination-guided RGBT object detection with inter-and intra-modality fusion. IEEE Transactions on Instrumentation and Measurement, 2023, 72: 1−13 doi: 10.1109/tim.2023.3251414 [6] Hu K, He Y, Li Y, et al. Ei2det: Edge-guided illumination-aware interactive learning for visible-infrared object detection. IEEE Transactions on Circuits and Systems for Video Technology, 2025, 35(7): 7101−7115 doi: 10.1109/TCSVT.2025.3539625 [7] Zhang X, Zhang X, Wang J, et al. TFDet: Target-aware fusion for RGB-T pedestrian detection. IEEE Transactions on Neural Networks and Learning Systems, 2024, 36(7): 13276−13290 doi: 10.1109/tnnls.2024.3443455 [8] Guan D, Cao Y, Yang J, et al. Exploiting fusion architectures for multispectral pedestrian detection and segmentation. Applied optics, 2018, 57(18): D108−D116 doi: 10.1364/AO.57.00D108 [9] Zhang L, Zhu X, Chen X, et al. Weakly aligned cross-modal learning for multispectral pedestrian detection. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. Seoul, Korea: IEEE, 2019. 5127−5137 [10] Zhang L, Liu Z, Zhu X, et al. Weakly aligned feature fusion for multimodal object detection. IEEE Transactions on Neural Networks and Learning Systems, 2025, 36(3): 4145−4159 doi: 10.1109/TNNLS.2021.3105143 [11] Chen Y, Guan Y, Shao Z. Real-Time Multispectral Pedestrian Detection with Weakly Aligned Cross-Modal Learning. In: Proceedings of the 2023 IEEE International Conference on Real-Time Computing and Robotics. Datong, China: IEEE, 2023. 829−834 [12] Bao W, Huang M, Hu J, et al. Attention-guided multi-modal and multi-scale fusion for multispectral pedestrian detection. In: Proceedings of the Chinese Conference on Pattern Recognition and Computer Vision. Cham, Switzerland: Springer International Publishing, 2022. 382−393 [13] Yang X, Qian Y, Zhu H, et al. BAANet: Learning bi-directional adaptive attention gates for multispectral pedestrian detection. In: Proceedings of the 2022 International Conference on Robotics and Automation. Philadelphia, USA: IEEE, 2022. 2920−2926 [14] Fu H, Wang S, Duan P, et al. Lraf-net: Long-range attention fusion network for visible-infrared object detection. IEEE Transactions on Neural Networks and Learning Systems, 2023, 35(10): 13232−13245 doi: 10.1109/tnnls.2023.3266452 [15] Cao Y, Bin J, Hamari J, et al. Multimodal object detection by channel switching and spatial attention. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. Vancouver, Canada: IEEE, 2023. 403−411 [16] Peng P, Xu T, Huang B, et al. HAFNet: hierarchical attentive fusion network for multispectral pedestrian detection. Remote Sensing, 2023, 15(8): 2041 doi: 10.3390/rs15082041 [17] Chen K, Liu J, Zhang H. IGT: Illumination-guided RGB-T object detection with transformers. Knowledge-Based Systems, 2023, 268: 110423 doi: 10.1016/j.knosys.2023.110423 [18] Yan Z, Dong Z, Zhang L. Visible and infrared images fusion for object detection: a survey. In: Proceedings of the International Conference on Algorithms, High Performance Computing, and Artificial Intelligence. Yinchuan, China: SPIE, 2023. 967−975 [19] Sun Y, Meng Y, Wang Q, et al. Visible and infrared image fusion for object detection: a survey. In: Proceedings of the International Conference on Image, Vision and Intelligent Systems. Singapore: Springer Nature Singapore, 2023. 236−248 [20] 王元喆, 梁腾飞, 曾宇乔, 金一, 李浥东. 多光谱目标检测综述. 信息与控制, 2024, 53(03): 287−301 doi: 10.13976/j.cnki.xk.2024.3238Wang Yuan-Zhe, Liang Teng-Fei, Zeng Yu-Qiao, Jin Yi, Li Yi-Dong. Overview of multispectral object detection. Information and Control, 2024, 53(03): 287−301 doi: 10.13976/j.cnki.xk.2024.3238 [21] 朱自文, 宋晓鸥, 崔巍, 岂峰利. 可见光−红外图像融合的目标检测综述. 计算机工程与应用, 2025, 61(17): 17−32 doi: 10.3778/j.issn.1002-8331.2501-0206Zhu Zi-Wen, Song Xiao-Ou, Cui Wei, Qi Feng-Li. Review of Visible and Infrared Image Fusion for Intelligent Object Detection. Computer Engineering and Applications, 2025, 61(17): 17−32 doi: 10.3778/j.issn.1002-8331.2501-0206 [22] Bie Q, Wang X, Xu X, et al. Visible-infrared cross-modal pedestrian detection: a summary. Journal of Image and Graphics, 2023, 28(5): 1287−1307 [23] Hwang S, Park J, Kim N, et al. Multispectral pedestrian detection: benchmark dataset and baseline. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Boston, USA: IEEE, 2015. 1037−1045 [24] 张天泷, 耿远超, 廖予祯, 许党朋. 多光谱目标检测算法及相关数据集综述. 强激光与粒子束, 2025, 37(05): 5−22 doi: 10.11884/HPLPB202537.240370Zhang Tian-Long, Geng Yuan-Chao, Liao Yu-Zhen, Xu Dang-Peng. A review of multispectral target detection algorithms and related datasets. Hign Power Laser and Particle Beams, 2025, 37(05): 5−22 doi: 10.11884/HPLPB202537.240370 [25] Zitova B, Flusser J. Image registration methods: a survey. Image and vision computing, 2003, 21(11): 977−1000 doi: 10.1016/S0262-8856(03)00137-9 [26] Maintz J B A, Viergever M A. A survey of medical image registration. Medical image analysis, 1998, 2(1): 1−36 doi: 10.1016/S1361-8415(01)80026-8 [27] Xie J, Nie J, Ding B, et al. Cross-modal local calibration and global context modeling network for RGB-infrared remote-sensing object detection. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2023, 16: 8933−8942 doi: 10.1109/JSTARS.2023.3315544 [28] Chen N, Xie J, Nie J, et al. Attentive alignment network for multispectral pedestrian detection. In: Proceedings of the 31st ACM International Conference on Multimedia. Ottawa, Canada: ACM, 2023. 3787−3795 [29] Shi Y, Li G, Shen Z, et al. BANet: Enhancing Weakly Aligned Multimodal Object Detection via Balanced Bidirectional Alignment Network. IEEE Transactions on Geoscience and Remote Sensing, 2026, 64: 1−14 doi: 10.1109/tgrs.2026.3674946 [30] Yuan M, Wang Y, Wei X. Translation, scale and rotation: cross-modal alignment meets RGB-infrared vehicle detection. In: Proceedings of the European Conference on Computer Vision. Cham, Switzerland: Springer Nature Switzerland, 2022. 509−525 [31] Yuan M, Shi X, Wang N, et al. Improving RGB-infrared object detection with cascade alignment-guided transformer. Information Fusion, 2024, 105: 102246 doi: 10.1016/j.inffus.2024.102246 [32] Song K, Xue X, Wen H, et al. Misaligned visible-thermal object detection: A drone-based benchmark and baseline. IEEE Transactions on Intelligent Vehicles, 2024, 9(11): 7449−7460 doi: 10.1109/TIV.2024.3398429 [33] Yu C, Shin Y. A Cross-Modality Feature Adaptive Interaction Approach for RGB-Infrared Object Detection in Aerial Imagery. IEEE Transactions on Geoscience and Remote Sensing, 2026, 64: 1−15 [34] Xiao H, Zhuang J, Yang B, et al. SAFF: A Spatially-Aware Fusion Framework for Effective and Efficient Aerial Object Detection. IEEE Transactions on Geoscience and Remote Sensing, 2026, 64: 1−13 doi: 10.1109/tgrs.2026.3674467 [35] Fu H, Yuan J, Zhong G, et al. CF-Deformable DETR: an end-to-end alignment-free model for weakly aligned visible-infrared object detection. In: Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence. Jeju, South Korea: IJCAI, 2024. 758−766 [36] Tian C, Zhou Z, Huang Y, et al. Cross-modality proposal-guided feature mining for unregistered RGB-thermal pedestrian detection. IEEE Transactions on Multimedia, 2024, 26: 6449−6461 doi: 10.1109/TMM.2024.3350926 [37] Chen C, Qi J, Liu X, et al. Weakly misalignment-free adaptive feature alignment for UAVs-based multimodal object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle, USA: IEEE, 2024. 26836−26845 [38] Zhao Z, Zhang W, Xiao Y, et al. Reflectance-Guided Progressive Feature Alignment Network for All-Day UAV Object Detection. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 1−15 doi: 10.1109/tgrs.2025.3574963 [39] Jin G, Zhao T, Yan J, et al. Contextually-guided state space fusion for misaligned multi-spectral object detection. In: Proceedings of the 33rd ACM International Conference on Multimedia. Dublin, Ireland: ACM, 2025. 2526−2535 [40] Guan D, Cao Y, Yang J, et al. Fusion of multispectral data through illumination-aware deep neural networks for pedestrian detection. Information Fusion, 2019, 50: 148−157 doi: 10.1016/j.inffus.2018.11.017 [41] Li C, Song D, Tong R, et al. Illumination-aware faster R-CNN for robust multispectral pedestrian detection. Pattern Recognition, 2019, 85: 161−171 doi: 10.1016/j.patcog.2018.08.005 [42] Zhou K, Chen L, Cao X. Improving multispectral pedestrian detection by addressing modality imbalance problems. In: Proceedings of the European Conference on Computer Vision. Cham, Switzerland: Springer International Publishing, 2020. 787−803 [43] Hu Q, Yu H, Zhou Z, et al. IAF-RTDETR: Illumination Evaluation-Driven Multimodal Object Detection Network for Infrared-Visible Dual-Source Fusion. Electronics, 2026, 15(6): 1332 doi: 10.3390/electronics15061332 [44] Liu T, Lam K M, Zhao R, et al. Deep cross-modal representation learning and distillation for illumination-invariant pedestrian detection. IEEE Transactions on Circuits and Systems for Video Technology, 2021, 32(1): 315−329 doi: 10.1109/tcsvt.2021.3060162 [45] Fu L, Gu W, Ai Y, et al. Adaptive spatial pixel-level feature fusion network for multispectral pedestrian detection. Infrared Physics & Technology, 2021, 116: 103770 doi: 10.1016/j.infrared.2021.103770 [46] Zhang H, Fromont E, Lefèvre S, et al. Guided attentive feature fusion for multispectral pedestrian detection. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. Virtual Event: IEEE, 2021. 72−80 [47] Kim J U, Park S, Ro Y M. Uncertainty-guided cross-modal learning for robust multispectral pedestrian detection. IEEE Transactions on Circuits and Systems for Video Technology, 2021, 32(3): 1510−1523 doi: 10.1109/tcsvt.2021.3076466 [48] Cao B, Sun Y, Zhu P, et al. Multi-modal gated mixture of local-to-global experts for dynamic image fusion. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris, France: IEEE, 2023. 23555−23564 [49] Cai W, Li Z, Dong J, et al. Modality balancing mechanism for RGB-infrared object detection in aerial image. In: Proceedings of the Chinese Conference on Pattern Recognition and Computer Vision. Singapore: Springer Nature Singapore, 2023. 81−93 [50] Lee W Y, Jovanov L, Philips W. Multimodal pedestrian detection based on cross-modality reference search. IEEE Sensors Journal, 2024, 24(10): 17291−17306 doi: 10.1109/JSEN.2024.3386709 [51] Xing Y, Yang S, Wang S, et al. MS-DETR: Multispectral pedestrian detection transformer with loosely coupled fusion and modality-balanced optimization. IEEE Transactions on Intelligent Transportation Systems, 2024, 25(12): 20628−20642 doi: 10.1109/TITS.2024.3450584 [52] Dang M, Liu G, Zhao J, et al. DDFD: diffusion-based denoising fusion for object detection in infrared-visible images. In: Proceedings of the 33rd ACM International Conference on Multimedia. Dublin, Ireland: ACM, 2025. 1452−1461 [53] Sun X, Yu Y, Cheng Q. Low-rank multimodal remote sensing object detection with frequency filtering experts. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 1−14 doi: 10.1109/tgrs.2024.3446814 [54] Liu Y, Jiang W. Frequency mining and complementary fusion network for RGB-infrared object detection. IEEE Geoscience and Remote Sensing Letters, 2024, 21: 1−5 doi: 10.1109/lgrs.2024.3448493 [55] Zhu H, Dong W, Yang L, et al. WaveMamba: wavelet-driven Mamba fusion for RGB-infrared object detection. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. Honolulu, USA: IEEE, 2025. 11219−11229 [56] Wang L, Bao Z, Lu D. GEM-YOLO: A Lightweight and Real-Time RGBT Object Detector with Gated Multimodal Fusion. Sensors, 2026, 26(7): 2035 doi: 10.3390/s26072035 [57] He X, Tang C, Zou X, et al. Multispectral object detection via cross-modal conflict-aware learning. In: Proceedings of the 31st ACM International Conference on Multimedia. Ottawa, Canada: ACM, 2023. 1465−1474 [58] Hu Y, Shi L, Yao L, et al. Dual attention feature fusion for visible-infrared object detection. In: Proceedings of the International Conference on Artificial Neural Networks. Cham, Switzerland: Springer Nature Switzerland, 2023. 53−65 [59] Zhou M, Li T, Qiao C, et al. Dmm: Disparity-guided multispectral mamba for oriented object detection in remote sensing. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 1−13 doi: 10.1109/tgrs.2025.3578309 [60] Bao W, Huang M, Hu J, et al. Dual dynamic cross-modal interaction network for multimodal remote sensing object detection. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 1−13 doi: 10.1109/tgrs.2025.3530085 [61] Li T, Li S, Li S, et al. SAM-guided semantic knowledge fusion for visible-infrared object detection. In: Proceedings of the 33rd ACM International Conference on Multimedia. Dublin, Ireland: ACM, 2025. 8835−8844 [62] Zhao T, Yuan M, Jiang F, et al. Removal then selection: A coarse-to-fine fusion perspective for RGB-infrared object detection. IEEE Transactions on Intelligent Transportation Systems, 2025, 27(2): 2504−2519 doi: 10.1109/tits.2025.3638627 [63] Li X, Chen S, Tian C, et al. M2FNet: mask-guided multi-level fusion for RGB-T pedestrian detection. IEEE Transactions on Multimedia, 2024, 26: 8678−8690 doi: 10.1109/TMM.2024.3381377 [64] Pan C, Jiang Q, Zheng H, et al. DANet: A Dual-Branch Framework with Diffusion-Integrated Autoencoder for Infrared-Visible Image Fusion. IEEE Transactions on Instrumentation and Measurement, 2025, 74: 1−13 [65] Cao Y, Guan D, Huang W, et al. Pedestrian detection with unsupervised multispectral feature learning using deep neural networks. Information Fusion, 2019, 46: 206−217 doi: 10.1016/j.inffus.2018.06.005 [66] Guan D, Luo X, Cao Y, et al. Unsupervised domain adaptation for multispectral pedestrian detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. Long Beach, USA: IEEE, 2019. 0−0 [67] Do D P, Kim T, Na J, et al. D3T: distinctive dual-domain teacher zigzagging across RGB-thermal gap for domain-adaptive object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle, USA: IEEE, 2024. 23313−23322 [68] Zhang H, Fromont E, Lefèvre S, et al. Deep active learning from multispectral data through cross-modality prediction inconsistency. In: Proceedings of the 2021 IEEE International Conference on Image Processing. Anchorage, USA: IEEE, 2021. 449−453 [69] Kim J, Kim H, Kim T, et al. MLPD: Multi-label pedestrian detector in multispectral domain. IEEE Robotics and Automation Letters, 2021, 6(4): 7846−7853 doi: 10.1109/LRA.2021.3099870 [70] Wang Y, Wei S, Xu S, et al. Confidence-driven Unimodal Interference Removal for Enhanced Multimodal Object Detection. IEEE Transactions on Circuits and Systems for Video Technology, 2025, 35(11): 11041−11053 doi: 10.1109/TCSVT.2025.3578340 [71] Yang S, Xing Y, Zhang S, et al. On Modality Incomplete Infrared-Visible Object Detection: An Architecture Compatibility Perspective. arXiv preprint arXiv: 2511.06406, 2025 [72] Kim M, Joung S, Park K, et al. Unpaired cross-spectral pedestrian detection via adversarial feature learning. In: Proceedings of the 2019 IEEE International Conference on Image Processing. Taipei, Taiwan: IEEE, 2019. 1650−1654 [73] Liu Y, Zhang L. Multimodal decomposed distillation with instance alignment and uncertainty compensation for thermal object detection. In: Proceedings of the 33rd ACM International Conference on Multimedia. Dublin, Ireland: ACM, 2025. 2294−2303 [74] Kieu M, Bagdanov A D, Bertini M, et al. Task-conditioned domain adaptation for pedestrian detection in thermal imagery. In: Proceedings of the European Conference on Computer Vision. Cham, Switzerland: Springer International Publishing, 2020. 546−562 [75] Medeiros H R, Pena F A G, Aminbeidokhti M, et al. HalluciDet: hallucinating RGB modality for person detection through privileged information. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. Waikoloa, USA: IEEE, 2024. 1444−1453 [76] Medeiros H R, Aminbeidokhti M, Peña F A G, et al. Modality translation for object detection adaptation without forgetting prior knowledge. In: Proceedings of the European Conference on Computer Vision. Cham, Switzerland: Springer Nature Switzerland, 2024. 51−68 [77] Ouyang J, Jin P, Wang Q. Multimodal feature-guided pretraining for RGB-T perception. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2024, 17: 16041−16050 doi: 10.1109/JSTARS.2024.3454054 [78] Azeem A, Li Z, Siddique A, et al. Unified multimodal fusion transformer for few shot object detection for remote sensing images. Information Fusion, 2024, 111: 102508 doi: 10.1016/j.inffus.2024.102508 [79] Zhang Y, Zeng W, Jin S, et al. When pedestrian detection meets multi-modal learning: generalist model and benchmark dataset. In: Proceedings of the European Conference on Computer Vision. Cham, Switzerland: Springer Nature Switzerland, 2024. 430−448 [80] Zhang L, Liu Z, Zhang S, et al. Cross-modality interactive attention network for multispectral pedestrian detection. Information Fusion, 2019, 50: 20−29 doi: 10.1016/j.inffus.2018.09.015 [81] Cao Y, Luo X, Yang J, et al. Locality guided cross-modal feature aggregation and pixel-level fusion for multispectral pedestrian detection. Information Fusion, 2022, 88: 1−11 doi: 10.1016/j.inffus.2022.06.008 [82] Li Q, Zhang C, Hu Q, et al. Confidence-aware fusion using dempster-shafer theory for multispectral pedestrian detection. IEEE Transactions on Multimedia, 2022, 25: 3420−3431 doi: 10.1109/tmm.2022.3160589 [83] Qingyun F, Zhaokui W. Cross-modality attentive feature fusion for object detection in multispectral remote sensing imagery. Pattern Recognition, 2022, 130: 108786 doi: 10.1016/j.patcog.2022.108786 [84] Wang Q, Chi Y, Shen T, et al. Improving RGB-infrared object detection by reducing cross-modality redundancy. Remote Sensing, 2022, 14(9): 2020 doi: 10.3390/rs14092020 [85] Li R, Xiang J, Sun F, et al. Multiscale cross-modal homogeneity enhancement and confidence-aware fusion for multispectral pedestrian detection. IEEE Transactions on Multimedia, 2023, 26: 852−863 doi: 10.1109/tmm.2023.3272471 [86] Shen J, Chen Y, Liu Y, et al. ICAFusion: Iterative cross-attention guided feature fusion for multispectral object detection. Pattern Recognition, 2024, 145: 109913 doi: 10.1016/j.patcog.2023.109913 [87] Lan X, Liu S, Zhang Z, et al. DCSANet: dual cross-channel and spatial attention make RGB-T object detection better. In: Proceedings of the 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems. Abu Dhabi, United Arab Emirates: IEEE, 2024. 12552−12558 [88] Kang X, Yin H, Duan P. Global-local feature fusion network for visible-infrared vehicle detection. IEEE Geoscience and Remote Sensing Letters, 2024, 21: 1−5 [89] Hu S, Bonardi F, Bouchafa S, et al. Rethinking self-attention for multispectral object detection. IEEE Transactions on Intelligent Transportation Systems, 2024, 25(11): 16300−16311 doi: 10.1109/TITS.2024.3412417 [90] Deevi S A, Lee C, Gan L, et al. RGB-X object detection via scene-specific fusion modules. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. Waikoloa, USA: IEEE, 2024. 7366−7375 [91] Sun X, Zhu Y, Huang H. Specificity-guided cross-modal feature reconstruction for RGB-infrared object detection. IEEE Transactions on Intelligent Transportation Systems, 2024, 26(1): 950−961 doi: 10.1109/tits.2024.3495028 [92] Xiong Z, Yao Z, Liu X, et al. Efficient Multispectral Object Detection with attentive feature aggregation leveraging zero-shot implicit illumination guidance. Information Fusion, 2025, 118: 102939 doi: 10.1016/j.inffus.2025.102939 [93] Zheng S, Junfeng L, Zeng J. MCAFNet: Multiscale cross-modality adaptive fusion network for multispectral object detection. Digital Signal Processing, 2025, 159: 104996 doi: 10.1016/j.dsp.2025.104996 [94] Li C, Peng X. Joint Transformer and Mamba fusion for multispectral object detection. Image and Vision Computing, 2025, 156: 105468 doi: 10.1016/j.imavis.2025.105468 [95] Qian J, Qiao B, Zhang Y, et al. DACFusion: Dual Asymmetric Cross-Attention guided feature fusion for multispectral object detection. Neurocomputing, 2025, 635: 129913 doi: 10.1016/j.neucom.2025.129913 [96] Li T, Ye M, Wu T, et al. Pseudo visible feature fine-grained fusion for thermal object detection. In: Proceedings of the Computer Vision and Pattern Recognition Conference. Nashville, USA: IEEE, 2025. 6710−6719 [97] González A, Fang Z, Socarras Y, et al. Pedestrian detection at day/night time with visible and FIR cameras: A comparison. Sensors, 2016, 16(6): 820 doi: 10.3390/s16060820 [98] FLIR Systems. (2018). FREE Teledyne FLIR Thermal Dataset for Algorithm Training [99] Zhang H, Fromont E, Lefèvre S, et al. Multispectral fusion for object detection with cyclic fuse-and-refine blocks. In: Proceedings of the 2020 IEEE International Conference on Image Processing. Virtual Event: IEEE, 2020. 276−280 [100] Takumi K, Watanabe K, Ha Q, et al. Multispectral object detection for autonomous vehicles. In: Proceedings of the Thematic Workshops of ACM Multimedia. Mountain View, USA: ACM, 2017. 35−43 [101] Jia X, Zhu C, Li M, et al. LLVIP: a visible-infrared paired dataset for low-light vision. In: Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops. Virtual Event: IEEE, 2021. 3496−3504 [102] Liu J, Fan X, Huang Z, et al. Target-aware dual adversarial learning and a multi-scenario multi-modality benchmark to fuse infrared and visible for object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans, USA: IEEE, 2022. 5802−5811 [103] Shin U, Park J, Kweon I S. Deep depth estimation from thermal image. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver, Canada: IEEE, 2023. 1043−1053 [104] Franchi G, Hariat M, Yu X, et al. InfraParis: a multi-modal and multi-task autonomous driving dataset. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. Waikoloa, USA: IEEE, 2024. 2973−2983 [105] Chen Z, Qian Y, Yang X, et al. AMFD: Distillation via adaptive multimodal fusion for multispectral pedestrian detection. IEEE Transactions on Multimedia, 2025, 27: 8298−8310 doi: 10.1109/TMM.2025.3604937 [106] Razakarivony S, Jurie F. Vehicle detection in aerial imagery: A small target detection benchmark. Journal of Visual Communication and Image Representation, 2016, 34: 187−203 doi: 10.1016/j.jvcir.2015.11.002 [107] Drone-based RGB-infrared cross-modality vehicle detection via uncertainty-aware learning. IEEE Transactions on Circuits and Systems for Video Technology, 2022, 32(10): 6700−6713 [108] Sun Y, Cao B, Zhu P, et al. DetFusion: a detection-driven infrared and visible image fusion network. In: Proceedings of the 30th ACM International Conference on Multimedia. Lisbon, Portugal: ACM, 2022. 4003−4011 [109] Zhu Y, Sun X, Wang M, et al. Multi-modal feature pyramid transformer for rgb-infrared object detection. IEEE Transactions on Intelligent Transportation Systems, 2023, 24(9): 9984−9995 doi: 10.1109/TITS.2023.3266487 [110] Li Q, Zhang C, Hu Q, et al. Stabilizing multispectral pedestrian detection with evidential hybrid fusion. IEEE Transactions on Circuits and Systems for Video Technology, 2023, 34(4): 3017−3029 doi: 10.1109/tcsvt.2023.3306870 [111] Guo J, Gao C, Liu F, et al. DAMSDet: dynamic adaptive multispectral detection transformer with competitive query selection and adaptive feature fusion. In: Proceedings of the European Conference on Computer Vision. Cham, Switzerland: Springer Nature Switzerland, 2024. 464−481 [112] Wang H, Wang C, Fu Q, et al. Cross-modal oriented object detection of UAV aerial images based on image feature. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 1−21 doi: 10.1109/tgrs.2024.3367934 [113] Shang X, Li N, Li D, et al. CCLDet: A cross-modality and cross-domain low-light detector. IEEE Transactions on Intelligent Transportation Systems, 2025, 26(3): 3284−3294 doi: 10.1109/TITS.2024.3522086 [114] Yuan M, Wei X. C2former: Calibrated and complementary transformer for rgb-infrared object detection. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 1−12 doi: 10.1109/tgrs.2024.3376819 [115] Cheng X, Geng K, Wang Z, et al. SLBAF-Net: Super-Lightweight bimodal adaptive fusion network for UAV detection in low recognition environment. Multimedia Tools and Applications, 2023, 82(30): 47773−47792 doi: 10.1007/s11042-023-15333-w [116] Liu K, Li T, Peng D. Aerial image object detection based on RGB-Infrared multi-branch progressive fusion. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 1−14 [117] Kim J U, Park S, Ro Y M. Towards versatile pedestrian detector with multisensory-matching and multispectral recalling memory. In: Proceedings of the 36th AAAI Conference on Artificial Intelligence. Virtual Event: AAAI, 2022. 1157−1165 [118] Dasgupta K, Das A, Das S, et al. Spatio-contextual deep network-based multimodal pedestrian detection for autonomous driving. IEEE transactions on intelligent transportation systems, 2022, 23(9): 15940−15950 doi: 10.1109/TITS.2022.3146575 [119] Li C, Song D, Tong R, et al. Multispectral pedestrian detection via simultaneous detection and segmentation. In: Proceedings of the British Machine Vision Conference. Newcastle, UK: BMVA Press, 2018. 225.1−225.12 [120] Kim T, Shin S, Yu Y, et al. Causal mode multiplexer: a novel framework for unbiased multispectral pedestrian detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle, USA: IEEE, 2024. 26784−26793 [121] Jang J, Park C, Kim H, et al. Multispectral object detection enhanced by cross-modal information complementary and cosine similarity channel resampling modules. In: Proceedings of the 2025 IEEE/CVF Winter Conference on Applications of Computer Vision. Tucson, USA: IEEE, 2025. 9437−9446 [122] Xiang X, Zhou G, Niu B, et al. Infrared-visible image fusion meets object detection: Towards unified optimization for multimodal perception. Remote Sensing, 2025, 17(21): 3637 doi: 10.3390/rs17213637 [123] Ahmed M, El-Sheimy N, Leung H. Dual-modal approach for ship detection: Fusing synthetic aperture radar and optical satellite imagery. Sensors, 2025, 25(2): 329 doi: 10.3390/s25020329 [124] Gao G, Wang M, Zhang X, et al. DEN: A new method for SAR and optical image fusion and intelligent classification. IEEE Transactions on Geoscience and Remote Sensing, 2024, 63: 1−18 [125] Qi Y, Yang S, Chen J, et al. A Modality Alignment and Fusion-Based Method for Around-the-Clock Remote Sensing Object Detection. Sensors, 2025, 25(16): 4964 doi: 10.3390/s25164964 [126] Choi J D, Kim M Y. A sensor fusion system with thermal infrared camera and LiDAR for autonomous vehicles and deep learning based object detection. ICT Express, 2023, 9(2): 222−227 doi: 10.1016/j.icte.2021.12.016 [127] Li C, Ni J, Luo Y, et al. Cross-Modality Fusion of Visible Light, Infrared, and SAR Images Under Few-Shot Conditions for Target Recognition. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2025, 19: 3657−3672 doi: 10.1109/jstars.2025.3649648 -
计量
- 文章访问数: 30
- HTML全文浏览量: 13
- 被引次数: 0
下载: