Audio-visual Segmentation Network Combining State Space Models and Dynamic Focal Loss
-
摘要: 视听目标分割任务旨在依据音频信号对视频中的发声物体进行像素级分割, 但在处理复杂动态场景时, 现有方法面临音视频特征交互不充分与像素级样本分布失衡两大挑战.为此, 提出结合状态空间模型与动态焦点损失的视听分割网络MambaAVS, 构建跨时间音频查询生成与跨模态音视频交互两大核心模块.首先, 利用跨时间音频查询生成机制提取关键频域信息以增强音频表征; 随后, 通过跨模态音视频交互模块动态捕捉音视频模态间的长距离依赖与语义对应关系, 实现特征的深度对齐与交互.此外, 针对训练过程中存在的难易样本分布不均与正负比例悬殊问题, 提出动态焦点损失, 通过自适应调节难易样本权重与梯度表征, 有效缓解了模型优化偏差.在 AVSBench-Object 与 AVSBench-Semantic 数据集上的实验结果表明, 所提 MambaAVS 在单声源、多声源及视听语义分割数据集上的mIoU比分别达到82.86%、62.56%和36.79%; 与同类先进方法 AVSegFormer 相比, MambaAVS在3个数据集上的精度分别提升了0.80%、4.20%和0.13%, 验证了其在复杂视听场景下的有效性.Abstract: Audio-visual segmentation aims to perform pixel-level segmentation of sounding objects in videos by leveraging audio-visual cues. However, existing approaches still suffer from two major challenges in complex dynamic scenarios: insufficient cross-modal feature interaction and severe pixel-level class imbalance. To address these issues, this paper proposes MambaAVS, an audio-visual segmentation network built upon a state space model and optimized with a dynamic focal loss. Specifically, MambaAVS introduces two key modules: Cross-Temporal Audio Query Generation and Cross-Modal Audio-Visual Interaction. First, the Cross-Temporal Audio Query Generation module captures discriminative frequency-domain representations across temporal dimensions to enhance audio representations. Then, the Cross-Modal Audio-Visual Interaction module dynamically models long-range cross-modal dependencies and semantic correspondences between audio and visual modalities, enabling effective feature alignment and interaction. Furthermore, to alleviate the optimization bias caused by the imbalanced distribution of hard/easy samples and the severe foreground-background pixel imbalance during training, a dynamic focal loss is proposed to adaptively adjust sample weights and gradient contributions. Extensive experiments on the AVSBench-Object and AVSBench-Semantic benchmarks demonstrate that MambaAVS achieves mIoU scores of 82.86%, 62.56%, and 36.79% on the Single Sound Source, Multiple Sound Source, and Audio-Visual Semantic Segmentation tasks, respectively. Compared with the state-of-the-art method AVSegFormer, MambaAVS improves mIoU by 0.80%, 4.20%, and 0.13% on the three benchmarks, respectively, demonstrating its effectiveness in complex audio-visual scenarios.
-
Key words:
- audio-visual segmentation /
- state space model /
- feature interaction /
- dynamic focal loss
-
表 1 以 ResNet-50 为骨干网络, 所提 MambaAVS 与不同 AVS 方法的分割性能统计
Table 1 Using ResNet-50 as the backbone network, a comparison of the segmentation performance of the proposed MambaAVS with various AVS methods
方法 骨干网络 S4 MS3 AVSS F-score mIoU(%) F-score mIoU(%) F-score mIoU(%) MSSL [5] ResNet-50 66.30 44.89 36.30 26.13 — — 3DC [35] ResNet-50 75.90 57.10 50.30 36.92 21.60 17.17 LVS [36] ResNet-50 51.00 37.94 33.00 29.45 — — SST [37] ResNet-50 80.10 66.29 57.20 42.57 — — AVSegFormer [38] ResNet-50 85.21 75.08 62.89 50.37 31.26 26.71 MambaAVS(本文方法) ResNet-50 86.30 76.37 63.75 51.95 33.83 28.59 表 2 不同骨干网络设置下, 所提 MambaAVS 与不同 AVS 方法的分割性能统计
Table 2 Performance statistics for the proposed MambaAVS and various AVS methods across different backbone networks
方法 骨干网络 S4 MS3 AVSS F-score mIoU(%) F-score mIoU(%) F-score mIoU(%) AOT [39] Swin-B – – – – 31.00 25.40 LGVT [40] Swin-T 87.30 74.94 59.30 40.71 – – AVSBench [1] PVTv2 87.90 78.74 64.50 54.00 35.20 29.77 AVSC [26] PVTv2 88.60 81.29 65.70 59.50 – – ECMVAE [2] PVTv2 90.10 81.74 70.80 57.84 – – AVSegFormer [38] PVTv2 89.90 82.06 69.30 58.36 42.00 36.66 CQFormer [41] PVTv2 91.18 83.62 72.67 60.99 43.01 38.09 DiffusionAVS [42] PVTv2 90.30 81.51 71.20 59.62 43.00 38.10 MambaAVS(本文方法) PVTv2 90.78 82.86 73.49 62.56 43.04 36.79 表 3 使用与不使用 MambaAVS 和 DFL 组件时的性能比较
Table 3 Performance comparison with and without the MambaAVS and DFL components
方法 骨干网络 S4 MS3 参数量(M) 推理速度(FPS) F-score mIoU(%) F-score mIoU(%) 基线 ResNet-50 85.82 75.41 62.89 50.37 150.90 152.22 基线+ MambaAVS ResNet-50 85.82 76.03 64.71 52.82 154.68 113.04 基线+ DFL ResNet-50 86.07 75.97 63.62 51.73 150.90 152.22 基线+ MambaAVS + DFL ResNet-50 86.30 76.37 63.75 51.95 154.68 113.04 表 4 MambaAVS 组件的消融实验
Table 4 Ablation experiment of the MambaAVS component
方法 骨干网络 S4 MS3 F-score mIoU(%) F-score mIoU(%) 基线 ResNet-50 85.21 75.08 62.89 50.37 基线 + CTAQG ResNet-50 85.53 75.42 63.22 50.44 基线 + CAVII ResNet-50 85.60 75.70 63.72 51.20 MambaAVS ResNet-50 85.82 76.03 64.71 52.82 表 5 DFL 组件的消融实验
Table 5 Ablation experiment of the DFL component
方法 骨干网络 S4 MS3 F-score mIoU(%) F-score mIoU(%) 基线 ResNet-50 85.31 74.93 64.05 51.15 基线 +$\xi_{a}$ ResNet-50 85.90 75.89 63.67 50.95 基线 +$\xi_{\text{g}}$ ResNet-50 86.02 75.77 62.71 51.98 基线 +$\xi_{\text{l}}$ ResNet-50 85.99 75.79 64.29 51.98 DFL ResNet-50 86.07 75.97 63.62 51.73 -
[1] J. Zhou, J. Wang, J. Zhang, W. Sun, J. Zhang, S. Birchfield, D. Guo, L. Kong, M. Wang, Y. Zhong, Audio–visual segmentation, in: European Conference on Computer Vision, Springer, 2022, pp. 386–403. [2] Y. Mao, J. Zhang, M. Xiang, Y. Zhong, Y. Dai, Multimodal Variational Auto-encoder based Audio-Visual Segmentation, in: 2023 IEEE/CVF International Conference on Computer Vision (ICCV), IEEE Computer Society, Los Alamitos, CA, USA, 2023, pp. 954–965. doi: 10.1109/ICCV51070.2023.00094. [3] R. Arandjelovic, A. Zisserman, Look, listen and learn, in: Proceedings of the IEEE international conference on computer vision, 2017, pp. 609–617. [4] T. Liang, G. Lin, L. Feng, Y. Zhang, F. Lv, Attention is not enough: Mitigating the distribution discrepancy in asynchronous multimodal sequence fusion, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 8148–8156. [5] R. Qian, D. Hu, H. Dinkel, M. Wu, N. Xu, W. Lin, Multiple sound sources localization from coarse to fine, in: A. Vedaldi, H. Bischof, T. Brox, J.-M. Frahm (Eds.), Computer Vision – ECCV 2020, Springer International Publishing, Cham, 2020, pp. 292–308. [6] A. Nagrani, S. Yang, A. Arnab, A. Jansen, C. Schmid, C. Sun. Attention bottlenecks for multimodal fusion. Advances in neural information processing systems, 2021, 34: 14200−14213 [7] S. Gong, Y. Zhuge, L. Zhang, Y. Wang, P. Zhang, L. Wang, H. Lu, Avs-mamba: Exploring temporal and multi-modal mamba for audiovisual segmentation, IEEE Transactions on Multimedia (2025). [8] 王 锋, 银莹, 王佳炎, 唐勇, 李胜, 赵静, “基于高斯 泼溅的轻量级重建场景分割方法, ”计算机学报, vol. 48, no. 5, 2025.F. Wang, Y. Yin, J. Wang, Y. Tang, S. Li, J. Zhao, Object segmentation in 3d reconstructed scenes based on gaussian splatting, Chinse Joural of Computers 48 (5) (2025 [9] J. Lin, Z. Xiao, X. Wei, P. Duan, X. He, R. Dian, Z. Li, S. Li. Click-pixel cognition fusion network with balanced cut for interactive image segmentation. IEEE Transactions on Image Processing, 2024, 33: 177−190 doi: 10.1109/TIP.2023.3338003 [10] 林家丞, 陈嘉俊, 李智勇, 王耀南. 基于语义概念关联的参考多目标跟 踪方法. 自动化学报, 2025, 51(12): 2664−2678 doi: 10.16383/j.aas.c250118J. Lin, J. Chen, Z. Li, Y. Wang. Semantic conceptual association-based method for referring multi-object tracking. Acta Automatica Sinica, 2025, 51(12): 2664−2678 doi: 10.16383/j.aas.c250118 [11] J. Lin, J. Chen, K. Peng, X. He, Z. Li, R. Stiefelhagen, K. Yang, Echotrack: Auditory referring multi-object tracking for autonomous driving, IEEE Transactions on Intelligent Transportation Systems 25 (11) (2024) 18964–18977. [12] 刘袁缘, 刘树阳, 刘云娇, 袁雨晨, 唐厂, 罗威. 提示学习在 计算机视觉中的分类、应用及展望. 自动化学报, 2025, 51(5): 1021−1040 doi: 10.16383/j.aas.c240177Y. Liu, S. Liu, Y. Liu, Y. Yuan, C. Tang, W. Luo. The classification, applications, and prospects of prompt learning in computer vision. Acta Automatica Sinica, 2025, 51(5): 1021−1040 doi: 10.16383/j.aas.c240177 [13] K. Sofiiuk, I. Petrov, O. Barinova, A. Konushin, f-brs: Rethinking backpropagating refinement for interactive segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 8623–8632. [14] 封筠, 张天, 史屹琛, 王辉, 胡晶晶. 融合双阶段特征与 transformer 编码的交互式图像分割. 计算机辅助设计 与图形学学报, 2024, 36(6): 831−843 doi: 10.3724/SP.J.1089.2024.19922J. Feng, T. Zhang, Y. Shi, H. Wang, J. Hu. Interactive image segmentation based on fusion of two-stage feature and transformer encoder. Journal of Computer-Aided Design & Computer Graphics, 2024, 36(6): 831−843 doi: 10.3724/SP.J.1089.2024.19922 [15] A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al., Segment anything, in: Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 4015–4026. [16] D. Liu, H. Zhang, F. Wu, Z.-J. Zha, Learning to assemble neural module tree networks for visual grounding, in: Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 4673–4682. [17] B. Miao, M. Bennamoun, Y. Gao, A. Mian, Spectrum-guided multi-granularity referring video object segmentation, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 920–930. [18] D. Wu, W. Han, T. Wang, X. Dong, X. Zhang, J. Shen, Referring multi-object tracking, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 14633–14642. [19] M. Cheng, Y. Sun, L. Wang, X. Zhu, K. Yao, J. Chen, G. Song, J. Han, J. Liu, E. Ding, J. Wang, Vista: Vision and scene text aggregation for cross-modal retrieval, in: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 5174–5183. doi: 10.1109/CVPR52688.2022.00512. [20] Y. Wang, W. Liu, G. Li, J. Ding, D. Hu, X. Li, Prompting segmentation with sound is generalizable audio-visual source localizer, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, 2024, pp. 5669–5677. [21] J. Chen, J. Lin, G. Zhong, H. Fu, K. Nai, K. Yang, Z. Li. Expression prompt collaboration transformer for universal referring video object segmentation. Knowledge-Based Systems, 2025, 311: 113006 doi: 10.1016/j.knosys.2025.113006 [22] Y. Ling, Y. Li, Z. Gan, J. Zhang, M. Chi, Y. Wang, Transavs: End-to-end audio-visual segmentation with transformer, in: ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2024, pp. 7845–7849. [23] K. Li, Z. Yang, L. Chen, Y. Yang, J. Xiao, Catr: Combinatorial-dependence audio-queried transformer for audio-visual video segmentation, in: Proceedings of the 31st ACM International Conference on Multimedia, MM ’23, Association for Computing Machinery, New York, NY, USA, 2023, p. 1485–1494. doi: 10.1145/3581783.3611724. [24] J. Liu, C. Ju, C. Ma, Y. Wang, Y. Wang, Y. Zhang, Audio-aware query-enhanced transformer for audio-visual segmentation, arXiv preprint arXiv: 2307.13236 (2023). [25] S. Huang, H. Li, Y. Wang, H. Zhu, J. Dai, J. Han, W. Rong, S. Liu, Discovering sounding objects by audio queries for audio visual segmentation, arXiv preprint arXiv: 2309.09501 (2023). [26] C. Liu, P. P. Li, X. Qi, H. Zhang, L. Li, D. Wang, X. Yu, Audio-visual segmentation by exploring cross-modal mutual semantics, in: Proceedings of the 31st ACM International Conference on Multimedia, 2023, pp. 7590–7598. [27] Q. Yang, X. Nie, T. Li, P. Gao, Y. Guo, C. Zhen, P. Yan, S. Xiang, Cooperation does matter: Exploring multi-order bilateral relations for audio-visual segmentation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 27134–27143. [28] J. Liu, Y. Wang, C. Ju, C. Ma, Y. Zhang, W. Xie, Annotation-free audio-visual segmentation, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 5604–5614. [29] M. M. Islam, G. Bertasius, Long movie clip classification with state-space video models, in: European Conference on Computer Vision, Springer, 2022, pp. 87–104. [30] A. Gu, T. Dao, Mamba: Linear-time sequence modeling with selective state spaces (2024). [31] L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, X. Wang, Vision mamba: Effcient visual representation learning with bidirectional state space model, arXiv preprint arXiv: 2401.09417 (2024). [32] Y. Yang, C. Ma, J. Yao, Z. Zhong, Y. Zhang, Y. Wang, Remamber: Referring image segmentation with mamba twister, in: European Conference on Computer Vision, Springer, 2024, pp. 108–126. [33] K. Zeng, H. Shi, J. Lin, S. Li, J. Cheng, K. Wang, Z. Li, K. Yang, Mambamos: Lidar-based 3d moving object segmentation with motion-aware state space model, in: Proceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 1505–1513. [34] J. Lin, J. Chen, K. Yang, A. Roitberg, S. Li, Z. Li, S. Li, Adaptiveclick: Click-aware transformer with adaptive focal loss for interactive image segmentation, IEEE Transactions on Neural Networks and Learning Systems 36 (3) (2025) 5759–5773. doi: 10.1109/TNNLS.2024.3378295. [35] S. Mahadevan, A. Athar, A. Osep, S. Hennen, L. Leal-Taixé, B. Leibe, Making a case for 3d convolutions for object segmentation in videos, BMVC abs/2008.11516 (2020). [36] H. Chen, W. Xie, T. Afouras, A. Nagrani, A. Vedaldi, A. Zisserman, Localizing visual sounds the hard way, in: 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 16862–16871. doi: 10.1109/CVPR46437.2021.01659. [37] B. Duke, A. Ahmed, C. Wolf, P. Aarabi, G. W. Taylor, Sstvos: Sparse spatiotemporal transformers for video object segmentation, 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021) 5908–5917. [38] S. Gao, Z. Chen, G. Chen, W. Wang, T. Lu. Avsegformer: Audio-visual segmentation with transformer. Proceedings of the AAAI Conference on Artificial Intelligence, 2024, 38(11): 12155−12163 doi: 10.1609/aaai.v38i11.29104 [39] Z. Yang, Y. Wei, Y. Yang. Associating objects with transformers for video object segmentation. Advances in Neural Information Processing Systems, 2021, 34: 2491−2502 [40] J. Zhang, J. Xie, N. Barnes, P. Li, Learning generative vision transformer with energy-based latent space for saliency prediction, in: M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, J. W. Vaughan (Eds.), Advances in Neural Information Processing Systems, Vol. 34, Curran Associates, Inc., 2021, pp. 15448–15463. [41] Y. Lv, Z. Liu, X. Chang. Consistency-queried transformer for audio-visual segmentation. IEEE Transactions on Image Processing, 2025, 34: 2616−2627 doi: 10.1109/TIP.2025.3563076 [42] Y. Mao, J. Zhang, M. Xiang, Y. Lv, D. Li, Y. Zhong, Y. Dai. Contrastive conditional latent diffusion for audio-visual segmentation. IEEE Transactions on Image Processing, 2025, 34: 4108−4119 doi: 10.1109/TIP.2025.3580269 -
计量
- 文章访问数: 13
- HTML全文浏览量: 3
- 被引次数: 0
下载: