• 中文核心
  • EI
  • 中国科技核心
  • Scopus
  • CSCD
  • 英国科学文摘

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

结合状态空间模型与动态焦点损失的视听分割网络

胡丙齐 林家丞 杨观赐

胡丙齐, 林家丞, 杨观赐. 结合状态空间模型与动态焦点损失的视听分割网络. 自动化学报, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250753
引用本文: 胡丙齐, 林家丞, 杨观赐. 结合状态空间模型与动态焦点损失的视听分割网络. 自动化学报, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250753
Hu Bing-Qi, Lin Jia-Cheng, Yang Guan-Ci. Audio-visual segmentation network combining state space models and dynamic focal loss. Acta Automatica Sinica, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250753
Citation: Hu Bing-Qi, Lin Jia-Cheng, Yang Guan-Ci. Audio-visual segmentation network combining state space models and dynamic focal loss. Acta Automatica Sinica, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250753

结合状态空间模型与动态焦点损失的视听分割网络

doi: 10.16383/j.aas.c250753 cstr: 32138.14.j.aas.c250753
基金项目: 国家自然科学基金(62373116), 贵州省基础研究计划面上项目(黔科合基础MS [2026] 079), 贵州省基础研究(自然科学)青年引导项目(黔科合基础QN [2025] 022)资助
详细信息
    作者简介:

    胡丙齐:贵州大学助理实验师. 2021年获得贵州大学硕士学位.主要研究方向为视频分割, 零样本学习. E-mail: bqhu@gzu.edu.cn

    林家丞:贵州大学特聘教授. 2025年获得湖南大学博士学位.主要研究方向为具身机器人的场景理解与多模态融合认知. 本文通信作者. E-mail: jcheng_lin@hnu.edu.cn

    杨观赐:贵州大学教授. 2012 年获得中国科学院计算机软件与理论专业博士学位. 主要研究方向为多模态融合认知与智能机器人, 智能机器人技能学习. E-mail: gcyang@gzu.edu.cn

  • 中图分类号: Y

Audio-visual Segmentation Network Combining State Space Models and Dynamic Focal Loss

Funds: Supported by National Natural Science Foundation of China (Grant Nos. 62373116), the Guizhou Provincial Science and Technology Projects (Grant No. QKHJC MS [2026] 079), the Youth Guidance Project of Guizhou Provincial Basic Research Program (Natural Science) (Grant No. Qiankehe Jichu QN[2025]022)
More Information
    Author Bio:

    HU Bing-Qi Assistant laboratory instructor at Guizhou University. He received his Master degree from Guizhou University in 2021. His research interests include video segmentation and zero-shot learning

    LIN Jia-Cheng Specially professor at Guizhou University. He received his Ph. D. degree from Hunan University in 2025. His research interest include embodied scene understanding and multimodal fusion cognition for robots. Corresponding author of this paper

    YANG Guan-Ci Professor at Guizhou University. He received his Ph.D. degree in computer software and theory from the Chinese Academy of Sciences in 2012. His research interests include multimodal fusion cognition for multimodal fusion cognition and intelligent robots, skill learning for intelligent robots

  • 摘要: 视听目标分割任务旨在依据音频信号对视频中的发声物体进行像素级分割, 但在处理复杂动态场景时, 现有方法面临音视频特征交互不充分与像素级样本分布失衡两大挑战.为此, 提出结合状态空间模型与动态焦点损失的视听分割网络MambaAVS, 构建跨时间音频查询生成与跨模态音视频交互两大核心模块.首先, 利用跨时间音频查询生成机制提取关键频域信息以增强音频表征; 随后, 通过跨模态音视频交互模块动态捕捉音视频模态间的长距离依赖与语义对应关系, 实现特征的深度对齐与交互.此外, 针对训练过程中存在的难易样本分布不均与正负比例悬殊问题, 提出动态焦点损失, 通过自适应调节难易样本权重与梯度表征, 有效缓解了模型优化偏差.在 AVSBench-Object 与 AVSBench-Semantic 数据集上的实验结果表明, 所提 MambaAVS 在单声源、多声源及视听语义分割数据集上的mIoU比分别达到82.86%、62.56%和36.79%; 与同类先进方法 AVSegFormer 相比, MambaAVS在3个数据集上的精度分别提升了0.80%、4.20%和0.13%, 验证了其在复杂视听场景下的有效性.
  • 图  1  所提MambaAVS示意图

    Fig.  1  Schematic diagram of the proposed MambaAVS

    图  2  融合状态空间模型与动态焦点损失的视听目标分割框架

    Fig.  2  A audio-visual object segmentation framework integrating state space models and dynamic focal loss

    图  3  不同数据集上MambaAVS的分割结果

    Fig.  3  Segmentation results of MambaAVS on different datasets

    图  4  使用或不使用MambaAVS组件时Grad-CAM可视化结果

    Fig.  4  Grad-CAM visualization results with and without the MambaAVS component

    图  5  失败案例分析

    Fig.  5  Analysis of failure cases

    表  1  以 ResNet-50 为骨干网络, 所提 MambaAVS 与不同 AVS 方法的分割性能统计

    Table  1  Using ResNet-50 as the backbone network, a comparison of the segmentation performance of the proposed MambaAVS with various AVS methods

    方法 骨干网络 S4 MS3 AVSS
    F-score mIoU(%) F-score mIoU(%) F-score mIoU(%)
    MSSL [5] ResNet-50 66.30 44.89 36.30 26.13
    3DC [35] ResNet-50 75.90 57.10 50.30 36.92 21.60 17.17
    LVS [36] ResNet-50 51.00 37.94 33.00 29.45
    SST [37] ResNet-50 80.10 66.29 57.20 42.57
    AVSegFormer [38] ResNet-50 85.21 75.08 62.89 50.37 31.26 26.71
    MambaAVS(本文方法) ResNet-50 86.30 76.37 63.75 51.95 33.83 28.59
    下载: 导出CSV

    表  2  不同骨干网络设置下, 所提 MambaAVS 与不同 AVS 方法的分割性能统计

    Table  2  Performance statistics for the proposed MambaAVS and various AVS methods across different backbone networks

    方法 骨干网络 S4 MS3 AVSS
    F-score mIoU(%) F-score mIoU(%) F-score mIoU(%)
    AOT [39] Swin-B 31.00 25.40
    LGVT [40] Swin-T 87.30 74.94 59.30 40.71
    AVSBench [1] PVTv2 87.90 78.74 64.50 54.00 35.20 29.77
    AVSC [26] PVTv2 88.60 81.29 65.70 59.50
    ECMVAE [2] PVTv2 90.10 81.74 70.80 57.84
    AVSegFormer [38] PVTv2 89.90 82.06 69.30 58.36 42.00 36.66
    CQFormer [41] PVTv2 91.18 83.62 72.67 60.99 43.01 38.09
    DiffusionAVS [42] PVTv2 90.30 81.51 71.20 59.62 43.00 38.10
    MambaAVS(本文方法) PVTv2 90.78 82.86 73.49 62.56 43.04 36.79
    下载: 导出CSV

    表  3  使用与不使用 MambaAVS 和 DFL 组件时的性能比较

    Table  3  Performance comparison with and without the MambaAVS and DFL components

    方法 骨干网络 S4 MS3 参数量(M) 推理速度(FPS)
    F-score mIoU(%) F-score mIoU(%)
    基线 ResNet-50 85.82 75.41 62.89 50.37 150.90 152.22
    基线+ MambaAVS ResNet-50 85.82 76.03 64.71 52.82 154.68 113.04
    基线+ DFL ResNet-50 86.07 75.97 63.62 51.73 150.90 152.22
    基线+ MambaAVS + DFL ResNet-50 86.30 76.37 63.75 51.95 154.68 113.04
    下载: 导出CSV

    表  4  MambaAVS 组件的消融实验

    Table  4  Ablation experiment of the MambaAVS component

    方法 骨干网络 S4 MS3
    F-score mIoU(%) F-score mIoU(%)
    基线 ResNet-50 85.21 75.08 62.89 50.37
    基线 + CTAQG ResNet-50 85.53 75.42 63.22 50.44
    基线 + CAVII ResNet-50 85.60 75.70 63.72 51.20
    MambaAVS ResNet-50 85.82 76.03 64.71 52.82
    下载: 导出CSV

    表  5  DFL 组件的消融实验

    Table  5  Ablation experiment of the DFL component

    方法 骨干网络 S4 MS3
    F-score mIoU(%) F-score mIoU(%)
    基线 ResNet-50 85.31 74.93 64.05 51.15
    基线 +$\xi_{a}$ ResNet-50 85.90 75.89 63.67 50.95
    基线 +$\xi_{\text{g}}$ ResNet-50 86.02 75.77 62.71 51.98
    基线 +$\xi_{\text{l}}$ ResNet-50 85.99 75.79 64.29 51.98
    DFL ResNet-50 86.07 75.97 63.62 51.73
    下载: 导出CSV
  • [1] J. Zhou, J. Wang, J. Zhang, W. Sun, J. Zhang, S. Birchfield, D. Guo, L. Kong, M. Wang, Y. Zhong, Audio–visual segmentation, in: European Conference on Computer Vision, Springer, 2022, pp. 386–403.
    [2] Y. Mao, J. Zhang, M. Xiang, Y. Zhong, Y. Dai, Multimodal Variational Auto-encoder based Audio-Visual Segmentation, in: 2023 IEEE/CVF International Conference on Computer Vision (ICCV), IEEE Computer Society, Los Alamitos, CA, USA, 2023, pp. 954–965. doi: 10.1109/ICCV51070.2023.00094.
    [3] R. Arandjelovic, A. Zisserman, Look, listen and learn, in: Proceedings of the IEEE international conference on computer vision, 2017, pp. 609–617.
    [4] T. Liang, G. Lin, L. Feng, Y. Zhang, F. Lv, Attention is not enough: Mitigating the distribution discrepancy in asynchronous multimodal sequence fusion, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 8148–8156.
    [5] R. Qian, D. Hu, H. Dinkel, M. Wu, N. Xu, W. Lin, Multiple sound sources localization from coarse to fine, in: A. Vedaldi, H. Bischof, T. Brox, J.-M. Frahm (Eds.), Computer Vision – ECCV 2020, Springer International Publishing, Cham, 2020, pp. 292–308.
    [6] A. Nagrani, S. Yang, A. Arnab, A. Jansen, C. Schmid, C. Sun. Attention bottlenecks for multimodal fusion. Advances in neural information processing systems, 2021, 34: 14200−14213
    [7] S. Gong, Y. Zhuge, L. Zhang, Y. Wang, P. Zhang, L. Wang, H. Lu, Avs-mamba: Exploring temporal and multi-modal mamba for audiovisual segmentation, IEEE Transactions on Multimedia (2025).
    [8] 王 锋, 银莹, 王佳炎, 唐勇, 李胜, 赵静, “基于高斯 泼溅的轻量级重建场景分割方法, ”计算机学报, vol. 48, no. 5, 2025.

    F. Wang, Y. Yin, J. Wang, Y. Tang, S. Li, J. Zhao, Object segmentation in 3d reconstructed scenes based on gaussian splatting, Chinse Joural of Computers 48 (5) (2025
    [9] J. Lin, Z. Xiao, X. Wei, P. Duan, X. He, R. Dian, Z. Li, S. Li. Click-pixel cognition fusion network with balanced cut for interactive image segmentation. IEEE Transactions on Image Processing, 2024, 33: 177−190 doi: 10.1109/TIP.2023.3338003
    [10] 林家丞, 陈嘉俊, 李智勇, 王耀南. 基于语义概念关联的参考多目标跟 踪方法. 自动化学报, 2025, 51(12): 2664−2678 doi: 10.16383/j.aas.c250118

    J. Lin, J. Chen, Z. Li, Y. Wang. Semantic conceptual association-based method for referring multi-object tracking. Acta Automatica Sinica, 2025, 51(12): 2664−2678 doi: 10.16383/j.aas.c250118
    [11] J. Lin, J. Chen, K. Peng, X. He, Z. Li, R. Stiefelhagen, K. Yang, Echotrack: Auditory referring multi-object tracking for autonomous driving, IEEE Transactions on Intelligent Transportation Systems 25 (11) (2024) 18964–18977.
    [12] 刘袁缘, 刘树阳, 刘云娇, 袁雨晨, 唐厂, 罗威. 提示学习在 计算机视觉中的分类、应用及展望. 自动化学报, 2025, 51(5): 1021−1040 doi: 10.16383/j.aas.c240177

    Y. Liu, S. Liu, Y. Liu, Y. Yuan, C. Tang, W. Luo. The classification, applications, and prospects of prompt learning in computer vision. Acta Automatica Sinica, 2025, 51(5): 1021−1040 doi: 10.16383/j.aas.c240177
    [13] K. Sofiiuk, I. Petrov, O. Barinova, A. Konushin, f-brs: Rethinking backpropagating refinement for interactive segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 8623–8632.
    [14] 封筠, 张天, 史屹琛, 王辉, 胡晶晶. 融合双阶段特征与 transformer 编码的交互式图像分割. 计算机辅助设计 与图形学学报, 2024, 36(6): 831−843 doi: 10.3724/SP.J.1089.2024.19922

    J. Feng, T. Zhang, Y. Shi, H. Wang, J. Hu. Interactive image segmentation based on fusion of two-stage feature and transformer encoder. Journal of Computer-Aided Design & Computer Graphics, 2024, 36(6): 831−843 doi: 10.3724/SP.J.1089.2024.19922
    [15] A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al., Segment anything, in: Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 4015–4026.
    [16] D. Liu, H. Zhang, F. Wu, Z.-J. Zha, Learning to assemble neural module tree networks for visual grounding, in: Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 4673–4682.
    [17] B. Miao, M. Bennamoun, Y. Gao, A. Mian, Spectrum-guided multi-granularity referring video object segmentation, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 920–930.
    [18] D. Wu, W. Han, T. Wang, X. Dong, X. Zhang, J. Shen, Referring multi-object tracking, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 14633–14642.
    [19] M. Cheng, Y. Sun, L. Wang, X. Zhu, K. Yao, J. Chen, G. Song, J. Han, J. Liu, E. Ding, J. Wang, Vista: Vision and scene text aggregation for cross-modal retrieval, in: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 5174–5183. doi: 10.1109/CVPR52688.2022.00512.
    [20] Y. Wang, W. Liu, G. Li, J. Ding, D. Hu, X. Li, Prompting segmentation with sound is generalizable audio-visual source localizer, in: Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, 2024, pp. 5669–5677.
    [21] J. Chen, J. Lin, G. Zhong, H. Fu, K. Nai, K. Yang, Z. Li. Expression prompt collaboration transformer for universal referring video object segmentation. Knowledge-Based Systems, 2025, 311: 113006 doi: 10.1016/j.knosys.2025.113006
    [22] Y. Ling, Y. Li, Z. Gan, J. Zhang, M. Chi, Y. Wang, Transavs: End-to-end audio-visual segmentation with transformer, in: ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2024, pp. 7845–7849.
    [23] K. Li, Z. Yang, L. Chen, Y. Yang, J. Xiao, Catr: Combinatorial-dependence audio-queried transformer for audio-visual video segmentation, in: Proceedings of the 31st ACM International Conference on Multimedia, MM ’23, Association for Computing Machinery, New York, NY, USA, 2023, p. 1485–1494. doi: 10.1145/3581783.3611724.
    [24] J. Liu, C. Ju, C. Ma, Y. Wang, Y. Wang, Y. Zhang, Audio-aware query-enhanced transformer for audio-visual segmentation, arXiv preprint arXiv: 2307.13236 (2023).
    [25] S. Huang, H. Li, Y. Wang, H. Zhu, J. Dai, J. Han, W. Rong, S. Liu, Discovering sounding objects by audio queries for audio visual segmentation, arXiv preprint arXiv: 2309.09501 (2023).
    [26] C. Liu, P. P. Li, X. Qi, H. Zhang, L. Li, D. Wang, X. Yu, Audio-visual segmentation by exploring cross-modal mutual semantics, in: Proceedings of the 31st ACM International Conference on Multimedia, 2023, pp. 7590–7598.
    [27] Q. Yang, X. Nie, T. Li, P. Gao, Y. Guo, C. Zhen, P. Yan, S. Xiang, Cooperation does matter: Exploring multi-order bilateral relations for audio-visual segmentation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 27134–27143.
    [28] J. Liu, Y. Wang, C. Ju, C. Ma, Y. Zhang, W. Xie, Annotation-free audio-visual segmentation, in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 5604–5614.
    [29] M. M. Islam, G. Bertasius, Long movie clip classification with state-space video models, in: European Conference on Computer Vision, Springer, 2022, pp. 87–104.
    [30] A. Gu, T. Dao, Mamba: Linear-time sequence modeling with selective state spaces (2024).
    [31] L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, X. Wang, Vision mamba: Effcient visual representation learning with bidirectional state space model, arXiv preprint arXiv: 2401.09417 (2024).
    [32] Y. Yang, C. Ma, J. Yao, Z. Zhong, Y. Zhang, Y. Wang, Remamber: Referring image segmentation with mamba twister, in: European Conference on Computer Vision, Springer, 2024, pp. 108–126.
    [33] K. Zeng, H. Shi, J. Lin, S. Li, J. Cheng, K. Wang, Z. Li, K. Yang, Mambamos: Lidar-based 3d moving object segmentation with motion-aware state space model, in: Proceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 1505–1513.
    [34] J. Lin, J. Chen, K. Yang, A. Roitberg, S. Li, Z. Li, S. Li, Adaptiveclick: Click-aware transformer with adaptive focal loss for interactive image segmentation, IEEE Transactions on Neural Networks and Learning Systems 36 (3) (2025) 5759–5773. doi: 10.1109/TNNLS.2024.3378295.
    [35] S. Mahadevan, A. Athar, A. Osep, S. Hennen, L. Leal-Taixé, B. Leibe, Making a case for 3d convolutions for object segmentation in videos, BMVC abs/2008.11516 (2020).
    [36] H. Chen, W. Xie, T. Afouras, A. Nagrani, A. Vedaldi, A. Zisserman, Localizing visual sounds the hard way, in: 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 16862–16871. doi: 10.1109/CVPR46437.2021.01659.
    [37] B. Duke, A. Ahmed, C. Wolf, P. Aarabi, G. W. Taylor, Sstvos: Sparse spatiotemporal transformers for video object segmentation, 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021) 5908–5917.
    [38] S. Gao, Z. Chen, G. Chen, W. Wang, T. Lu. Avsegformer: Audio-visual segmentation with transformer. Proceedings of the AAAI Conference on Artificial Intelligence, 2024, 38(11): 12155−12163 doi: 10.1609/aaai.v38i11.29104
    [39] Z. Yang, Y. Wei, Y. Yang. Associating objects with transformers for video object segmentation. Advances in Neural Information Processing Systems, 2021, 34: 2491−2502
    [40] J. Zhang, J. Xie, N. Barnes, P. Li, Learning generative vision transformer with energy-based latent space for saliency prediction, in: M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, J. W. Vaughan (Eds.), Advances in Neural Information Processing Systems, Vol. 34, Curran Associates, Inc., 2021, pp. 15448–15463.
    [41] Y. Lv, Z. Liu, X. Chang. Consistency-queried transformer for audio-visual segmentation. IEEE Transactions on Image Processing, 2025, 34: 2616−2627 doi: 10.1109/TIP.2025.3563076
    [42] Y. Mao, J. Zhang, M. Xiang, Y. Lv, D. Li, Y. Zhong, Y. Dai. Contrastive conditional latent diffusion for audio-visual segmentation. IEEE Transactions on Image Processing, 2025, 34: 4108−4119 doi: 10.1109/TIP.2025.3580269
  • 加载中
计量
  • 文章访问数:  13
  • HTML全文浏览量:  3
  • 被引次数: 0
出版历程
  • 收稿日期:  2025-12-29
  • 录用日期:  2026-04-12
  • 网络出版日期:  2026-07-28

目录

    /

    返回文章
    返回