• 中文核心
  • EI
  • 中国科技核心
  • Scopus
  • CSCD
  • 英国科学文摘

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

图像数据集知识效能评估方法

丁昕苗 李好男 刘雨帆 刘彤飞 张克卿 原春锋 李兵 王坚 胡卫明

丁昕苗, 李好男, 刘雨帆, 刘彤飞, 张克卿, 原春锋, 李兵, 王坚, 胡卫明. 图像数据集知识效能评估方法. 自动化学报, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250457
引用本文: 丁昕苗, 李好男, 刘雨帆, 刘彤飞, 张克卿, 原春锋, 李兵, 王坚, 胡卫明. 图像数据集知识效能评估方法. 自动化学报, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250457
Ding Xin-Miao, Li Hao-Nan, Liu Yu-Fan, Liu Tong-Fei, Zhang Ke-Qing, Yuan Chun-Feng, Li Bing, Wang Jian, Hu Wei-Ming. Knowledge efficiency evaluation method for image datasets. Acta Automatica Sinica, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250457
Citation: Ding Xin-Miao, Li Hao-Nan, Liu Yu-Fan, Liu Tong-Fei, Zhang Ke-Qing, Yuan Chun-Feng, Li Bing, Wang Jian, Hu Wei-Ming. Knowledge efficiency evaluation method for image datasets. Acta Automatica Sinica, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250457

图像数据集知识效能评估方法

doi: 10.16383/j.aas.c250457 cstr: 32138.14.j.aas.c250457
基金项目: 北京市自然科学基金(JQ24022), 国家自然科学基金(61876100, 62072286, 62372451, 62192785, 62372082, 62192782, 62532015), 中国人工智能学会−蚂蚁集团联合研究基金(CAAI-MYJJ2024-02), 中国科协青年人才托举工程(2024QNRC001)资助
详细信息
    作者简介:

    丁昕苗:山东工商学院信息与电子工程学院教授. 2013年获得中国矿业大学(北京)博士学位. 主要研究方向为图像与视频理解, 机器学习和网络安全. E-mail: dingxinmiao@126.com

    李好男:山东工商学院信息与电子工程学院硕士研究生. 主要研究方向为机器学习, 模式识别. E-mail: 2023420059@sdtbu.edu.cn

    刘雨帆:中国科学院自动化研究所副研究员. 2023年获得中国科学院大学博士学位. 主要研究方向为计算机视觉, 模型压缩和视频理解. 本文通信作者. E-mail: liuyufan@ia.ac.cn

    刘彤飞:中国科学院自动化研究所博士研究生. 主要研究方向为数据集压缩. E-mail: liutongfei@ia.ac.cn

    张克卿:中国科学院自动化研究所博士研究生. 主要研究方向为大语言模型安全, 知识蒸馏. E-mail: zhangkeqing2025@ia.ac.cn

    原春锋:中国科学院自动化研究所研究员. 2010年获得中国科学院自动化研究所博士学位. 主要研究方向为模式识别, 计算机视觉. E-mail: yuanchunfeng@ia.ac.cn

    李兵:中国科学院自动化研究所研究员. 2009年获得北京交通大学博士学位. 主要研究方向为视频理解, 颜色恒常性, 视觉显著性, 多实例学习和网络内容安全. E-mail: libing@ia.ac.cn

    王坚:中国科学院自动化研究所工程师. 主要研究方向为机器学习, 计算机视觉. E-mail: wangjian@ia.ac.cn

    胡卫明:中国科学院自动化研究所研究员. 1998年获得浙江大学博士学位. 主要研究方向为视觉运动分析, 网络不良信息识别和网络入侵检测. E-mail: huweiming@ia.ac.cn

Knowledge Efficiency Evaluation Method for Image Datasets

Funds: Supported by Beijing Natural Science Foundation (JQ24022), National Natural Science Foundation of China (61876100, 62072286, 62372451, 62192785, 62372082, 62192782, 62532015), CAAI-Ant Group Research Fund (CAAI-MYJJ2024-02), and Young Elite Scientists Sponsorship Program by CAST (2024QNRC001)
More Information
    Author Bio:

    DING Xin-Miao Professor at the School of Information and Electronic Engineering, Shandong Technology and Business University. She received her Ph.D. degree from China University of Mining and Technology, Beijing in 2013. Her research interests include image and video understanding, machine learning, and internet security

    LI Hao-Nan Master student at the School of Information and Electronic Engineering, Shandong Technology and Business University. His research interests include machine learning and pattern recognition

    LIU Yu-Fan Associate researcher at the Institute of Automation, Chinese Academy of Sciences. She received her Ph.D. degree from University of Chinese Academy of Sciences in 2023. Her research interests include computer vision, model compression, and video understanding. Corresponding author of this paper

    LIU Tong-Fei Ph.D. candidate at the Institute of Automation, Chinese Academy of Sciences. His main research interest is dataset compression

    ZHANG Ke-Qing Ph.D. candidate at the Institute of Automation, Chinese Academy of Sciences. His research interests include large language model safety and knowledge distillation

    YUAN Chun-Feng Researcher at the Institute of Automation, Chinese Academy of Sciences. She received her Ph.D. degree from the Institute of Automation, Chinese Academy of Sciences in 2010. Her research interests include pattern recognition and computer vision

    LI Bing Researcher at the Institute of Automation, Chinese Academy of Sciences. He received his Ph.D. degree from Beijing Jiaotong University in 2009. His research interests include video understanding, color constancy, visual saliency, multi-instance learning, and web content security

    WANG Jian Engineer at the Institute of Automation, Chinese Academy of Sciences. His research interests include machine learning and computer vision

    HU Wei-Ming Researcher at the Institute of Automation, Chinese Academy of Sciences. He received his Ph.D. degree from Zhejiang University in 1998. His research interests include visual motion analysis, recognition of web objectionable information, and network intrusion detection

  • 摘要: 当前数据交易过程面临价值评估难题, AI模型训练缺乏数据选择标准. 其原因在于不同数据集在信息密度、质量结构与任务适配性方面呈现出显著差异, 而现有方法多聚焦于数据质量或模型性能的单一维度, 缺乏对数据集信息密度与多维价值特性的系统性量化框架, 限制了其在任务导向场景下数据集的科学选择与高效利用. 为此, 提出数据集知识效能评估框架, 即一种综合衡量数据集有效知识含量及其在特定任务中适应能力的量化方法, 并基于知识提炼、知识质量分析与知识任务适配构建三维评估框架: 1) 知识提炼: 将无损数据集蒸馏算法引入效能评估, 通过知识密度比量化数据集有效知识承载能力; 2) 知识质量分析: 构建多维质量指标体系, 深入分析和评估数据的知识质量; 3) 知识任务适配: 创新性提出任务驱动的评估适配机制, 通过契合度分析与权重优化实现对特定任务适应性的评估. 在CIFAR-10/100、VGGFace2、ImageNet-1K和NABirds等典型数据集上的实验结果表明, 所提方法能够较为准确地评估知识密度与任务适配性, 为数据集选择与应用提供理论依据和技术支撑.
  • 图  1  知识效能评估框架

    Fig.  1  Knowledge efficiency evaluation framework

    图  2  蒸馏前后数据集每类IPC变化

    Fig.  2  Per-class IPC changes of datasets before and after distillation

    图  3  语义层面标注错误

    Fig.  3  Semantic-level annotation errors

    图  4  数据集类内多样性统计

    Fig.  4  Intra-class diversity statistics of the datasets

    图  5  数据集灰度均值与标准差分布

    Fig.  5  Distribution of grayscale mean and standard deviation of the datasets

    图  6  数据集信息熵分布

    Fig.  6  Information entropy distribution of the datasets

    表  1  不同数据评估范式的对比与本文方法定位

    Table  1  Comparison of different data evaluation paradigms and positioning of our method

    评估范式 评估对象 核心问题与边界 知识承载 多维质量 任务适配
    数据质量评估 数据本身 关注数据是否规范及单一维度是否达标, 通常难以刻画数据集整体知识结构
    数据价值评估 经济层 关注数据的经济或业务收益, 但强依赖具体场景与收益假设, 可复用性弱
    样本贡献度 单样本 刻画单样本对训练目标的边际贡献, 计算代价高且难形成数据集层整体结论
    迁移适配预测 表征层 评估预训练表征的可迁移性, 主要反映模型与任务匹配, 难解释数据本体与质量结构
    DKEE (本文) 数据集层 统一刻画数据集中可被提炼的有效知识含量, 并在不同任务场景下进行多维质量分析,
    从而支持训练前的数据选择与治理决策
    下载: 导出CSV

    表  2  蒸馏前后数据集整体规模知识密度对比

    Table  2  Overall-scale knowledge density comparison of datasets before and after distillation

    数据集 类别数$ K $ 原始样本 蒸馏后样本 $ g(D,\; T) $
    MNIST[60] 10 60000 250 0.004
    CIFAR-10[3] 10 50000 8750 0.175
    CIFAR-100[3] 100 50000 7500 0.150
    TinyImageNet[61] 200 100000 8000 0.080
    ImageNet[4] 1000 1300000 1100000 0.846
    NABirds[65] 555 44400 33300 0.750
    Places365[62] 365 1788500 1168000 0.653
    VGGFace2[63] 9131 3287160 2739300 0.833
    下载: 导出CSV

    表  3  数据集完整性与标签一致性检测结果

    Table  3  Dataset completeness and label consistency detection results

    数据集总图像数重复图像数有效率(%)正确标签数标签一致性(%)
    MNIST700000100.0070000100.00
    CIFAR-10600000100.0060000100.00
    CIFAR-100600002699.9660000100.00
    TinyImageNet1300000100.00130000100.00
    Places36522049600100.002204960100.00
    VGGFace233112860100.003311286100.00
    ImageNet14311670100.001431167100.00
    NABirds485620100.0048562100.00
    下载: 导出CSV

    表  4  数据集噪声检测结果

    Table  4  Dataset noise detection results

    数据集 总样本数 正常样本数 噪声样本数 噪声比例(%)
    MNIST 60000 42238 17762 29.60
    CIFAR 10 50000 33078 16922 33.84
    CIFAR 100 50000 33269 16731 33.46
    TinyImageNet 100000 89462 10538 10.54
    Places365 1803460 1801968 1492 0.08
    VGGFace2 3141890 3140174 1716 0.05
    ImageNet 1281167 1279738 1429 0.12
    NABirds 23912 6021 17891 74.82
    下载: 导出CSV

    表  5  数据集类间多样性测量结果

    Table  5  Dataset inter-class diversity measurement results

    数据集 类间多样性均值
    MNIST 0.004162
    CIFAR-10 0.002321
    CIFAR-100 0.009125
    TinyImageNet 0.004166
    Places365 0.011322
    ImageNet 0.012894
    VGGFace2 0.029938
    NABirds 0.009781
    下载: 导出CSV

    表  6  经典模型在不同数据集上的测试准确率(%)

    Table  6  Test accuracy of classic models on different datasets (%)

    模型 MNIST CIFAR-10 CIFAR-100 TinyImageNet Places365 VGGFace2 ImageNet NABirds
    VGG16[55] 99.20 93.25 73.50 56.80 55.00 89.65 71.50 74.90
    ResNet18[54] 99.50 93.02 77.10 58.70 54.70 91.20 69.80 76.21
    ResNet50[54] 99.60 93.62 79.20 61.50 55.20 92.75 76.15 79.55
    DenseNet121[56] 99.70 95.04 79.30 64.20 56.20 93.10 74.90 80.30
    MobileNetV2[58] 99.40 94.50 75.30 59.10 54.00 90.50 71.80
    EfficientNet-B0[57] 99.70 96.00 80.10 66.00 57.00 94.00 77.10 63.70
    PyramidNet[59] 99.75 97.25 83.54 68.00 58.70 95.50 82.00
    平均准确率 99.55 94.67 78.29 62.04 55.83 92.39 74.75 74.93
    —表示该数据集未公开基准结果
    下载: 导出CSV

    表  7  歧义性、领域偏移与泄露比测量结果(%)

    Table  7  Measurement results of ambiguity domain shift and leakage ratio (%)

    数据集 歧义性 领域偏移 泄露比
    MNIST 99.88 0.120 4.99
    CIFAR-10 97.75 0.170 5.15
    CIFAR-100 99.04 0.210 4.49
    TinyImageNet 95.36 0.010 8.74
    Places365 99.69 0.040 73.35
    ImageNet 99.64 0.020 62.82
    VGGFace2 99.85 0.036 78.43
    NABirds 99.67 0.370 2.39
    下载: 导出CSV

    表  8  预训练模型在不同数据集上的迁移准确率(%)

    Table  8  Transfer accuracy of pre-trained models on different datasets (%)

    源模型 目标数据集
    MNIST CIFAR-10 CIFAR-100 TinyImageNet Places365 ImageNet VGGFace2 NABirds LFW[64] CUB[66]
    MNIST 83.68 7.42 10.98 0 0 0 0 0 0.81 0.78
    CIFAR-10 25.19 91.90 72.81 25.82 61.65 79.41 52.47 65.00 1.59 0.67
    CIFAR-100 5.30 8.35 76.92 10.17 31.55 56.00 23.96 42.66 2.66 0.86
    TinyImageNet 1.39 3.91 20.52 44.14 26.26 62.34 20.83 36.87 4.01 1.85
    Places365 0.45 0.68 2.51 8.81 45.60 39.30 8.78 26.19 5.18 2.59
    ImageNet 0.13 0.29 1.50 7.85 27.61 69.15 5.56 38.71 8.80 28.39
    VGGFace2 0.22 0.26 0.33 1.28 9.64 23.59 98.00 17.95 98.46 1.69
    NABirds 0.24 0.33 0.42 1.35 2.60 26.85 1.59 73.63 18.29 80.84
    下载: 导出CSV

    表  9  不同任务类型下的数据集适配效能

    Table  9  Dataset adaptation efficiency under different task types

    数据集 通用任务$ V_{\text{task}}(D,\; T) $ 特定任务$ V_{\text{task}}(D,\; T) $
    MNIST 0.420 0.420
    CIFAR-10 0.439 0.439
    CIFAR-100 0.456 0.456
    TinyImageNet 0.519 0.519
    Places365 0.537 0.537
    ImageNet 0.546 0.546
    VGGFace2 0.420 0.477
    NABirds 0.510 0.536
    下载: 导出CSV

    A1  NABirds数据集中困难样本剔除前后的准确率对比

    A1  Accuracy comparison before and after removing hard samples from NABirds dataset

    实验设置 分类准确率(%)
    原始数据集 56.0
    剔除20% 困难样本 44.0
    下载: 导出CSV

    B1  基于ResNet18的跨数据集迁移分类准确率(%)

    B1  Cross-dataset transfer classification accuracy based on ResNet18 (%)

    源模型 目标数据集
    MNIST CIFAR-10 CIFAR-100 TinyImageNet Places365 ImageNet VGGFace2 NABirds
    MNIST 99.36 17.37 3.31 1.10 0.27 0.10 0.29 0.26
    CIFAR-10 44.11 81.13 14.64 4.57 0.73 0.29 0.32 0.48
    CIFAR-100 62.41 67.56 79.26 17.80 2.34 1.29 0.30 0.42
    TinyImageNet 31.38 55.80 33.59 99.36 17.97 16.04 2.38 1.52
    Places365 91.60 68.84 39.47 41.81 52.17 32.92 16.50 3.69
    ImageNet 94.64 77.74 52.80 58.05 38.80 63.61 25.24 22.14
    VGGFace2 94.59 78.20 52.56 57.83 38.91 63.51 25.65 0.23
    NABirds 54.28 27.35 8.23 3.50 4.47 1.81 1.23 22.41
    下载: 导出CSV

    C1  COCO数据集在通用分类任务下的多维评估结果

    C1  Multi-dimensional evaluation results of COCO dataset for general classification task

    指标 结果
    知识密度 95%
    完整度 100%
    标签一致性 100%
    噪声比 93.08%
    类内多样性 0.034485
    类间多样性 0.01
    均值 0.0018
    方差 0.0022
    0.7269
    基线模型性能 36.28%
    歧义性 0.9984
    领域偏移 0.059
    泄露比 11.03%
    知识效能 4.47
    下载: 导出CSV

    D1  单次评估流程的耗时统计(h)

    D1  Time cost statistics of single evaluation process (h)

    数据集 知识提炼 质量分析
    MNIST 1 0.5
    CIFAR-10 2 0.5
    TinyImageNet 8 1
    Places365 15 15
    VGGFace2 30 20
    ImageNet 20 20
    NABirds 5 1
    下载: 导出CSV
  • [1] 国家发展改革委, 国家数据局, 教育部, 财政部, 金融监管总局, 中国证监会. 关于促进数据产业高质量发展的指导意见 [Online], available: https://www.gov.cn/zhengce/zhengceku/202412/content_6995430.htm, 2025-09-01

    National Development and Reform Commission, National Data Administration, Ministry of Education, Ministry of Finance, State Administration for Financial Regulation, China Securities Regulatory Commission. Guiding opinions on promoting the high-quality development of the data industry [Online], available: https://www.gov.cn/zhengce/zhengceku/202412/content_6995430.htm, September 1, 2025
    [2] 国家发展改革委, 国家数据局, 中央网信办, 工业和信息化部, 公安部, 市场监管总局. 关于完善数据流通安全治理更好促进数据要素市场化价值化的实施方案 [Online], available: https://www.ndrc.gov.cn/xxgk/zcfb/tz/202501/t20250115_1395692.html, 2025-09-01

    National Development and Reform Commission, National Data Administration, Cyberspace Administration of China, Ministry of Industry and Information Technology, Ministry of Public Security, State Administration for Market Regulation. Implementation plan for improving data circulation security governance and better promoting the marketization and valorization of data elements [Online], available: https://www.ndrc.gov.cn/xxgk/zcfb/tz/202501/t20250115_1395692.html, September 1, 2025
    [3] Krizhevsky A, Hinton G. Learning Multiple Layers of Features From Tiny Images, Technical Report, Department of Computer Science, University of Toronto, Canada, 2009.
    [4] Deng J, Dong W, Socher R, Li L J, Li K, Li F F. ImageNet: A large-scale hierarchical image database. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Miami, USA: IEEE, 2009. 248−255
    [5] Hinton G, Vinyals O, Dean J. Distilling the knowledge in a neural network. arXiv preprint arXiv: 1503.02531, 2015.
    [6] Romero A, Ballas N, Kahou S E, Chassang A, Gatta C, Bengio Y. FitNets: Hints for thin deep nets. arXiv preprint arXiv: 1412.6550, 2014.
    [7] Zagoruyko S, Komodakis N. Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer. In: Proceedings of the 5th International Conference on Learning Representations. Toulon, France: OpenReview.net, 2017.
    [8] Tung F, Mori G. Similarity-preserving knowledge distillation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Seoul, South Korea: IEEE, 2019. 1365−1374
    [9] Liu Y F, Cao J J, Li B, Yuan C F, Hu W M, Li Y X. Knowledge distillation via instance relationship graph. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Long Beach, USA: IEEE, 2019. 7089−7097
    [10] Tian Y L, Krishnan D, Isola P. Contrastive representation distillation. arXiv preprint arXiv: 1910.10699, 2019.
    [11] Heo B, Kim J, Yun S, Park H, Kwak N, Choi J Y. A comprehensive overhaul of feature distillation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Seoul, South Korea: IEEE, 2019. 1921−1930
    [12] Liu Y F, Cao J J, Li B, Hu W M, Ding J T, Li L, et al. Cross-architecture knowledge distillation. International Journal of Computer Vision, 2024, 132(8): 2798−2824 doi: 10.1007/s11263-024-02002-0
    [13] Yang Z Y, Liu J G, Huang J, He X D, Mei T, Xu C L. Cross-modal contrastive distillation for instructional activity anticipation. In: Proceedings of the 26th International Conference on Pattern Recognition (ICPR). Montreal, Canada: IEEE, 2022. 5002−5009
    [14] Yu J H, Yang L J, Xu N, Yang J C, Huang T. Slimmable neural networks. In: Proceedings of the International Conference on Learning Representations. New Orleans, USA: OpenReview.net, 2019.
    [15] Liu Y F, Cao J J, Li B, Hu W M, Maybank S. Learning to explore distillability and sparsability: A joint framework for model compression. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(3): 3378−3395 doi: 10.1109/tpami.2022.3185317
    [16] Liu Y F, Cao J J, Bai W M, Li B, Hu W M. Learning from the raw domain: Cross modality distillation for compressed video action recognition. In: Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Rhodes Island, Greece: IEEE, 2023. 1−5
    [17] Fang Z Y, Wang J F, Hu X W, Wang L J, Yang Y Z, Liu Z C. Compressing visual-linguistic model via knowledge distillation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Montreal, Canada: IEEE, 2021. 1408−1418
    [18] Krizhevsky A, Sutskever I, Hinton G E. ImageNet classification with deep convolutional neural networks. In: Proceedings of the 26th International Conference on Neural Information Processing Systems. Lake Tahoe, Nevada: Curran Associates Inc., 2012. 1097−1105
    [19] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, et al. Attention is all you need. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. Long Beach, USA: Curran Associates Inc., 2017. 6000−6010
    [20] Wang T Z, Zhu J Y, Torralba A, Efros A A. Dataset distillation. In: Proceedings of the International Conference on Learning Representations. New Orleans, USA: OpenReview.net, 2019.
    [21] Zhao B, Mopuri K R, Bilen H. Dataset condensation with gradient matching. arXiv preprint arXiv: 2006.05929, 2020.
    [22] Nguyen T, Chen Z R, Lee J. Dataset meta-learning from kernel ridge-regression. In: Proceedings of the International Conference on Learning Representations. Vienna, Austria: OpenReview.net, 2021.
    [23] Nguyen T, Novak R, Xiao L C, Lee J. Dataset distillation with infinitely wide convolutional networks. In: Proceedings of the 35th International Conference on Neural Information Processing Systems. Virtual Event: Curran Associates Inc., 2021. Article No. 397
    [24] Zhao B, Bilen H. Dataset condensation with distribution matching. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). Waikoloa, USA: IEEE, 2023. 6503−6512
    [25] Zhao G L, Li G B, Qin Y P, Yu Y Z. Improved distribution matching for dataset condensation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE, 2023. 7856−7865
    [26] Cazenavette G, Wang T Z, Torralba A, Efros A A, Zhu J Y. Dataset distillation by matching training trajectories. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans, USA: IEEE, 2022. 10708−10717
    [27] Yin Z Y, Xing E, Shen Z Q. Squeeze, recover and relabel: Dataset condensation at ImageNet scale from a new perspective. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. New Orleans, USA: Curran Associates Inc., 2023. Article No. 3219
    [28] Zhang X, Du J W, Liu P, Zhou J T. Breaking class barriers: Efficient dataset distillation via inter-class feature compensator. In: Proceedings of the 13th International Conference on Learning Representations. Singapore: OpenReview.net, 2025.
    [29] Chen M Y, Du J W, Huang B, Wang Y, Zhang X B, Wang W. Influence-guided diffusion for dataset distillation. In: Proceedings of the 13th International Conference on Learning Representations. Singapore: OpenReview.net, 2025.
    [30] Qi D, Li J, Gao J Y, Dou S G, Tai Y, Hu J L. Towards universal dataset distillation via task-driven diffusion. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, USA: IEEE, 2025. 10557−10566
    [31] Guo Z Y, Wang K, Cazenavette G, Li H, Zhang K P, You Y. Towards lossless dataset distillation via difficulty-aligned trajectory matching. In: Proceedings of the 12th International Conference on Learning Representations. Vienna, Austria: OpenReview.net, 2024.
    [32] Shao S T, Zhou Z K, Chen H R, Shen Z Q. Elucidating the design space of dataset condensation. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. Vancouver, Canada: Curran Associates Inc., 2024. Article No. 3146
    [33] Schioppa A. Efficient sketches for training data attribution and studying the loss landscape. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. Vancouver, Canada: Curran Associates Inc., 2024. Article No. 1190
    [34] Liu T F, Liu Y F, Li B, Hu W M, Li Y M, Ma C G. Noise-optimized distribution distillation for dataset condensation. In: Proceedings of the 33rd ACM International Conference on Multimedia. Dublin, Ireland: ACM, 2025. 10352−10360
    [35] Liu P, Du J W. The evolution of dataset distillation: Toward scalable and generalizable solutions. arXiv preprint arXiv: 2502.05673, 2025.
    [36] Northcutt C G, Athalye A, Mueller J. Pervasive label errors in test sets destabilize machine learning benchmarks. In: Proceedings of the 35th Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1). Virtual Event: OpenReview.net, 2021.
    [37] Kang B Y, Xie S N, Rohrbach M, Yan Z C, Gordo A, Feng J S, et al. Decoupling representation and classifier for long-tailed recognition. In: Proceedings of the 8th International Conference on Learning Representations. Addis Ababa, Ethiopia: OpenReview.net, 2020.
    [38] Uddin S, Lu H H. Dataset meta-level and statistical features affect machine learning performance. Scientific Reports, 2024, 14(1): Article No. 1670 doi: 10.1038/s41598-024-51825-x
    [39] Ho T K. A data complexity analysis of comparative advantages of decision forest constructors. Pattern Analysis & Applications, 2002, 5(2): 102−112 doi: 10.1007/s100440200009
    [40] Nguyen C, Hassner T, Seeger M, Archambeau C. LEEP: A new measure to evaluate transferability of learned representations. In: Proceedings of the 37th International Conference on Machine Learning. Virtual Event: PMLR, 2020. 7294−7305
    [41] You K C, Liu Y, Wang J M, Long M S. LogME: Practical assessment of pre-trained models for transfer learning. In: Proceedings of the 38th International Conference on Machine Learning. Virtual Event: PMLR, 2021. 12133−12143
    [42] Swayamdipta S, Schwartz R, Lourie N, Wang Y Z, Hajishirzi H, Smith N A, et al. Dataset cartography: Mapping and diagnosing datasets with training dynamics. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Virtual Event: ACL, 2020. 9275−9293
    [43] Dai M Y, Raffiee A H, Jain A, Correa J. Evaluating transferability in retrieval tasks: An approach using MMD and kernel methods. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, USA: IEEE, 2024. 22390−22400
    [44] 李娟娟, 秦蕊, 丁文文, 王戈, 王坛, 王飞跃. 基于web3的去中心化自治组织与运营新框架. 自动化学报, 2023, 49(5): 985−998 doi: 10.16383/j.aas.c220753

    Li Juan-Juan, Qin Rui, Ding Wen-Wen, Wang Ge, Wang Tan, Wang Fei-Yue. A new framework for web3-powered decentralized autonomous organizations and operations. Acta Automatica Sinica, 2023, 49(5): 985−998 doi: 10.16383/j.aas.c220753
    [45] 张辉, 颜星雨, 毛建旭, 别克扎提•巴合提, 杜瑞, 王耀南. 面向源网荷的智能化数据协同推断技术研究综述. 自动化学报, 2025, 51(11): 2387−2411 doi: 10.16383/j.aas.c250203

    Zhang Hui, Yan Xing-Yu, Mao Jian-Xu, Biekezhati Baheti, Du Rui, Wang Yao-Nan. A review of intelligent data collaborative inference techniques for source-grid-load systems. Acta Automatica Sinica, 2025, 51(11): 2387−2411 doi: 10.16383/j.aas.c250203
    [46] 欧阳丽炜, 王帅, 袁勇, 倪晓春, 王飞跃. 智能合约: 架构及进展. 自动化学报, 2019, 45(3): 445−457 doi: 10.16383/j.aas.c180586

    Ouyang Li-Wei, Wang Shuai, Yuan Yong, Ni Xiao-Chun, Wang Fei-Yue. Smart contracts: Architecture and research progresses. Acta Automatica Sinica, 2019, 45(3): 445−457 doi: 10.16383/j.aas.c180586
    [47] Hotelling H. Analysis of a complex of statistical variables into principal components. Journal of Educational Psychology, 1993, 24(6): 417−441 doi: 10.1037/h0071325
    [48] van der Maaten L, Hinton G. Visualizing data using t-SNE. Journal of Machine Learning Research, 2008, 9: 2579−2605
    [49] Ester M, Kriegel H P, Sander J, Xu X W. A density-based algorithm for discovering clusters in large spatial databases with noise. In: Proceedings of the 2nd International Conference on Knowledge Discovery and Data Mining. Portland, USA: AAAI Press, 1996. 226−231
    [50] Lin J H. Divergence measures based on the Shannon entropy. IEEE Transactions on Information Theory, 1991, 37(1): 145−151 doi: 10.1109/18.61115
    [51] Radford A, Kim J W, Hallacy C, Ramesh A, Goh G, Agarwal S, et al. Learning transferable visual models from natural language supervision. In: Proceedings of the 38th International Conference on Machine Learning. Virtual Event: PMLR, 2021. 8748−8763
    [52] Kingma D P, Welling M. Auto-encoding variational Bayes. In: Proceedings of the International Conference on Learning Representations. Banff, Canada: OpenReview.net, 2014.
    [53] Kullback S, Leibler R A. On information and sufficiency. The Annals of Mathematical Statistics, 1951, 22(1): 79−86
    [54] He K M, Zhang X Y, Ren S Q, Sun J. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas, USA: IEEE, 2016. 770−778
    [55] Simonyan S, Zisserman A. Very deep convolutional networks for large-scale image recognition. In: Proceedings of the 3rd International Conference on Learning Representations. San Diego, USA: OpenReview.net, 2015.
    [56] Huang G, Liu Z, van der Maaten L, Weinberger K Q. Densely connected convolutional networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu, USA: IEEE, 2017. 2261−2269
    [57] Tan M X, Le Q. EfficientNet: Rethinking model scaling for convolutional neural networks. In: Proceedings of the 36th International Conference on Machine Learning. Long Beach, USA: PMLR, 2019. 6105−6114
    [58] Sandler M, Howard A, Zhu M L, Zhmoginov A, Chen L C. MobileNetV2: Inverted residuals and linear bottlenecks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City, USA: IEEE, 2018. 4510−4520
    [59] Han D, Kim J, Kim J. Deep pyramidal residual networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu, USA: IEEE, 2017. 6307−6315
    [60] Lecun Y, Bottou L, Bengio Y, Haffner P. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 1998, 86(11): 2278−2324 doi: 10.1109/5.726791
    [61] Le Y, Yang X. Tiny ImageNet Visual Recognition Challenge, Technical Report, Stanford University, USA, 2015.
    [62] López-Cifuentes A, Escudero-Viñolo M, Bescós J, García-Martín Á. Semantic-aware scene recognition. Pattern Recognition, 2020, 102: Article No. 107256 doi: 10.1016/j.patcog.2020.107256
    [63] Cao Q, Shen L, Xie W D, Parkhi O M, Zisserman A. VGGFace2: A dataset for recognising faces across pose and age. In: Proceedings of the 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2018). Xi'an, China: IEEE, 2018. 67−74
    [64] Huang G B, Ramesh M, Berg T, Learned-Miller E. Labeled faces in the wild: A database for studying face recognition in unconstrained environments. In: Proceedings of the Workshop on Faces in ‘Real-Life’ Images: Detection, Alignment, and Recognition. Marseille, France: ECCV 2008 Workshop, 2008. 1−11
    [65] van Horn G, Branson S, Farrell R, Haber S, Barry J, Ipeirotis P. Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Boston, USA: IEEE, 2015. 595−604
    [66] Wah C, Branson S, Welinder P, Perona P, Belongie S. The Caltech-UCSD Birds-200-2011 Dataset, Technical Report CNS-TR-2011-001, California Institute of Technology, USA, 2011.
  • 加载中
计量
  • 文章访问数:  76
  • HTML全文浏览量:  85
  • 被引次数: 0
出版历程
  • 收稿日期:  2025-09-18
  • 录用日期:  2026-03-04
  • 网络出版日期:  2026-08-04

目录

    /

    返回文章
    返回