• 中文核心
  • EI
  • 中国科技核心
  • Scopus
  • CSCD
  • 英国科学文摘

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

针对低比特Transformer的无约束训练后量化

江佳浩 尹鹏 王轩瀚 曾鹏鹏 宋井宽

江佳浩, 尹鹏, 王轩瀚, 曾鹏鹏, 宋井宽. 针对低比特Transformer的无约束训练后量化. 自动化学报, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250702
引用本文: 江佳浩, 尹鹏, 王轩瀚, 曾鹏鹏, 宋井宽. 针对低比特Transformer的无约束训练后量化. 自动化学报, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250702
Jiang Jia-Hao, Yin Peng, Wang Xuan-Han, Zeng Peng-Peng, Song Jing-Kuan. Constraint-free post-training quantization for low-bit transformer. Acta Automatica Sinica, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250702
Citation: Jiang Jia-Hao, Yin Peng, Wang Xuan-Han, Zeng Peng-Peng, Song Jing-Kuan. Constraint-free post-training quantization for low-bit transformer. Acta Automatica Sinica, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250702

针对低比特Transformer的无约束训练后量化

doi: 10.16383/j.aas.c250702 cstr: 32138.14.j.aas.c250702
基金项目: 教育部学科突破先导项目(JYB2025XDXM103), 国家自然科学基金(62425208, U22A2097, 62402094, 62502076)资助
详细信息
    作者简介:

    江佳浩:同济大学电子与信息工程学院硕士研究生. 主要研究方向为具身智能和计算机视觉. E-mail: jjhao@tongji.edu.cn

    尹鹏:腾讯科技(深圳)有限公司算法工程师. 2025年获得电子科技大学硕士学位. 主要研究方向为深度学习和模型量化. E-mail: raymond2yp@gmail.com

    王轩瀚:同济大学计算机科学与技术学院研究员. 主要研究方向为以人为中心的视觉智能和多模态具身智能.本文通信作者. E-mail: wxuanhan@hotmail.com

    曾鹏鹏:同济大学计算机科学与技术学院研究员. 主要研究方向为多模态理解, 计算机视觉和具身智能. E-mail: is.pengpengzeng@gmail.com

    宋井宽:同济大学计算机科学与技术学院教授. 主要研究方向为多模态学习和具身智能. E-mail: jingkuan.song@gmail.com

Constraint-free Post-training Quantization for Low-bit Transformer

Funds: Supported by Fundamental and Interdisciplinary Disciplines Breakthrough Plan of the Ministry of Education of China (JYB2025XDXM103) and National Natural Science Foundation of China (62425208, U22A2097, 62402094, 62502076)
More Information
    Author Bio:

    JIANG Jia-Hao Master student at the College of Electronic and Information Engineering, Tongji University. His research interests include embodied intelligence and computer vision

    YIN Peng Algorithm engineer at Tencent Technology (Shenzhen) Co., Ltd. He received his master degree from University of Electronic Science and Technology of China in 2025. His research interests include deep learning and model quantization

    WANG Xuan-Han Researcher at the School of Computer Science and Technology, Tongji University. His research interests include human-centered visual intelligence and multimodal embodied intelligence. Corresponding author of this paper

    ZENG Peng-Peng Researcher at the School of Computer Science and Technology, Tongji University. His research interests include multimodal understanding, computer vision, and embodied intelligence

    SONG Jing-Kuan Professor at the School of Computer Science and Technology, Tongji University. His research interests include multimodal learning and embodied intelligence

  • 摘要: Transformer模型具有庞大的参数规模和计算开销, 难以直接部署于资源受限的边缘设备, 限制了其在开放环境中的实际应用.低比特量化能够显著降低模型存储需求并提升推理效率, 是实现Transformer模型轻量化部署的重要技术路径.然而, 现有量化方法通常受固定量化区间约束, 在处理非线性激活产生的长尾分布时, 难以有效平衡裁剪误差与舍入误差, 导致模型性能明显下降. 为此, 提出一种面向低比特Transformer模型的无约束训练后量化方法(CFQuant), 通过建模神经网络激活值的概率分布实现自适应量化. 首先, CFQuant在校准过程中自适应估计激活分布密度, 并通过迭代搜索最小化量化误差. 其次, 设计高效尺度偏移算法动态调整模型分布, 以减小校准阶段与推理阶段之间的分布偏移, 进一步提升量化稳定性. 此外, 构建基于查找表的矩阵乘法加速机制, 将低比特推理中的浮点乘法转换为预计算查找操作, 从而降低推理计算开销.最后, 通过在多类Transformer模型以及视觉、语言和多模态任务上的大量实验, 充分验证了所提方法的灵活性和有效性.
  • 图  1  不同分布的量化误差对比

    Fig.  1  Comparison of quantization errors for different distributions

    图  2  舍入误差和裁剪误差

    Fig.  2  Rounding error and clipping error

    图  3  保留不同激活值为全精度的性能比较

    Fig.  3  Performance comparison when different activations are retained at full precision

    图  4  无约束训练后量化方法整体架构

    Fig.  4  Overall architecture of CFQuant

    图  5  CFQuant对量化误差和模型性能的影响

    Fig.  5  Effect of CFQuant on quantization error and model performance

    图  6  不同校准样本数量下的Top-1准确率对比

    Fig.  6  Top-1 accuracy comparison with different calibration sample sizes

    表  1  不同方法在ImageNet-1k数据集上分类任务的性能对比(%)

    Table  1  Performance comparison of different methods on classification tasks on ImageNet-1k dataset (%)

    方法 比特数(W/A) ViT-T ViT-S ViT-B DeiT-T DeiT-S DeiT-B Swin-T Swin-S Swin-B
    全精度 32/32 75.47 81.39 84.54 72.21 79.85 81.80 81.37 83.23 85.27
    PTQ4ViT[19] 4/4 16.53 42.57 30.69 36.96 34.08 64.39 73.06 76.09 74.02
    APQ-ViT[33] 4/4 17.56 47.95 41.41 47.94 43.55 67.48 77.15 76.48
    RepQ-ViT[20] 4/4 26.72 65.05 68.48 57.43 69.03 75.61 72.37 79.45 80.74
    CFQuant 4/4 55.41 70.95 73.72 62.57 72.57 78.02 73.80 80.65 81.84
    下载: 导出CSV

    表  2  不同方法在MS-COCO数据集上目标检测和实例分割任务的性能对比(%)

    Table  2  Performance comparison of different methods on object detection and instance segmentation tasks on the MS-COCO dataset (%)

    方法 比特数(W/A) Mask R-CNN Cascade Mask R-CNN
    Swin-T Swin-S Swin-T Swin-S
    边界框AP 掩码AP 边界框AP 掩码AP 边界框AP 掩码AP 边界框AP 掩码AP
    全精度 32/32 46.0 41.6 48.5 43.3 50.4 43.7 51.9 45.0
    PTQ4ViT[19] 4/4 6.9 7.0 26.7 26.6 14.7 13.5 0.5 0.5
    APQ-ViT[33] 4/4 23.7 22.6 44.7 40.1 27.2 24.4 47.7 41.1
    RepQ-ViT[20] 4/4 36.1 36.0 44.2 40.2 47.0 41.4 49.3 43.1
    CFQuant 4/4 36.6 36.5 43.4 40.6 47.7 42.0 50.1 43.6
    下载: 导出CSV

    表  3  在多个零样本任务上LLaMA系列模型不同方法的性能对比(%)

    Table  3  Performance comparison of different methods on multiple zero-shot tasks across LLaMA models (%)

    模型 方法 比特数(W/A) PIQA ARC-e WinoGrande BoolQ ARC-c HellaSwag 平均值
    LLaMA-7B全精度16/1677.4752.4867.0773.0841.4673.0064.09
    SmoothQuant[15]4/449.8030.4048.0049.1025.8027.4038.41
    OmniQuant[36]4/466.1545.2053.4363.5131.1456.4452.65
    AffineQuant[60]4/469.3742.5555.3363.7331.9157.6553.42
    CFQuant4/467.9143.9055.3664.4532.1957.1453.49
    LLaMA-13B全精度16/1679.1059.8970.3168.0144.4576.2166.33
    SmoothQuant[15]4/461.0439.1851.0661.8030.8052.2949.36
    OmniQuant[36]4/469.6947.3955.8062.8433.1058.9654.63
    AffineQuant[60]4/466.3243.9054.7064.1029.6156.8852.58
    CFQuant4/471.2045.6156.3964.1733.3559.5155.04
    下载: 导出CSV

    表  4  在多个零样本任务上Qwen3系列模型不同方法的性能对比(%)

    Table  4  Performance comparison of different methods on multiple zero-shot tasks across Qwen3 models (%)

    模型 方法 比特数(W/A) PIQA ARC-e WinoGrande BoolQ ARC-c HellaSwag 平均值
    Qwen3-8B全精度16/1677.8080.9367.7286.6156.6674.9274.11
    SmoothQuant[15]4/873.4271.0162.1679.1144.3251.6863.62
    PrefixQuant[61]4/871.9870.5064.0982.2646.9365.0966.81
    CFQuant4/876.1771.1365.8279.7947.6171.5668.68
    SmoothQuant[15]4/450.5325.5152.1739.5825.7326.7236.71
    PrefixQuant[61]4/456.4238.5548.0745.6926.7934.8241.72
    CFQuant4/472.7471.8464.0978.8147.8767.9367.21
    Qwen3-14B全精度16/1679.8282.8772.9389.3360.2478.8277.34
    SmoothQuant[15]4/451.2225.8150.1838.4726.5125.7936.33
    PrefixQuant[61]4/457.8943.8151.3056.7329.1039.8346.44
    CFQuant4/465.0758.7156.9178.8440.1049.3958.17
    下载: 导出CSV

    表  5  在WikiText-2和C4数据集上语言生成任务的性能比较

    Table  5  Performance comparison on language generation tasks on WikiText-2 and C4 datasets

    模型 方法 比特数(W/A) WikiText-2 $ \downarrow $ C4 $ \downarrow $
    LLaMA-7B全精度16/165.687.08
    SmoothQuant[15]4/425.2532.32
    OmniQuant[36]4/411.2614.51
    CFQuant4/410.3613.91
    LLaMA-13B全精度16/165.096.61
    SmoothQuant[15]4/440.0547.18
    OmniQuant[36]4/410.8713.78
    CFQuant4/410.1413.28
    Qwen3-8B全精度16/169.7213.29
    SmoothQuant[15]4/4$ 3.36\times10^{4} $$ 2.29\times10^{4} $
    PrefixQuant[61]4/4155.78119.93
    CFQuant4/413.7817.26
    Qwen3-14B全精度16/168.6412.01
    SmoothQuant[15]4/4$ 2.16\times10^{5} $$ 1.99\times10^{5} $
    PrefixQuant[61]4/4187.28129.05
    CFQuant4/432.3035.76
    下载: 导出CSV

    表  6  在MMBench数据集上Qwen3-VL-4B模型的 不同方法性能对比(%)

    Table  6  Performance comparison of different methods for the Qwen3-VL-4B model on the MMBench dataset (%)

    方法比特数(W/A)MMBench-ENMMBench-CN
    全精度16/1683.1380.80
    AWQ[62]4/1681.5879.26
    SmoothQuant[15]4/875.3175.00
    PrefixQuant[61]4/877.0977.79
    CFQuant4/879.4178.72
    下载: 导出CSV

    表  7  在ImageNet-1k数据集上查找表额外内存开销的消融实验

    Table  7  Ablation study of the additional lookup-table memory overhead on the ImageNet-1k dataset

    模型 CFQuant 模型尺寸(MB) Top-1准确率(%)
    DeiT-S × 11.18 68.27
    11.19 (+0.1%) 72.57 (+4.30)
    Swin-S × 24.78 79.03
    24.80 (+0.2%) 80.65 (+1.62)
    下载: 导出CSV

    表  8  在ImageNet-1k数据集上不同激活函数 量化策略的消融实验

    Table  8  Ablation study of quantization strategies for different activation functions on the ImageNet-1k dataset

    模型 Softmax GeLU Top-1准确率(%)
    DeiT-S × × 68.27
    × 68.82 (+0.55)
    × 71.33 (+3.06)
    72.57 (+4.30)
    Swin-S × × 79.03
    × 79.28 (+0.25)
    × 80.36 (+1.33)
    80.65 (+1.62)
    下载: 导出CSV

    表  9  在ImageNet-1k数据集上高效尺度偏移算法的 消融实验

    Table  9  Ablation studies of efficient scale-shift algorithm on the ImageNet-1k dataset

    模型ESATop-1准确率(%)
    DeiT-S×71.08
    72.57 (+1.49)
    Swin-S×78.93
    80.65 (+1.72)
    下载: 导出CSV

    表  10  在ImageNet-1k数据集上不同量化方法校准时间的对比

    Table  10  Comparison of calibration time for different quantization methods on the ImageNet-1k dataset

    模型 方法 Top-1准确率(%) 耗时(min)
    DeiT-S 全精度 79.85
    PTQ4ViT[19] 34.08 3.61
    RepQ-ViT[20] 69.03 3.12
    CFQuant 72.57 3.28
    Swin-S 全精度 83.23
    PTQ4ViT[19] 76.09 8.46
    RepQ-ViT[20] 79.45 6.87
    CFQuant 80.65 7.19
    下载: 导出CSV

    表  11  在ImageNet-1k数据集上的推理效率对比

    Table  11  Comparison on inference efficiency on the ImageNet-1k dataset

    方法模型尺寸(MB)位运算量(G)延迟(s)
    全精度89.444 74110.39
    RepQ-ViT[20]11.181412.67
    CFQuant11.183464.33
    CFQuant $ _{+\text{MM-LUT}} $11.191442.80
    下载: 导出CSV
  • [1] Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X H, Unterthiner T, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In: Proceedings of the 9th International Conference on Learning Representations. Virtual Event: OpenReview.net, 2021.
    [2] Dai Y, Chen X J, Wang X H, Pang M H, Gao L L, Shen H T. ReSParser: Fully convolutional multiple human parsing with representative sets. IEEE Transactions on Multimedia, 2024, 26: 1384−1394 doi: 10.1109/TMM.2023.3281070
    [3] Zhang S S, Roller S, Goyal N, Artetxe M, Chen M Y, Chen S H, et al. OPT: Open pre-trained Tansformer language models. arXiv preprint arXiv: 2205.01068, 2022.
    [4] Li S S, Xu X, Meng W X, Song J K, Peng C, Shen H T. Mitigating hallucinations in large vision-language models via reasoning uncertainty-guided refinement. IEEE Transactions on Multimedia, 2025, 27: 7380−7391 doi: 10.1109/TMM.2025.3599076
    [5] Touvron H, Martin L, Stone K, Albert P, Almahairi A, Babaei Y, et al. LLaMA 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv: 2307.09288, 2023.
    [6] Zitkovich B, Yu T H, Xu S C, Xu P, Xiao T, Xia F, et al. RT-2: Vision-language-action models transfer web knowledge to robotic control. In: Proceedings of the Conference on Robot Learning. Atlanta, USA: PMLR, 2023. 2165−2183
    [7] 李浩然, 陈宇辉, 崔文博, 刘卫恒, 刘锴, 周明才, 等. 面向具身操作的视觉-语言-动作模型综述. 自动化学报, 2026, 52(1): 18−51

    Li Hao-Ran, Chen Yu-Hui, Cui Wen-Bo, Liu Wei-Heng, Liu Kai, Zhou Ming-Cai, et al. Survey of vision-language-action models for embodied manipulation. Acta Automatica Sinica, 2026, 52(1): 18−51
    [8] Qu D L, Song H M, Chen Q Z, Wang D, Yao Y Q, Ye X Y, et al. SpatialVLA: Exploring spatial representations for visual-language-action model. In: Proceedings of the Robotics: Science and Systems 2025. Los Angeles, USA: University of Southern California, 2025. Article No. 011
    [9] Liu Z, Lin Y T, Cao Y, Hu H, Wei Y X, Zhang Z, et al. Swin Tansformer: Hierarchical vision Tansformer using shifted windows. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Montreal, Canada: IEEE, 2021. 9992−10002
    [10] 王文晟, 谭宁, 黄凯, 张雨浓, 郑伟诗, 孙富春. 基于大模型的具身智能系统综述. 自动化学报, 2025, 51(1): 1−19

    Wang Wen-Sheng, Tan Ning, Huang Kai, Zhang Yu-Nong, Zheng Wei-Shi, Sun Fu-Chun. Embodied intelligence systems based on large models: A survey. Acta Automatica Sinica, 2025, 51(1): 1−19
    [11] Gu Y X, Dong L, Wei F R, Huang M L. MiniLLM: Knowledge distillation of large language models. In: Proceedings of the 12th International Conference on Learning Representations. Vienna, Austria: OpenReview.net, 2024.
    [12] Chen G B, Choi W, Yu X, Han T, Chandraker M. Learning efficient object detection models with knowledge distillation. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. Long Beach, USA: Curran Associates, Inc., 2017. 742−751
    [13] Ashkboos S, Croci M L, do Nascimento M G, Hoefler T, Hensman J. SliceGPT: Compress large language models by deleting rows and columns. In: Proceedings of the 12th International Conference on Learning Representations. Vienna, Austria: OpenReview.net, 2024.
    [14] Sun M J, Liu Z, Bair A, Kolter J Z. A simple and effective pruning approach for large language models. In: Proceedings of the 12th International Conference on Learning Representations. Vienna, Austria: OpenReview.net, 2024.
    [15] Xiao G X, Lin J, Seznec M, Wu H, Demouth J, Han S. SmoothQuant: Accurate and efficient post-training quantization for large language models. In: Proceedings of the 40th International Conference on Machine Learning. Honolulu, USA: PMLR, 2023. 38087−38099
    [16] Liu Y J, Yang H R, Dong Z, Keutzer K, Du L, Zhang S H. NoisyQuant: Noisy bias-enhanced post-training activation quantization for vision Tansformers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver, Canada: IEEE, 2023. 20321−20330
    [17] Yin P, Zhu X S, Song J K, Gao L L, Shen H T. SI-BiViT: Binarizing vision Tansformers with spatial interaction. In: Proceedings of the 32nd ACM International Conference on Multimedia. Melbourne, Australia: ACM, 2024. 8169−8178
    [18] Li Y H, Gong R H, Tan X, Yang Y, Hu P, Zhang Q, et al. BRECQ: Pushing the limit of post-training quantization by block reconstruction. In: Proceedings of the 9th International Conference on Learning Representations. Virtual Event: OpenReview.net, 2021.
    [19] Yuan Z H, Xue C H, Chen Y Q, Wu Q, Sun G Y. PTQ4ViT: Post-training quantization for vision Tansformers with twin uniform quantization. In: Proceedings of the 17th European Conference on Computer Vision. Tel Aviv, Israel: Springer, 2022. 191−207
    [20] Li Z K, Xiao J R, Yang L W, Gu Q Y. RepQ-ViT: Scale reparameterization for post-training quantization of vision Tansformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Paris, France: IEEE, 2023. 17181−17190
    [21] Jacob B, Kligys S, Chen B, Zhu M L, Tang M, Howard A, et al. Quantization and training of neural networks for efficient integer-arithmetic-only inference. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City, USA: IEEE, 2018. 2704−2713
    [22] Cai J Y, Takemoto M, Nakajo H. A deep look into logarithmic quantization of model parameters in neural networks. In: Proceedings of the 10th International Conference on Advances in Information Technology. Bangkok, Thailand: ACM, 2018. Article No. 6
    [23] Lloyd S. Least squares quantization in PCM. IEEE Transactions on Information Theory, 1982, 28(2): 129−137
    [24] Max J. Quantizing for minimum distortion. IRE Transactions on Information Theory, 1960, 6(1): 7−12
    [25] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, et al. Attention is all you need. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. Long Beach, USA: Curran Associates, Inc., 2017. 6000−6010
    [26] Devlin J, Chang M W, Lee K, Toutanova K. BERT: Pre-training of deep bidirectional Tansformers for language understanding. In: Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Minneapolis, USA: Association for Computational Linguistics, 2019. 4171−4186
    [27] Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I. Language models are unsupervised multitask learners [Online], available: https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf, July 3, 2026
    [28] Touvron H, Cord M, Douze M, Massa F, Sablayrolles A, Jégou H. Training data-efficient image Tansformers & distillation through attention. In: Proceedings of the 38th International Conference on Machine Learning. Virtual Event: PMLR, 2021. 10347−10357
    [29] Zhang Z Z, Zhang H, Zhao L, Chen T, Arik S Ö, Pfister T. Nested hierarchical Tansformer: Towards accurate, data-efficient and interpretable visual understanding. In: Proceedings of the 36th AAAI Conference on Artificial Intelligence. Virtual Event: AAAI Press, 2022. 3417−3425
    [30] Strudel R, Garcia R, Laptev I, Schmid C. Segmenter: Transformer for semantic segmentation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Montreal, Canada: IEEE, 2021. 7242−7252
    [31] Brown T B, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, et al. Language models are few-shot learners. In: Proceedings of the 34th International Conference on Neural Information Processing Systems. Vancouver, Canada: Curran Associates, Inc., 2020. Article No. 159
    [32] Sun Y X, Liu R K, Bai H L, Bao H, Zhao K, Li Y N, et al. FlatQuant: Flatness matters for LLM quantization. In: Proceedings of the 42nd International Conference on Machine Learning. Vancouver, Canada: PMLR, 2025. 57587−57613
    [33] Ding Y F, Qin H T, Yan Q H, Chai Z H, Liu J J, Wei X L, et al. Towards accurate post-training quantization for vision Tansformer. In: Proceedings of the 30th ACM International Conference on Multimedia. Lisbon, Portugal: ACM, 2022. 5380−5388
    [34] Jeon Y, Lee C, Cho E, Ro Y. Mr.BiQ: Post-training non-uniform quantization based on minimizing the reconstruction error. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans, USA: IEEE, 2022. 12319−12328
    [35] Frantar E, Ashkboos S, Hoefler T, Alistarh D. OPTQ: Accurate quantization for generative pre-trained Tansformers. In: Proceedings of the 11th International Conference on Learning Representations. Kigali, Rwanda: OpenReview.net, 2023.
    [36] Shao W Q, Chen M Z, Zhang Z Y, Xu P, Zhao L R, Li Z Q, et al. OmniQuant: Omnidirectionally calibrated quantization for large language models. In: Proceedings of the 12th International Conference on Learning Representations. Vienna, Austria: OpenReview.net, 2024.
    [37] Liu Z C, Zhao C S, Fedorov I, Soran B, Choudhary D, Krishnamoorthi R, et al. SpinQuant: LLM quantization with learned rotations. In: Proceedings of the 13th International Conference on Learning Representations. Singapore: OpenReview.net, 2025.
    [38] Oh S, Sim H, Kim J, Lee J. Non-uniform step size quantization for accurate post-training quantization. In: Proceedings of the 17th European Conference on Computer Vision. Tel Aviv, Israel: Springer, 2022. 658−673
    [39] Lee E H, Miyashita D, Chai E, Murmann B, Wong S S. LogNet: Energy-efficient neural networks using logarithmic computation. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). New Orleans, USA: IEEE, 2017. 5900−5904
    [40] Lin Y, Zhang T Y, Sun P Q, Li Z, Zhou S C. FQ-ViT: Post-training quantization for fully quantized vision Tansformer. In: Proceedings of the 31st International Joint Conference on Artificial Intelligence. Vienna, Austria: IJCAI, 2022. 1173−1179
    [41] Berger T. Rate-distortion theory. Wiley Encyclopedia of Telecommunications. Hoboken: John Wiley & Sons, 2003.
    [42] Yao Z W, Dong Z, Zheng Z C, Gholami A, Yu J L, Tan E, et al. HAWQ-V3: Dyadic neural network quantization. In: Proceedings of the 38th International Conference on Machine Learning. Virtual Event: PMLR, 2021. 11875−11886
    [43] Song H Y, Dharmapurikar S, Turner J, Lockwood J. Fast Hash table lookup using extended bloom filter: An aid to network processing. ACM SIGCOMM Computer Communication Review, 2005, 35(4): 181−192
    [44] Deng J, Dong W, Socher R, Li L J, Li K, Fei-Fei L. ImageNet: A large-scale hierarchical image database. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Miami, USA: IEEE, 2009. 248−255
    [45] Lin T Y, Maire M, Belongie S, Hays J, Perona P, Ramanan D, et al. Microsoft COCO: Common objects in context. In: Proceedings of the 13th European Conference on Computer Vision. Zurich, Switzerland: Springer, 2014. 740−755
    [46] Bisk Y, Zellers R, le Bras R, Gao J F, Choi Y. PIQA: Reasoning about physical commonsense in natural language. In: Proceedings of the 34th AAAI Conference on Artificial Intelligence. New York, USA: AAAI Press, 2020. 7432−7439
    [47] Clark P, Cowhey I, Etzioni O, Khot T, Sabharwal A, Schoenick C, et al. Think you have solved question answering? Try ARC, the AI2 reasoning challenge. arXiv preprint arXiv: 1803.05457, 2018.
    [48] Clark C, Lee K, Chang M W, Kwiatkowski T, Collins M, Toutanova K. BoolQ: Exploring the surprising difficulty of natural yes/no questions. In: Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Minneapolis, USA: Association for Computational Linguistics, 2019. 2924−2936
    [49] Zellers R, Holtzman A, Bisk Y, Farhadi A, Choi Y. HellaSwag: Can a machine really finish your sentence? In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence, Italy: Association for Computational Linguistics, 2019. 4791−4800
    [50] Sakaguchi K, le Bras R, Bhagavatula C, Choi Y. WinoGrande: An adversarial Winograd schema challenge at scale. Communications of the ACM, 2021, 64(9): 99−106
    [51] Merity S, Xiong C, Bradbury J, Socher R. Pointer sentinel mixture models. In: Proceedings of the 5th International Conference on Learning Representations. Toulon, France: OpenReview.net, 2017.
    [52] Raffel C, Shazeer N, Roberts A, Lee K, Narang S, Matena M, et al. Exploring the limits of transfer learning with a unified text-to-text Tansformer. The Journal of Machine Learning Research, 2020, 21(1): Article No. 140
    [53] Liu Y, Duan H D, Zhang Y H, Li B, Zhang S Y, Zhao W B, et al. MMBench: Is your multi-modal model an all-around player? In: Proceedings of the 18th European Conference on Computer Vision. Milan, Italy: Springer, 2025. 216−233
    [54] Yang A, Li A F, Yang B S, Zhang B C, Hui B Y, Zheng B, et al. Qwen3 technical report. arXiv preprint arXiv: 2505.09388, 2025.
    [55] Bai S, Cai Y X, Chen R Z, Chen K Q, Chen X H, Cheng Z S, et al. Qwen3-VL technical report. arXiv preprint arXiv: 2511.21631, 2025.
    [56] Zheng X Y, Li Y Y, Chu H R, Feng Y, Ma X D, Wang Z N, et al. An empirical study of Qwen3 quantization. Visual Intelligence, 2025, 4(1): Article No. 11
    [57] Liu Z H, Wang Y H, Han K, Zhang W, Ma S W, Gao W. Post-training quantization for vision Tansformer. In: Proceedings of the 35th International Conference on Neural Information Processing Systems. Virtual Event: Curran Associates, Inc., 2021. Article No. 2152
    [58] Li R D, Wang Y, Liang F, Qin H W, Yan J J, Fan R. Fully quantized network for object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Long Beach, USA: IEEE, 2019. 2805−2814
    [59] Wightman R. PyTorch Image Models [Online], available: https://github.com/huggingface/pytorch-image-models, July 3, 2026
    [60] Ma Y X, Li H X, Zheng X W, Ling F, Xiao X F, Wang R, et al. AffineQuant: Affine transformation quantization for large language models. In: Proceedings of the 12th International Conference on Learning Representations. Vienna, Austria: OpenReview.net, 2024.
    [61] Chen M Z, Liu Y, Wang J H, Bin Y, Shao W Q, Luo P. PrefixQuant: Eliminating outliers by prefixed tokens for large language models quantization. IEEE Transactions on Pattern Analysis and Machine Intelligence, DOI: 10.1109/TPAMI.2026.3711802
    [62] Lin J, Tang J M, Tang H T, Yang S, Chen W M, Wang W C, et al. AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration. In: Proceedings of the 7th Annual Conference on Machine Learning and Systems. Santa Clara, USA: MLSys, 2024.
  • 加载中
计量
  • 文章访问数:  13
  • HTML全文浏览量:  8
  • 被引次数: 0
出版历程
  • 收稿日期:  2025-12-04
  • 录用日期:  2026-06-24
  • 网络出版日期:  2026-08-31

目录

    /

    返回文章
    返回