• 中文核心
  • EI
  • 中国科技核心
  • Scopus
  • CSCD
  • 英国科学文摘

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

交互式教学内容生成大模型的构建与自动化评估

张若华 刘涛 卢烨 宋思宇 刘文涛 宋宇 周爱民 郝昊

张若华, 刘涛, 卢烨, 宋思宇, 刘文涛, 宋宇, 周爱民, 郝昊. 交互式教学内容生成大模型的构建与自动化评估. 自动化学报, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c260145
引用本文: 张若华, 刘涛, 卢烨, 宋思宇, 刘文涛, 宋宇, 周爱民, 郝昊. 交互式教学内容生成大模型的构建与自动化评估. 自动化学报, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c260145
Zhang Ruo-Hua, Liu Tao, Lu Ye, Song Si-Yu, Liu Wen-Tao, Song Yu, Zhou Ai-Min, Hao Hao. Construction and automated evaluation of a large language model for interactive teaching content generation. Acta Automatica Sinica, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c260145
Citation: Zhang Ruo-Hua, Liu Tao, Lu Ye, Song Si-Yu, Liu Wen-Tao, Song Yu, Zhou Ai-Min, Hao Hao. Construction and automated evaluation of a large language model for interactive teaching content generation. Acta Automatica Sinica, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c260145

交互式教学内容生成大模型的构建与自动化评估

doi: 10.16383/j.aas.c260145 cstr: 32138.14.j.aas.c260145
基金项目: 国家自然科学基金(62306174)资助
详细信息
    作者简介:

    张若华:华东师范大学上海智能教育研究院、计算机科学与技术学院硕士研究生. 主要研究方向为教育大模型和模型高效微调. E-mail: 51275901086@stu.ecnu.edu.cn

    刘涛:华东师范大学上海智能教育研究院、计算机科学与技术学院硕士研究生. 主要研究方向为教育大模型和模型评测. E-mail: 51275901030@stu.ecnu.edu.cn

    卢烨:华东师范大学上海智能教育研究院博士研究生. 主要研究方向为教育大模型、强化学习和演化计算. E-mail: 51275901116@stu.ecnu.edu.cn

    宋思宇:华东师范大学上海智能教育研究院、计算机科学与技术学院博士研究生. 主要研究方向为教育大模型、强化学习和自进化. E-mail: siyusong00@gmail.com

    刘文涛:华东师范大学上海智能教育研究院、计算机科学与技术学院与上海创智学院联合培养博士研究生. 主要研究方向为超长上下文大模型和多模态大模型. E-mail: wtliu@stu.ecnu.edu.cn

    宋宇:华东师范大学上海智能教育研究院研究员. 主要研究方向为智能课堂教学评价、人机交互和多智能体协同. E-mail: ysong@mail.ecnu.edu.cn

    周爱民:华东师范大学上海智能教育研究院、计算机科学与技术学院教授, 上海创智学院全时导师. 主要研究方向为大语言模型、智能体系统和演化优化与学习. E-mail: amzhou@cs.ecnu.edu.cn

    郝昊:华东师范大学上海智能教育研究院副研究员. 主要研究方向为教育大模型、智能体设计和无梯度优化. 本文通信作者. E-mail: hhao@mail.ecnu.edu.cn

  • 中图分类号: Y

Construction and Automated Evaluation of a Large Language Model for Interactive Teaching Content Generation

Funds: Supported by National Natural Science Foundation of China (62306174)
More Information
    Author Bio:

    Zhang Ruo-Hua Master student at the Shanghai Institute of AI for Education and the School of Computer Science and Technology, East China Normal University. His research interests include educational large models and efficient fine-tuning

    Liu Tao Master student at the Shanghai Institute of AI for Education and the School of Computer Science and Technology, East China Normal University. His research interests include educational large models and model evaluation

    Lu Ye Ph.D. candidate at the Shanghai Institute of AI for Education, East China Normal University. His research interests include educational large models, reinforcement learning, and evolutionary computation

    Song Si-Yu Ph.D. candidate at the Shanghai Institute of AI for Education and the School of Computer Science and Technology, East China Normal University. His research interests include educational large models, reinforcement learning, and self-evolution

    Liu Wen-Tao Ph.D. candidate jointly trained by the Shanghai Institute of AI for Education and the School of Computer Science and Technology at East China Normal University, and Shanghai Innovation Institute. His research interests include ultra-long-context large models and multimodal large models

    Song Yu Researcher at the Shanghai Institute of AI for Education, East China Normal University. His research interests include intelligent classroom teaching evaluation, human-AI interaction, and multi-agent collaboration

    Zhou Ai-Min Professor at the Shanghai Institute of AI for Education and the School of Computer Science and Technology, East China Normal University, and a full-time mentor at Shanghai Innovation Institute. His research interests include large models, agent systems, and evolutionary optimization and learning

    Hao Hao Associate researcher at the Shanghai Institute of AI for Education, East China Normal University. His research interests include educational large models, agent design, and gradient-free optimization. Corresponding author of this paper

  • 摘要: 当前教育大模型主要依赖文本对话开展知识讲解, 缺乏直观且动态的交互式呈现能力, 从而在复杂概念的可视化与过程性理解支持方面存在一定局限. 为此, 提出一种面向富媒体交互生成的教育大模型构建方法及其配套评测框架. 首先, 依据多学科教学大纲并结合细粒度知识点, 构建包含高质量交互式超文本标记语言(hypertext markup language, HTML)动画以及教学代码的指令数据集, 以缓解该方向高质量数据稀缺的问题. 其次, 构建“监督微调 + 强化学习”的两阶段训练策略: 设计融合代码可执行性与教学相关性的混合奖励模型, 并通过强化学习进一步对齐模型生成内容的教育价值与代码鲁棒性, 从而获得具备交互式HTML教学素材生成能力的领域专用大模型. 最后, 针对交互式教学内容缺乏自动化评测体系的现状, 构建涵盖代码正确性、渲染成功率与教学交互性等维度的指标体系, 并实现相应的自动评估管线. 实验结果表明, 所提模型在交互式辅导内容生成的准确性与交互性方面优于主流开源与闭源基座模型. 相关模型、数据与评测工具将公开发布, 以支持教育人工智能研究与应用的进一步发展.
    1)  11 Qwen3-VL详情见https://huggingface.co/Qwen/Qwen3-VL; Qwen3-Coder详情见https://qwenlm.github.io/blog/qwen3-coder/.
  • 图  1  传统静态资源与生成式交互仿真中的认知加工过程对比

    Fig.  1  Comparison of cognitive processing in static resources and generative interactive simulations

    图  2  融合富媒体交互的生成式教育大模型总体架构

    Fig.  2  Overall architecture of the generative educational large model with rich-media interaction

    图  3  基于多模态反馈的迭代修正闭环

    Fig.  3  Closed-loop iterative refinement based on multimodal feedback

    图  4  级联评估与稀疏视觉采样策略

    Fig.  4  Cascaded evaluation and sparse visual sampling strategy

    图  5  多模态奖励模型计算流程图

    Fig.  5  Computational flowchart of the multimodal reward model

    图  6  Edu-Eval基准测试集的学科、学段与知识点分布

    Fig.  6  Subject, grade-band, and knowledge-point distribution of the Edu-Eval benchmark dataset

    图  7  不同模型的主实验结果

    Fig.  7  Main experimental results of different models

    图  8  模型能力多维评估雷达图

    Fig.  8  Radar chart of multidimensional model capabilities

    图  9  “牛顿与经典力学”交互式教学内容生成案例

    Fig.  9  Case of interactive teaching content generation for Newton and classical mechanics

    表  1  模型配置与训练开销

    Table  1  Model configurations and training costs

    角色 模型名称 总参数 / 激活参数 任务
    策略模型 Qwen3-30B-A3B-Coder 30.5 B / 2.4 B 生成交互式代码
    评估模型 Qwen3-VL-235B 235 B / 22 B 多模态联合评分
    下载: 导出CSV

    表  2  不同模型在Edu-Eval测试集上的性能对比

    Table  2  Performance comparison of different models on the Edu-Eval test set

    模型 总分 语法层(L1) 逻辑层(L2) 视觉层(L3) 交互层(L4)
    DeepSeek-V3.243.934.865.070.346.0
    Gemini 3 Pro48.944.066.759.455.0
    Qwen3-235B-Instruct52.845.571.067.956.4
    Qwen3-Coder-30B-Instruct55.046.965.364.364.8
    IE-code63.149.661.474.977.7
    下载: 导出CSV

    表  3  不同训练阶段与数据规模的消融实验结果

    Table  3  Ablation study on training stages and data scale

    模型 总分 视觉层(L3) 交互层(L4)
    基座模型55.064.364.8
    Direct-RL (无SFT)53.968.962.2
    IE-code-5k-SFT47.257.358.7
    IE-code-5k-RL57.672.070.7
    IE-code-10k-SFT46.354.958.3
    IE-code-10k-RL63.174.977.7
    下载: 导出CSV

    表  4  匿名A/B人类偏好评估结果

    Table  4  Anonymous A/B human preference evaluation results

    评价结果 样本级结果数 比例
    IE-code更好 13 65.0%
    Qwen3-Coder-30B-Instruct更好 3 15.0%
    差不多 2 10.0%
    都不合格 1 5.0%
    跳过 1 5.0%
    下载: 导出CSV
  • [1] Zhao W X, Zhou K, Li J Y, Tang T Y, Dong Z C, Hou Y P, et al. A survey of large language models. Frontiers of Computer Science, 2026, 20: Article No. 2012627 doi: 10.1007/s11704-026-60308-3
    [2] Google DeepMind. A new era of intelligence with Gemini 3[Online], available: https://blog.google/products-and-platforms/products/gemini/gemini-3/, July 29, 2026
    [3] Yang A, Li A F, Yang B S, Zhang B C, Hui B Y, Zheng B, et al. Qwen3 technical report. arXiv preprint arXiv: 2505.09388, 2025.
    [4] Chen Z, Liu T Q, Tian M, Tong Q, Luo W Q, Liu Z T. Advancing Math Reasoning in Language Models: The impact of problem-solving data, data synthesis methods, and training stages. arXiv preprint arXiv: 2501.14002, 2025.
    [5] Sweller J, van Merriënboer J J G, Paas F. Cognitive architecture and instructional design: 20 years later. Educational Psychology Review, 2019, 31(2): 261−292 doi: 10.1007/s10648-019-09465-5
    [6] Skulmowski A, Xu K M. Understanding cognitive load in digital and online learning: A new perspective on extraneous cognitive load. Educational Psychology Review, 2022, 34(1): 171−196 doi: 10.1007/s10648-021-09624-7
    [7] Risko E F, Gilbert S J. Cognitive offloading. Trends in Cognitive Sciences, 2016, 20(9): 676−688 doi: 10.1016/j.tics.2016.07.002
    [8] Rutten N, van Joolingen W R, van der Veen J T. The learning effects of computer simulations in science education. Computers & Education, 2012, 58(1): 136−153 doi: 10.1016/j.compedu.2011.07.017
    [9] Schnotz W, Rasch T. Enabling, facilitating, and inhibiting effects of animations in multimedia learning: Why reduction of cognitive load can have negative results on learning. Educational Technology Research and Development, 2005, 53(3): 47−58 doi: 10.1007/BF02504797
    [10] Wang X Y, Chen Y Y, Yuan L F, Zhang Y Z, Li Y Z, Peng H, et al. Executable code actions elicit better LLM agents. In: Proceedings of the 41st International Conference on Machine Learning. Vienna, Austria: Proceedings of Machine Learning Research, 2024, 235: 50208−50232
    [11] DeepSeek-AI, Zhu Q H, Guo D Y, Shao Z H, Yang D J, Wang P Y, et al. DeepSeek-Coder-V2: Breaking the barrier of closed-source models in code intelligence. arXiv preprint arXiv: 2406.11931, 2024.
    [12] Qwen Team. Qwen3-Coder: Agentic coding in the world[Online], available: https://qwenlm.github.io/blog/qwen3-coder/, July 29, 2026
    [13] Si C L, Zhang Y Z, Li R, Yang Z Y, Liu R B, Yang D Y. Design2Code: Benchmarking multimodal code generation for automated front-end engineering. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Albuquerque, USA: Association for Computational Linguistics, 2025. 3956−3974
    [14] Yang J, Prabhakar A, Narasimhan K, Yao S Y. InterCode: Standardizing and benchmarking interactive coding with execution feedback. Advances in Neural Information Processing Systems, 2023, 36: 23826−23854 doi: 10.52202/075280-1035
    [15] Shao Z H, Wang P Y, Zhu Q H, Xu R X, Song J X, Bi X, et al. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv: 2402.03300, 2024.
    [16] Qwen Team. Qwen3-VL-235B-A22B-Instruct[Online], available: https://huggingface.co/Qwen/Qwen3-VL-235B-A22B-Instruct, July 29, 2026
    [17] Kasneci E, Sessler K, Küchemann S, Bannert M, Dementieva D, Fischer F, et al. ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 2023, 103: Article No. 102274 doi: 10.1016/j.lindif.2023.102274
    [18] Kohnke L, Moorhouse B L, Zou D. ChatGPT for language teaching and learning. RELC Journal, 2023, 54(2): 537−550 doi: 10.1177/00336882231162868
    [19] Trinh T H, Wu Y H, Le Q V, He H, Luong T. Solving olympiad geometry without human demonstrations. Nature, 2024, 625(7995): 476−482 doi: 10.1038/s41586-023-06747-5
    [20] Hui B Y, Yang J, Cui Z Y, Yang J X, Liu D Y H, Zhang L, et al. Qwen2.5-Coder technical report. arXiv preprint arXiv: 2409.12186, 2024.
    [21] Zhou S Y, Xu F F, Zhu H, Zhou X H, Lo R, Sridhar A, et al. WebArena: A realistic web environment for building autonomous agents. arXiv preprint arXiv: 2307.13854, 2023.
    [22] Meng Y, Xia M Z, Chen D Q. SimPO: Simple preference optimization with a reference-free reward. Advances in Neural Information Processing Systems, 2024, 37: 124198−124235 doi: 10.52202/079017-3946
    [23] Zhang K C, Li G, Li J, Dong Y H, Li J, Jin Z. Focused-DPO: Enhancing code generation through focused preference optimization on error-prone points. In: Findings of the Association for Computational Linguistics: ACL 2025. Vienna, Austria: Association for Computational Linguistics, 2025. 9578−9591
    [24] Sheng G M, Zhang C, Ye Z L F, Wu X B, Zhang W, Zhang R, et al. HybridFlow: A flexible and efficient RLHF framework. In: Proceedings of the Twentieth European Conference on Computer Systems. Rotterdam, The Netherlands: Association for Computing Machinery, 2025. 1279−1297
    [25] Rodriguez J, Zhang H T, Puri A, Pramanik R, Feizi A, Wichmann P, et al. Rendering-aware reinforcement learning for vector graphics generation. Advances in Neural Information Processing Systems, 2025, 38: 60496−60534
    [26] Gao L, Schulman J, Hilton J. Scaling laws for reward model overoptimization. In: Proceedings of the 40th International Conference on Machine Learning. Honolulu, USA: Proceedings of Machine Learning Research, 2023, 202: 10835−10866
    [27] Zheng L M, Chiang W L, Sheng Y, Zhuang S Y, Wu Z H, Zhuang Y H, et al. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems, 2023, 36: 46595−46623 doi: 10.52202/075280-2020
    [28] Madaan A, Tandon N, Gupta P, Hallinan S, Gao L Y, Wiegreffe S, et al. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems, 2023, 36: 46534−46594 doi: 10.52202/075280-2019
    [29] Shinn N, Cassano F, Gopinath A, Narasimhan K, Yao S Y. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 2023, 36: 8634−8652 doi: 10.52202/075280-0377
    [30] Mayer R E, Fiorella L. Principles for reducing extraneous processing in multimedia learning: Coherence, signaling, redundancy, spatial contiguity, and temporal contiguity principles. In: Mayer R E, ed. The Cambridge Handbook of Multimedia Learning. 2nd ed. New York, USA: Cambridge University Press, 2014. 279−315
  • 加载中
计量
  • 文章访问数:  10
  • HTML全文浏览量:  5
  • 被引次数: 0
出版历程
  • 收稿日期:  2026-02-27
  • 录用日期:  2026-06-24
  • 网络出版日期:  2026-08-18

目录

    /

    返回文章
    返回