Construction and Automated Evaluation of Large Language Model for Interactive Teaching Content Generation
-
摘要: 当前教育大语言模型主要依赖文本对话开展知识讲解, 缺乏直观且动态的交互式呈现能力, 从而在复杂概念的可视化与过程性理解支持方面存在一定局限. 为此, 提出一种面向富媒体交互生成的教育大语言模型构建方法及其配套评测框架. 首先, 依据多学科教学大纲并结合细粒度知识点, 构建包含高质量交互式超文本标记语言(HTML)动画以及教学代码的指令数据集, 以缓解高质量数据稀缺的问题. 其次, 构建“监督微调 + 强化学习”的两阶段训练策略: 设计融合代码可执行性与教学相关性的混合奖励模型, 并通过强化学习进一步对齐模型生成内容的教育价值与代码鲁棒性, 从而获得具备交互式HTML教学素材生成能力的领域专用大语言模型. 最后, 针对交互式教学内容缺乏自动化评测体系的现状, 构建涵盖代码正确性、渲染成功率与教学交互性等维度的指标体系, 并实现相应的自动评估管线. 实验结果表明, 所提模型在交互式教学内容生成的准确性与交互性方面优于主流开源与闭源基座模型. 相关模型、数据与评测工具将公开发布, 以支持教育人工智能研究与应用的进一步发展.Abstract: Existing educational large language models (LLMs) primarily rely on text-based dialogue for knowledge explanation, lacking intuitive and dynamic interactive presentation capabilities. Consequently, they face certain limitations in visualizing complex concepts and supporting process-oriented understanding. To address this gap, we propose a method for constructing educational LLMs for rich-media interactive generation, along with a supporting evaluation framework. First, guided by multi-disciplinary curricula and fine-grained knowledge points, we construct an instruction dataset containing high-quality interactive hypertext markup language (HTML) animations and teaching code, alleviating the scarcity of high-quality data. Second, we develop a two-stage training strategy——supervised fine-tuning followed by reinforcement learning (RL). Specifically, we design a hybrid reward model that integrates code executability and pedagogical relevance, and apply RL to further align the educational value of model-generated content with code robustness, resulting in a domain-specific LLM capable of producing interactive HTML teaching materials. Finally, to address the current absence of an automated evaluation system for interactive teaching content, we establish a multi-dimensional metric suite covering code correctness, rendering success rate, and teaching interactivity, and implement an automated evaluation pipeline. Experimental results show that the proposed model outperforms mainstream open-source and closed-source base models in both the accuracy and interactivity of generated interactive teaching content. The related models, datasets, and evaluation tools will be publicly released to support the further development of educational artificial intelligence research and applications.
-
表 1 模型配置
Table 1 Model configurations
角色 模型名称 总参数/激活参数 任务 策略模型 Qwen3-30B-A3B-Coder 30.5 B/2.4 B 生成交互式代码 评估模型 Qwen3-VL-235B 235 B/22 B 多模态联合评分 表 2 不同模型在Edu-Eval测试集上的性能对比
Table 2 Performance comparison of different models on the Edu-Eval test set
模型 总分 语法层(L1) 逻辑层(L2) 视觉层(L3) 交互层(L4) DeepSeek-V3.2 43.9 34.8 65.0 70.3 46.0 Gemini 3 Pro 48.9 44.0 66.7 59.4 55.0 Qwen3-235B-Instruct 52.8 45.5 71.0 67.9 56.4 Qwen3-30B-A3B-Coder 55.0 46.9 65.3 64.3 64.8 IE-code 63.1 49.6 61.4 74.9 77.7 表 3 不同训练阶段与数据规模的消融实验结果
Table 3 Ablation study results on different training stages and data scale
模型 总分 视觉层(L3) 交互层(L4) 基座模型 55.0 64.3 64.8 Direct-RL (无SFT) 53.9 68.9 62.2 IE-code-5k-SFT 47.2 57.3 58.7 IE-code-5k-RL 57.6 72.0 70.7 IE-code-10k-SFT 46.3 54.9 58.3 IE-code-10k-RL 63.1 74.9 77.7 表 4 匿名A/B人类偏好评估结果
Table 4 Anonymous A/B human preference evaluation results
评价结果 样本级结果数 比例 IE-code更好 13 65.0% Qwen3-30B-A3B-Coder更好 3 15.0% 差不多 2 10.0% 都不合格 1 5.0% 跳过 1 5.0% -
[1] Zhao W X, Zhou K, Li J Y, Tang T Y, Dong Z C, Hou Y P, et al. A survey of large language models. Frontiers of Computer Science, 2026, 20: Article No. 2012627 doi: 10.1007/s11704-026-60308-3 [2] Google DeepMind. A new era of intelligence with Gemini 3 [Online], available: https://blog.google/products-and-platforms/products/gemini/gemini-3/, July 29, 2026 [3] Yang A, Li A F, Yang B S, Zhang B C, Hui B Y, Zheng B, et al. Qwen3 technical report. arXiv preprint arXiv: 2505.09388, 2025. [4] Chen Z, Liu T Q, Tian M, Tong Q, Luo W Q, Liu Z T. Advancing math reasoning in language models: The impact of problem-solving data, data synthesis methods, and training stages. arXiv preprint arXiv: 2501.14002, 2025. [5] Sweller J, van Merriënboer J J G, Paas F. Cognitive architecture and instructional design: 20 years later. Educational Psychology Review, 2019, 31(2): 261−292 doi: 10.1007/s10648-019-09465-5 [6] Skulmowski A, Xu K M. Understanding cognitive load in digital and online learning: A new perspective on extraneous cognitive load. Educational Psychology Review, 2022, 34(1): 171−196 doi: 10.1007/s10648-021-09624-7 [7] Risko E F, Gilbert S J. Cognitive offloading. Trends in Cognitive Sciences, 2016, 20(9): 676−688 doi: 10.1016/j.tics.2016.07.002 [8] Rutten N, van Joolingen W R, van der Veen J T. The learning effects of computer simulations in science education. Computers & Education, 2012, 58(1): 136−153 doi: 10.1016/j.compedu.2011.07.017 [9] Schnotz W, Rasch T. Enabling, facilitating, and inhibiting effects of animations in multimedia learning: Why reduction of cognitive load can have negative results on learning. Educational Technology Research and Development, 2005, 53(3): 47−58 doi: 10.1007/BF02504797 [10] Wang X Y, Chen Y Y, Yuan L F, Zhang Y Z, Li Y Z, Peng H, et al. Executable code actions elicit better LLM agents. In: Proceedings of the 41st International Conference on Machine Learning. Vienna, Austria: Proceedings of Machine Learning Research, 2024, 235: 50208−50232 [11] DeepSeek-AI, Zhu Q H, Guo D Y, Shao Z H, Yang D J, Wang P Y, et al. DeepSeek-Coder-V2: Breaking the barrier of closed-source models in code intelligence. arXiv preprint arXiv: 2406.11931, 2024. [12] Qwen Team. Qwen3-Coder: Agentic coding in the world [Online], available: https://qwenlm.github.io/blog/qwen3-coder/, July 29, 2026 [13] Si C L, Zhang Y Z, Li R, Yang Z Y, Liu R B, Yang D Y. Design2Code: Benchmarking multimodal code generation for automated front-end engineering. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Albuquerque, USA: Association for Computational Linguistics, 2025. 3956−3974 [14] Yang J, Prabhakar A, Narasimhan K, Yao S Y. InterCode: Standardizing and benchmarking interactive coding with execution feedback. Advances in Neural Information Processing Systems, 2023, 36: 23826−23854 doi: 10.52202/075280-1035 [15] Shao Z H, Wang P Y, Zhu Q H, Xu R X, Song J X, Bi X, et al. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv: 2402.03300, 2024. [16] Qwen Team. Qwen3-VL-235B-A22B-Instruct [Online], available: https://huggingface.co/Qwen/Qwen3-VL-235B-A22B-Instruct, July 29, 2026 [17] Kasneci E, Sessler K, Küchemann S, Bannert M, Dementieva D, Fischer F, et al. ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 2023, 103: Article No. 102274 doi: 10.1016/j.lindif.2023.102274 [18] Kohnke L, Moorhouse B L, Zou D. ChatGPT for language teaching and learning. RELC Journal, 2023, 54(2): 537−550 doi: 10.1177/00336882231162868 [19] Trinh T H, Wu Y H, Le Q V, He H, Luong T. Solving olympiad geometry without human demonstrations. Nature, 2024, 625(7995): 476−482 doi: 10.1038/s41586-023-06747-5 [20] Hui B Y, Yang J, Cui Z Y, Yang J X, Liu D Y H, Zhang L, et al. Qwen2.5-Coder technical report. arXiv preprint arXiv: 2409.12186, 2024. [21] Zhou S Y, Xu F F, Zhu H, Zhou X H, Lo R, Sridhar A, et al. WebArena: A realistic web environment for building autonomous agents. arXiv preprint arXiv: 2307.13854, 2023. [22] Meng Y, Xia M Z, Chen D Q. SimPO: Simple preference optimization with a reference-free reward. Advances in Neural Information Processing Systems, 2024, 37: 124198−124235 doi: 10.52202/079017-3946 [23] Zhang K C, Li G, Li J, Dong Y H, Li J, Jin Z. Focused-DPO: Enhancing code generation through focused preference optimization on error-prone points. In: Findings of the Association for Computational Linguistics: ACL 2025. Vienna, Austria: Association for Computational Linguistics, 2025. 9578−9591 [24] Sheng G M, Zhang C, Ye Z L F, Wu X B, Zhang W, Zhang R, et al. HybridFlow: A flexible and efficient RLHF framework. In: Proceedings of the Twentieth European Conference on Computer Systems. Rotterdam, The Netherlands: Association for Computing Machinery, 2025. 1279−1297 [25] Rodriguez J, Zhang H T, Puri A, Pramanik R, Feizi A, Wichmann P, et al. Rendering-aware reinforcement learning for vector graphics generation. Advances in Neural Information Processing Systems, 2025, 38: 60496−60534 [26] Gao L, Schulman J, Hilton J. Scaling laws for reward model overoptimization. In: Proceedings of the 40th International Conference on Machine Learning. Honolulu, USA: Proceedings of Machine Learning Research, 2023, 202: 10835−10866 [27] Zheng L M, Chiang W L, Sheng Y, Zhuang S Y, Wu Z H, Zhuang Y H, et al. Judging LLM-as-a-judge with MT-Bench and Chatbot arena. Advances in Neural Information Processing Systems, 2023, 36: 46595−46623 doi: 10.52202/075280-2020 [28] Madaan A, Tandon N, Gupta P, Hallinan S, Gao L Y, Wiegreffe S, et al. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems, 2023, 36: 46534−46594 doi: 10.52202/075280-2019 [29] Shinn N, Cassano F, Gopinath A, Narasimhan K, Yao S Y. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 2023, 36: 8634−8652 doi: 10.52202/075280-0377 [30] Mayer R E, Fiorella L. Principles for reducing extraneous processing in multimedia learning: Coherence, signaling, redundancy, spatial contiguity, and temporal contiguity principles. The Cambridge Handbook of Multimedia Learning. New York: Cambridge University Press, 2014. 279−315 -
计量
- 文章访问数: 97
- HTML全文浏览量: 53
- 被引次数: 0
下载: