重新思考LLM辅导员的脚手架:基准测试与现实部署的交互错位

Rethinking Scaffolding in LLM Tutors: The Interactional Mismatch Between Benchmarks and Real-World Deployments

精选理由

这篇论文用9490个对话数据证明,AI辅导系统在测试中的表现和实际使用差很多,学生根本不吃那套引导。做教育AI的值得看看。

AI 摘要

该论文引入了一个评估管道,包含两个指标——聊天机器人脚手架和学生吸收率,并在9个数据集(共9490个对话)上应用,涵盖AI导师基准测试和现实部署。分析发现,基准测试假设高脚手架、高学生吸收率环境,但现实中的学生整体吸收率较低,经常绕过聊天机器人的教学框架。论文认为,绕过脚手架不一定有害,反而常突显了聊天机器人的教学框架与学生目标之间的不匹配。未来基准测试应评估聊天机器人如何导航多样化的学习情境和学生驱动的交互模式。

原文 · arXiv cs.AI

Rethinking Scaffolding in LLM Tutors: The Interactional Mismatch Between Benchmarks and Real-World Deployments

A central pedagogical value evaluated in AI tutor benchmarks is scaffolding: guiding students through graduated steps toward a solution. Alignment and evaluation methods for embedding scaffolding behaviour into chatbots, however, rest on an implicit assumption: that students will take up the scaffolding and engage in the conversation. To examine whether this assumption holds, we introduce an evaluation pipeline around two metrics - Chatbot Scaffolding and Student Uptake - and apply them across nine datasets of 9,490 chats, spanning AI tutor benchmarks and real-world deployments of educational chatbots. Our analysis reveals that while benchmarks assume a high-scaffolding, high-student-uptake environment, students in real-world settings exhibit lower levels of uptake overall - frequently bypassing the chatbot's pedagogical framing to drive the interaction toward their own learning goals at little interpersonal cost. We argue that bypassing scaffolding is not necessarily detrimental; rather, it frequently highlights a mismatch between a chatbot's pedagogical framing and the student's learning goals. To meaningfully evaluate the effectiveness of a chatbot's assistance, future benchmarks must move beyond the assumption that students will simply take up the scaffolding, and instead evaluate how these chatbots navigate diverse learning contexts and student-driven interaction patterns.