论文

Sherpa:用多轮强化学习让 LLM 学会因材施教

Sherpa: Teaching LLMs to Teach Adaptively

精选理由

用 LLM 扮演不同类型的学生,让教师模型通过强化学习直接优化学生成绩,教学得分从 52.5% 拉到 79.2%,做 AI 教育的可以看看。

Sherpa 是一个多轮强化学习框架,用 LLM 模拟多种有不同学习偏好的学生原型,训练教师模型根据学生个体学习结果调整教学策略。经过 Sherpa 训练后,被指导学生在所有原型上的表现平均提升 20.5 个百分点。在 MathTutorBench 评测中,教学综合得分从 52.5% 提高到 79.2%。人工评估中,训练后的教师模型在 79.6% 的成对比较中优于基座模型。

原文 · arXiv cs.AI

Sherpa: Teaching LLMs to Teach Adaptively

Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria that specify what good teaching looks like. However, these signals are often not grounded in individual student learning outcomes, where effective teaching strategies can vary substantially across learners. To address this, we introduce Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with LLMs conditioned on distinct learning preferences and trains a teacher model to adapt its instruction by directly maximizing their learning outcomes. Teacher LLMs trained with Sherpa improve instructed students' performance across all archetypes by an average of 20.5 percentage points. Under MathTutorBench's evaluation, Sherpa raises the overall pedagogy score from 52.5% to 79.2%, indicating better teaching responses. Our human studies show that the trained teacher is preferred over the base model in 79.6% of pairwise comparisons. Together, Sherpa trains LLM teachers to adapt to diverse simulated students and become better aligned with human teachers, paving the road towards AI tutors teaching real students.