Papers
arxiv:2610.08778

Sherpa: Teaching LLMs to Teach Adaptively

Published on Oct 6
· Submitted by
Yanzhe Zhang
on Oct 7
Authors:
,
,

Abstract

Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria that specify what good teaching looks like. However, these signals are often not grounded in individual student learning outcomes, where effective teaching strategies can vary substantially across learners. To address this, we introduce Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with LLMs conditioned on distinct learning preferences and trains a teacher model to adapt its instruction by directly maximizing their learning outcomes. Teacher LLMs trained with Sherpa improve instructed students' performance across all archetypes by an average of 20.5 percentage points. Under MathTutorBench's evaluation, Sherpa raises the overall pedagogy score from 52.5% to 79.2%, indicating better teaching responses. Our human studies show that the trained teacher is preferred over the base model in 79.6% of pairwise comparisons. Together, Sherpa trains LLM teachers to adapt to diverse simulated students and become better aligned with human teachers, paving the road towards AI tutors teaching real students.

Community

Paper author Paper submitter

Sherpa trains LLMs to become adaptive teachers by optimizing directly for student learning outcomes rather than predefined notions of good teaching. Using multi-turn RL with diverse simulated student archetypes, the teacher learns to infer different students’ needs and adapt its teaching strategy accordingly. Sherpa improves student performance by 20.5 percentage points on average, raises MathTutorBench pedagogy scores from 52.5% to 79.2%, and is preferred over the base model by human teachers in 79.6% of comparisons.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.08778
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.08778 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.08778 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.08778 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.