StudentBench: AI and human tutoring yield equivalent GRE learning gains
Abstract
Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). The StudentBench platform is freely available at https://studentbench.org.
Community
We introduce StudentBench, a suite of AI teaching evaluations and a public platform for measuring how well AI helps people learn. The paper compares AI tutoring, expert human tutoring, and a no tutoring control on Quantitative and Verbal GRE learning gains, and evaluates AI-generated lesson plans and practice problems.
Dataset: https://huggingface.co/datasets/handshake-ai-research/studentbench
Project: https://studentbench.org
Code: https://github.com/Handshake-AI-Research/studentbench
This is the first work to establish and measure the cost to augment a student’s learning by one percentage point.
Over 1 trillion USD is being spent teaching machines to improve machines.
StudentBench shows the value of teaching machines to improve humans.
across 2000+ students, AI tutoring was statistically equivalent to expert human tutoring on GRE learning gains (p=.015).
this work establishes an important precedent for the democratization of education -- a student who has access to an LLM can improve their GRE score for one penny.
We provide five new kinds of tutoring evaluation leaderboards where we separate 13 LLM-based AI tutors (and we can clearly see which models do best at which capabilities).
Get this paper in your agent:
hf papers read 2609.28470 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper