Papers
arxiv:2609.28470

StudentBench: AI and human tutoring yield equivalent GRE learning gains

Published on Sep 23
· Submitted by
Curtis Northcutt
on Sep 24
Authors:
,
,
,
,
,

Abstract

Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). The StudentBench platform is freely available at https://studentbench.org.

Community

Paper submitter

We introduce StudentBench, a suite of AI teaching evaluations and a public platform for measuring how well AI helps people learn. The paper compares AI tutoring, expert human tutoring, and a no tutoring control on Quantitative and Verbal GRE learning gains, and evaluates AI-generated lesson plans and practice problems.

Dataset: https://huggingface.co/datasets/handshake-ai-research/studentbench
Project: https://studentbench.org
Code: https://github.com/Handshake-AI-Research/studentbench

Paper submitter

This is the first work to establish and measure the cost to augment a student’s learning by one percentage point.

Over 1 trillion USD is being spent teaching machines to improve machines.
StudentBench shows the value of teaching machines to improve humans.

Paper submitter

across 2000+ students, AI tutoring was statistically equivalent to expert human tutoring on GRE learning gains (p=.015).

Paper submitter

this work establishes an important precedent for the democratization of education -- a student who has access to an LLM can improve their GRE score for one penny.

Paper submitter

We provide five new kinds of tutoring evaluation leaderboards where we separate 13 LLM-based AI tutors (and we can clearly see which models do best at which capabilities).

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.28470
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.28470 in a model README.md to link it from this page.

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.28470 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.