Research
StudentBench: AI and human tutoring yield equivalent GRE learning gains
Overview Research area: AI for education / human intelligence augmentation; large-scale randomized evaluation of LLM tutoring versus human tutoring. Technical level: Intermediate. The conceptual frami

- arXiv
- 2609.28470
- Published
- 2026-09-23
- Authors
- Curtis Northcutt, Inaara Hasmani, Kevin Feng, Trevor Khangi, Andreas Plesner, Jonas Mueller
AI summary
Overview
Research area: AI for education / human intelligence augmentation; large-scale randomized evaluation of LLM tutoring versus human tutoring.
Technical level: Intermediate. The conceptual framing is accessible to a general audience, but the paper relies on equivalence testing (two one-sided tests), ANCOVA, item response theory, and Bradley–Terry preference models.
Scope: StudentBench is a suite of AI teaching evaluations and a public platform that tests whether general-purpose LLMs, with minimal software scaffolding, produce GRE learning gains statistically equivalent to expert human tutoring.
What This Paper Is About
Most AI progress is measured by model capability benchmarks, not by whether a model actually improves what a human can do. This paper asks whether autonomous, general-purpose LLMs can teach as effectively as expert human tutors, using the GRE as a standardized testbed. The authors built StudentBench to collect real student learning data at scale and to separate AI tutors along lesson planning, practice-problem creation, conversational pedagogy, cost, and engagement.
Key Contributions
- StudentBench platform and dataset: A public AI teaching evaluation suite and platform that collected over 175,000 student–AI messages, with open-sourced data and code. It is freely available at https://studentbench.org.
- Learning-gain equivalence study: A randomized comparison of AI tutoring (13 AI tutors), human tutoring, and no-tutoring control across 2,383 human participants and 2,469 Quantitative and Verbal sessions, using GRE pre-tests and post-tests written by former ETS and Kaplan GRE exam creators.
- Three evaluations of AI teaching plus five leaderboards: Expert pairwise rubric review of lesson plans and practice problems (2,028 pairwise evaluations from 51 expert human tutors, based on 381 student pre-tests), plus transcript-based analysis of conversational pedagogy, yielding leaderboards across lesson planning, practice-problem creation, conversational pedagogy, cost, and engagement.
- Cost-per-learning-gain metric: A proposed standardized reporting unit, USD per percentage point of learning gain, used to compare AI inference cost against a $75 per hour human tutoring reference.
Main Findings
- AI tutoring matches human tutoring on pooled GRE gains: Pooled AI tutoring and human tutoring produced statistically equivalent learning gains (p = .015, ±0.25 pooled-SD bounds), and the result also held under tighter ±0.20-SD bounds (p = .023, 1.59 d.f.). The combined AI − human difference was −0.58 percentage points (90% CI [−2.18, 1.03]). Equivalence at the ±0.25-SD margin was also established for Quantitative but not Verbal.
- Best AI tutors beat humans in most domains: In five of the seven GRE domains, the best performing AI tutor surpassed the human tutor on average, and a different tutor led each domain. In Quantitative, the best AI tutors had higher mean gains than humans in three of the four domains; in Verbal, human mean gains remained above the pooled AI mean in all three domains.
- Large gains over no tutoring: After adjustment for pre-test score, the AI − control difference was 6.86 percentage points in Quantitative (95% CI [4.02, 9.69]) and 5.47 in Verbal ([2.46, 8.47]); across both sections the difference was 6.15 percentage points ([4.08, 8.21]), about 1.5–2 more correct answers out of 27.
- Human-level gains at 918 times lower cost: Gemma 4 31B achieved GRE learning gains equivalent to expert human tutoring (p = .044) at 918 times lower cost per percentage point of learning gain ($0.0052 for AI versus $4.81 for human). Its mean inference cost was $0.067 for the entire tutoring session covering lesson planning, practice-problem creation, and one hour of interactive tutoring.
- Six AI tutors passed individual equivalence tests: Six AI tutors passed individual equivalence tests against human tutoring in the cost analysis, with Gemma 4 31B used as the 1× cost reference.
- Cost range across tutors: Across the 12 AI tutors, mean inference cost ranged from $0.067 per session for Gemma 4 31B to $21.24 per session for GPT-5.5 Pro. All AI tutors on the Pareto frontier had a mean cost per session under $5. Gemini 3.5 Flash had a similar learning gain to GPT-5.5 Pro while costing 20 times less.
- Faster replies track with engagement and learning: Lower AI latency was strongly associated with higher student engagement (Spearman ρ = −0.81, p = .0056). In 1,137 Quantitative AI-tutoring sessions, faster replies were associated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002); these associations were not statistically significant in Verbal.
- Reply-time range: AI reply time ranged from 1.9 seconds for GPT-5.4 mini to 31.0 seconds for GPT-5.5 Pro across the 12 AI tutors.
- Anthropic models lead expert-rated teaching materials: Anthropic models, particularly Opus, performed strongly across all three leaderboards. Non-overlapping confidence intervals separated most model pairs for lesson planning and practice-problem creation.
- AI tutors cluster by model family on pedagogy: In conversational pedagogy, AI tutors clustered by model family with no overlap between family ranges, suggesting pedagogical characteristics may reflect company-wide training practices.
- Expert-preferred practice problems draw fewer student disputes: AI tutors scoring highly on practice-problem design received fewer student answer disputes in Quantitative and Combined (Spearman ρ = 0.71 and 0.77; p = .035 and p = .033). GPT-5.4 mini had the highest observed flag rate, with students disputing approximately 11% of answered Quantitative practice problems.
- Gemini tutors led on high-performing students: Using item response theory to estimate Quantitative pre-test proficiency, among students in the top 25%, four of the five highest mean gains came from Gemini tutors, led by Gemini 3.5 Flash.
Methodology in Plain English
Students were recruited through Handshake's platform of 25 million fellows; invitations went to over 70,000 randomly selected students across majors, primarily aged 18 to 23, who were paid $50 for completing a session independent of test performance. Each session followed three steps: a 27-question Quantitative or Verbal pre-test with eligibility and effort screening plus five minutes reviewing incorrect answers; one hour in an assigned condition (AI tutoring, human tutoring, or no tutoring control, where control students watched GRE-unrelated educational videos); and a post-test with different questions matched in length and concepts. Quantitative allowed 47 minutes and Verbal 41 minutes. Assessment forms P and Q were randomly counterbalanced so half the students got P as pre-test and Q as post-test and half the reverse, preventing test-difficulty differences from being mistaken for learning gains. AI tutors received two low-guidance prompts (one for lesson planning and problem creation, one for interactive tutoring) and never had access to post-test questions. Learning gain was defined as post-test minus pre-test percentage-correct, with unanswered questions counted as incorrect. Comparisons used ANCOVA to adjust for pre-test score, and AI−human equivalence used two one-sided tests at α = 0.05 with CR2 covariance and Satterthwaite degrees of freedom to account for shared tutors and repeated students. Data quality filters excluded sessions for incompleteness, low effort, rapid responses, below-random performance, frequent tab switching, and pre-test scores above 24/27. In the second study, 51 expert human tutors completed 2,028 pairwise reviews of AI-generated lesson plans and embedded practice problems from 381 student pre-tests, answering eight comparison questions per pair (five on lesson planning, three on practice problems), fitted with centered Bradley–Terry models. Conversational pedagogy was scored using six fixed text and turn-order rules applied to 1,971 AI and 135 human transcripts.
Why This Matters
- Impact on research: The paper argues that most LLM benchmarks measure model capability, not the ability to augment human capability, and it supplies a new evaluation paradigm and public infrastructure for measuring real learning gains on real exams with unaided post-tests. It also frames AI tutoring as the first condition for what the authors call recursive human self-improvement.
- Real-world applications:
- Affordable one-on-one test preparation, with AI tutoring on the Pareto frontier costing under $5 per session.
- Democratized access to individualized instruction, addressing Bloom's two-sigma problem of making one-on-one tutoring broadly affordable.
- Cost planning for institutions, using the proposed USD-per-percentage-point-of-learning-gain unit.
- Tutoring design guidance: latency appears to influence learning indirectly through engagement and practice, making it a first-class design consideration.
- Industry relevance: Mean AI inference cost ranged from $0.067 to $21.24 per session across 12 tutors, and Gemini 3.5 Flash matched GPT-5.5 Pro's learning gain at 20 times lower cost, showing that cost-to-learning ratios differ markedly across providers. The finding that pedagogical behavior clusters by model family suggests provider-level training practices matter for educational use. Handshake AI funded the work.
Future Directions
- Longitudinal retention studies: The authors measured immediate learning gains only and explicitly hope their findings encourage studies of whether gains persist over months.
- Testing outside the study population: The study used English-literate adults with Handshake platform access recruited for a paid study; deployment across languages, devices, and educational settings was not tested.
- Adding a solo-practice control: No condition had students work through GRE practice problems on their own, so the added benefit of interacting with an AI tutor over practice alone is not separately measured.
- Extending to other exams: The authors suggest AI tutoring may also help students prepare for other standardized exams such as the SAT and MCAT.
Target Audience
Readers who benefit most are AI and education researchers evaluating whether LLMs genuinely improve human capability, learning-science and psychometrics researchers interested in equivalence testing and expert-rated teaching evaluations, and education technology builders and institutional decision-makers weighing tutoring cost against measured learning gain. The paper is also useful to policy audiences interested in access to individualized instruction, though casual readers should note the study population was adults with English literacy who had platform access and were paid for participation.
Authors’ abstract
Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). The StudentBench platform is freely available at https://studentbench.org.