Skip to content
AI.info

Industry Transformation

AI in Education: Tutors, Cheating, and What the Evidence Actually Shows

Bloom's two-sigma tutoring claim, Khanmigo and Duolingo Max, Anthropic's July 2026 free-teacher launch, Turnitin's bias, blue books' return, and what RCTs in Nigeria, Harvard, and Turkey actually found about AI tutoring.

AI in Education: Tutors, Cheating, and What the Evidence Actually Shows

Gabriele Masetti ·

The Two-Sigma Promise, and Why It's Been Oversold

Every pitch for an AI tutor eventually cites Benjamin Bloom. In a 1984 paper, the University of Chicago education researcher described dissertation work by his students Joanne Anania and Joseph Burke showing that one-on-one tutoring combined with mastery learning pushed the average tutored student two standard deviations above students in a conventional 30-person classroom — better than 98% of their peers. Bloom called it "the two-sigma problem": if tutoring works this well, why can't schools deliver it to everyone?

Tutoring effect sizes: Bloom's famous claim versus what meta-analyses actually measured.

The number is real, but the field has spent four decades failing to reproduce it. A 1982 meta-analysis by Peter Cohen, James Kulik, and Chen-Lin Kulik put the average tutoring effect closer to 0.33 standard deviations — about 13 percentile points. A 2020 meta-analysis of 96 randomized tutoring studies found an average effect of 0.37 standard deviations, roughly 14 percentile points, and not a single study in the set reached two full sigmas.

Bloom's number describes an upper bound under near-ideal conditions, not a typical outcome. That gap between the founding legend and the measured reality is the right frame for everything AI-tutoring companies now claim, and for the backlash against AI-assisted cheating that has followed them into schools.

The Products Already in Classrooms

Khan Academy's Khanmigo, built on OpenAI models, costs individual users $4 a month or $44 a year, with one subscription covering up to ten children in a household. Since a Microsoft-brokered deal announced in May 2024 gave Khan Academy access to Azure OpenAI Service, every verified US K-12 teacher gets Khanmigo free. Districts pay separately: base reporting tools start around $5 per student per year, an "Enterprise Starter" tier runs $10, and a student-facing tutoring add-on runs about $15 per student annually. Palm Beach County's rollout reached all 58 high schools in a district of more than 179,000 students, at a cost of roughly $1.2 million.

Duolingo Max, launched in 2023 and originally powered by GPT-4, sits above the standard Super subscription at $29.99 a month or $168 a year, adding a live "Video Call" AI conversation partner and scripted "Roleplay" scenarios. Its differentiation narrowed sharply in January 2026, when "Explain My Answer" — previously Max's signature grammar-explanation feature — became free for every Duolingo user, leaving the paid tier to justify itself mainly through conversation practice for exam prep like TOEFL or IELTS.

Universities have moved faster and spent more. Arizona State became OpenAI's first higher-education partner in January 2024, initially rolling out ChatGPT Enterprise; by 2025 the partnership expanded so that every ASU student, faculty member, and staff member got ChatGPT Edu at no individual cost, on OpenAI's current models, supporting nearly 700 projects through the university's AI Innovation Challenge.

California State University went bigger still: a systemwide contract reported at roughly $17 million, since renewed at $13 million a year for three years, giving ChatGPT Edu to more than 470,000 students and 63,000 faculty and staff — OpenAI's largest higher-ed deal. The rollout has not been smooth.

A CSU-wide survey of over 94,000 students and staff found 65% of students and 59% of faculty skeptical that AI has actually benefited education, 80% of students uncomfortable submitting AI-generated work as their own, and 52% of faculty reporting a negative effect on their teaching.

Anthropic Enters K-12, and Doubles Down on Higher Ed

Anthropic's push into education split into two moves a year apart. In April 2025, Claude for Education launched for universities, with the London School of Economics, Northeastern, and Champlain College as first campus-wide partners and free access for any US .edu Pro user. Its centerpiece, Learning mode, is built to walk students through reasoning steps via Socratic-style prompting rather than hand them finished answers — a direct response to the "answer machine" critique of chatbot tutoring.

On July 14, 2026, Anthropic launched Claude for Teachers, giving verified US K-12 educators free access to premium Claude features, a library of teaching skills, and lesson-planning tools tied to academic standards in all 50 states through a partner called Learning Commons. Teachers who sign up by June 30, 2027 lock in a full year of free access. Anthropic developed FERPA-aligned data terms jointly with the American Federation of Teachers, and lined up partnerships with the Gates Foundation, Teach For America, and a pilot in Detroit Public Schools.

Coverage in Chalkbeat and Education Week framed the launch as the clearest sign yet that AI companies see teachers, not just students, as the entry point into American classrooms — and some critics quoted in that coverage worried about a handful of vendors shaping curricula nationwide through free access rather than public procurement.

The Cheating Panic, By the Numbers

The anxiety driving all this positioning is not imagined. Pew Research Center found that the share of US teens who say they've used ChatGPT for schoolwork rose from 13% in 2023 to 26% by late 2024, with Black and Hispanic teens (31% each) more likely to report use than white teens (22%). Teens themselves draw sharp lines: 54% say it's fine to use a chatbot for researching a topic, but only 18% think it's acceptable for writing essays.

Faculty perceive the shift as worse than students do. A July 2024 survey covered by Inside Higher Ed found 68% of instructors expected generative AI to hurt academic integrity, while 47% of more than 2,000 surveyed students admitted cheating had gotten easier. An Elon University and American Association of Colleges and Universities survey released in January 2026 found 78% of faculty believe AI-driven cheating is rising, with 57% saying it has risen "a lot."

A separate poll of higher-education leaders found 59% believed cheating had increased since generative AI became widely available, but only 11% of chief technology officers said their institution had anything resembling a comprehensive AI strategy to respond.

Why Detection Tools Can't Be Trusted

Schools' first instinct was to fight software with software, and that instinct has largely backfired. Turnitin advertises under a 1% false-positive rate for documents containing more than 20% AI-generated text. Independent testing tells a different story for the population most likely to be wrongly accused: a 2023 Stanford study (Liang, Yuksekgonul, Mao, Wu, and Zou, published in Patterns) ran seven popular GPT detectors against TOEFL essays written by non-native English speakers and found an average false-positive rate of 61.3%.

One detector alone flagged 97.8% of those essays as AI-generated, and all seven detectors unanimously misclassified 19.8% of them, while the same tools performed close to flawlessly on essays from US eighth-graders. The mechanism is straightforward: non-native writers tend to use simpler, more predictable sentence structures, which is exactly the low-"perplexity" pattern detectors are trained to flag as machine-written.

Group flagged by AI detectors False-positive rate
Non-native English speakers (TOEFL essays, Stanford 2023 study) 61.3% (average)
Black students (Common Sense Media, 2024) ~20%
White students ~7%
Latino students ~10%

The bias isn't limited to language background. Common Sense Media surveyed roughly 1,045 teens and parents in 2024 and found Black students' work was falsely flagged as AI-generated about 20% of the time, compared with 7% for white students and 10% for Latino students — a gap researchers tied partly to detectors misreading African American Vernacular English as a statistical anomaly. Put those two findings together and the tools schools bought to enforce integrity turn out to systematically punish the students least responsible for the underlying behavior.

Blue Books Are Back

Faced with detectors that don't work and chatbots that do, instructors reached for a solution older than the personal computer: the handwritten exam. Reporting from Axios, U.S. News, and the trade press describes blue-book sales climbing as much as 80% at some universities, with Texas A&M, the University of Florida, and UC Berkeley among the campuses citing surging demand for in-class, pen-and-paper testing.

The logic is simple — a student writing by hand in a proctored room has no path to a chatbot mid-exam — and some instructors report the forced slowdown improves the visible quality of student reasoning. Critics counter that timed handwriting disadvantages students with disabilities and multilingual writers who process information differently, and that determined cheaters will find workarounds regardless of exam format.

What the Randomized Trials Actually Show

Strip away the marketing and the panic, and the controlled research tells a more precise, more useful story — one where scaffolding is the variable that decides everything.

A World Bank-supported randomized controlled trial in Benin City, Nigeria gave roughly 800 secondary students six weeks of after-school tutoring — 12 sessions, twice a week — using Microsoft Copilot running on GPT-4. The intervention lifted English scores by 0.23 standard deviations at a cost of about $48 per student, a result the researchers say outperformed 80% of rigorously evaluated education programs and delivered gains comparable to roughly two years of typical schooling.

A Harvard RCT led by Gregory Kestin, published in Scientific Reports in June 2025, put 194 introductory physics students through identical material delivered either by a purpose-built AI tutor or an active-learning classroom session. The AI-tutored group's estimated effect size landed between 0.73 and 1.3 standard deviations over the classroom condition (p < 0.00000001), finishing in a median of 49 minutes against roughly 60 for the in-class group, while reporting higher engagement and motivation. The tutor was deliberately engineered around known learning science — active questioning, small steps, scaffolded hints, timely feedback — rather than left as an open chat window.

That engineering detail is the crux of the field's most sobering result. Hamsa Bastani, Osbert Bastani, and colleagues ran an RCT across ninth-, tenth-, and eleventh-grade math classes in a Turkish high school, published in PNAS in 2025 as "Generative AI without guardrails can harm learning." Students given open GPT-4 access during practice sessions improved practice scores by 48% to 127%.

Once the tool was removed for unassisted exams, those same students scored 17% worse than a no-AI control group — the chatbot had let them get right answers without building the underlying skill. A second version of the tool, restricted to Socratic hints designed with teacher input rather than direct answers, avoided that harm entirely.

Stanford's National Student Support Accelerator ran a complementary RCT with 900 tutors and 1,800 K-12 students in a Southern school district, testing a real-time coaching tool called Tutor CoPilot that suggests better questions and explanations to human tutors mid-session. It did not touch students directly — it made adult tutors better. Students working with the district's lowest-rated tutors saw their pass rate jump from 56% to 65%, nearly closing the gap with the 66% pass rate students got from top-rated tutors.

Study Result
Benin City World Bank RCT (Microsoft Copilot/GPT-4 tutoring) English scores +0.23 SD, ~$48/student
Harvard physics RCT (Kestin, 2025) Effect size 0.73-1.3 SD over classroom
Turkish PNAS RCT (Bastani et al.) Practice scores +48% to +127%; unassisted exam scores -17% vs control
Stanford Tutor CoPilot RCT Lowest-rated tutors' pass rate: 56% -> 65%

The Clearer Win, and What Comes Next

Line those studies up and a pattern holds across every one: AI tutoring built with scaffolding — hints instead of answers, active questioning, a human in the loop — produces real, often large learning gains. AI access with no structure, handed to a novice as an answer key, produces short-term score inflation and a longer-term skills deficit once the tool disappears. The Bastani team said as much explicitly: restricting the model to teacher-approved hints was what separated a helpful tutor from a crutch.

The steadier payoff so far is on the adult side of the classroom. Gallup and the Walton Family Foundation surveyed 2,232 US public school teachers between March 18 and April 11, 2025, and found 6 in 10 had used AI tools during the school year; those using AI weekly reported saving about six hours a week on lesson planning, grading, and materials creation — roughly six weeks of time across a 37.4-week school year.

UNESCO's 2023 Guidance for Generative AI in Education and Research, still the reference framework cited by ministries drafting national policy, recommends a minimum age for unsupervised chatbot use and mandatory data-privacy protections, alongside its 2024 AI competency framework for students built around human-centered thinking and ethics rather than tool fluency alone.

None of this resolves into a tidy verdict on "AI in schools" as a single thing, because it isn't one. It's a scaffolded Socratic tutor in a Harvard physics course, an unrestricted chat window in a Turkish algebra class, and a $13-million-a-year enterprise contract at Cal State, and the evidence treats those three cases completely differently. The policy fights over blue books and detection software are, in effect, arguments about which of those three a given classroom is going to become.

Teachers who sign up for Anthropic's free program have until June 30, 2027 to lock in their year of access — the next real test of which model wins is already running.

Explore

More articles