Skip to content
AI.info

The Pulse

Pangram Puts Colleges in an AI-Detection Arms Race

About 1,400 North American colleges have purchased Turnitin’s AI detector, while Pangram is gaining attention for its low reported false-positive rate. Professors and universities are split between expanding detection tools and redesigning

Pangram Puts Colleges in an AI-Detection Arms Race

AI.info Team ·

About 1,400 colleges and universities in North America have purchased Turnitin’s AI-detection tool, and the company says it found at least some AI writing in nearly half of student submissions at U.S. universities during the past academic year. The numbers show how quickly automated authorship checks have moved from experimental software to routine campus infrastructure. They also explain why professors now face a conflict between tools that promise evidence and research that warns against treating a detector’s score as proof.

The shift is unfolding around Pangram, a newer detector that has become a reference point in arguments over whether AI-written prose can be identified reliably. Turnitin says its detector produces false positives in less than 1 percent of cases, but also misses about 15 percent of AI-written passages. At a university processing tens of thousands of papers, even a small error rate can put hundreds of students into disciplinary conversations.

The Atlantic reported the figures and interviewed professors, detection companies and university officials as colleges reassess how to police AI-assisted writing.

Turnitin’s reach makes small error rates consequential

Turnitin’s AI checker arrived as an add-on to its widely used plagiarism-detection service. Annie Chechitelli, Turnitin’s chief product officer, told The Atlantic that roughly 1,400 North American institutions have bought the detection product. The company says the tool is designed to give students the benefit of the doubt, which helps explain its reported false-negative rate but does not eliminate the consequences of false positives.

Vanderbilt University illustrated that problem when it disabled Turnitin’s AI checker in 2023. The university estimated that a 1 percent false-positive rate applied to roughly 75,000 student papers could wrongly flag as many as 750 submissions. A detector score cannot establish who wrote a paper, but it can still shape the first conversation between a student and an instructor.

Timothy Paustian, a biology professor at the University of Wisconsin at Madison, tested early detectors after ChatGPT became publicly available. He found them “comically bad,” according to The Atlantic, and later tried placing invisible prompts in assignments that chatbots might answer differently from students.

Pangram claims a narrower margin of error

Pangram has attracted attention by reporting a far lower false-positive rate than older tools. The company claims that internal testing produced one false positive for every 25,000 cases. A 2025 working paper from researchers at the University of Chicago found that Pangram almost never misclassified long passages of human writing, although the researchers found that shorter samples could still produce errors.

The University of Chicago now makes Pangram available to College instructors, but its policy treats the result as supporting information rather than a verdict. The university says instructors and its Office of College Community Standards should use the software within a transparent, human-led process that can include questions about a student’s subject knowledge, drafting history, research process, editing tools and document version history.

Pangram’s own product materials say the detector analyzes patterns in writing produced by models including ChatGPT, Gemini, Grok, Llama and Claude. The company says it benchmarks against 26 models and reports 99.7 percent accuracy on its tested AI-generated samples. Those figures come from Pangram’s own testing, while the University of Chicago research examined detector performance under its own study design.

Students can test the detector as they revise

The central problem is that Pangram is available to students as well as instructors. A student can submit a draft, see the score, revise the prose and submit it again. That turns the detector into a feedback mechanism for both sides of the dispute: professors use it to identify suspicious work, while students can use it to make AI-generated writing less detectable.

Marc Watkins, a lecturer at the University of Mississippi who studies AI’s effect on education, described the result bluntly: “Basically it’s an arms race.” The Atlantic reported that students can keep modifying AI-generated essays until Pangram returns a result of “100-percent human,” while Turnitin’s closed system is harder to probe through repeated trial and error.

The competition has also created a market for humanizer tools that rewrite AI-generated prose. Research published in 2026 found that adversarial rewriting can reduce detector performance without necessarily producing text that readers would recognize as machine-written. The technical contest therefore shifts whenever a detector improves: one side updates its classifier, and the other changes the text it is trying to conceal.

Faculty are caught between suspicion and policy

Many instructors remain skeptical of detectors but lack an easy replacement. “I think most faculty feel entirely overwhelmed by this,” Watkins told The Atlantic. “It’s a mess.” That pressure is strongest in large courses where instructors must evaluate hundreds of papers and cannot reconstruct every student’s drafting process from memory.

Some universities are choosing not to expand detection. An MIT working group recently recommended against relying on AI detectors, warning that they could create “an atmosphere of distrust between instructors and students.” Indiana University’s Kelley School of Business has barred professors from using AI detectors and instead advises them to design assignments around process, reasoning and authentic engagement.

The policy debate is also moving toward conditional acceptance of AI. Indiana’s guidance says students are already expected to use AI responsibly in their careers. A Harvard dean recently urged the college to leave the “AI-detection business” and consider accepting or encouraging AI use in writing-intensive courses, according to The Atlantic.

The score cannot settle authorship

The dispute over Pangram shows why detection results are becoming evidence in a conversation rather than a final answer. A low false-positive rate may make a tool more useful, but it does not resolve questions about AI-assisted editing, mixed authorship or a student’s permitted use of software. A paper can contain a student’s ideas, an AI-assisted revision and a detector score that describes only one part of that process.

University of Chicago’s policy reflects that distinction. Instructors may use Pangram, but they are directed to examine the student’s command of the subject, research records, outside assistance and writing history. That approach puts more work on faculty while limiting the chance that a percentage becomes a punishment by itself.

For colleges, the immediate choice is not simply whether to buy a detector. It is whether automated scores should govern discipline, prompt a human review or push courses toward assignments that make drafting visible. As Pangram and its competitors improve, the campus arms race is producing better detection tools and better evasion tools at the same time.

Source

The Atlantic

Explore

More articles