The Pulse
Epoch AI Labels 9 of 15 AI Benchmarks Flawed
Epoch AI has reviewed 15 external AI benchmarks and found substantive problems in nine of them. Its audit flags scoring errors, weak evaluation setups, inconsistent versions and conditions that can distort model comparisons.

AI.info Team ·
Epoch AI has classified nine of the 15 external AI benchmarks it reviewed as flawed, exposing substantial weaknesses in several widely used tests for measuring model capability. The review list, updated through September 12, includes SWE-bench Verified, Humanity’s Last Exam, Terminal-Bench 4.0.0, DeepSWE v1.1 and the Berkeley Function Calling Leaderboard v4.
The finding does not mean that every score from those benchmarks is worthless. Epoch defines a flawed benchmark as one with at least one serious defect that affects how results should be interpreted. The organization says its reviews concern specific benchmark versions, and a later release can be considered separately if developers correct the problems.
Epoch published its review framework and results in its Benchmark Reviews Documentation. The initial set contains four benchmarks labeled Verified, nine labeled Flawed and two classified as Not enough information.
Fifteen reviews, nine negative verdicts
The flawed group spans software engineering, health, general knowledge, writing and tool use. Epoch marked Berkeley Function Calling Leaderboard v4 and HealthBench Professional as flawed on September 10, followed by DeepSWE v1.1 on September 7 and Terminal-Bench 4.0.0 on September 4.
Earlier reviews flagged SWE-bench Verified and SWE-Bench Pro, two prominent tests of software-engineering performance. Humanity’s Last Exam, Lech Mazur Writing and TextQuests also received flawed verdicts. The two remaining entries, CritPt and FrontierCode, were not judged because Epoch said it did not have enough information to support either a Verified or Flawed label.
Four benchmarks passed the organization’s minimum standard: ExploitBench v0.1, SimpleQA Verified, PostTrainBench v1.1 and WeirdML v2. Epoch’s published table records the review dates and verdicts for all 15 benchmarks.
What Epoch counts as a serious defect
Epoch’s rubric separates whether a benchmark can be inspected from whether its scoring produces a meaningful result. A review can proceed when all tasks and scoring logic are available, or when a representative sample can be examined and the evaluation settings are fully disclosed. If those conditions are not met, the benchmark receives a Not enough information verdict rather than a quality judgment.
For benchmarks that can be reviewed, Epoch looks for errors that change scores or undermine the stated capability being measured. Its examples include tasks that are impossible to answer as written, overly strict or incorrect scorers that produce false negatives, lax scoring that permits reward hacking, and shortcuts that allow a model to retrieve an answer from the evaluation environment.
The standard also covers benchmark consistency, model elicitation and evaluation bias. Results can fail the consistency test when the scorer, instructions or ground truth change without a version update. Epoch also examines whether token budgets, time limits, tool access, sandboxes and scaffolds give models a fair chance to perform.
The 20 percent threshold
Epoch’s default scoring threshold is quantitative: at least 20 percent of an inspected sample must contain errors, or a single issue must corrupt grading at scale. For benchmarks with more than 50 tasks, reviewers sample 50 questions, expanding to 100 when the observed error rate falls between 15 and 25 percent. Benchmarks with 50 or fewer available tasks are assessed in full.
The threshold applies to the minimum standard for a Verified label. A benchmark can still carry limitations without being classified as flawed, but a defect that crosses the threshold ends the review with a Flawed verdict. Epoch says it stops once it has identified sufficient errors and does not claim that its first review has found every possible problem.
Epoch’s labels are version-specific. A benchmark developer can publish a corrected release, but the label attached to the version reviewed by Epoch does not automatically change.
Why benchmark conditions matter
Epoch’s documentation places particular emphasis on the conditions surrounding a test, not only on its questions. A model’s result can change when developers alter the scaffold, resource limits, system prompts, tool access or termination rules. The organization cites earlier analysis estimating that changing the scaffold produced differences of up to 11 percent for GPT-5 and up to 15 percent for Kimi K2 Thinking.
Those conditions complicate comparisons between model developers. A leaderboard may appear to rank models by capability while also reflecting different prompts, agent frameworks, token budgets or environmental permissions. When the benchmark does not disclose those settings, outside researchers have less ability to reproduce or interpret the result.
Epoch also warns that agentic environments can create reward-hacking opportunities. A model may achieve a high score through a shortcut, exploit or environment-level weakness rather than by demonstrating the capability the benchmark claims to measure.
A review system built for public scrutiny
Epoch says it selected the initial sample to cover different error types and research domains. It plans to prioritize benchmarks with broad reach, high citation counts, inclusion in recent system cards, safety relevance or limited coverage in its existing reviews. Tests that require specialist knowledge may take longer because the organization expects to bring in outside subject-matter experts.
The organization does not review benchmarks it created because of conflicts of interest. It also says it contacts benchmark developers at least one day before publication and will publish a developer response in full if the developer wants one included.
The immediate result is a more qualified reading of benchmark leaderboards. Nine of the 15 reviewed tests contain problems Epoch considers serious enough to affect interpretation, while two others cannot yet be assessed from the available information. For anyone comparing models, the benchmark name alone is no longer enough: the version, scoring rules and evaluation conditions now matter just as much as the final percentage.