The Pulse
UK AISI Publishes Evaluation Results Through EvalEval
The UK AI Security Institute is sharing verified evaluation results, methods and configuration information through EvalEval’s open infrastructure.

AI.info Team ·
The UK AI Security Institute is using EvalEval’s infrastructure to openly share evaluation results, giving researchers access to verified findings alongside the context and configuration information needed to interpret them.
The announcement, published on September 22, 2026, marks a new phase in the collaboration between AISI and the EvalEval Coalition. AISI is making publicly reported evaluation methods and findings available through Evaluation Cards where appropriate. The open platform brings evaluation results and the information needed to interpret them into a common structure.
AISI and EvalEval previously collaborated on research that began at a joint workshop alongside NeurIPS 2025. Feedback from the institute also helped shape the Every Eval Ever schema, the shared reporting format behind the project.
“We are excited about this adoption and look forward to further standardising and sharing evaluations with AISI and other AI evaluation organisations.”
— EvalEval Coalition, research community
Five benchmarks, six frontier models
The release includes verified results, context and configuration information for five benchmarks used in the paper’s main experiment: HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0.
The results cover six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. AISI is also sharing results from two related cyber evaluations, Cyber CTFs and The Last Ones, which use a different, partially overlapping set of models.
The data accompany AISI’s paper, How Inference Compute Shapes Frontier LLM Evaluation. The study examines how benchmark performance depends on inference-time compute and the protocol used to run an evaluation.
One example in the release concerns Humanity’s Last Exam. The published material shows performance changing with both evaluation protocol and inference compute. When models received correctness feedback from an oracle after each attempt, they continued to solve additional tasks as token use increased.
Why evaluation setup changes the result
Evaluation results are often reported across different formats, platforms and outlets, sometimes without enough information to reproduce them. Prompting, tool access, token budgets, feedback rules, scoring procedures and the amount of inference compute can all affect what a measured result means.
EvalEval’s platform is intended to preserve those details rather than reduce an evaluation to a single number. Evaluation Cards combine benchmark metadata, evaluation-run data and model metadata, creating records that help users understand whether apparently similar results were produced under meaningfully different conditions.
Repeating a large evaluation can be expensive or impossible. Public configuration data cannot replace a rerun, but it can help researchers examine individual studies more closely, compare protocols and understand how setup choices may influence reported performance.
EvalEval wants a shared record of AI testing
The Every Eval Ever project provides the reporting schema, while Evaluation Cards presents results and surrounding metadata in a browsable form. The coalition is asking model developers to report verified evaluation results, evaluation developers to submit benchmark and run data using the Every Eval Ever schema, and evaluation, governance and policy researchers to explore the cards by benchmark or model.
The arrangement gives AISI’s work a place in a wider collection of evaluations rather than leaving each result inside a standalone paper or institutional report. Researchers can compare results for a benchmark or model across different setups and examine where differences may reflect protocol choices.
For AISI, the release offers a public record of selected evaluation work without claiming that every internal method or result can be published. The announcement says the shared data include the five benchmarks from the paper’s main experiment and two related cyber evaluations.
The immediate result is a set of frontier-model evaluations with surrounding information that allows other researchers and practitioners to inspect the conditions behind the scores. EvalEval says that as more evaluators adopt the Every Eval Ever schema, open comparisons can support broader and more reliable meta-research.