UK AISI Shares Evaluation Results via EvalEval Infrastructure
The UK AI Security Institute is publishing verified benchmark results for frontier models using EvalEval's Every Eval Ever schema and Evaluation Cards, advancing reproducible…

- AISI is releasing results, context and configuration data for HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro and Terminal-Bench 2.0 covering Claude Opus 4, 4.5, 4.6 plus GPT-5, 5.2 and 5.4 from the paper's main experiment.
- The effort builds on earlier AISI-EvalEval collaboration from a NeurIPS 2025 workshop that informed the Every Eval Ever schema and complements AISI tools such as OptStop and HiBayES.
- Open sharing of transcript-level details supports analysis of how inference compute and evaluation protocols influence performance, offering reference points for the broader evaluation ecosystem.
Advancing Reproducible Evaluations
Evaluation results for advanced AI systems are often shared in inconsistent formats that omit critical details needed for reproduction. Replicating expensive runs can be impractical. The EvalEval Coalition created the Every Eval Ever schema and Evaluation Cards to standardize benchmark metadata, run data and model information.
This September 2026 announcement details how the UK AI Security Institute is applying that infrastructure to make its publicly reported methods and findings available through Evaluation Cards where appropriate. The work extends prior joint research begun at a NeurIPS 2025 workshop and aligns with AISI's projects on evaluation efficiency, statistical rigor and standardization.
Released Benchmarks and Models
The shared data include verified results for the five benchmarks featured in the main experiment of the paper "How Inference Compute Shapes Frontier LLM Evaluation": HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro and Terminal-Bench 2.0.
These cover six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. Separate results from two cyber evaluations, Cyber CTFs and The Last Ones, use a partially overlapping set of models. The paper examines how benchmark scores vary with inference-time compute and protocol choices.
Insights from Open Data
The source illustrates how performance on Humanity's Last Exam changes with evaluation protocol and compute budget, presenting cumulative shares of tasks solved under different conditions including oracle feedback. Such examples highlight the value of detailed reporting.
Transcript-level transparency enables closer study and cross-ecosystem comparison. Releases with full setup information provide verified reference points, especially when other reports lack comparable context. Wider adoption of the schema could strengthen meta-research on evaluation practices.
Broader Context and Next Steps
The EvalEval Coalition continues developing tools to improve evaluation science and documentation of applicability. AISI's mission focuses on scientific understanding of advanced AI risks to support government policy.
The announcement encourages contributions from model developers, benchmark creators and researchers to expand standardized sharing. Further reading includes the AISI paper and related technical reports on OptStop, HiBayES and the Every Eval Ever schema.


