UK AISI Publishes AI Evaluation Results Through EvalEval Cards
AISI is using EvalEval’s Every Eval Ever schema and Evaluation Cards platform to share benchmark methods, context and results for frontier-model evaluations.

The UK AI Security Institute is using EvalEval infrastructure to publish AI evaluation results in a common format, Hugging Face’s blog carried from the EvalEval Coalition, turning an earlier research collaboration into a public reproducibility project.
The release is aimed at a practical problem in model assessment: evaluation results are often scattered across papers, platforms and reporting styles, while rerunning the same tests can be too expensive for many researchers.
EvalEval’s approach puts methods, findings and context into a shared structure so readers can inspect how a result was produced instead of treating a benchmark score as a standalone number.
The collaboration builds on work that began around a joint workshop alongside NeurIPS 2025.
Feedback from AISI helped shape Every Eval Ever, the schema behind EvalEval’s reporting system, and the new phase applies that structure through Evaluation Cards, an open platform for organising evaluation results and the metadata needed to interpret them.
AISI’s initial Evaluation Cards entry packages the benchmark material as a reusable record rather than a narrative-only announcement.
For each of the five main experiment benchmarks, the card set brings together the checked result, the run context and the configuration choices that a reader would need before comparing the finding with another study.
The underlying AISI paper examines the link between inference-time compute, evaluation protocol and frontier large-language-model benchmark outcomes.
The model set named for the main results spans Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4.
Separate cyber benchmark material is also attached for Cyber CTFs and The Last Ones, with model coverage that partly overlaps the main experiment but is not identical.
The source highlights Humanity’s Last Exam as one example of why protocol details matter.
Performance changes when models receive correctness feedback from an oracle after each attempt; in that setup, models continued solving additional tasks as token use increased.
The benchmark record is therefore not only a score, but a description of conditions such as token budget, feedback and evaluation setup.
That level of disclosure supports comparison across a wider evaluation ecosystem.
When other reports omit setup details, openly released Evaluation Cards can act as reference points for understanding how model results may differ because of experimental choices rather than capability alone.
AISI’s adoption also connects to its existing evaluation infrastructure work, including OptStop for efficiency, HiBayES for statistical rigor, and standardisation efforts in transcript analysis and capability elicitation.
EvalEval frames the shared infrastructure as a way to diagnose gaps in evaluation reporting and make model assessments easier to reproduce, verify and compare.
The EvalEval Coalition says its broader projects connect benchmark descriptions, individual run details and model information in records that readers can interpret together.
The next test for the effort is whether more evaluation organisations adopt the Every Eval Ever schema, giving researchers a larger base for meta-research on how advanced AI systems are measured.




















