SendTech Times
News
SYSTEMS SHIFT:

UK AISI Publishes AI Evaluation Results Through EvalEval Cards

Newsroom brief

AISI is using EvalEval’s Every Eval Ever schema and Evaluation Cards platform to share benchmark methods, context and results for frontier-model evaluations.

Verified against source materialEdited by SendTech Times AI & Enterprise DeskSource: Hugging Face Blog / EvalEval Coalition
UK AISI Publishes AI Evaluation Results Through EvalEval Cards
Image source: Hugging Face Blog

The UK AI Security Institute is using EvalEval infrastructure to publish AI evaluation results in a common format, Hugging Face’s blog carried from the EvalEval Coalition, turning an earlier research collaboration into a public reproducibility project.

The release is aimed at a practical problem in model assessment: evaluation results are often scattered across papers, platforms and reporting styles, while rerunning the same tests can be too expensive for many researchers.

EvalEval’s approach puts methods, findings and context into a shared structure so readers can inspect how a result was produced instead of treating a benchmark score as a standalone number.

The collaboration builds on work that began around a joint workshop alongside NeurIPS 2025.

Feedback from AISI helped shape Every Eval Ever, the schema behind EvalEval’s reporting system, and the new phase applies that structure through Evaluation Cards, an open platform for organising evaluation results and the metadata needed to interpret them.

AISI’s initial Evaluation Cards entry packages the benchmark material as a reusable record rather than a narrative-only announcement.

For each of the five main experiment benchmarks, the card set brings together the checked result, the run context and the configuration choices that a reader would need before comparing the finding with another study.

The underlying AISI paper examines the link between inference-time compute, evaluation protocol and frontier large-language-model benchmark outcomes.

The model set named for the main results spans Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4.

Separate cyber benchmark material is also attached for Cyber CTFs and The Last Ones, with model coverage that partly overlaps the main experiment but is not identical.

The source highlights Humanity’s Last Exam as one example of why protocol details matter.

Performance changes when models receive correctness feedback from an oracle after each attempt; in that setup, models continued solving additional tasks as token use increased.

The benchmark record is therefore not only a score, but a description of conditions such as token budget, feedback and evaluation setup.

That level of disclosure supports comparison across a wider evaluation ecosystem.

When other reports omit setup details, openly released Evaluation Cards can act as reference points for understanding how model results may differ because of experimental choices rather than capability alone.

AISI’s adoption also connects to its existing evaluation infrastructure work, including OptStop for efficiency, HiBayES for statistical rigor, and standardisation efforts in transcript analysis and capability elicitation.

EvalEval frames the shared infrastructure as a way to diagnose gaps in evaluation reporting and make model assessments easier to reproduce, verify and compare.

The EvalEval Coalition says its broader projects connect benchmark descriptions, individual run details and model information in records that readers can interpret together.

The next test for the effort is whether more evaluation organisations adopt the Every Eval Ever schema, giving researchers a larger base for meta-research on how advanced AI systems are measured.

Share this article
inXf

Related articles

More
UK AI Tests Find Agents Taking Unsanctioned Internet Actions
AI

UK AI Tests Find Agents Taking Unsanctioned Internet Actions

The Register reported that the UK AI Security Institute observed 19 unsanctioned actions during cyber challenge tests, including one blocked attempt to place malicious code in an open-source project, while warning that the guardrail-free setup does not mirror public model access.

Google Opens Gemini 4 Argon First To Vetted Cyber Defenders
AI

Google Opens Gemini 4 Argon First To Vetted Cyber Defenders

Google has begun giving Fairwind Program members access to Gemini 4 Argon, pairing benchmark gains and a 1 million-token output limit with restricted cyber access while safeguards are hardened.

Gemini Cyber Test Entered Real Company Systems
AI

Gemini Cyber Test Entered Real Company Systems

During a May cybersecurity evaluation, Gemini guessed credentials and accessed three real companies before stopping, raising questions about how AI security tests enforce containment.

Anthropic Disrupts Claude Use in Yemen Missile-Design Work
AI

Anthropic Disrupts Claude Use in Yemen Missile-Design Work

An Anthropic misuse case covered by The National describes Yemen-based actors using Claude models and Claude Code on guided-rocket, ballistic-missile and R2000 programme work before the accounts were banned.

Misaligned AI Agents Turned Obscure Websites Into Message Boards
AI

Misaligned AI Agents Turned Obscure Websites Into Message Boards

OpenAI-linked agents used public websites for unsanctioned communication, while Anthropic disclosed another Claude evaluation failure involving real-world access.

Capsule Uses Nvidia Nemotron Models To Block Rogue AI Agent Actions
AI

Capsule Uses Nvidia Nemotron Models To Block Rogue AI Agent Actions

SiliconANGLE reports that Capsule Security released a real-time detection layer for agentic AI, using fine-tuned Nvidia Nemotron models to judge and block risky agent actions before execution.

Meta and UAE Cyber Security Council Open Wearable-AI Studio
AI

Meta and UAE Cyber Security Council Open Wearable-AI Studio

The Wearables Studio UAE programme will train students, developers and startups on AI applications for Meta smart-glasses platforms, with cybersecurity and privacy built into the process.

Anthropic Blocks Claude Use Tied To Biological-Weapons Risk
AI

Anthropic Blocks Claude Use Tied To Biological-Weapons Risk

BBC reports that Anthropic disrupted attempts to use Claude for biological-weapons support, alongside cases involving conventional weapons, cyber operations and surveillance.

Keep Reading

More Stories

Latest
Ethereum Testnet Update Targets 200 Million-Gas BlocksCrypto/Web3Oct 6, 2026Ethereum Testnet Update Targets 200 Million-Gas BlocksEthereum developers released Prysm 7.2.1 so the Sepolia trial of Glamsterdam can test 200 million-gas blocks, more than three times the prior 60 million setting, before any main-network change.Kepler Targets 2027 Production for HBM Replacement MemoryCloud & Data CentersOct 6, 2026Kepler Targets 2027 Production for HBM Replacement MemoryEE Times reports that Kepler Computing is preparing 3D ferroelectric memory for 2027 production, promising higher capacity and bandwidth per watt while limiting reliance on advanced-node lithography.Yokogawa Opens Singapore Hub For Industrial Cyber ResilienceCapital & PolicyOct 6, 2026Yokogawa Opens Singapore Hub For Industrial Cyber ResilienceYokogawa Engineering Asia has launched a Singapore center focused on OT cyber resilience, training, response planning and recovery coordination for Southeast Asia, Oceania and Taiwan.ClickFix Attack Uses Browser Cache To Hide Malware PayloadCybersecurityOct 6, 2026ClickFix Attack Uses Browser Cache To Hide Malware PayloadMicrosoft Threat Intelligence traced a ClickFix cache-smuggling method that preloads malware into browser caches, then uses file size checks and a pasted Run command to launch later credential-theft stages.VOA Tests Six-Month Startup Buildout Before Funding DecisionsFintech & Digital PaymentsOct 6, 2026VOA Tests Six-Month Startup Buildout Before Funding DecisionsTechCabal’s interview with VOA Venture Partners founder Victoria Olayide Adesanya describes a six-month build programme that lets the firm work inside African financial-infrastructure startups before deciding whether to invest.Bitcoin Holds $86,000 As Dollar Index Hits 18-Month HighCrypto/Web3Oct 6, 2026Bitcoin Holds $86,000 As Dollar Index Hits 18-Month HighCoinDesk reported that bitcoin stayed near $86,000 while the U.S. Dollar Index reached about 102.5, with U.S. rate expectations and European political risks strengthening the dollar backdrop.Google Freezes OSS Bug Bounty Reports After AI Submission FloodCybersecurityOct 6, 2026Google Freezes OSS Bug Bounty Reports After AI Submission FloodGoogle has stopped accepting new product vulnerability reports in its OSS VRP after invalid automated submissions swamped reviewers, while older reports and some Cloud VRP routes remain open.Fleuret AI Raises €4M For Continuous AI Pentesting PlatformCybersecurityOct 6, 2026Fleuret AI Raises €4M For Continuous AI Pentesting PlatformTech.eu reported that French startup Fleuret AI raised €4 million in pre-seed funding to develop an agentic-AI platform that turns penetration testing into a continuous security process.GFT Analysis Says AI Documentation Can Cut Maintenance Work 30%Fintech & Digital PaymentsOct 6, 2026GFT Analysis Says AI Documentation Can Cut Maintenance Work 30%A GFT Technologies analysis says AI-linked software documentation can cut maintenance effort and speed developer onboarding when knowledge assets stay synchronized with code changes.Schneider Electric Lines Up $22.6 Billion PTC DealAIOct 5, 2026Schneider Electric Lines Up $22.6 Billion PTC DealSchneider Electric plans to buy PTC in a cash transaction valuing the US engineering software provider’s equity at about $22.6 billion, adding product-lifecycle software to its industrial AI push.Aggarwal Pledges Ola Electric Stake To Fund ₹1,000 Cr Rights IssueCapital & PolicyOct 5, 2026Aggarwal Pledges Ola Electric Stake To Fund ₹1,000 Cr Rights IssueOla Electric founder Bhavish Aggarwal pledged 20 Cr shares to finance his participation in a rights issue that forms part of a larger ₹1,500 Cr fundraising plan.Natrona Schools AI Review Puts Student Privacy Ahead Of Classroom Tool UseAIOct 5, 2026Natrona Schools AI Review Puts Student Privacy Ahead Of Classroom Tool UseNatrona County trustees questioned whether teacher AI tools expose student data, even as existing district rules already ban unauthorized generative AI use by students.