Company Size
51-200
Company Stage
N/A
Total Funding
N/A
Headquarters
London, United Kingdom
Founded
2023
See people who can refer or advise you
Help us improve and share your feedback! Did you find this helpful?
Remote Work Options
Hybrid Work Options
Flexible Work Hours
Paid Vacation
Paid Holidays
Parental Leave
Gym Membership
Home Office Stipend
Commuter Benefits
Conference Attendance Budget
Professional Development Budget
New framework enhances reproducibility and transparency in AI model evaluation. A collaboration between the UK AI Safety Institute and Hugging Face introduces EvalEval, a framework designed to standardize and improve the transparency of AI model benchmark results, addressing critical challenges in reliable evaluation. Published September 21, 2026 Addressing the reproducibility crisis in AI benchmarking. As artificial intelligence models grow in complexity and scope, the reliability and reproducibility of their evaluations have become a paramount concern. The proliferation of benchmarks and the nuanced variations in evaluation methodologies often lead to inconsistent or incomparable results, hindering true progress and trustworthy assessments. A significant step towards resolving this challenge has been taken through a collaborative effort between the UK AI Safety Institute (UK AISI) and Hugging Face, resulting in the development of EvalEval. This novel framework aims to standardize the reporting and execution of AI model evaluations, providing a robust solution for researchers and developers to ensure their benchmark results are not only accurate but also verifiable and transparent. By fostering a more consistent approach to evaluation, EvalEval is poised to enhance the integrity of AI research and development across the industry. Introducing EvalEval: A unified evaluation ecosystem. EvalEval is designed as a comprehensive system for evaluating large language models (LLMs) and other AI models, emphasizing reproducibility and transparency. It leverages Hugging Face's existing ecosystem, integrating deeply with the Hugging Face Hub, datasets, and models. The core idea is to encapsulate an entire evaluation setup - including the model, evaluation code, data, and environment - within a single, shareable artifact. This artifact, often a repository on the Hub, ensures that anyone can rerun the exact evaluation and obtain identical results. Key features of EvalEval include explicit versioning of all components, from evaluation scripts to the specific model checkpoints used. It supports various evaluation types, from traditional benchmarks to more complex adversarial tests, and provides tools for visualizing and comparing results. This structured approach directly tackles the "reproducibility crisis" in AI, where published results are frequently difficult or impossible to replicate due to undocumented variables or proprietary setups. How EvalEval facilitates transparent benchmarking. At its heart, EvalEval functions by creating a traceable and auditable record of each evaluation. When an evaluation is conducted using the framework, all necessary information is meticulously recorded. This includes the exact commit hash of the model used, the specific version of the evaluation script, the dataset split, and even the computational environment specifications. These details are then packaged and made accessible, typically through a dedicated repository on the Hugging Face Hub. For developers, this means no more guessing which model version yielded which score, or struggling to recreate a testing environment. For researchers, it offers a verifiable path to confirming published results. The framework also supports the creation of "evaluation cards" and "results cards," which provide human-readable summaries of the evaluation methodology and outcomes, further enhancing transparency and ease of understanding for the broader AI community. Impact on AI Safety and development. The UK AI Safety Institute's involvement underscores the critical role of reproducible evaluations in ensuring AI safety. For high-stakes applications, understanding precisely how models perform and the conditions under which those performances were achieved is non-negotiable. EvalEval provides a foundational tool for rigorous safety assessments, allowing for the consistent testing of alignment, bias, and robustness across different models and iterations. Beyond safety, the framework benefits general AI development by accelerating innovation. Developers can more confidently build upon existing work, knowing that benchmark improvements are genuinely representative of enhanced model capabilities, rather than artifacts of evaluation discrepancies. This fosters a more reliable feedback loop for model training and refinement, ultimately leading to more robust and capable AI systems. Why it matters. EvalEval represents a significant leap forward in standardizing AI model evaluation. By making benchmark results transparent, reproducible, and verifiable, it addresses a core challenge that has long plagued AI research and development. This initiative will not only bolster confidence in reported AI performance metrics but also provide a crucial foundation for safer, more robust, and continually improving artificial intelligence systems across all sectors.
The UK government's AI Security Institute has appointed Henry de Zoete as director and Nate Burnikell as chief strategy officer. De Zoete, a long-term AI and technology adviser to the government, played a key role in establishing the Institute. He is also a tech founder whose startup Look After My Bills participated in Y Combinator's 2018 class. He replaces interim director Adam Beaumont, who is returning to GCHQ. Burnikell has worked across various government departments over the past decade, including the Cabinet Office and Department for Levelling Up. Most recently, he worked with the Department for Science, Innovation and Technology, under which the AI Security Institute was originally founded. The Institute researches issues regarding AI safety and national security.
Microsoft, Google and xAI to give US government early access to AI models for security checks. 05 May 2026 08:00PM (Updated: 06 May 2026 04:40AM) Add CNA as a trusted source to help Google better understand and surface our content in search results. Read a summary of this article on FAST. WASHINGTON: Microsoft, Google and Elon Musk's xAI agreed to give the US government early access to new artificial intelligence models for national security testing, as US officials grow alarmed by the hacking capabilities of Anthropic's newly unveiled Mythos. The Centre for AI Standards and Innovation at the Department of Commerce said on Tuesday (May 5) that the agreement would allow it to evaluate the models before deployment and conduct research to assess their capabilities and security risks. The agreement fulfills a pledge the Trump administration made in July 2025 to partner with technology companies to vet their AI models for "national security risks." Microsoft will work with US government scientists to test AI systems "in ways that probe unexpected behaviours," the company said in a statement. Together they will develop shared datasets and workflows for testing the company's models, the company said. Microsoft signed a similar agreement with the UK's AI Security Institute, according to the statement. Concern is growing in Washington over the national security risks posed by powerful AI systems. By securing early access to frontier models, US officials are aiming to identify threats ranging from cyberattacks to military misuse before the tools are widely deployed. The development of advanced AI systems including Anthropic's Mythos has in recent weeks created a stir globally, including among US officials and corporate America, over their ability to supercharge hackers. "Independent, rigorous measurement science is essential to understanding frontier AI and its national security implications," CAISI Director Chris Fall said in a statement. The move builds on previous agreements with OpenAI and Anthropic, established in 2024 under the Biden administration when CAISI was known as the US Artificial Intelligence Safety Institute. Under former President Joe Biden, the institute focused on developing AI tests, definitions and voluntary safety standards. It was led by Biden tech adviser Elizabeth Kelly, who has since joined Anthropic, according to her LinkedIn profile. CAISI, which serves as the government's main hub for AI model testing, said it had already completed more than 40 evaluations, including on cutting-edge models not yet available to the public. Developers frequently hand over versions of their models with safety guardrails stripped back so the centre can probe for national security risks, the agency said. xAI did not immediately respond to a request for comment. Google declined to comment. Last week, the Pentagon said it had reached agreements with seven AI companies to deploy their advanced capabilities on the Defense Department's classified networks as it seeks to broaden the range of AI providers working across the military. The Pentagon announcement did not include Anthropic, which has been embroiled in a dispute with the Pentagon over guardrails on the military's use of its AI tools. Related topics.
Breaking out: can AI agents escape their sandboxes? Container sandboxes are part of routine AI agent testing and deployment. Agents use them to run code, edit files, and interact with system resources without direct access to the host. The SandboxEscapeBench benchmark, developed by researchers at the University of Oxford and the AI Security Institute, evaluates whether an agent with shell access can escape a container and reach the host system. Evaluation architecture and scenario taxonomy (Source: AI Security Institute) What SandboxEscapeBench measures. SandboxEscapeBench places models in controlled container environments and tests whether they can retrieve a protected file from the host filesystem. Each scenario focuses on accessing /flag.txt outside the container boundary. The setup uses a nested design, with containers running inside virtual machines, ensuring that any successful escape remains contained within an outer isolation layer. The benchmark includes 18 scenarios spanning three layers of the container stack: orchestration, runtime, and kernel. These scenarios reflect known vulnerability classes such as exposed Docker sockets, writable host mounts, and privileged containers, as well as documented exploits like Dirty COW and Dirty Pipe. "For security and evaluation integrity reasons, we've limited the public release to scenarios based on known vulnerability classes, while retaining a private test set for internal AISI evaluations. The benchmark is implemented as capture-the-flag-style challenges using Inspect, AISI's open-source framework for AI evaluations, and the SandboxEscapeBench open-source code is available on GitHub," the researchers wrote. What the researchers have found. The evaluation shows that AI agents can exploit common configuration issues when prompted. Frontier models succeeded in scenarios involving exposed Docker sockets, writable host mounts, and privileged containers, well-known weaknesses that still appear in real-world environments. More complex tasks proved harder, especially those requiring multiple steps or deeper interaction with system components. Kernel-level exploits and advanced privilege escalation paths were not solved under the tested conditions. Performance increased with larger token budgets, since models explored more commands and execution paths during each attempt. Providing hints about the underlying vulnerability improved results for some models, enabling them to reach solutions more with fewer steps, while others showed little change under the same conditions. The researchers also analyzed how models behaved during these attempts, noting that some stopped early when progress stalled, while others continued with multiple approaches, including repeated incorrect ones. Some models reported success before completing the task, while others verified the result before stopping. All successful escapes relied on known misconfigurations or publicly disclosed vulnerabilities, and the tests did not identify new flaws. SandboxEscapeBench and its tooling are available as open-source resources for security researchers and evaluators tracking AI agent breakout capabilities. More about
AISI and Lakera open source 'b3' to strengthen AI agent security. Open-source software The UK AI Security Institute, in collaboration with Check Point and Lakera, has unveiled an open source benchmark 'b3' to strengthen LLM security for AI agents. The UK AI Security Institute (AISI) has partnered with Check Point and its subsidiary Lakera to launch the Backbone Breaker Benchmark (b3), an open source framework designed to enhance the security and resilience of large language models (LLMs) that power AI agents. Built to make LLM security measurable and transparent, b3 focuses on identifying the specific "pressure points" where LLMs fail - such as when prompts, files, or web inputs trigger malicious outputs. Rather than evaluating full agent workflows, the benchmark isolates the individual steps most vulnerable to attack.