AI Policy Accelerator Projects
Through our 2026 AI Policy Accelerator, our teams produced the following projects.
​Building the UK’s AI Assurance Infrastructure
The UK cannot compete with the US, China, or EU on frontier compute or regulatory market power, but it can build a cost-effective advantage at the "deployment layer" by connecting existing institutions (NPL, BSI, UKAS, AI Standards Hub) into an Independent Verification Organisation ecosystem that certifies how AI systems perform once deployed. The paper recommends five low-cost, largely legislation-free steps—starting with formalising the AI Standards Hub's role and building toward mandatory incident reporting—to establish this infrastructure before the EU and US-based standards efforts lock in the global benchmark first.
Evaluation Dependence in Frontier AI Safety Frameworks
Anthropic’s April 2026 decision to withhold Claude Mythos Preview reveals that frontier AI safety frameworks are entirely evaluation-dependent, but the actual gating decision was made by senior human judgement outside the formal risk framework. Evaluations themselves are demonstrably unreliable, with scores swinging 5-20 points from scaffolding changes alone, and sandbagging detection missing up to 36% of covert cases. The paper argues frontier AI Governance needs defense-in-depth: multiple independent safety mechanisms that don’t all fail on the same input, since evaluations currently have no fallback when they themselves fail.
AI Proliferation and Middle Powers: Preparation and Response Mechanisms
As AI access broadens, middle powers (unlike the US and China, which dominate frontier development) face proliferation risks (cyber-attacks, CBRN threats, influence operations) they have little control over, making national-level preparedness a more strategic priority than contesting the capability imbalance. The paper offers a three-part framework to help middle powers detect, escalate, and respond to these risks, based on their capacity, governance, and infrastructure posture.
Could Underwater Data Centers Pose a Risk to AI Treaty Verification?
Underwater data centers are technically appealing for treaty evasion but face major construction and maintenance hurdles that make them unlikely to be used, though buildout is detectable while an operational one might evade detection long-term. This non-zero risk still warrants developing detection capabilities now.
Near-Term Feasibility of Monitoring and Verification Methods in the Compute Governance Of Artificial Intelligence
Since compute is the most tractable lever for regulating frontier AI (as opposed to data or algorithms), this paper evaluates six proposed compute governance monitoring and verification methods against both technical feasibility (accuracy, evasion-resistance) and institutional feasibility (cost, enforceability), rating each as near-, medium-, or long-term deployable. It highlights near-term methods and their requirements for deployment, aiming to give policymakers a more critical, balanced basis for evaluating compute governance options.
