top of page

When Frontier AI Goes Closed: Why AISI Needs Model-Agnostic Assurance

Sep 14
7 min read




Anthropic’s refusal to provide pre-release access exposes the risks of relying on privileged model access for frontier AI assessment.

 


Summary


Anthropic’s decision not to provide Claude Mythos 5.1 to the UK AI Security Institute (AISI) for pre-release testing has been discussed mainly as a problem of voluntary cooperation, regulation and geopolitics. It is also a technical warning. Some of the strongest methods for understanding and controlling advanced models depend on privileged access to model weights, activations, training or fine-tuning. That access is valuable, but it cannot be assumed. If the UK wants a durable national capability for frontier AI assurance, it also needs model-agnostic tests that remain useful when an evaluator can see only inputs, outputs, tool traces and externally observable behaviour.


The problem is not only whether a company cooperates


Mythos 5.1 is a useful case because it breaks an expectation that had become normal: a leading developer would give AISI early access to a high-capability model so that independent testing could take place before wider release. The immediate policy question is how the UK should respond when that access is withheld. The deeper engineering question is whether the assurance process itself has been designed to remain effective when the level of access changes.

This matters beyond one company. Advanced models will increasingly be developed under different national laws, export controls, security restrictions and commercial strategies. Some developers may provide full technical cooperation; others may offer only a restricted API, a hosted service, or access after deployment. In a period of geopolitical tension, a foreign government may also limit what can be shared with a UK body. No technical method can solve a complete refusal to provide any access at all. However, evaluation methods that require less privileged access reduce the number of ways in which independent scrutiny can be blocked.


What AISI contributes


AISI has an important role in the UK’s AI assurance landscape. Its stated mission is to equip governments with a scientific understanding of the risks posed by advanced AI, and it evaluates leading systems before and after release across areas linked to national security and public safety. Pre-deployment access gives AISI an unusual opportunity to test capabilities, safeguards and failure modes before they affect users at scale. That capability should be protected.

But independence is stronger when the evidence does not depend on a single access route. A robust assurance system should be able to use deep internal access when it is available, while still producing meaningful evidence when it is not.


What the public Model Transparency work tells us


The table below is based only on publicly available information and should not be read as an official AISI organisation chart. Public material does, however, show a clear concentration of expertise around model transparency, white-box control, model organisms, internal representations, steering and oversight.

Person

Public evidence of involvement

What the visible work suggests

Joseph Bloom

Publicly describes himself as leading AISI’s Model Transparency workstream; recent work covers auditing, monitorability, scheming and loss of oversight.

Strong use of interpretability and internal model signals, alongside broader oversight.

Jordan Taylor

Author on AISI White Box Control work and the 2026 Loss of Oversight report.

White-box detection plus broader oversight signals, including behaviour and model internals.

Sid Black

AISI Research Engineer; contributor to White Box Control and recent Model Transparency work on reward hacking and model behaviour.

Internal representations, model organisms, training and chain-of-thought behaviour.

Satvik Golechha

AISI researcher publicly working on training and interpreting model organisms of misalignment within the Model Transparency team.

Training-based and interpretability methods that benefit from open or privileged model access.

Thomas Read

Contributor to recent Model Transparency work on evaluation awareness, steering and the Loss of Oversight report.

Behavioural evaluation combined with steering and access to open-weight models.

Alan Cooney

Current Head of Autonomous Systems & Control at AISI and a contributor to the earlier White Box Control programme.

Empirical control and internal-access methods; best treated as adjacent to, rather than a confirmed current member of, Model Transparency.

Important qualification: it would be inaccurate to say that AISI’s interpretability or oversight work is exclusively white-box. AISI explicitly studies black-box, white-box and training-based methods. The point is that the publicly visible Model Transparency and White Box Control work has a strong centre of gravity around techniques that become more powerful when internal or training access is available.

AISI’s own sandbagging research makes this dependency especially clear. In its auditing game, black-box methods had little success in that particular setting, while more promising approaches used internal activations or additional training. AISI also notes that these stronger methods require deeper access than external evaluators are often given. That finding should not be interpreted as a reason to give up on model-agnostic testing. It identifies the research gap that now matters most.


Why model-agnostic assurance is strategically necessary


White-box methods should remain a major part of AISI’s research. Access to activations, representations and weights can expose evidence that may never appear clearly in a model’s final answer. The mistake would be to allow those methods to become a single point of failure for national assurance.

For explainability, model-agnostic testing does not promise to reconstruct a model’s true internal reasoning. It should instead aim for something testable: whether changes to prompts, context, retrieved evidence, tools or environmental conditions lead to stable and attributable changes in behaviour. Perturbation tests, counterfactual tests, local surrogate explanations, consistency checks and repeated sampling can provide external evidence about what influences a system and where its behaviour becomes unstable. These explanations are not ground truth about internal cognition, but they can still support audit, comparison and risk assessment.

The same principle extends beyond explainability. Safety, security and alignment can be tested through adversarial interaction, capability elicitation, red teaming, tool-use monitoring, sandboxed system tests, robustness checks, uncertainty measurement, cross-model comparisons and incident-focused evaluation. For agentic systems, an evaluator can examine actions, tool calls, state changes and failure propagation even when the model itself remains closed. No single test is sufficient, but multiple independent forms of evidence can provide a stronger assurance case than one privileged technique used in isolation.


A two-track strategy for AISI


The practical response is not to choose between white-box and black-box evaluation. AISI should maintain two parallel capabilities. The first is privileged assurance: use weights, activations, fine-tuning and internal signals whenever developers provide them. The second is sovereign model-agnostic assurance: maintain test suites whose minimum requirement is only the level of access available to an external user or regulator.

Every evaluation could state its access requirement and how its confidence changes as access is reduced. This would create an assurance ladder, from public interface testing through richer telemetry and finally to full internal access. A model should not become effectively untestable simply because a developer declines to share its internals. The quality of evidence may decrease, but the evaluation capability should degrade gracefully rather than disappear.


The lesson from Mythos 5.1

The Anthropic episode should therefore trigger two responses. One is legal and diplomatic: determine what access the UK should expect from companies offering powerful AI systems and how that expectation can be supported through agreements, standards or regulation. The other is scientific: invest in evaluation methods that remain useful when cooperation is partial.

AISI’s work on model transparency is valuable precisely because it shows how much can be learned from deeper access. The next step is to make sure the UK is not dependent on that access. The long-term objective should be an assurance system that can evaluate advanced AI under both cooperative and non-cooperative conditions. In a world where frontier models are strategic assets, model-agnostic testing is not a second-best option. It is part of technical sovereignty.


Building model-agnostic assurance at DEIS


This is also the direction we are exploring at the Dependable Intelligent Systems Research Centre. Our aim is to develop assurance methods that can test the explainability and dependability of AI systems without assuming access to model weights, internal activations or training data.

One example is XWhy, an approach we are developing to examine explainability from a model-agnostic perspective. Rather than trying to reconstruct the hidden reasoning process of a model, the objective is to test whether the explanation provided for its behaviour is supported by externally observable evidence. This can include systematically changing the input, context, evidence or operating conditions and measuring whether the model's decisions and explanations change in a consistent and attributable way.

This distinction is important. An explanation can sound convincing while having little connection to the factors that actually influence a system's behaviour. A model-agnostic assurance approach should therefore test explanations rather than simply accept or visualise them. Questions such as whether an explanation is stable, whether the identified factors are genuinely influential, whether similar inputs produce consistent explanations, and whether counterfactual changes lead to the expected behavioural response can all be examined without privileged access to the model.

We see this as part of a broader approach to dependable intelligent systems. Explainability should not only help a human understand an individual prediction. It should also become an observable property that can be tested, challenged and monitored as part of an assurance process. For organisations such as AISI, this creates an important complementary capability. White-box interpretability can provide valuable evidence when internal access is available. Model-agnostic approaches such as XWhy can provide another layer of evidence when that access is restricted. The long-term goal should therefore not be to decide which form of interpretability is superior, but to build an assurance framework in which evidence remains available across different levels of model access. This is the research direction we are pursuing: moving from explaining AI systems towards testing whether their explanations can be trusted.


Related Sources

Editorial note: This draft deliberately uses original wording and argumentation rather than reproducing the Responsible AI UK article. The public-team table is a best-effort snapshot from public sources, not an official AISI staff list.

 
 
 

Comments


bottom of page