Essay  ·  Artificial Intelligence  ·  Security  ·  Governance

The Vault
& the Virus

Why closing access to frontier AI produces the precise outcomes it claims to prevent — and what the history of every dangerous technology tells us about why.

There is a position, held with genuine conviction by serious people, that the most powerful AI systems should be locked away — accessible only to vetted institutions, audited deployments, and trusted partners. The argument sounds like prudence. It is, on examination, closer to its opposite: a strategy that concentrates risk, destroys defensive infrastructure, and has now, in the case of Anthropic's own Mythos model, produced a textbook demonstration of its own failure on the day it was announced.

The Claim Under Scrutiny Closing access to frontier AI models makes the world safer by preventing dangerous capabilities from reaching malicious actors.

IThe Dual-Use Fallacy

The closure argument begins with a premise that sounds obvious: powerful tools, in the wrong hands, cause harm. Therefore restrict the tools. The problem is that this logic, applied consistently, would have locked away Linux, strong cryptography, the internet, and the printing press. Every one of those technologies was dual-use. Every one enabled both extraordinary human progress and genuine harm. None were made safer by restriction; all became more useful and more manageable through the development of norms, institutions, and defensive infrastructure that only open access made possible.

What the closure argument actually proposes is not safety — it is a shift of power. Closing a general-purpose AI capability does not make the capability unavailable; it makes the capability available only to the organisation that controls it. That is a political and economic outcome dressed in the language of safety. The question it refuses to answer is: safer for whom, and compared to what?

On the Nature of Knowledge

Nuclear material is dangerous at rest. Enriched uranium in a container is a threat independent of anyone's knowledge or intent. AI is dangerous only in combination with human agency — it is pure knowledge, and knowledge has a different relationship to containment than matter does. You can build a fence around uranium. You cannot build a fence around an idea that thousands of researchers worldwide are independently converging on.

IITwo Claims, One Argument

The closure debate is routinely framed as open weights versus closed weights. That framing obscures the more important distinction. There are actually two separate claims being made, and they carry entirely different burdens of proof.

The first claim is an intellectual property claim: we built this model, therefore the weights are proprietary. This requires no special justification. It is the standard position of every closed-source software company. Nobody disputes it as a category. A lab that trains a model and chooses not to release the weights is making the same decision as a company that ships a binary without publishing the source. The weights are their IP.

The second claim is a governance claim: this capability is so dangerous that the public must not be permitted to interact with it at all, and we are the appropriate authority to make and enforce that determination. This claim is categorically different. It is not a licensing decision. It is an assertion of regulatory power that the organisation was never granted, over a public that never consented to the arrangement. The burden of proof required to justify restricting access is an order of magnitude higher than the burden required to protect IP — and it must clear that bar without appeal to the same commercial interests that benefit from the restriction.

The conflation of these two claims is doing most of the work in the closure argument. Ordinary closed models attract no particular criticism because they are making the first claim only. Mythos attracted a qualitatively different response because it made the second claim: not just that the weights were proprietary, but that society should be prevented from interacting with the capability entirely. That is not a software decision. It is a political one, and it demands political scrutiny.

The Burden of Proof Asymmetry

"We own this model" requires only that you built it. "Society may not access this model" requires that you demonstrate the danger is real, that your judgement is trustworthy, that no alternative governance structure exists, and that your commercial interest in the restriction does not contaminate the safety argument. The second claim has never cleared that bar. It has simply been asserted, with confidence, by the organisations that benefit most from it being accepted.

IIIThe Bottleneck Is Never the Book

In 1994, a seventeen-year-old named David Hahn built a functional neutron source in his mother's garden shed in Michigan. He used publicly available chemistry textbooks, mail-order materials, and correspondence with regulatory bodies who did not realise what he was building. He got surprisingly far. The limiting factor was not knowledge — it was isotope acquisition. The physical substrate was the bottleneck, not the synthesis of understanding.

This is the correct model for thinking about AI and dangerous capability. A group capable of acquiring nuclear materials or synthesising dangerous chemicals is also capable of finding the knowledge to use them, with or without AI. The population that AI meaningfully assists is not the determined state-level actor — it is the opportunistic lower-skilled one. And the question then becomes whether AI raises the defensive capability for that same population equally. The evidence suggests it does: AI-assisted fraud detection, anomaly identification, and security auditing all benefit from the same general capability that improves the attacker's toolkit.

AI is an accelerator, not an originator. It reduces the friction of synthesising existing knowledge. So does a skilled Google search, a good research librarian, or a PhD supervisor. The difference is speed and accessibility — a difference of degree, not of kind. The public debate has consistently mistaken acceleration for invention of new danger, and that category error is doing enormous work in the closure argument.

IVThe Geopolitical Boomerang

The closure argument has a structural failure that becomes visible the moment you trace its geopolitical consequences. Full closure — banning export of open weights, restricting high-end compute, cutting off foreign cloud access — does not prevent capable adversaries from developing AI. It converts AI development from a commercial race into a national survival mandate. The one-time capital cost of building a sovereign parallel ecosystem, however large, becomes trivially justifiable against the framing of existential strategic deficit.

Once that fixed cost is paid and the parallel infrastructure is operational, the original gatekeeper loses one hundred percent of their leverage. The capability was not stopped; it was replicated entirely outside any sphere of visibility, auditing, or diplomatic influence. This is not a theoretical prediction — it is the documented history of nuclear technology, space technology, and semiconductor manufacturing. The Soviet nuclear programme reached capability four years after the most aggressively secured scientific effort in history. Chinese orbital capability developed regardless of ITAR. Huawei's Ascend chip series accelerated in direct response to NVIDIA export restrictions.

Full Closure: The Complete Failure Sequence

Full closure forces sovereign adaptation at high one-time cost. Once that cost is paid, control is lost entirely and the capability now exists in an ecosystem completely outside diplomatic reach. Partial closure is worse: the capability leaks anyway through distillation and API exploitation, but the public never gains the direct experience needed to develop accurate threat models. You get the proliferation without the inoculation.

The compute restriction argument follows the same arc. The assumption that restricting high-end GPU exports prevents capable actors from developing frontier AI ignores that training does not strictly require data centre hardware — it requires time and electricity. Consumer-grade GPUs are inefficient for training large models, but inefficiency is a cost, not a barrier. Meanwhile the restrictions raise barriers for universities, open researchers, and the civil society organisations building defensive infrastructure — the population least likely to cause harm and most likely to produce safety research that eventually benefits everyone.

VThe Framing Is the Signal

There is a mechanism in the closure debate that receives almost no attention: the announcement itself is the proliferation event. Before a lab declares that its model is too dangerous to release due to its exceptional offensive capability, that model is one of many research artefacts in a crowded field. After the announcement, it is the confirmed location of the most strategically valuable capability on the planet, certified by the people who built it.

The security framing does not distribute fear evenly. It concentrates trust — in the developers. It converts a positive-sum capability race into a zero-sum strategic competition. And it provides the moral and strategic justification for every extreme acquisition measure an adversary might consider: state-sponsored theft, insider recruitment, supply chain compromise. A foreign intelligence service that might have deprioritised AI lab penetration now has a board-approved mandate and unlimited budget, courtesy of the press release that announced the danger.

The Better Framing

Consider the difference between two announcements. The first: "Our model is so dangerous it can hack critical infrastructure — we will not release it." The second: "Our model found previously unknown bugs in widely-deployed production systems — here are the CVEs, here are the patches, here is the API." Same underlying capability. The first certifies a high-value target and mobilises every adversary with strategic motivation. The second distributes the defensive benefit, collapses the acquisition incentive, and earns trust through demonstrated outcomes rather than claimed authority.

The practical consequence is that closing a model while announcing its offensive capability does not protect anyone from the capability — it ensures the people who most need to defend against it have no access to it, while every motivated adversary now has a certified reason to acquire it through whatever channel is available. The announcement does more for the attacker's intelligence picture than most espionage operations could achieve independently.

VIThe Inoculation Argument

In September 1987, a worker at a radiotherapy clinic in Goiânia, Brazil, broke open an abandoned caesium-137 teletherapy unit. The glowing blue powder inside looked magical. It was shared through a neighbourhood. Four people died. Dozens were contaminated. Hundreds required monitoring. The harm came not from radiological knowledge being too widely distributed — it came from its being too narrowly distributed. The people handling the material did not know what they were handling because the knowledge of what caesium-137 looked like and what it could do had not reached them.

The lesson the nuclear safety community drew from Goiânia was not further restriction. It was wider distribution of detection awareness, reporting norms, and public health education. Known-signature awareness saves lives. Opacity kills. The same principle applies to AI capability: a population with no direct experience of what AI-generated phishing, synthetic media, or automated social engineering actually looks like is a population that cannot defend against it.

There is a subtler failure mode that the controlled-demo approach introduces. Showing users demonstrations of AI-generated attacks produces a mental model anchored to the demonstrated capability. That model becomes their filter. The actual adversarial deployment is optimised against exactly that filter. The demo does not show you the threat — it shows you a safely bounded version that produces false confidence about recognising the real one. A user who has spent time with an open model understands its capability range from direct experience. That knowledge is durable and calibrated in a way no awareness training produces.

VIIThe Accountability Gap

There is a version of the closure argument that deserves to be taken seriously: perhaps closed labs, despite their incentive problems, are simply more careful than the distributed chaos of open deployment. The historical record does not support this. Closed labs face simultaneous pressure to attract investment, maintain competitive advantage, avoid regulation, and project safety credibility. These incentives are nearly perfectly designed to produce motivated underreporting of dangerous capabilities. A lab that discovers its model has unexpected offensive properties has strong financial incentives to classify that internally and remediate quietly rather than disclose. There is no mandatory adverse event reporting system, no equivalent of the aviation near-miss database that has made commercial flight so safe precisely because failures are shared rather than buried.

Safety research is also more productive on open models. Interpretability, alignment, and robustness work all require access to weights. If safety research is concentrated inside the same labs building frontier models, you get captured safety teams with commercial conflicts of interest rather than independent verification. The open source security community has demonstrated for decades that public codebases receive more scrutiny, find more bugs, and develop more robust defensive tooling than proprietary equivalents. There is no reason to expect AI to differ.

VIIIEternalBlue

In 2017, a group calling itself the Shadow Brokers leaked a collection of offensive cyber tools from the United States National Security Agency. Among them was EternalBlue — an exploit targeting a critical vulnerability in Microsoft's SMB protocol. The NSA had known about the vulnerability for years. Microsoft had no patch. The global infrastructure had no inoculation. Within weeks, WannaCry and NotPetya used EternalBlue to cause an estimated ten billion dollars in damage. Hospitals diverted patients. Shipping infrastructure froze. Manufacturing plants went dark.

The NSA's logic for keeping EternalBlue secret was identical to the AI closure argument at every step. We found something powerful. The public is not ready. We can control this better than the ecosystem can. The security framing justifies the secrecy. Each assumption failed in sequence. Closure created a high-value target that attracted sophisticated theft. The theft delivered the capability to criminal operations with zero safety investment and zero defensive accompaniment. The public was completely unprepared because the inoculation had never been distributed. And the organisation that hoarded the capability bore none of the cost while hospitals and shipping companies paid it in full.

The responsible disclosure norm the security community developed in the years since is not idealism. It is hard empirical learning from exactly this outcome: find the vulnerability, notify the developer privately, agree on a patch timeline, then publish. The ecosystem gets the fix before the exploit is weaponised. EternalBlue is the complete worked example of what happens when the closure argument wins. The AI governance conversation is being asked to learn the same lesson without paying the same tuition.

IXThe Mythos Proof

On April 7th, 2026, Anthropic announced Claude Mythos Preview — described as sufficiently capable at offensive cybersecurity that public access was not appropriate. It was made available to a small number of trusted partners under Project Glasswing. On the same day, a group accessed the model by guessing its URL from familiarity with Anthropic's URL conventions for other models. They have been using it since.

The breach did not require a zero-day. It required pattern recognition.

A second exposure followed weeks earlier, when source maps for Claude Code were accidentally published due to a missing entry in an .npmignore file. Static analysis tools have been able to catch this class of error for years. These are not novel attack surfaces — they are the ordinary texture of misconfigured builds and predictable naming schemes. The same category of failure accounts for the majority of production breaches across the industry: exposed environment variables, open S3 buckets, default credentials left in place, API keys committed to public repositories. The zero-day is the exception. The misconfiguration is the rule.

What makes the Mythos case structurally interesting is not the embarrassment of the breach. It is what the breach implies about the model's actual deployment. Grant, for the sake of argument, that Mythos performs exactly as claimed — that it can identify critical vulnerabilities in production systems. The most natural application of that capability is continuous audit of the organisation's own infrastructure. The URL access and the .npmignore exposure suggest one of three things: the model was not running on Anthropic's own systems, it was running and not acting on its findings, or it was running and findings were not reaching the people responsible for remediation. None of these is a reassuring interpretation — and critically, none of them depends on whether the capability claim is true. The indictment holds either way.

Detection is not remediation. A finding that does not reach the right person, at the right time, through a process that produces a fix, contributes nothing to security regardless of how sophisticated the underlying model is.

This is the limitation the closure argument consistently underweights. The gap between what a model can identify and what an organisation actually patches is where most security failures live. The responsible disclosure community has a name for this interval: patch lag. The vulnerability exists. The knowledge of it exists. The breach happens in the space between discovery and remediation — or because remediation never happened at all.

Closing public access to the model shifts none of that. It leaves the gap intact while ensuring that the organisations best positioned to audit their own systems for the same class of failure — the ones without dedicated security teams — remain without the tool. The capability that was declared too dangerous for public access demonstrably did not protect the organisation that declared it. The lesson is not that the capability is overstated. It is that capability without process is not security.

XThe Patch That Never Shipped

There is a simpler way to state what closed access actually means in practice. When a capability exists and bad actors gain access to it before defenders are ready, that is not an argument against open access. It is the permanent condition of security. The question is always what compresses the time between exposure and defense.

The entire responsible disclosure pipeline — researcher finds vulnerability, notifies vendor privately, patch ships, CVE publishes — is premised on offensive knowledge reaching the defensive side at the same time it reaches attackers, or sooner. That pipeline has been hard-won over decades. It works because both sides have access to the same information. Closing access breaks the pipeline at the source: defenders cannot patch what they cannot see.

The "bad actors move first" concern is not a special property of open AI. It is the default state of every security domain. The security community's answer was never restriction — it was infrastructure: bug bounties, CVE databases, open threat intelligence feeds, public vulnerability research. All of it depends on the assumption that broad access to offensive knowledge produces faster and more robust defense than restricted access does. The alternative — hiding vulnerabilities and hoping adversaries don't find them independently — has a name. It is called the NSA's EternalBlue strategy. The outcome is documented in the previous section.

Closing access to avoid the scenario where bad actors move before defenders is equivalent to arguing that CVEs should not be published because attackers read them too. Technically true. Net effect: the sophisticated attacker finds the vulnerability anyway; the defender never gets the patch. The only population that closed access reliably protects against is the unsophisticated opportunistic actor — and that is precisely the population most capable of being defended against through the open ecosystem that access enables.

XIWhat Remains

Every argument for closure ultimately rests on the same assumption: that safety comes from reducing access to capability.

The historical record suggests the opposite. Safety emerges when defensive capacity grows at least as quickly as offensive capacity. The mechanism differs across domains, but the pattern repeats. Aviation became safe because incidents were shared. Cybersecurity became safer because vulnerabilities were disclosed and patched. Public health became safer because warning signs, reporting systems, and detection knowledge were distributed widely enough that ordinary people could recognise danger before it became catastrophe.

The lesson is not that dangerous capabilities are harmless. It is that dangerous capabilities do not remain confined to the institutions that discover them. They diffuse. Sometimes through publication. Sometimes through theft. Sometimes through independent rediscovery. Sometimes through simple persistence by motivated adversaries. The mechanism changes. The outcome rarely does.

Once that reality is accepted, the question changes. The relevant policy problem is no longer how to prevent capability from spreading indefinitely. The relevant policy problem is how to ensure that defensive infrastructure spreads faster than offensive use.

Closed weights are a legitimate intellectual property choice. No organisation is required to publish what it builds. But the stronger claim—that society itself should be denied access because a capability is too dangerous—requires demonstrating that restriction produces more security than preparation. The evidence presented so far points in the opposite direction.

The Mythos episode is instructive not because it was uniquely embarrassing, but because it compressed the entire closure argument into a single case study. A capability was declared too dangerous for public access. Access leaked anyway. The people most likely to benefit from defensive exposure were excluded. The people most motivated to obtain the capability immediately gained a reason to pursue it. The capability diffused. The inoculation did not.

What remains, then, is not a choice between safety and openness. It is a choice between two security models. The first assumes that capability can be successfully concentrated indefinitely. The second assumes that capability will eventually spread and focuses on preparing society for that reality. History has been unusually consistent about which model survives contact with reality.

Closing a physics textbook does not make the physics go away. It ensures that the people who most need to understand it — the engineers building defences, the researchers finding fixes, the public developing instincts — remain in the dark while the people with the worst intentions find their own copy. The vault metaphor is seductive because it implies control. What it actually describes, as the record now shows, is a target. And targets get found. They always have.