Nonproliferation is the wrong approach to AI misuse
- title
- Nonproliferation is the wrong approach to AI misuse
- type
- summary
- summary
- Helen Toner on why fixed dangerous capabilities cannot be kept from bad actors, and why the frontier-to-open lag should be used as an adaptation buffer
- tags
- ai, open-weights, ai-safety, policy, security
- created
- 2026-09-14
- updated
- 2026-09-14
Helen Toner published this on 5 April 2025, the third of three posts launching her Substack, Rising Tide. It is aimed at her own side: people in AI safety who think catastrophic misuse is a real prospect and conclude that the only defense is to stop dangerous models from spreading. She accepts the premises and rejects the conclusion. Nathan Lambert's open-source-ai-reading-list files it under cyber risks and open models, as the argument that "you cannot expect to control access to AI at a certain threshold (e.g. open weight models) and need to prepare society to tackle risks downstream of available intelligence". It is the oldest of the three pieces in that group, and the two Joshua Saxe posts beside it push its reasoning further than she does.
Four premises accepted, one rejected
Toner sets the argument out in five steps. AI is getting better at coding and scientific R&D. At some point that could make large attacks, on critical infrastructure or with novel bioweapons, much easier. There are enough bad actors that attacks will be attempted. Society cannot reliably stop them, and one could be catastrophic. So the only answer is hardcore nonproliferation.
She agrees with the first four. The fifth is the one she calls "probably wrong": a serious nonproliferation regime would be both ineffective and damaging to freedom. Her main example is a paper by Dan Hendrycks, Eric Schmidt and Alexandr Wang (published at nationalsecurity.ai), which she calls an unusually clear statement of the view. Its nonproliferation section has three parts: compute security (tracking where advanced chips are), information security (locking down misusable weights) and AI safeguards. That section is explicitly about keeping catastrophic capability from malicious actors, not about a great-power race.
Frontier targets and fixed-capability targets
The argument turns on a cost curve she takes from Paul Scharre (CNAS, 2024). Training the best model keeps getting more expensive. But once a given level of capability exists, reproducing it gets cheaper quickly. That gives policy two different kinds of target. A frontier target, such as tracking progress toward AGI, can focus on the biggest and most compute-hungry systems. A fixed-capability target is a specific ability, like helping a novice hack the power grid or make smallpox. That ability "will rapidly become cheaper and more widely available" however expensive it was at first. Bioweapons and infrastructure hacking are fixed-capability problems.
Her analogy is a counterfactual about nuclear weapons. Suppose the uranium and enrichment needed for a tactical bomb kept falling. Controls on enrichment plants would work at first, but eventually anyone with a plot of land, and the trace uranium in its soil, would need watching. The regime would become more invasive and less effective at the same time. Nuclear technology never went that way. AI does: "it's often only a couple of years between when you need a world-class computing cluster to do something and when you can run it on a high-end gaming chip on your laptop."
Where the open-source argument stops too early
The critique is not only aimed at nonproliferation advocates. Toner says the open-source community, whom she describes as AI nonproliferation's main opponents, tend to think the argument ends once control is shown to be draconian and futile. It does not. If AI that can be catastrophically misused is coming, something else has to deal with it. That sets her apart from the pure-rebuttal pieces in the vault, such as arguments-against-open-source-ai, which treats the danger claim as mostly made in bad faith.
The adaptation buffer
Her alternative is to treat the gap between frontier models and widely available models as an adaptation buffer. The window opens when it becomes clear which dangerous capability is coming and closes when that capability is widespread. The work in between is making society resilient to it. That can mean using unreleased frontier models for defense, but most of what she proposes involves no AI at all.
For cyber she suggests a large capacity-building push to prepare critical infrastructure operators for AI-driven attacks, incentives to build defensive tools on the best models so new capabilities are in deployed defensive software before attackers get them, and economy-wide plans for responding to and recovering from a major AI-driven attack. For biosecurity she suggests screening access to key services and materials (like the DNA synthesis screening rules in Biden's AI executive order), pathogen-agnostic medical countermeasures such as a universal flu vaccine and broad-spectrum antivirals, the ability to scale vaccine production fast, PPE stockpiles, and better disease surveillance. She says she is not an expert in either field and offers these as the kind of measure she means, not a policy agenda.
There are two ways to widen the buffer, and she supports both. One is limited, targeted nonproliferation that delays wide availability. She approves of the status quo, where frontier models stay private and open equivalents follow months or years later, and calls that "precautionary friction". The other is better forecasting of specific dangerous capabilities. She praises the uplift trials labs had begun running for bioweapons assistance, and says that for helping novices make known pathogens, the buffer has already started.
Scope and caveats
The argument does not cover frontier targets. There, some of the Hendrycks proposals look reasonable to her. Chip export controls cannot stop someone putting eight GPUs in a basement, but they can limit who builds clusters of hundreds of thousands of chips. Firmware-level chip features are a poor anti-terrorism tool but might help verify a US-China agreement. It also does not cover loss-of-control risk from the models themselves. Humanity has long experience adapting to dangerous tools in bad actors' hands, and none with machines smarter than people.
She does not promise the strategy works. With several misuse threats arriving at once alongside economic and geopolitical upheaval, she doesn't "particularly like our chances". The buffer could be too short, a new technique could suddenly make already-released models better hackers, or infrastructure operators could prove impossible to change. The fix still has to be "multiple ambitious, large-scale projects with a real sense of urgency".
The post ends with a diagnosis of the debate. When she presses nonproliferation advocates, many turn out to worry more about AI takeover and to expect progress so fast that nonproliferation would only matter briefly. What they actually want is a year or two of delay on open releases, protection of US models from theft by China, and time for defenders to adapt, which is close to her own view. They argue in terms of misuse because hackers and bioterrorists sound more respectable than takeover, and end up appearing to demand a permanent, invasive regime they do not intend. Some in the safety coalition have argued for totalitarian surveillance as a long-term answer, she notes, citing Bostrom's "Vulnerable World" paper, and anyone who does not mean that should not imply it.
How it fits with the others
a-safe-path-to-open-weights (July 2026) is her framework put into practice by a lab: staged, defender-first access is precautionary friction with a plan for using the window. openai-huggingface-incident-autonomous-hacking and national-ai-cybersecurity-policy (July and August 2026) share her conclusion that restriction cannot contain a fixed capability, but go further. Saxe argues restriction should be light even during the buffer, because open frontier weights are already available to attackers and defensive diffusion is what counts. The disagreement is over how much the delay itself is worth.
In 2026 the buffer has a measured size. Lambert's list puts the open-closed gap at roughly four to six months, with the leading open models from Chinese labs; open-closed-model-gap and how-far-behind-are-open-models cover the measurements. That is much shorter than the "months or years" lag Toner describes approvingly, and it is set by labs outside US policy reach. The post does not consider that the delay she values could come from someone other than the US frontier labs releasing on their own schedule.
- A Safe Path to Open Weights
- Detecting and countering misuse of AI: September 2026
- The Arguments Against Open Source AI are Very Bad
- We urgently need a coherent national AI cybersecurity policy
- Open-Source AI & Open Models Reading List
- The OpenAI/Huggingface incident; how we should manage the imminent arrival of autonomous hacking too cheap to meter
- 6 months to live for open models
- On the Societal Impact of Open Foundation Models
- Some Simple Economics of Open versus Closed AI