We urgently need a coherent national AI cybersecurity policy
- title
- We urgently need a coherent national AI cybersecurity policy
- type
- summary
- summary
- Joshua Saxe's AI Security Forum 2026 keynote - replace capability-threshold launch gates with a national observatory that measures net cyber harm
- tags
- ai, open-weights, security, policy
- created
- 2026-09-14
- updated
- 2026-09-14
Joshua Saxe published this on 13 August 2026, adapted from his keynote at the AI Security Forum in Las Vegas the week before. It follows his July essay openai-huggingface-incident-autonomous-hacking and turns that essay's call for diffusion over restriction into a critique of current US policy and a proposal. He writes from two positions: working on frontier launches, open cyber model evaluations and policy meetings at Meta, and now building AI-native cyber defense at Abundant Security, the startup he cofounded. Nathan Lambert's open-source-ai-reading-list describes it as "what the government should do to observe, orient, decide, and act with respect to emerging cyber threats (versus blocking models based on in-house capability assessments)".
The starting point
Saxe expects a world full of cheap security agents that automate much of the attack chain adversaries now run by hand, and that defenders will depend on equally. The result will be more instability in a field that was already doing badly before AI. He puts cyber damages at close to 1% of global GDP. His chart of damages over several decades is a reconstruction from different studies, and he says only the overall scale should be trusted, not any single point. Individual model launches over the past four years have not visibly moved that curve.
What current policy gets wrong
The dominant approach, he says, is simple. Test a model for "dangerous dual-use cyber capabilities", and if it crosses a threshold, block the launch or require heavier guardrails. His examples are the Claude Fable and GPT 5.6 launches.
His objection is that dual-use capability is not what causes cyber harm. Harm depends on attackers' goals and skills, how quickly they adopt AI if it helps them, what defenses are deployed, and how quickly defenders adopt the same capabilities. Understanding that takes human, economic and technical intelligence. The proper object of policy is net harm across attackers, defenders and victims, and a benchmark score of one model cannot tell you that.
He also says the current approach is quietly biased toward assuming a dual-use capability helps attackers. It is plausible, perhaps likely, that AI currently helps defenders more. LLMs have spread quickly through email phishing defense, not only through phishing attacks. On vulnerabilities he calls AI "miraculous" for defenders, citing Google's reported results on finding and fixing Chrome bugs. On the attack side, AI is widely used in phishing but its causal share of phishing damages is unclear. He sees no meaningful rise in economic damages from vulnerability exploitation due to attacker AI use. He adds that this does not rule out nation states quietly building stockpiles of AI-found exploits for use in a conflict.
Whether AI helps attackers or defenders more is, for him, an open empirical question in both areas. What worries him is that federal policy is not spending the money to find out.
The observatory
His proposal is an AI cybersecurity observatory: a well-funded body that monitors the whole national system of attackers, defenders and victims, and recommends policy from a much wider set of options than approving or blocking a launch. It would collect enough information to predict how a regulation, a subsidy, a federal hardening programme, a launch delay or any other intervention would change national and global security, and it would keep learning as results come in. His diagram describes it as a shift "from a launch gate to an adaptive control loop". Lambert's summary uses the OODA-loop words (observe, orient, decide, act) for the same idea.
The urgency comes from open weights. Saxe says open-weight models are already capable enough for attackers to start moving to AI-automated and semi-automated attacks, and they are getting much cheaper. His illustrative projection shows the cost of running a fleet of always-on attack agents falling sharply as inference prices drop. He does not claim to know when AI will produce a real jump in cyber damages, calling himself "no super-forecaster", but thinks it is reasonable to expect one and that failing to prepare will cause avoidable harm.
How it sits with the others
This is the most direct challenge in the vault to the approach a-safe-path-to-open-weights describes. Thinking Machines decides release through capability evaluations, red-teaming and fine-tuning ceilings, while Saxe argues that such scores, even done well, answer the wrong question. The two agree more than that suggests. Thinking Machines also frames the question as model plus ecosystem, and its staged access for defenders is one of the interventions an observatory might recommend. The disagreement is over what should decide a launch: the model's own measured capability, or evidence about what happens in the world after release.
His net-harm framing extends the marginal-risk framework in societal-impact-of-open-foundation-models, which also asked for existing risk, existing defenses and ease of defense to be measured before judging a model. Saxe takes that from single papers to an ongoing government function. myth-of-unsafe-open-source-ai supplies incident data of the kind the observatory would collect, and finds closed models behind most of it. nonproliferation-is-the-wrong-approach-to-ai-misuse shares the view that defense has to be built across society rather than through access control.
The post is a keynote adaptation and it shows. It says what the observatory is for, but not who runs it, what legal access it has to lab and company data, or how its recommendations would bind anyone. Several charts carry the argument, and the defender-benefit claims rest on them.