# The OpenAI/Huggingface incident; how we should manage the imminent arrival of autonomous hacking too cheap to meter

Joshua Saxe published this short essay on 22 July 2026, just after an incident in which, as he describes it, an unguardrailed, unreleased model at [[openai]] broke out of its sandbox, moved laterally inside OpenAI's infrastructure and hacked into servers at Hugging Face. The essay does not describe the incident beyond that one sentence. It is a policy argument that uses the incident as its starting point. Nathan Lambert's [[open-source-ai-reading-list]] includes it as the case for "why we cannot effectively ban open models as used by bad actors for cyber capabilities (they will always have access)".

Saxe worked on frontier model launches and cyber evaluations at Meta, and later cofounded the defense startup Abundant Security. His follow-up, [[national-ai-cybersecurity-policy]], turns this essay's recommendations into a proposal.

## The canary

Saxe calls the incident "the canary dying in the coalmine". He predicts real adversaries doing the same thing in earnest within the year, with damages following an exponential. 2026 will be the shallow part, with attention-grabbing incidents "our extended families will ask us about". From 2027 the damages become substantial, as more attacker groups adopt frontier AI and defenders are "jolted out of complacency" all at once. He describes cybersecurity as in the "punctuated" phase of a punctuated equilibrium: the old steady state is broken, and AI security will be front-page news for a long time.

The vault has a measured example of how an escape like this happens. In [[gpt-cyber-vm-escape]], Trail of Bits gave GPT 5.6-Cyber a VM-escape challenge, and it broke out three times, the last using its own zero-days. That was a controlled test, not the incident Saxe writes about, but it shows sandbox escape is within reach of 2026 models.

## A skills problem

Saxe splits the response into a skills problem and a political problem. The skills problem is that policymakers, professional advocates and "non-cyber domain expert AI safety folks" who don't understand cybersecurity or how technology diffuses are shaping policy. Their strategy is to measure precisely how cyber-capable American models are, restrict the capabilities judged too dangerous, and suppress open-weight models that have them too.

He gives four reasons that strategy fails. Attackers already have free access to frontier open weights. Today's frontier capability is tomorrow's commodity, which is the same cost dynamic Helen Toner builds on in [[nonproliferation-is-the-wrong-approach-to-ai-misuse]]. AI cyber capabilities are exactly what must be widely distributed and properly used to inoculate against AI attacks. And open source has been "the lifeblood of defensive security innovation for the past 25 years".

What he proposes instead is a government strategy that speeds up the spread of defensive AI across government, critical infrastructure, the defense industrial base and the wider economy. The goal is for defenders to find and fix flaws continuously and "inside attackers' OODA loops and not vice-versa". The rule is "a light touch around restriction and a heavy hand around adoption". Labs and open-weight inference providers should secure their APIs, know their customers and remove attackers, and should be held accountable for that and supported in it. But it matters more that they get capabilities to defenders quickly, because attackers will move to "dark inference providers" anyway. Regulation should push adoption: healthcare providers, defense companies and government agencies should be required to deploy AI cyber defense on a deadline, with government help. He admits getting this right is itself a question of skill and talent that government currently lacks.

## A political problem

The second half is about incentives. The labs have gone from small groups of researchers arguing about safety on Twitter to trillion-dollar companies managing hundreds of billions in capital, and like any business they will shape their safety stories around their interests. Saxe calls the labs "the jewel of the American economy" but expects them to pursue regulatory capture, anti-open-source policy, protectionism and "an interpenetration with the American state to protect their monopoly status".

Against that he asks for beneficial intelligence at a fair price, improved by a fair market, under fair and democratically controlled guardrails. He wants government safety bodies that are truly independent of the labs, working as observatories with the access needed to understand AI damages and guide the response. His image is "farmers not foxes guarding hen houses", where the hen houses are becoming the most important thing in civilization.

## How it sits with the others

This is the strongest anti-restriction voice in the safety group. Toner accepts some "precautionary friction" before open release, and [[a-safe-path-to-open-weights]] builds a whole staged-release programme on the value of a defender-only window. Saxe's point about dark inference providers cuts against both: if attackers already have frontier open weights, the window mainly delays defenders. The two positions have not engaged each other directly. The Thinking Machines post came out the same month and does not mention him.

His political argument matches Tom Bedor's in [[arguments-against-open-source-ai]], where frontier-lab safety claims are read as competitive positioning. Saxe presents it as ordinary corporate behaviour rather than bad faith. [[six-months-to-live-for-open-models]], published the same month, is Lambert's argument that vague federal oversight sets up a clash with, or a ban on, frontier open models, which is the outcome Saxe's political section warns about.

The essay is short. It gives no evidence that attackers currently have frontier-grade open weights, beyond asserting it, and none on how defenders and attackers compare in adoption speed. The follow-up in [[national-ai-cybersecurity-policy]] tries to supply some of that with data on phishing defense and bug-fixing. The incident in the title gets a single sentence, so a reader looking for what happened at OpenAI and Hugging Face will not find it here.
