Open-Source AI & Open Models Reading List

The ATOM Project: American Truly Open Models

title
The ATOM Project: American Truly Open Models
type
summary
summary
Nathan Lambert's 2025 memo arguing the US lost open-model leadership to China and needs several 10,000-GPU labs building fully open models
tags
ai, open-weights, open-source, policy, china, llm, ai-lab
created
2026-09-14
updated
2026-09-14

Nathan Lambert published the ATOM Project at atomproject.ai on August 4, 2025. It is a memo with a sign-on form: a statement of the problem, a funding recommendation, and a call for researchers, companies and philanthropies to back it. The page has kept moving since. Its charts were refreshed through March 2026, and a list of American open model builders was added on November 23, 2025. In his open-source-ai-reading-list Lambert files it as the piece on why the US needs to invest in open models for fundamental research in the face of competition from China. The data behind it was later expanded into atom-report-open-model-ecosystem, which deliberately makes no policy recommendations and points back to this page for them.

"Truly open" is doing work in the name. The memo separates open-weight models, where only the weights are released, from open source models that also ship training data, intermediate checkpoints, base models, training code and a permissive license. It wants the second kind.

The argument

The starting claim is historical. American AI leadership was built by being the hub of open research: Google invented the Transformer and shared it, and academia and industry worked closely enough that tens of thousands of students and researchers could build on each other's results. Lambert's view is that this playbook is now the default at Chinese companies and increasingly neglected at American ones. Leading US scientists have moved into closed labs, and "closed labs make closed research". He is careful not to blame those labs for doing good work privately; his point is that they cover only a few ideas, spend most of their resources on the next model, and leave the 2-, 5- and 10-year research questions to an open community that now runs mostly on Chinese models.

The concrete example is reinforcement learning with verifiable rewards. It was the most active research area after the first reasoning models, and most of that work was done on Alibaba's Qwen because Qwen was the strongest open base for math and code.

The adoption numbers in 2025

The memo tells the adoption story through Meta's Llama. When Llama 2 came out, its models drew about 60M downloads against 10M for the comparable early Qwen models, with half as many models. By Llama 3 and Qwen 2.5 in 2024 the lead had shrunk to 20M, both families past 120M. By August 2025 the leading US and Chinese providers each had around 300M total downloads on Hugging Face, with China growing faster and Europe near 100M. Derivatives told the same story: Chinese models accounted for 10-30% of new fine-tunes in early 2024, Qwen alone for more than 40% by mid-2025, while Llama's share fell from nearly 50% in fall 2024 to 15%. The updated charts add that Qwen passed Llama as the most-downloaded family in September 2025.

On performance, Llama 3 led every openly available Chinese model in July 2024. By the August 2025 publication date the top 10 open models on LMArena were all Chinese, as were the top 3 open models on Artificial Analysis.

The later ATOM Report revises some of these figures with a cleaner method. It puts Meta's derivative peak at 44% in August 2024, falling to 11% by February 2026, where the memo said nearly 50% falling to 15%. The report is the newer and better-documented source.

Lambert's diagnosis is structural rather than a verdict on Llama 4. China has "5 amazing open labs", the US had Meta, and "we are running Meta in a race against 5 other Chinese runners, and then complain when it doesn't win every race every time." Licenses made it worse: Llama's terms needed legal review to know whether a use was allowed, while Qwen and other Chinese models adopted plain licenses borrowed from open source software practice.

Security, values and dumping

A shorter section argues from national interest. Adoption of Chinese models by American companies has been slow in some sectors because of worries about backdoors and insecure generated code. Lambert admits this is hard to measure and rests on private conversations. There is also concern that the models' outputs are censored, with SpeechMap.ai cited as evidence. And he draws a parallel between Chinese labs racing to release cheap open models and the PRC's history of dumping state-subsidized exports below cost.

That parallel sits awkwardly with the rest of the memo. The American models it asks for are also supposed to be free, cheap to run and modifiable, which is the same offer the dumping comparison treats as suspect when a Chinese lab makes it.

What it asks for

The recommendation is specific. The US should maintain multiple labs, each with 10,000 or more H100-class GPUs, dedicated to training open models. The PRC had at least five labs releasing open models at or beyond the best US open model. Concentration matters: Lambert argues that spreading the same money across many small projects will not produce leading models, while a few focused efforts can learn from each other without growing so large that they slow down.

Each center should build two things. The first is frontier open models, which in 2025 meant 100 to 600+ billion parameters in a mixture-of-experts architecture. The second is a family of smaller sizes, from something that runs on a phone up to cloud scale, because most tasks do not need the largest model and researchers need to find the minimum size that solves a given problem. Top models are trained by teams of fifty to a few hundred people. The often-quoted $5 million cost of DeepSeek V3 is called misleading, as the deepseek authors themselves acknowledged; 10,000 GPUs is the entry point that allows fast experiments alongside a large training run.

Money should come from private companies, philanthropies and government. Programs like the National AI Research Resource help broaden access but are not concentrated enough, and the target is frontier open models within 6-12 months. The White House AI Action Plan of 2025 is read as a turning point in how Washington saw open models, with benefits outweighing measured risks.

American builders as of November 2025

The list added in November makes a narrower point than the memo: the US has about as many labs releasing good models as China's roughly 20, but American labs mostly release smaller models under more restrictive licenses, so their impact is muted.

Ai2's olmo is described as the open source leader, with Olmo 3 32B Think called the best fully open language model ever made, always under Apache 2.0. Nvidia's Nemotron is "arguably the open leader in the U.S. after Llama 4", with Nemotron Nano 9B v2 matching Chinese models of its size and more open data releases. openai's gpt-oss-120b, its first open-weights language model since GPT-2, is called incredibly important. Google's Gemma, IBM's Granite, Microsoft's Phi, Hugging Face's SmolLM, Liquid AI, Moondream, ServiceNow's Apriel, the Arcee/Datology/Prime Intellect Trinity models and Stanford's Marin community models fill out the list, mostly with small dense models. Reflection AI appears with no model at all, only the remark that raising $2B to do this is convincing enough. The four watched most closely were Ai2, Nvidia, Arcee and Reflection.

The "unclear" group is telling. meta-platforms gets "crickets since the Llama 4 fiasco", xAI's Grok has promised releases that have not been useful, Reka has gone quiet, and the Institute for Foundation Models is based in the US but marketed its K2-Think release as the "UAE DeepSeek". Cohere, Mistral and AI21 are named as belonging to a wider Western list.

Later context

Two things in the vault pick up threads from this memo. Reflection AI still had no public model in June 2026, when, according to six-months-to-live-for-open-models, its representative argued at a White House meeting that open models should be exempt from a capability-review framework. And Lambert's own later writing moves away from the China-threat framing used here. banning-open-source-ai-would-be-a-mistake, written with Kevin Xu in 2026, warns that using China as a pretext to regulate open source will backfire, and argues that transparency makes open models more secure. The memo raised backdoor and censorship concerns as reasons to build at home; the op-ed argues those concerns do not justify restricting the models. These positions can both be held, but the emphasis has clearly shifted.

Source gaps

The clipped page is incomplete. The FAQ has three category headings (Open Models, AI Ecosystem, Miscellaneous) with no questions or answers captured. The signatories list is empty, and the seven charts appear only as titles and captions, so the figures quoted above come from the memo's prose rather than the plots.