The ATOM Report: Measuring the Open Language Model Ecosystem
- title
- The ATOM Report: Measuring the Open Language Model Ecosystem
- type
- summary
- summary
- Lambert and Brand's 2026 adoption study of ~1.5K open models - China passed the US in downloads in mid-2025, Qwen dominates, plus a size-normalized metric
- tags
- ai, open-weights, llm, china, research, ai-lab
- created
- 2026-09-14
- updated
- 2026-09-14
The ATOM Report is an April 2026 arXiv paper (2604.07190) by Nathan Lambert and Florian Brand of Interconnects AI. It is the data half of atom-project-american-truly-open-models: the methods first used to argue for that project were expanded into a measurement study, and the report says outright that it makes no policy recommendations and sends readers to the ATOM Project essay for those. Lambert's open-source-ai-reading-list marks it optional and describes it as the general summary of US versus China model adoption.
The scope is narrower than "open source AI". The authors track about 1.5K mainline open language models released since ChatGPT, models that take text in and produce text out. Embedding models, image diffusion models and small classifiers are excluded, because those can dominate download and fine-tune counts while having little effect on how the technology advances.
How it measures
Hugging Face downloads are the primary signal. A download is any HTTP request to a model file, including programmatic pulls. For history, Hugging Face supplied monthly data for six organizations (Meta, Qwen, Mistral, DeepSeek, Google, Microsoft) from November 2023 through July 10, 2025, filtered for spikes with a 2.5x interquartile range rule. From July 2025 the authors run their own daily scraper over every public model and splice its monthly deltas onto the filtered baseline. Derivatives come from the base_model tag, counting only children of tracked models with more than five downloads and excluding GGUF and MLX re-uploads, which would otherwise inflate fine-tune counts for small popular models.
Three other sources fill in. OpenRouter shared the top 10 open models per month by tokens served, which captures use rather than downloads but undercounts organizations whose traffic is spread across many models. Arena's style-controlled Elo supplies human preference, with scores before May 19, 2025 shifted up 59.2 points for a platform recalibration, and each region's best score held monotonic so the frontier only advances. The Artificial Analysis Intelligence Index supplies benchmark performance, with a linear fit per region.
Models are sorted into seven size buckets: under 1B, 1-5B, 7-9B, 10-50B, 50-100B, 100-250B and 250B+. The 7-9B bucket is split out because Llama 8B, Mistral 7B and Qwen 7B cluster there, and it alone takes 33.8% of all downloads; models under 10B take about 75%. mixture-of-experts models count by total parameters, so DeepSeek-R1 (671B total, 37B active) is in 250B+. Region is the headquarters of the releasing organization, which the authors call an imperfect proxy. Europe is just Mistral and Hugging Face.
The limitations section is candid. ModelScope and other hosts are not counted. A company that downloads a model once and serves it from its own infrastructure counts as one download, and so do clouds that cache weights. In the other direction, CI pipelines, bots and repeated pulls inflate counts, especially for small models and for small models with new architectures that get used as test fixtures. The authors' defense is that every usage signal they have correlates strongly with the download data.
Adoption by region
The data shows three eras, each with one region above 50% of downloads and fine-tunes. Europe led first, entirely on Mistral 7B and Mixtral 8x7B. The US took over with Llama 3, helped by its range of sizes. After DeepSeek V3, R1 and Qwen3, China took the lead and has widened it since. China overtook the US in late July 2025.
By March 2026 tracked downloads reached 2.04 billion, six times the 339M of March 2025. China went from 97M to 1.15B in that year (11.9x), the US from 177M to 723M (4.1x), and the EU from 65M to 163M (2.5x). The gap between China and the US grew from 23M in August 2025 to 428M.
Derivatives moved further. China's share of new fine-tunes went from 10% in November 2023 to 70% in February 2026, while the EU fell from a 58% peak in January 2024 to 4%. Inference moved furthest: Chinese models' share of open-model tokens on OpenRouter rose from 2.8% to 72.7% in 14 months, ending January 2026.
Performance followed the same direction with a smaller margin. On Arena, China's best open model passed the US in December 2024 on DeepSeek V3 and has led since. On Artificial Analysis, China's best open score quadrupled in 18 months to 68.4 by January 2026. The authors' explanation for why a modest performance lead turns into a large adoption lead is that people picking an open model default to the best one.
Adoption by organization
Qwen is the center of the story. It passed Llama in cumulative downloads in September 2025, 325.4M to 323.7M, and reached 942.1M by March 2026, nearly double Llama's 476.0M. As a base for fine-tunes it led much earlier, from June 2024, and its share of new derivatives rose from 1% in January 2024 to 69% in February 2026, while Meta's peaked at 44% in August 2024 and fell to 11%. Qwen 2.5, released in September 2024, is identified as the turning point.
Small models drive the gap. In February 2026 Qwen generated 153.6M monthly downloads against 71.2M combined for the next eight organizations, a figure the authors admit is inflated by the Qwen3.5 launch that month; December 2025 shows the same pattern at 87.5M against 61.3M. Six small Qwen3 models, 0.6B through 8B out of 66 in the family, drew 32.9M downloads in February, about equal to the 32.8M from Zhipu AI, MiniMax, Mistral, Moonshot AI, NVIDIA and openai combined.
deepseek plays a different role. It holds 47% of downloads in the 250B+ bucket, where Qwen holds 44% of the sub-10B market, and those large models are the one class where Qwen does not lead. DeepSeek passed Mistral in cumulative downloads only in January 2026 (128.2M to 124.3M), but its inference share is much larger than its downloads: V3 and R1 took up to 75.6% of open-model tokens on OpenRouter in June 2025 and still 31.1% in January 2026.
The other leaders each have a short arc. Mistral led derivatives at 57% in December 2023 and never turned that into download growth. OpenAI's GPT-OSS models, released from September 2025, had monthly downloads above Mistral's entire historical portfolio by spring 2026. Meta is second in cumulative downloads but stalled after Llama 4, and its OpenRouter token share went from a 37.4% peak in January 2025 to zero by August 2025.
The labs known for frontier-scale Chinese MoE models, MiniMax, Moonshot AI and Z.ai, are far behind the leaders in adoption, though growing. OpenRouter also shows how quickly a model can appear without a download footprint: Xiaomi's MiMo-V2-Flash, a 309B MoE released in December 2025, went from zero to 27.2% of open-model tokens by January 2026. The American entrants tracked since August 2025 are small by comparison: NVIDIA at 30.7M downloads by late March 2026, Ai2's olmo at 14.8M, IBM's Granite at 8.6M, about 56M for all of them together against Qwen's 942.1M.
The Relative Adoption Metric
Raw downloads favor small models, since a 1.5B model routinely gets 10-50x the downloads of a 400B one. The report's one methodological contribution is the Relative Adoption Metric, which compares a model to its size class at the same age:
RAM(m, t) = D(m, t) / C10(b, t)
D(m, t) is model m's cumulative downloads t days after release, and C10(b, t) is the 10th-highest release-aligned download count in its size bucket at that age. Scores are computed at 7, 14, 30, 60, 90, 180 and 365 days, and anything above 1.0 beats the bucket's top-10 line. The 10th place is used instead of the mean because the distribution is so skewed that in the 1-5B bucket one breakout model moves the mean by about an order of magnitude. The cutoffs are held monotonic and the baseline is refreshed quarterly; the numbers here use the 2026-Q2 baseline.
The standout results are American. In the 100-250B bucket, GPT-OSS 120B scored 19.19x at 7 days and was still at 20.31x at 180 days, and Nemotron Super 120B opened at 16.96x and 19.38x at 7 and 14 days and held 10.96x at 60. The authors read this as evidence that US models remain relatively popular long after release. MiniMax M2.1 is the counter-example, falling from 4.16x at 7 days to 0.78x at 90.
In 250B+, GLM-5 reached 20.21x at 30 days and 7.43x at 90, while GLM 4.7 and DeepSeek V3.2 never crossed the top-10 line. Kimi K2.5 started slowly at 1.64x and climbed to 9.55x by 90 days. The Qwen3.5 launch shows why normalization matters for small models: Qwen3.5-4B hit 6.97x at 7 days, which sounds large until you see the 1-5B cutoff was only 24K downloads at 7 days and 1.5M at 60. DeepSeek OCR (3B) launched higher in that bucket, at 12.68x and 15.45x. The clearest Qwen breakout was Qwen3.5-35B-A3B at 10.44x after a week and 8.75x after two months.
Reading it carefully
Two numbers in the report do not reconcile. The methodology says the tracked models account for over 3 billion downloads across the study period, while the regional section says cumulative tracked downloads reached 2.04 billion by March 2026, which is exactly the sum of the three regions. The report does not explain the gap; models from organizations outside the three regions would be one explanation, but the text does not say so.
The RAM finding about American models rests on two models, GPT-OSS 120B and Nemotron Super, in a single size bucket, and on downloads, the metric the authors admit bots and CI inflate. It says these two launches were exceptional, which is narrower than saying US models are gaining ground; the regional totals point the other way.
The report's figures also update the numbers on the ATOM Project page, which was written in August 2025. That page put Meta's peak derivative share at nearly 50% and its current share at 15%; the report, with cleaner filtering and a later snapshot, says 44% and 11%. Where they differ, the report is the newer source. The two also disagree on where Hugging Face belongs: its November 2025 list of American builders includes SmolLM, while the report counts Hugging Face as European. Other views of the same question are on are-open-models-catching-up, how-far-behind-are-open-models and open-and-closed-models-are-on-different-exponentials, and kimi-k3 is a later release that falls outside the report's window.
Source gaps
The extraction is text only. All 17 figures survive only as captions, Appendix B's seven tables of top models per size bucket are empty headings, and Table 1 with exact RAM scores per milestone is missing. The figures quoted above come from the body text and captions. The reference list was skimmed, not summarized; Appendix A's related work points mainly at Longpre et al. on ecosystem participation and at the Foundation Model Transparency Index.