#open-weights
Wiki 52
- 2 OLMo 2 Furious AI2's OLMo 2 report, fully open 7B/13B/32B models on up to 6.6T tokens with a stability fix list, Dolmino mid-training and the Tülu 3 RLVR recipe
- 6 months to live for open models Lambert predicts a US ban or delay on frontier open weights by early 2027, calls Anthropic's distillation campaign regulatory capture, and proposes off-ramps
- A $500 RL Fine-Tune That Beat the Frontier A $500 GRPO fine-tune of a 9B open model beat every frontier config on catalog review, 68× cheaper
- A Safe Path to Open Weights Thinking Machines Lab's release framework for open weights - test the model, stage access, filter dangerous knowledge - applied to its Inkling models
- Are Open Models Catching Up? SemiAnalysis reruns era-specific benchmarks on open and closed LLMs and finds catch-up time roughly halving each era, from 19.7 to 4.8 months
- Banning Open Source AI Would Be A Mistake Lambert and Kevin Xu's 2026 op-ed defending open source AI on education, innovation and competition as Washington moves to regulate models
- China's Structural Advantage in Open Source AI Kevin Xu, after Ion Stoica, on why Chinese labs default to open weights - talent, data, academia-industry ties, and shared artifacts
- Chinese Open Source: A Definitive History Kevin Xu traces Chinese open source from Linux in 1994 through Alibaba, Kaiyuanshe, Huawei and the state to the DeepSeek-era AI labs
- Consent in Crisis: The Rapid Decline of the AI Data Commons Longpre et al.'s audit of 14,000 domains behind C4, RefinedWeb and Dolma - AI crawl restrictions surged in 2023-2024, cutting off the best web data
- DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, 10 points above previous DeepSeek V4 Flash Artificial Analysis puts DeepSeek V4 Flash 0731 at 50, one point under GPT-5.6 Luna for ~60% less per task and on the cost Pareto frontier
- From Open Source Software to Open Source Strategy Bill Gurley on open source as a corporate weapon, from Android and Kubernetes to autonomous vehicles and the fight over open-weight AI
- Frontiers in synthetic data Lambert's 2024 notes on synthetic data in post-training, from SFT on GPT-4 outputs to Gemini Flash being distilled from Pro
- GLM-5.2 is the step change for open agents Nathan Lambert on GLM-5.2 as the first open-weight model that works as a general coding agent, released while Claude Fable 5 was restricted
- GLM-5.3: How Chinese labs keep stride with the frontier Lambert on GLM-5.3 and why Chinese labs match US models without relying mainly on distillation, plus Z.ai's staged release for cyber capabilities
- How distillation is used today and what performance uplift it gives to open models Lambert's July 2026 note that distilled data seeds SFT for Chinese labs but matters less as RL grows, written against a Stratechery claim
- How far behind are open models? Håvard Tveit Ihle measures open-model lag on 17 benchmarks, finding 8-10 months on private ones, 4-6 on public, and a gap growing since R1
- How much does distillation really matter for Chinese LLMs? Lambert reads Anthropic's February 2026 disclosure against DeepSeek, Moonshot and MiniMax and argues distillation helps but is not decisive
- Interconnects (interconnects.ai) Nathan Lambert's newsletter on how AI models are trained, reasoning models and post-training, and the open-model race between the US and China
- Kimi K3 Raschka on Moonshot's 2.8T open-weight K3 — LatentMoE, KDA, attention residuals, NoPE everywhere
- Kimi K3: The open-weights escalation Nathan Lambert on Kimi K3 as the first true frontier open-weight model, what it says about Chinese labs, and why banning open weights backfires
- LLM Distillation Training one LLM on another's outputs, from SFT on generated text to on-policy KD, and the 2025-2026 fight over Chinese labs distilling US models
- LLMs: Intelligence vs. Cost Guido Imperiale replots Artificial Analysis's intelligence-vs-cost chart on a linear axis with OpenRouter and local-electricity prices
- Neutrino-1 8B Fermion Research's Apache-2.0 8B: a 3.88 GB ternary-coded container for H100, MacBook, and CPU
- Nonproliferation is the wrong approach to AI misuse Helen Toner on why fixed dangerous capabilities cannot be kept from bad actors, and why the frontier-to-open lag should be used as an adaptation buffer
- Notes from inside China's AI labs Nathan Lambert's May 2026 trip report from Chinese AI labs - student-heavy teams, less ego, Claude everywhere, in-house data, and too few Nvidia chips
- Olmo The Allen Institute for AI's family of fully open language models, OLMo (2024) to OLMo 2 to Olmo 3, released with data, code, checkpoints and logs
- Olmo 3 AI2's Olmo 3 report (Dec 2025), fully open 7B/32B Base, Think, Instruct and RL-Zero models with every stage's data, code and checkpoints released
- OLMo: Accelerating the Science of Language Models AI2's first OLMo release (Feb 2024), 1B and 7B models on 2T+ Dolma tokens with weights, data, code, logs and 500+ checkpoints under Apache 2.0
- On the Societal Impact of Open Foundation Models Kapoor, Bommasani et al.'s 2024 paper - five properties of open-weight models and a marginal-risk framework showing most misuse studies were incomplete
- Open and closed models are on different exponentials Nathan Lambert on coding agents proving users pay a premium for top closed models, while open models take the larger, slower diffusion market
- Open models in perpetual catch-up Nathan Lambert on why the roughly six-month gap between open and closed models holds steady, plus trends in adoption, specialization and China
- Open Source AI is the Path Forward Mark Zuckerberg's July 2024 letter releasing Llama 3.1 405B, arguing open models are better for developers, for Meta, and for safety against China
- Open-closed model gap How far the best open-weight LLMs trail the best closed ones - how it is measured, estimates from 4 to 10 months, and whether it closes
- Open-Source AI & Open Models Reading List Nathan Lambert's annotated list of the best writing on open models — why they exist, why China leads, the gap, cyber risk and distillation
- Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling EleutherAI's 2023 Pythia suite, 16 models from 70M to 12B trained on the Pile in one fixed order, with 154 checkpoints each for training-dynamics research
- Qwen4: The Architecture of the Future (Qwen3.8-Flash-Next) Habr piece from reseller gptunnel on Qwen3.8-Flash-Next, the open-weight preview of Qwen4's architecture; Qwen4 itself is not out
- Reasoning Prefills on Open Models, v1.1 A reasoning-prefill test where Qwen3.8 follows GPT-5.5 Pro's trace far more than other open models, read as a sign of GPT distillation
- Some Simple Economics of Open versus Closed AI Christian Catalini's a16z essay using innovation economics to argue open weights change where AI investment goes and who profits, not how much
- Telnyx Inference Inference provider serving open-weight models on its own GPUs via an OpenAI-compatible API
- The Arguments Against Open Source AI are Very Bad Tom Bedor's rebuttal to the frontier-lab case against open weights, using encryption export controls as the precedent
- The ATOM Project: American Truly Open Models Nathan Lambert's 2025 memo arguing the US lost open-model leadership to China and needs several 10,000-GPU labs building fully open models
- The ATOM Report: Measuring the Open Language Model Ecosystem Lambert and Brand's 2026 adoption study of ~1.5K open models - China passed the US in downloads in mid-2025, Qwen dominates, plus a size-normalized metric
- The distillation panic Lambert argues "distillation attacks" wrongly brands a standard training technique and warns US policy could end up banning Chinese open weights
- The Gradient of Generative AI Release: Methods and Considerations Irene Solaiman's 2023 framework placing AI releases on a six-level gradient from fully closed to fully open, with the tradeoffs and controls at each
- The Myth of unsafe Open Source AI Florian Brand's survey of third-party incident reports - real AI misuse in 2025-2026 ran mostly through closed models, except image abuse
- The OpenAI/Huggingface incident; how we should manage the imminent arrival of autonomous hacking too cheap to meter Joshua Saxe reads an unreleased OpenAI model hacking Hugging Face as the start of cheap autonomous hacking, and argues for diffusion over restriction
- The Z.ai Playbook ChinaTalk interviews Z.ai's Zixuan Li on GLM - why Zhipu open-sources, the coding plan, role-play and translation, release within hours
- US Scrutiny of Chinese Model Use House probes of Airbnb, Cursor and DoorDash over Chinese AI models, set against Western companies moving to Chinese open weights to cut costs
- We urgently need a coherent national AI cybersecurity policy Joshua Saxe's AI Security Forum 2026 keynote - replace capability-threshold launch gates with a national observatory that measures net cyber harm
- What comes next with open models Nathan Lambert on why open models should stop chasing the closed frontier and build small, specialized models that closed agents call as tools
- Why I build open language models Nathan Lambert's 2024 case for fully open LLMs beyond Meta's self-interest, and what building them at Ai2 looks like day to day
- Z.ai (Zhipu AI) Chinese lab behind the open-weight GLM models, from a 2021 Tsinghua paper to GLM-5.3, and how its release practice changed along the way
talks 1
- Nathan Lambert on China's AI Ecosystem and the Open Model Gap Lambert's 2025 recap of open models - Qwen overtaking Llama, a crowded Chinese field, and why the US needs funded, fully open models