Open-closed model gap
- title
- Open-closed model gap
- type
- concept
- summary
- How far the best open-weight LLMs trail the best closed ones - how it is measured, estimates from 4 to 10 months, and whether it closes
- tags
- ai, open-weights, benchmarks, llm, china
- sources
- open-source-ai-reading-list, are-open-models-catching-up, deepseek-v4-flash-0731-artificial-analysis, how-far-behind-are-open-models, open-models-in-perpetual-catch-up, open-and-closed-models-are-on-different-exponentials, glm-5-3-how-chinese-labs-keep-stride, the-distillation-panic, how-much-does-distillation-matter-for-chinese-llms, how-distillation-is-used-today, zai-playbook, lambert-china-ai-ecosystem-open-model-gap
- created
- 2026-09-14
- updated
- 2026-09-14
The open-closed model gap is the distance between the best LLM whose weights anyone can download and the best one available only through an API. It is usually stated as a time lag, "open models are N months behind", and sometimes as points on a benchmark index. In September 2026, Nathan Lambert's open-source-ai-reading-list gave the consensus as "roughly 4–6 months", said the gap had narrowed in recent years, and noted that the leading open models have all come from Chinese labs since about 2024. The sources behind that sentence do not agree with each other. Estimates run from about four months to about ten depending on the benchmarks used, and one analysis finds the gap widening while another finds it closing faster every era.
Measuring it
The simplest measure is a composite index. Artificial Analysis runs a fixed suite on every model (nine evaluations in Index v4.1, among them GDPval-AA, Terminal-Bench, Humanity's Last Exam and GPQA Diamond) and publishes a step chart of the best open-weights and proprietary scores over time. Epoch AI does the same with its Epoch Capabilities Index, which can be plotted by release date and filtered by accessibility. An index gives the gap in points. On 31 July 2026 Artificial Analysis had Claude Opus 5 (max) at 61 and the best open model, kimi-k3 (max), at 57 (deepseek-v4-flash-0731-artificial-analysis). Guido Imperiale's reading rule for the same index is that one point will not be noticed and five is substantial (intelligence-vs-cost-linear).
Points do not translate into months, so the more quoted figures come from matching levels. Håvard Tveit Ihle (how-far-behind-are-open-models) sets score thresholds on each of 17 benchmarks and measures, for each threshold, how long after the first closed model an open model first crossed it. SemiAnalysis (are-open-models-catching-up) splits LLM history into three eras, each with its own four benchmarks, and times how long each era's founding closed flagship stood before an open model passed it.
Two choices inside these methods change the answer. The first is which end of the interval gets the date. Ihle dates each gap at the open model's release, a backward-looking measure: how far back you have to go to find closed models as good as today's open ones. The forward-looking question, how long until open models reach today's closed frontier, can only be answered later. SemiAnalysis's per-era clock starts at the closed model's launch and only asks about that era's first closed model, not about the closed frontier that kept moving.
The second is that all of these count release dates. A lab's internal models are ahead of its public ones, and the lead varies. SemiAnalysis puts GPT-4's training-to-release delay at 218 days and Mythos's at 114 (are-open-models-catching-up). Z.ai's product lead said in November 2025 that GLM models are open-sourced "within a few hours" of finishing (zai-playbook). A release-date gap therefore measures what the public can use, and it understates how far ahead the American labs' internal models are. Lambert says as much in glm-5-3-how-chinese-labs-keep-stride.
Estimates over time
| When | Who | Method | Estimate |
|---|---|---|---|
| Nov 2025 | Nathan Lambert, talk at The Curve (lambert-china-ai-ecosystem-open-model-gap) | Judgement, Epoch AI plots | Open models "months" behind the frontier, American open models years behind; he calls the lag a useful safety buffer |
| Feb 2026 | Lambert, open-models-in-perpetual-catch-up | Artificial Analysis index | About 6 months and "holding steady"; the likely future is a 6–9 month lag |
| May 2026 | Ihle, how-far-behind-are-open-models | Threshold crossings, 17 benchmarks | 8–10 months on private benchmarks, 4–6 on public; smallest around R1 (Jan 2025) and growing since |
| Jul 2026 | Artificial Analysis, deepseek-v4-flash-0731-artificial-analysis | Intelligence Index v4.1 | 4 points: Claude Opus 5 at 61, Kimi K3 at 57 |
| Aug 2026 | SemiAnalysis, are-open-models-catching-up | Era composites, own runs | Catch-up time 19.7, 8.5, then 4.8 months per era; GPT-5.6 at 99.1 against Kimi K3 at 93.0 on the agentic composite |
| Sep 2026 | Lambert, open-source-ai-reading-list | Summary of the above | Roughly 4–6 months, narrowed in recent years |
Lambert's February post gives two figures for the same thing, "~6month gap holding steady" near the top and "lag the best closed models by 6-9months" further down. The table keeps both.
Where they disagree
On size, the estimates agree more than it first seems. SemiAnalysis's 4.8 months is measured on four recent public benchmarks, and it falls inside Ihle's 4–6 month band for public benchmarks. Lambert's 6 months comes from an index built on public evaluations. The reading list's 4–6 months is the public-benchmark number.
Ihle's private-benchmark result is where the size breaks from the rest: 8–10 months, close to double. His explanation is that open developers filter benchmark data less carefully and train toward the tests more. That fits Lambert's February claim that the Artificial Analysis index is "a bit unrepresentative of the true frontier", and his remark in how-much-does-distillation-matter-for-chinese-llms that Chinese labs approach benchmarks in a way that makes them "appear that they're a bit closer than they really are". SemiAnalysis concedes the point in practice: Kimi K3 beats Fable 5 on its composite, and the firm still uses Fable for its own work. Practitioners outside the benchmark debate say the same about long unsupervised tasks (local-ai-is-not-opus).
Not every bias favours open models, though. Ihle identifies several that inflate the measured gap: a winner's-curse effect from benchmarks testing more closed models, bugs in third-party hosting of open models, and closed-side contamination on FrontierMath (OpenAI funded it and has seen most problems) and ARC-AGI (its semi-private set passes through commercial APIs).
On direction, the estimates do disagree. SemiAnalysis sees catch-up time halving every era. Ihle sees the gap at its smallest in January 2025 and wider since, on both public and private benchmarks. Lambert saw it holding steady. The methods explain part of this. SemiAnalysis has one datapoint per era, its clock starts before open models have joined the era, and open models join each new era sooner, which shortens the clock without saying anything about the current distance. Ihle's curves fit many datapoints but depend on how accept/reject judgements were made, and those were made largely by Claude Opus 4.7 and loosened by hand. Neither analysis can settle the other. The reading list's "has reduced in recent years" is true on SemiAnalysis's numbers, but on Ihle's it is true only if "recent years" goes back to 2023.
Cost frontier and performance frontier
A model can trail on intelligence and still be the best buy. Artificial Analysis's Intelligence-vs-cost chart draws a Pareto line through the cheapest model at each score. At the end of July 2026 two open models sat on it: DeepSeek V4 Flash 0731, which scores 50 for about 60% less per task than GPT-5.6 Luna at 51, and Kimi K3 near the top. Nothing open sat at the performance end of the line, where Claude Opus 5 was (deepseek-v4-flash-0731-artificial-analysis). Imperiale pushes this further in intelligence-vs-cost-linear: GLM-5.3 scores 60 at $0.49 per task against Fable 5.1's 66 at $3.69, and he doubts most users would notice the difference.
Which frontier matters depends on the buyer. Lambert's argument in open-and-closed-models-are-on-different-exponentials is that coding-agent users have shown they will pay large margins for the best model, because their output is visibly higher with it, while enterprises adopt open models once one clears a task's threshold and then rarely replace it. Seen that way, the performance gap sets the price of the top closed models and the cost gap decides adoption across everything else. SemiAnalysis frames the same question as the fear that open models close enough to the frontier will turn the model layer into a commodity (are-open-models-catching-up). A footnote in arguments-against-open-source-ai attacks the other side of that question, the claim that non-frontier markets are "just the frontier minus n-months" and so belong to whoever holds the frontier.
Why the leading open models are Chinese
Until 2024 the open frontier was mostly Meta's Llama. The SemiAnalysis charts show Llama-2 opening the first era and Llama-3.1-405B passing GPT-3.5 Turbo in July 2024. Llama-4 Maverick in 2025 then "squashed" the momentum DeepSeek R1 had built, and Lambert said at The Curve that a better Llama 4 would have changed that summer's discussion of the gap. From July 2024 onward, Ihle's Chinese-only analysis gives essentially the same curves as the full analysis, with a few exceptions. So from then on a Chinese model was nearly always the first open model across each threshold. In August 2026 SemiAnalysis's agentic composite had the best American open models, Inkling at 51.1 and Nemotron 3 Ultra at 32.7, far behind Kimi K3 at 93.0.
Lambert's explanations in glm-5-3-how-chinese-labs-keep-stride are mostly organizational. Z.ai releases within days, so it spends the months American labs spend on pre-release testing climbing benchmarks. Its financing depends on public benchmark scores more directly than OpenAI's or Anthropic's does. Its flagship is narrower (text-only, focused on agentic coding), which makes post-training easier. A Chinese market for RL data is growing, much of it sold by American data companies, so Chinese labs may be buying the same environments as the frontier labs. And the team is skilled, close to Tsinghua, and probably more compute-efficient than the American labs. open-models-in-perpetual-catch-up adds that Chinese labs build on each other's published ideas more openly than Silicon Valley does. For the structural and historical case see chinas-structural-advantage-in-open-source-ai and notes-from-inside-chinas-ai-labs.
The political explanation is llm-distillation: that Chinese labs are close only because they train on the outputs of American models. Lambert grants that distillation happens, including API abuse to extract reasoning traces, and that it helps most when a lab enters a new domain. He argues it is not the main factor and has become less important as RL takes over post-training, because RL needs on-policy generations from the lab's own model. He also thinks Fable 5 came too late to be a teacher for Kimi K3 or GLM-5.2 (how-distillation-is-used-today, the-distillation-panic). SemiAnalysis is less dismissive and counts distillation among the reasons "nothing stays secret forever". Kevin Xu, quoted in the-distillation-panic, turns the argument around: labs that rely on distillation to stay close may never learn what it takes to lead, so cutting them off could help them in the long run.
Does it close?
Three positions run through the sources.
The closing case belongs to SemiAnalysis: each era's lead has lasted about half as long as the previous one, frontier labs now release every 51 days without building a durable margin, and research advances leak or get reverse-engineered.
The steady case is Lambert's "perpetual catch-up". American labs have far more compute, data and users, but in a log-linear scaling regime those advantages buy only thin margins, which is why open models stay closer than expected. Closed models keep reaching new, more valuable tasks, though, so open models keep trailing. In his view open models would win only through a fundamental change, such as a way to merge and share expert models or a hundredfold cut in training cost (open-models-in-perpetual-catch-up).
The widening case has two parts. The first is measurement: Ihle's curves have risen since January 2025. The second is strategy. Lambert expects closed labs to hold their best models back from APIs, to protect token supply, avoid distillation and keep the high-margin uses, which would widen the release-date gap even if internal capability did not change (open-and-closed-models-are-on-different-exponentials). At The Curve he framed the open question as whether the American compute advantage "kicks in now and the gap really rebuilds", or whether Chinese labs stay close behind (lambert-china-ai-ecosystem-open-model-gap). He expects open developers eventually to stop chasing Claude and GPT on the Artificial Analysis index and serve niches instead. If they do, the index gap will stop measuring anything open developers are trying to close.
One effect could cut the other way. If self-improvement loops at the labs come to depend on user data, a lab that releases in hours collects that data months earlier than one that tests for months, which would favour the Chinese labs (glm-5-3-how-chinese-labs-keep-stride).
- Nathan Lambert on China's AI Ecosystem and the Open Model Gap
- Are Open Models Catching Up?
- The Arguments Against Open Source AI are Very Bad
- DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, 10 points above previous DeepSeek V4 Flash
- DeepSeek
- How far behind are open models?
- How much does distillation really matter for Chinese LLMs?
- Kimi K3: The open-weights escalation
- Kimi K3
- LLM Distillation
- Nonproliferation is the wrong approach to AI misuse
- Open models in perpetual catch-up
- Open-Source AI & Open Models Reading List
- 6 months to live for open models
- What comes next with open models