DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, 10 points above previous DeepSeek V4 Flash
- title
- DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, 10 points above previous DeepSeek V4 Flash
- type
- summary
- summary
- Artificial Analysis puts DeepSeek V4 Flash 0731 at 50, one point under GPT-5.6 Luna for ~60% less per task and on the cost Pareto frontier
- tags
- ai, open-weights, benchmarks, llm, china, cost
- created
- 2026-09-14
- updated
- 2026-09-14
Artificial Analysis published this short results note on 31 July 2026, the day deepseek put a new checkpoint of its small V4 model on its API. DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index v4.1, ten points above the April release of V4 Flash and six above the much larger V4 Pro. Nathan Lambert's open-source-ai-reading-list uses it as the example for one claim: open models sit on the Pareto cost frontier even where they are not at the absolute performance frontier. The note has the numbers to show both halves of that.
Where 50 sits
The index is a composite of nine evaluations that Artificial Analysis runs itself: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience and AA-LCR. At 50, V4 Flash 0731 ties Gemini 3.6 Flash and sits one point behind GPT-5.6 Luna (max), GLM-5.2 (max) and Muse Spark 1.1 (xhigh). The open-weights leader is kimi-k3 (max) at 57. The note does not name the closed leader, but its chart does: Claude Opus 5 (max) at 61, Claude Fable 5 at 60 and GPT-5.6 Sol (max) at 59. On this index, at the end of July 2026, the best open model was four points behind the best closed one, and V4 Flash 0731 eleven.
The cost half
What gets the model onto the frontier is price. Even after OpenAI cut GPT-5.6 Luna's price by 80% that same day, V4 Flash 0731's cost per index task on DeepSeek's first-party API comes out about 60% lower than Luna (max), a model with one point more intelligence. Artificial Analysis names the main driver: DeepSeek discounts cache hits by about 98%, where most of the industry discounts them by 90%. List pricing is $0.14 per million input tokens and $0.28 per million output, with cache hits at $0.0028, all unchanged from the earlier V4 Flash. The model also got cheaper to run on the index without any price change, because it used about 206M output tokens against its predecessor's 234M.

The lower panel is the chart behind Lambert's example. The dotted Pareto line runs from MiMo-V2.5 near $0.01 per task, through V4 Flash 0731 at roughly $0.03 and GPT-5.6 Luna (max) at roughly $0.07, up to Kimi K3 (max) near $0.85 and Claude Opus 5 (max) near $2.30. Two open models are on that line, and the one at the top end is still four points short of Opus 5. The cost axis is logarithmic, which is the first thing intelligence-vs-cost-linear objects to: a linear axis would show a much larger dollar gap between the cheap end and the expensive end.
Nothing about the model's shape changed. It has the same architecture as V4 Flash, 284B total parameters with 13B active, a 1M-token context window, and text input and output only.
What improved
The biggest gain is agentic. On GDPval-AA v2, Artificial Analysis's Elo-rated evaluation of real-world work tasks, the model rises from 1189 to 1559. Once the weights are released, that will be the second-highest open-weights score, behind Kimi K3 (max) at 1687 and ahead of GLM-5.2 (max) at 1510. Terminal-Bench 2.1 rises 17 points to 79% and τ³-Banking 8 points to 31%. Every other evaluation in the index also went up: CritPt by 9 points to 17%, SciCode by 5 to 50%, Humanity's Last Exam by 5 to 37%, AA-LCR by 3 to 66% and GPQA Diamond by 1 to 91%.
The knowledge result is worth reading carefully, because the headline number moved and the underlying ability did not. The AA-Omniscience Index rose 7 points to -16, but accuracy stayed at 37%. All of the gain came from answering wrongly less often: the hallucination rate, the share of non-correct responses that were confident wrong answers rather than refusals, fell to 84%. Artificial Analysis reads the unchanged accuracy as what you would expect from a model whose parameter count did not change, which fits the idea that factual recall is bounded by model size (factual-capacity-scaling). An 84% hallucination rate is still near the bottom of the chart, comparable to GPT-5.6 Terra (max) at 85%.
Caveats for the open-model reading
At publication, V4 Flash 0731 was not yet an open-weights model. It was available only through DeepSeek's API, and Artificial Analysis says DeepSeek "is expected to release the model's full weights in the coming weeks". Using it as the open example on the cost frontier assumes those weights arrive with the same behaviour.
The cost advantage is also measured on DeepSeek's own API, and the note credits the cache discount for much of it. Imperiale's objection in intelligence-vs-cost-linear is that Artificial Analysis prices open models at their developer's API when third-party hosts are often cheaper. Here the first-party API is the cheap option, and a host without a comparable cache discount would not necessarily reproduce the 60% figure. For the broader question of how far open models trail, see open-closed-model-gap.
Source defects
The note gives the hallucination-rate drop as 12 points in the text and 11 points in a chart caption. The chart shows the predecessor at 96%, so 12 is correct. The image captions in the clip are also offset by one: each caption sits above the image it describes, not below it.