DeepSeek Gained 10 Points on the Top AI Benchmark Overnight, for Free. Part of the Score Is 0% Smarter.
On July 31, 2026, DeepSeek shipped an update called V4 Flash 0731: same 284B/13B parameter count, same 1M-token context, same $0.14/$0.28 per-million-token price as the model it replaced. Artificial Analysis tested it anyway. The Intelligence Index score went from 40 to 50, a bigger single-release jump than DeepSeek's own flagship Pro tier scores in total. Some of that gain is genuine: +370 Elo on the agentic-work benchmark, +17 points on real command-line tasks. But the honesty sub-score moved entirely because the model got better at saying “I don't know,” with raw accuracy flat at 37%, unchanged to the decimal. Here's the full breakdown of what actually moved, and why the same 10-point number now means two different things at once.
TL;DR
- 📈 +10 points, $0 more— DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, up from 40 for the April 2026 model it replaced. Same 284B total / 13B active parameters, same 1M-token context, same $0.14/$0.28 price per million input/output tokens.
- 🤖 The agentic gains are real— GDPval-AA v2 Elo jumped from 1189 to 1559 (+370, against a human baseline of 1000). Terminal-Bench 2.1 rose from 62% to 79%. It improved on every one of the index's evaluations, not just one lucky test.
- 🎭 The honesty score is a different animal— AA-Omniscience improved 7 points, but Artificial Analysis states plainly that the gain is “purely driven by a reduced hallucination rate, with overall accuracy unchanged.” Hallucination rate: 96% → 84%. Raw accuracy: 37% → 37%.
- 💰 It undercuts OpenAI in the same week OpenAI cut prices— Hours after OpenAI slashed GPT-5.6 Luna's price 80% (from $1/$6 to $0.20/$1.20 per million tokens), DeepSeek V4 Flash 0731's cost per task still ran roughly 60% below Luna's, for a nearly identical Intelligence Index score.
- 🏢 It beat DeepSeek's own flagship, for free— V4 Flash 0731's score of 50 is 6 points ahead of DeepSeek V4 Pro's 44, the more expensive tier in the same lineup.
An update that, on paper, shouldn't have moved anything
Model updates usually come with a reason attached: more parameters, more training tokens, a new architecture. DeepSeek V4 Flash 0731 has none of those. Artificial Analysis's write-up, published July 31, 2026, confirms the new release “shares identical architecture and pricing with the earlier DeepSeek V4 Flash.” Total parameters stay at 284B, active parameters at inference stay at 13B, the context window holds at 1M tokens, and the API price is untouched at $0.14 per million input tokens and $0.28 per million output tokens, with the same 98% cache-hit discount DeepSeek has offered since the original release.
| Spec | V4 Flash (Apr 2026) | V4 Flash 0731 (Jul 2026) |
|---|---|---|
| Total parameters | 284B | 284B |
| Active parameters | 13B | 13B |
| Context window | 1M tokens | 1M tokens |
| Input price | $0.14 / 1M tok | $0.14 / 1M tok |
| Output price | $0.28 / 1M tok | $0.28 / 1M tok |
| Intelligence Index | 40 | 50 |
Every row is identical except the last one. A model that costs the same to run scored 10 points higher on the industry's most-cited aggregate benchmark, which currently places Claude Opus 5 at the top around 63 and spans dozens of models from roughly 20 to that ceiling. A 10-point jump at zero cost is not a rounding error on that scale; it's the gap between a mid-tier model and one landing one point behind GPT-5.6 Luna, OpenAI's own flagship-efficient tier.
Where the gains are real: agentic work and coding, provably up
Split the Intelligence Index into its component evaluations and most of the movement is unambiguous capability, not noise. GDPval-AA v2, Artificial Analysis's benchmark for real-world agentic work, re-baselines its Elo scale so that 1000 equals average human performance on the underlying tasks. DeepSeek V4 Flash 0731 scores 1559 on it, up from 1189 for the model it replaced, a 370-point jump in a single release, which Artificial Analysis notes will land as the second-highest open-weights score once the weights are public, behind Kimi K3 (1687) and ahead of GLM-5.2 (1510).
GDPval-AA v2: agentic real-world work, re-baselined to human = 1000
Source: Artificial Analysis, Jul 31, 2026
Terminal-Bench 2.1, which scores an agent's ability to complete real command-line tasks, moved from 62% to 79%, a 17-point gain. And the improvement wasn't concentrated in one favorable test: every single evaluation in the index rose, including τ³-Bench Banking (+8 points), CritPt graduate-physics reasoning (+9 points), SciCode scientific coding (+5 points), Humanity's Last Exam (+5 points), AA-LCR long-context reading (+3 points), and GPQA Diamond (+1 point). It even got more efficient doing it: total output tokens used to run the full Intelligence Index battery fell 12%, from roughly 234 million to 206 million.
| Evaluation | V4 Flash (Apr) | V4 Flash 0731 (Jul) | Change |
|---|---|---|---|
| GDPval-AA v2 (Elo) | 1189 | 1559 | +370 |
| Terminal-Bench 2.1 | 62% | 79% | +17 pts |
| τ³-Bench Banking | 23% | 31% | +8 pts |
| CritPt | 8% | 17% | +9 pts |
| SciCode | 45% | 50% | +5 pts |
| Humanity's Last Exam | 32% | 37% | +5 pts |
| AA-LCR | 63% | 66% | +3 pts |
| GPQA Diamond | 90% | 91% | +1 pt |
None of that reads like an artifact. A model that completes almost 1 in 5 more real terminal tasks and gains 370 points of agentic-work Elo, at unchanged parameter count and price, is a genuinely better tool for the jobs those benchmarks approximate. If the story stopped here, it would just be a good week for DeepSeek.
The score that isn't what it looks like
It doesn't stop there, because part of that 10-point Intelligence Index gain flows through AA-Omniscience, Artificial Analysis's test for whether a model knows the boundary of its own knowledge, and the mechanism behind that particular sub-score is not the same as the mechanism behind GDPval-AA v2 or Terminal-Bench. Artificial Analysis is explicit about it: DeepSeek V4 Flash 0731's AA-Omniscience Index improves from -23 to -16, a 7-point gain, but the write-up states the improvement is “purely driven by a reduced hallucination rate, with overall accuracy (percentage correct) unchanged.” The hallucination rate itself falls from 96% to 84%, an 11-point improvement (Artificial Analysis's summary line puts it at 12 points falling to 84%; both figures appear in the source and describe the same delta). Raw accuracy holds at 37% in both versions, “consistent,” the article notes, “with the model being unchanged in size at 284B total parameters.”
AA-Omniscience: what actually moved
Source: Artificial Analysis, Jul 31, 2026
That is a real and useful skill. A model that says “I don't know” instead of confidently inventing an answer is safer to deploy in anything user-facing, and calibration has been a genuine, hard-won research problem for years. But it is a categorically different achievement from “the model knows more.” Zero new facts were learned. What changed is the model's willingness to guess, and that distinction gets fully absorbed into a single composite number the moment it enters the Intelligence Index, indistinguishable on the leaderboard from the very real capability gains sitting right next to it in GDPval-AA v2 and Terminal-Bench.
This isn't a knock on Artificial Analysis's methodology, which documents the split transparently in its own write-up. It's a warning about how that write-up gets consumed: as a single delta, “+10 points,” stripped of the breakdown, in a tweet, a Slack message, or a routing config comment. The composite score doesn't carry the asterisk with it once it leaves the source.
Why this matters more than a benchmark curiosity
The timing sharpens the stakes. OpenAI cut GPT-5.6 Luna's API price by 80% the same week, from $1 and $6 per million input/output tokens down to $0.20 and $1.20, repositioning it as the aggressive low-cost option in its lineup. DeepSeek V4 Flash 0731 lands one Intelligence Index point behind Luna's max configuration (50 vs. 51), but Artificial Analysis calculates that even after that price cut, DeepSeek's first-party cost per task still runs roughly 60% lower than Luna's, driven in part by DeepSeek's 98% cache-hit discount against a typical industry discount closer to 90%.
Cost per task, relative: DeepSeek V4 Flash 0731 vs. GPT-5.6 Luna, after Luna's price cut
Source: Artificial Analysis, Jul 31, 2026; OpenAI pricing update, Jul 30, 2026
Inside DeepSeek's own lineup, the comparison is even starker. V4 Flash 0731's score of 50 now sits 6 points above DeepSeek V4 Pro's 44, the company's more expensive tier, meaning a customer who upgrades to the pricier model this week gets a worse Intelligence Index score than the free update to the cheap one. None of that 6-point gap is in question, it's the same benchmark family, same methodology, two DeepSeek models tested the same way. But teams that route traffic by Intelligence Index score, whether by hand or through an automatic router, are about to treat a partly-calibration-driven number as equivalent to GDPval-AA v2's 370-point, clearly-capability jump. For anything where a model's job is to know facts, not just refuse gracefully when it doesn't, that distinction changes which model is actually the better pick.
This isn't a DeepSeek-only story
The broader pattern predates this release. Anthropic's Opus 4.7 moved from 80.8% to 87.6% on SWE-Bench Verified largely through post-training refinement rather than a base model size increase, and small open models like VibeThinker-3B have shown large reasoning-benchmark gains built entirely on top of an unchanged base checkpoint. 2026 has been the year reinforcement-learning post-training, not pretraining scale, produces the visible leaderboard movement. DeepSeek V4 Flash 0731 is the clearest single data point yet that this shift cuts two ways: post-training can teach a model to reason and act better, which is what GDPval-AA v2 and Terminal-Bench captured here, and it can separately teach a model to hedge better, which is what AA-Omniscience captured, and a composite benchmark score has no obligation to tell you which one you just watched happen.
How to check this yourself before trusting a leaderboard delta
Artificial Analysis's per-evaluation tables are public, and the fastest gut-check on any single-number benchmark claim is to pull the sub-scores instead of the composite. For pricing and model metadata specifically, OpenRouter's public API surfaces DeepSeek's current listings directly, which is useful for confirming a price claim like the one in this article hasn't drifted since publication:
curl -s "https://openrouter.ai/api/v1/models" \
| jq '.data[] | select(.id | test("deepseek.*v4.*flash"; "i")) | {
id,
context_length,
pricing: .pricing
}'When a benchmark write-up is available, the more important habit is reading past the composite score to the sub-benchmark table, specifically checking whether a claimed gain sits in a capability evaluation (agentic task completion, coding, tool use) or a calibration/honesty evaluation (refusal rate, hallucination rate). The two are not interchangeable inputs to a routing decision, even when they land in the same index.
What this means if you route models by score
DeepSeek V4 Flash 0731 is, on net, a better model than its predecessor at identical cost, and the agentic gains here are large enough to matter for anyone running coding or real-world-task agents. That part of the story is straightforwardly good news. But “10 points higher on the Intelligence Index” is no longer a single, legible fact the way it might have been two years ago, back when most benchmark movement came from bigger, more capable base models. In a year where post-training increasingly drives the visible score, the same delta can mean “meaningfully smarter,” “meaningfully more honest,” or some blend of both, and only reading the sub-scores tells you which one you're paying for.
If you're routing production traffic across models and want to see the sub-benchmark breakdown behind a score instead of trusting the composite number, MegaBrain routes every request through a single gateway with transparent, at-cost pricing across 500+ models, so a capability upgrade and a calibration upgrade show up as two different line items instead of one indistinguishable leaderboard delta. And if you want something watching new benchmark releases for exactly this kind of asterisk on a schedule, mhermesruns 24/7 on its own isolated VM and can poll a source like Artificial Analysis's article feed and flag the next time a headline number and its sub-scores tell two different stories.
MegaBrain Gateway
500+ models. One API. No markup.
Use in Claude Code, Cline, Cursor, or any coding agent.
Newsletter
Stay in the loop
Get the latest model comparisons and guides — no spam, unsubscribe anytime.