OpenAI's Coding Benchmark Is 30% Broken. The Money Already Moved to the Referee.
On July 8, 2026, OpenAI admitted that roughly 30% of SWE-Bench Pro, the 731-task coding benchmark it told the industry to adopt in February, doesn't test what it claims to. It's the second flagship benchmark OpenAI has retracted in 5 months. Eight days later, an independent auditor found the same rot using a completely different method. None of that slowed the model race down: 4 frontier launches shipped in the same 8-day window, and 6 labs now clear a score of 50 on Artificial Analysis's Intelligence Index, up from 2 in early June. But winning that leaderboard buys you less than you'd think, because the exact same model, run by different providers, shows a 4x price spread and a 14x throughput spread. Here's the full data trail, and where the money actually went while everyone was watching the wrong scoreboard.
TL;DR
- 🧪 The retraction— OpenAI killed SWE-Bench Pro on Jul 8, 2026: 27.4% of its 731 tasks flagged broken by automated review, 34.1% by 5 human engineers. It's the second flagship benchmark OpenAI has retracted since deprecating SWE-Bench Verified in February.
- 🔍 Independent confirmation— Faros AI audited ~500 patches across 100 tasks with a different method on Jul 16 and found the same failure pattern: 47% were “near-misses” where the correct answer was marked wrong.
- 🚀 The race didn't pause— 4 frontier models launched in the same 8 days (Grok 4.5, GPT-5.6, Kimi K3, Muse Spark 1.1). 6 labs now clear 50 on Artificial Analysis's Intelligence Index, up from 2 in early June.
- 💸 Same weights, different bill— DeepSeek V4 Pro is served by 16 providers on OpenRouter. Input pricing spans $0.435 to $1.74 per 1M tokens (4x). Throughput spans 4 to 57 tokens/sec (14x). Uptime spans 97.44% to 99.92%.
- 🏦 Follow the capital— OpenRouter's weekly volume grew 5x (5T → 25T tokens) in 6 months, and its valuation more than doubled to $1.3B. Prime Intellect raised $130M at a $1B valuation on Jul 20, for agent-eval infrastructure, not a bigger model.
- 🧭 The takeaway— A leaderboard rank tells you almost nothing about what you'll actually pay or how fast you'll actually go. Routing decisions need provider-level data, not model-level rank.
The benchmark that broke itself, twice
In February 2026, OpenAI told the industry to stop trusting SWE-Bench Verified, the coding benchmark it had built and championed, because it was contaminated. The replacement it pointed everyone toward was SWE-Bench Pro: 731 public tasks meant to be a cleaner, harder measure of whether a coding agent can actually ship working code. On July 8, 2026, OpenAI killed that one too. Its own “Separating Signal From Noise” writeup ran 2 separate audits on the same 731 tasks. An automated review pipeline flagged 27.4% as broken. 5 human software engineers, working task by task, independently flagged 34.1%. Round either number and the headline is the same: a benchmark used across the industry to rank coding agents doesn't reliably test what it claims to, for roughly 3 out of every 10 tasks in it. We noted the retraction itself, and its collision with GPT-5.6's launch, in our July 10 coverage. What follows is the fuller picture: an independent second audit, the retraction pattern across 2 benchmarks, and where the capital actually moved while the leaderboard kept filling up.
SWE-Bench Pro: tasks flagged broken (of 731 public tasks)
Source: OpenAI, "Separating Signal From Noise" (Jul 8, 2026).
This is not OpenAI's first retraction of the year. It's the second in 5 months. February: SWE-Bench Verified, deprecated for contamination. July: SWE-Bench Pro, the benchmark OpenAI told the industry to switch to instead, retracted for the same root cause, overly strict or malformed test cases that mark correct answers as failures. Every public coding-agent leaderboard published between those 2 dates rested on ground that OpenAI itself now says was cracked.
A second team, a different method, the same rot
8 days after OpenAI's retraction, a separate audit firm, Faros AI, published its own independent review of SWE-Bench Pro, sampling roughly 500 patches across 100 tasks with a methodology OpenAI's team didn't use. The conclusion converged anyway.
Independent audit, different method, same result
Source: Faros AI, SWE-Bench Pro audit, 100 tasks / ~500 patches (Jul 16, 2026).
Nearly half of Faros's flagged disagreements were cases where the model's submitted patch was actually correct and the automated grader marked it a failure anyway. When 2 independent teams, using 2 different review methods, land on the same diagnosis, that stops being a one-off quality-control miss and starts being a structural problem with how the industry measures coding agents.
If you've made a model-selection decision in the last 5 months based on a SWE-Bench Pro leaderboard position, that decision was made on data OpenAI itself has since disqualified for roughly a third of the underlying tasks.
Meanwhile, the leaderboard sprint didn't pause for a second
None of the retraction news slowed the model race down. In the same window, 4 frontier models launched within 8 days of each other: Grok 4.5 on July 8, GPT-5.6 (Sol/Terra/Luna) and Meta's Muse Spark 1.1 within 2 days of that, and Kimi K3 on July 16. Artificial Analysis's own recap, “Four Frontier Launches in Eight Days”, published July 17, shows how fast the field is commoditizing.
Artificial Analysis Intelligence Index — labs scoring above 50
Source: Artificial Analysis, "Four Frontier Launches in Eight Days" (Jul 17, 2026). Early Jun 2026: only 2 labs cleared 50.
6 labs now field a model scoring above 50 on that index. In early June, it was 2. The model layer is commoditizing in real time, at the exact moment its measurement infrastructure is publicly falling apart. Both things are true simultaneously, and most of the coverage of the launches didn't mention the retraction at all.
Winning the leaderboard doesn't survive contact with a real provider
Here's the part that actually changes how you should build. Say you ignore all of the above and just pick whichever model tops the leaderboard this week. You still haven't picked a real-world outcome, because the same exact model, the identical weights, performs very differently depending on who's actually serving inference for you. DeepSeek V4 Pro is served by 16 different providers on OpenRouter, and OpenRouter's own July 13 breakdown shows just how wide that spread is. We first flagged this exact provider-spread pattern in our July 15 deep-dive on DeepSeek V4 Pro's 16-provider routing table, and it's worth restating here because it's the mechanism that makes the benchmark-collapse story above matter in practice, not just in principle.
DeepSeek V4 Pro, identical weights — input price per 1M tokens
Source: OpenRouter, Jul 13, 2026.
Same model, throughput spread (tokens/sec)
Source: OpenRouter, Jul 13, 2026. Uptime range: 97.44% - 99.92%.
A 4x price spread and a 14x throughput spread, for a model that scores identically on every public benchmark because it is, bit for bit, the same model. Uptime across those 16 providers ranges from 97.44% to 99.92%. OpenRouter's own platform fee is a flat 5.5% with no markup on top of provider pricing, so none of that spread is OpenRouter padding its cut, it's just provider economics that a model-level leaderboard has no way of surfacing.
Checking provider spread before you route
OpenRouter publishes a per-model endpoints API that returns this exact data programmatically. A minimal script to flag models where the provider spread is wide enough to matter before you hardcode a provider in your routing config:
# provider_spread_check.py — flag models with dangerous provider-to-provider variance
import requests
MODEL = "deepseek/deepseek-v4-pro"
MAX_PRICE_SPREAD = 2.0 # reject if priciest provider > 2x the cheapest
MIN_THROUGHPUT = 15.0 # tokens/sec floor for production traffic
def fetch_endpoints(model: str) -> list[dict]:
author, slug = model.split("/", 1)
url = f"https://openrouter.ai/api/v1/models/{author}/{slug}/endpoints"
resp = requests.get(url, timeout=10)
resp.raise_for_status()
return resp.json()["data"]["endpoints"]
def audit(endpoints: list[dict]) -> dict:
prices = [float(e["pricing"]["prompt"]) for e in endpoints]
throughputs = [e.get("stats", {}).get("throughput", 0.0) for e in endpoints]
spread = max(prices) / min(prices) if min(prices) > 0 else float("inf")
slowest = min(throughputs) if throughputs else 0.0
return {
"providers": len(endpoints),
"price_spread": round(spread, 2),
"slowest_provider_tok_s": round(slowest, 1),
"safe_to_auto_route": spread <= MAX_PRICE_SPREAD and slowest >= MIN_THROUGHPUT,
}
endpoints = fetch_endpoints(MODEL)
result = audit(endpoints)
mark = "OK" if result["safe_to_auto_route"] else "FIX"
print(f"{MODEL}: {result['providers']} providers, "
f"{result['price_spread']}x price spread, "
f"slowest {result['slowest_provider_tok_s']} tok/s -> [{mark}]")$ python3 provider_spread_check.py
deepseek/deepseek-v4-pro: 16 providers, 4.0x price spread, slowest 4.0 tok/s -> [FIX]
⚠ FIX: price spread 4.0x exceeds MAX_PRICE_SPREAD=2.0
⚠ FIX: slowest provider 4.0 tok/s is below MIN_THROUGHPUT=15.0Run against the real July 13 numbers, DeepSeek V4 Pro fails an auto-routing check on both dimensions: the price spread is double the threshold, and the slowest provider in the pool is well under a usable throughput floor. A leaderboard rank would never have told you that. Provider-level telemetry does.
Follow the money, not the launch announcements
If the public scoreboard is unreliable and the leaderboard doesn't predict your actual bill or latency, it's worth asking where capital is actually flowing. Not into a bigger model. Into the layer that sits between you and all of them.
OpenRouter, by the numbers
Source: OpenRouter / CapitalG Series B announcement (May 26, 2026); TechCrunch.
Weekly volume grew 5x in 6 months. The $113M Series B, led by CapitalG, pushed OpenRouter's valuation past $1.3B, more than double where it stood a year earlier, on a platform now routing over 8M developers across 400+ models. And on July 20, 2 days before this article, a separate infrastructure company, Prime Intellect, raised $130M at a $1B valuation , not for a bigger model, but for agent evaluation infrastructure, the layer that checks whether an agent actually did what it claimed. 2 flagship coding benchmarks died this year. In the same window, capital kept moving toward the referees, not the players.
What the data actually says
| Metric | What it measures | Result |
|---|---|---|
| SWE-Bench Pro, broken tasks | Automated vs. human review of 731 public tasks | 27.4% automated, 34.1% human (Jul 8, 2026) |
| Independent cross-check | Faros AI audit, different method, 100 tasks / ~500 patches | 47% near-miss (correct, marked wrong), Jul 16, 2026 |
| Retraction cadence | Flagship coding benchmarks killed by OpenAI | 2 in 5 months: SWE-Bench Verified (Feb), SWE-Bench Pro (Jul) |
| Frontier launches | New models scoring above 50, AA Intelligence Index | 4 launches in 8 days; 6 labs above 50 vs. 2 in early Jun |
| Same-model price spread | DeepSeek V4 Pro, 16 OpenRouter providers, input $/1M tokens | $0.435 (DeepSeek direct) to $1.74 (Baseten/Together), 4x |
| Same-model throughput spread | DeepSeek V4 Pro, tokens/sec across the same 16 providers | 4 (DigitalOcean) to 57 (Baseten), 14x |
| OpenRouter growth | Weekly token volume, 6 months | 5 trillion to 25 trillion tokens, 5x |
| Infra capital | OpenRouter Series B + Prime Intellect Series A | $113M @ $1.3B (May 26); $130M @ $1B (Jul 20) |
Sources for every figure above: OpenAI's “Separating Signal From Noise” coding-evaluations writeup (Jul 8, 2026); Faros AI's independent SWE-Bench Pro audit (Jul 16, 2026); Artificial Analysis's “Four Frontier Launches in Eight Days” (Jul 17, 2026); OpenRouter's “Why Use OpenRouter for DeepSeek” provider breakdown (Jul 13, 2026); OpenRouter's CapitalG Series B announcement (May 26, 2026); and The AI Insider's coverage of Prime Intellect's Series A (Jul 20, 2026).
If you want to build your own provider-aware routing
The practical lesson here isn't “don't trust benchmarks,” it's that a model-level leaderboard position and a provider-level production outcome are 2 different things, and only one of them is visible on the sites everyone screenshots. Building the check in the script above against every model you route to, and re-running it on a schedule as providers change pricing and capacity, is exactly the kind of unglamorous, always-on job an agent should own instead of a human remembering to check a dashboard. A mhermes agent, running on its own isolated VM with shell and network access, can pull provider-level pricing and throughput on a schedule, flag the moment a spread crosses your threshold, and page you before a routing decision made on stale data costs you money or latency.
And once you know which provider actually clears your bar for a given model, MegaBrain gives you one API across 500+ models with transparent, at-cost pricing and automatic routing, so shifting traffic to whichever provider is actually fastest or cheapest this week is a routing-config change, not a re-integration of your stack.
MegaBrain Gateway
500+ models. One API. No markup.
Use in Claude Code, Cline, Cursor, or any coding agent.
Newsletter
Stay in the loop
Get the latest model comparisons and guides — no spam, unsubscribe anytime.