AI BenchmarksArtificial AnalysisOpenAIAnthropicAI EconomicsLLM Pricing

A Benchmark Update Erased a 3-Point Gap in 3 Days. At the New Tied Score, One Model Costs 57% Less.

On September 4, 2026, Artificial Analysis published Intelligence Index v4.2. Claude Fable 5.1 led it outright, GPT-6 Astra followed. On September 7, three days later, the same lab published v4.3: one benchmark swapped, the private-test-set weighting bumped from 40% to 45%, and Astra and Fable 5.1 came out tied at 53. Neither model was retrained, re-priced, or updated in between. Only the ruler moved. And at the score they now share, Astra's average cost per completed task is $3.26 against Fable 5.1's $7.63, 57% less for an identical result.

2026-09-09·13 min read

TL;DR

  • 📊 3 days, one ranking flip— v4.2 (Sep 4) had Claude Fable 5.1 leading GPT-6 Astra outright. v4.3 (Sep 7) ties them at 53. Neither model changed.
  • 🔧 What actually moved— Terminal-Bench upgraded v2.1 to v4.0, τ³-Banking replaced by AutomationBench-AA (a new 657-task Zapier collaboration), and the weight given to evaluations with private test sets rose from 40% to 45%.
  • 💰 The cost number underneath the tie— at the shared score of 53, Astra's average cost per Intelligence Index task is $3.26, Fable 5.1's is $7.63. 57% less for the same result.
  • 🖥️ Coding gap widened too— on the new Terminal-Bench v4.0, Astra scores 59.1% against Fable 5.1's 52.0% and Claude Opus 5's 49.0%, a 7-point lead that did not exist on the old benchmark version.
  • 🏭 Open-weight frontier held its ground— GLM-5.3-Flash (42) stayed the top open-weight score under the new methodology, ahead of Qwen3.8 2.4T A95B (40) and DeepSeek V4 Pro 0813 (36).
  • 📈 The real lesson isn't about these two models— it's that a leaderboard rank is a function of the test suite's current composition, not a fixed property of the model. Cost per completed task at the score you actually need is the number that survives a methodology update.

Two press releases, three days apart, from the same source

Artificial Analysis runs the Intelligence Index, a composite score built from evaluations spanning agentic work, coding, general knowledge, and scientific reasoning, and it is one of the most cited independent leaderboards in the industry. On September 4, 2026, it published v4.2, its own words: an “interim update” to keep pace with a fast-moving frontier, ahead of a full v5 release still in progress. That update added AA-Briefcase (an agentic knowledge-work evaluation with a private test set) and Surge AI's GDP.pdf (long-context document reasoning across 4,592 PDF pages), and retired GPQA Diamond for being saturated. Under that version, Claude Fable 5.1 led the index outright. GPT-6 Astra followed, described in Artificial Analysis's own summary as showing “a 4pt gain over GPT-5.6 Sol”, an improvement, but not a tie, and not a lead.

Three days later, on September 7, the same lab shipped v4.3. This time the change was narrower: Terminal-Bench, the evaluation that scores whether an agent can complete complex work through a real terminal, was upgraded from v2.1 to v4.0. τ³-Banking, a narrower agentic-workflow benchmark, was retired and replaced with AutomationBench-AA, a new evaluation built in collaboration with Zapier on a held-out set of 657 business workflow tasks across finance, HR, marketing, operations, sales, and support. And the share of the index's total weight coming from evaluations with private, unpublished test sets rose from 40% to 45%.

Intelligence Index Score: Same Two Models, 3 Days Apart

v4.2 (Sep 4) — Claude Fable 5.1Leads
v4.2 (Sep 4) — GPT-6 AstraTrails
v4.3 (Sep 7) — Claude Fable 5.153
v4.3 (Sep 7) — GPT-6 Astra53

Source: Artificial Analysis, 'Announcing Artificial Analysis Intelligence Index v4.2,' Sep 4, 2026 ('Claude Fable 5.1 leads the Index, followed by...GPT-6 Astra'); 'Announcing the Artificial Analysis Intelligence Index v4.3,' Sep 7, 2026 (both score 53).

Nobody retrained either model in that window. Nobody re-priced them. The only thing that changed between the two headlines was which tasks Artificial Analysis was averaging together and how heavily it weighted the ones neither lab can see in advance.

What the reweighting actually rewarded

The composite category weights themselves held steady across the update: Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20%. What changed sat one level down, inside those categories. Terminal-Bench v4.0 tests an agent across software, machine learning, science, operations, security, hardware, and media, run three times per task with average pass@1 reported, a harder and broader bar than v2.1. On that new version, GPT-6 Astra (max) scores 59.1%, ahead of Claude Fable 5.1 (max with fallback) at 52.0% and Claude Opus 5 (max) at 49.0%, and 19.2 percentage points ahead of GPT-5.6 Sol (max) at 39.9%. AutomationBench-AA, the new Zapier-built replacement for τ³-Banking, awards partial credit per completed objective but zeroes the task on any guardrail violation; there, Astra (max) scores 68.5%, ahead of Grok 4.6 (high) at 66.7% and GLM-5.3 (max) at 62.2%, completing every objective with zero guardrail violations on 41.6% of workflows against 32.1% for the nearest competitor.

Terminal-Bench v4.0: pass@1

GPT-6 Astra (max)59.1%
Claude Fable 5.1 (max, fallback)52.0%
Claude Opus 5 (max)49.0%
GPT-5.6 Sol (max)39.9%

Source: Artificial Analysis, 'Announcing the Artificial Analysis Intelligence Index v4.3,' Sep 7, 2026.

Both of the two new or upgraded evaluations happen to be exactly where Astra is strongest: agentic, tool-using, terminal- and workflow-shaped tasks, evaluated on private held-out sets it could not have optimized against. Swap the harness's hardest agentic tasks for a newer, tougher version and add a fresh business-automation suite nobody could game in advance, and a model built around agentic tool use closes a gap that existed on the older composite. That is not evidence of manipulation on Artificial Analysis's part, its own changelog is public and its stated goal is reducing saturation and gaming, not favoring a lab. It is evidence that “this model leads the leaderboard” is a claim about a specific, versioned test suite at a specific date, not a durable fact about the model.

The number underneath the tie

Here is the part that matters more than the rank. Artificial Analysis tracks an Intelligence-vs-Cost-per-Task Pareto frontier alongside the raw score, and at the tied score of 53, the two models are not tied on price. Astra's average cost per Intelligence Index task is $3.26. Fable 5.1's is $7.63. Same composite score, same test suite, 57% less spent per completed task.

# Cost per completed task, at an identical Intelligence Index score of 53
astra_cost_per_task = 3.26   # $, GPT-6 Astra (max)
fable_cost_per_task = 7.63   # $, Claude Fable 5.1 (max with fallback)

savings_pct = 1 - (astra_cost_per_task / fable_cost_per_task)
print(f"Astra costs {savings_pct:.0%} less per completed task at the same score")
# Astra costs 57% less per completed task at the same score
# Source: Artificial Analysis, "Announcing the Artificial Analysis Intelligence Index v4.3," Sep 7, 2026

Artificial Analysis's own framing of the frontier makes the pattern explicit: OpenAI occupies the majority of the cost-efficiency frontier across all five of Astra's reasoning effort levels, each the cheapest way to reach its respective intelligence level on the index. GLM-5.3-Flash (42), MiMo-V2.5-Pro (26), and Claude Fable 5.1 at its highest effort setting (53) round out the rest of that frontier. Four labs occupy it in total. Being tied for the top score and being the cheapest way to get that score are two separate claims, and only the second one changed hands cleanly, in either version of the index.

ModelIntelligence Index v4.3Cost per TaskDelta vs. Astra
GPT-6 Astra (max)53$3.26baseline
Claude Fable 5.1 (max, fallback)53$7.63+134% cost, same score
Claude Opus 5 (max)51not published in article2 points lower score
GLM-5.3-Flash42top open-weight, on frontierlower score, open weights

The open-weight frontier didn't move

One more data point the recalibration left untouched: GLM-5.3-Flash held its position as the strongest open-weight model on the index at 42, ahead of Qwen3.8 2.4T A95B at 40 and DeepSeek V4 Pro 0813 (max) at 36. Both the closed-weight tie at the top and the open-weight ranking below it moved for different reasons this cycle, one from a benchmark swap that favored a specific capability profile, the other from nothing changing at all. That contrast is itself the tell: when a score moves, check whether the model changed or the ruler did before treating the new number as news about the model.

Three days. One tied leaderboard position. Zero changes to either model's weights or price. If your routing logic, your procurement deck, or your “best model” blog post cites a single leaderboard score with no version number attached, it is already out of date the next time that lab ships an interim update, and interim updates are shipping roughly every few days right now.

What actually survives a methodology update

A composite intelligence score is a snapshot of a specific, versioned test suite. Cost per completed task, measured on the workload you actually run, is not immune to benchmark churn either, since it is usually reported per the same versioned suite, but it survives a reweighting better than a bare rank does, because it is denominated in dollars against a fixed unit of work rather than in points against a shifting composite. $3.26 versus $7.63 per completed Intelligence Index task tells you something actionable regardless of which version of the index produced the 53. “Tied for first” tells you something that can flip again in the next changelog, and per Artificial Analysis's own disclosure, this is the second interim update in a single week, with a full v5 release still coming.

The practical failure mode is the same one that shows up every time a lab or a benchmark publisher ships a headline number: hard-coding a routing or purchasing decision to a rank that is true only as of one dated snapshot, then never re-checking it once the version number ticks over. MegaBrain routes every call across 500+ models at zero markup and surfaces real, current cost-per-call data for your actual workload, not a leaderboard position frozen at whichever version was live the day you read the blog post, so when a benchmark update erases or opens a gap between two models overnight, that shows up as a number you can route around immediately, not a claim you have to remember to fact-check every time the test suite changes again.

Sign up at getmegabrain.com to route across GPT-6 Astra, Claude Fable 5.1, and 500+ other models with real, per-call cost visibility, benchmarked against the workload you are actually running, not the leaderboard snapshot from three days ago.

MegaBrain Gateway

500+ models. One API. No markup.

Use in Claude Code, Cline, Cursor, or any coding agent.

Try MegaBrain free →

Newsletter

Stay in the loop

Get the latest model comparisons and guides — no spam, unsubscribe anytime.