AI BenchmarksAI EconomicsAI AgentsAI InfrastructureArtificial Analysis

Anthropic Is Raising Sonnet 5's Price 50%. Its Own Benchmark Says Real Agent Cost Just Fell 20%.

On September 1, 2026, Claude Sonnet 5's API price jumps from $2/$10 to $3/$15 per million input/output tokens, a clean +50%on every token, per Anthropic's own pricing docs. Nine days before that change was even announced, Artificial Analysis published a very different number from the same company: Claude Opus 5 beating Claude Fable 5 by 146 Elo points on AA-Briefcase, its agentic knowledge-work benchmark, for 20% less per finished task. One lab, one quarter, two numbers pointing in opposite directions. Only one of them tells you what your agents will actually cost.

2026-07-30ยท15 min read

TL;DR

  • ๐Ÿ’ธ The hike โ€” Claude Sonnet 5 moves from $2/$10 to $3/$15 per million input/output tokens on Sep 1, 2026, a flat +50%. Confirmed directly on Anthropic's own pricing page.
  • ๐Ÿ“‰ The inversion โ€” On Jul 24, 2026, Claude Opus 5 beat Claude Fable 5 by 146 Elo on AA-Briefcase (1,720 vs 1,574, max effort) while costing $17.79/task vs $22.30/task, 20% less. At "high" effort, Opus 5 still beats Fable 5 for $10.41/task, under half the price.
  • ๐Ÿ“ The spread โ€” across every model AA-Briefcase tested, cost per finished task varies 800x: DeepSeek V4 Flash at $0.04/task, Claude Fable 5 averaging $31/task, same benchmark, same day.
  • ๐Ÿชค The trap โ€” Kimi K3 (Jul 21, 2026) prices at $3/$15 per million tokens, the exact sticker Sonnet 5 moves to. But it needs 83 turns and 120,000 output tokens to finish one AA-Briefcase task: $10.57/task for a 51% rubric pass rate, worse than Opus 5 at essentially the same price.
  • ๐Ÿงฎ The cross-check โ€” a completely different Artificial Analysis benchmark (the general Intelligence Index), same 8-day launch window, shows the identical pattern: a 5x cost-per-task spread (GPT-5.6 Luna $0.21 to GPT-5.6 Sol $1.04) that tracks nothing about the models' per-token rate cards.
  • ๐Ÿงญ The takeaway โ€” $/task = turns-to-completion ร— tokens-per-turn ร— $/token. The rate card is only the last term. Picking a model off its pricing page, without checking the other two, is how teams end up overpaying without knowing it.

The number every launch quotes

Every model that shipped in the last few weeks came with a price-per-million-token headline, and all of them pointed the same direction: down. Grok 4.5 landed at $2/$6. Meta's Muse Spark 1.1 came in at $1.25/$4.25. OpenAI's smallest GPT-5.6 variant, Luna, priced at $1/$6. Ten new model families landed on OpenRouter in the 22 days between July 1 and July 22, 2026 alone. Reading only that number, the story writes itself: AI is getting radically cheaper, month over month.

But an agent doesn't buy tokens. It buys finished tasks. Two models with wildly different per-token rate cards can cost exactly the same to finish a job, or ten times different, depending on how many turns and how many tokens each one burns getting there. Cost per task is the number that actually predicts an agent bill, and it's the one number that never makes it into a launch post. So we pulled it, from the primary sources that publish it.

Headline price per million tokens, July 2026 launches

Grok 4.5 โ€” input$2 / M tok
Muse Spark 1.1 โ€” input$1.25 / M tok
GPT-5.6 Luna โ€” input$1 / M tok
Sonnet 5 โ€” input (through Aug 31)$2 / M tok
Sonnet 5 โ€” input (from Sep 1)$3 / M tok

Sources: SpaceXAI, Meta, OpenAI launch pricing; Anthropic pricing docs.

The 800x spread nobody puts in a press release

Artificial Analysis built AA-Briefcase precisely to measure this: 91 real knowledge-work tasks across banking operations, product management, and heavy-industry strategy, multi-week projects built from roughly 2,000 source files, 3,500+ emails and 25,000+ Slack messages, designed with experts from Google, McKinsey, and BCG. Every submission is graded on rubric pass/fail checks plus head-to-head analytical-quality and presentation-quality comparisons, aggregated into a single Elo score. Cost per task is calculated from actual token usage split across input, cache-hit, cache-write, reasoning, and answer tokens, at each model's real pricing.

Across every model the benchmark has tested, cost to finish one task varies by more than 800x. DeepSeek V4 Flash finishes a task for 4 cents. Claude Fable 5 averages $31 for the same class of work.

AA-Briefcase: cost per finished task (log spread)

DeepSeek V4 Flash (max)$0.04 / task
Claude Fable 5 (avg)$31.00 / task

Source: artificialanalysis.ai/articles/aa-briefcase (Jun 18, 2026).

Opus 5 breaks the assumption

On July 24, 2026, Anthropic shipped Claude Opus 5, and it broke the pattern most people assume holds by default: that a smarter model costs more. At max effort, Opus 5 scored an AA-Briefcase Elo of 1,720, 146 points ahead of Fable 5's 1,574, for $17.79per task against Fable 5's $22.30. Twenty percent cheaper, and smarter, at the same time. Dial the effort setting down one notch and the gap widens further: Opus 5 on โ€œhighโ€ effort still outscores Fable 5 outright, for $10.41 per task, under half of Fable 5's price.

Opus 5 vs. Fable 5 on AA-Briefcase, max effort

Fable 5 โ€” Elo1,574
Opus 5 โ€” Elo1,720
Fable 5 โ€” cost/task$22.30
Opus 5 โ€” cost/task$17.79

Source: artificialanalysis.ai/articles/claude-opus-5-leader-agentic-knowledge-work (Jul 24, 2026).

Opus 5 on โ€œhighโ€ effort: still ahead of Fable 5's score, for $10.41/task, less than half the price. Intelligence and cost decoupled in the direction nobody markets: both got better at once.

The Kimi K3 trap

Three days before Opus 5 shipped, Moonshot AI had already released Kimi K3, priced at $3 input / $15 output per million tokens, the exact sticker Sonnet 5 moves to on September 1. Same rate card, dollar for dollar. If cost-per-token were the whole story, Kimi K3 and September's Sonnet 5 should land in the same place on cost per finished task.

They don't. Finishing one AA-Briefcase task takes Kimi K3 83 back-and-forth turns and 120,000 output tokens, averaging 56.4 minutes wall-clock per task. That works out to $10.57 per task, for a 51% rubric pass rate and an Elo of 1,543, while Opus 5, at essentially the same price per task, clears a meaningfully higher bar. Same rate card, different bill, because the rate card was never the whole equation.

ModelTurns/taskOutput tok/taskCost/taskRubric passElo
Kimi K383120,000$10.5751%1,543
Opus 5 (high)โ€”โ€”$10.41higher+32 vs Fable 5

The cross-check: a different benchmark, same pattern

This isn't an artifact of one unusual benchmark. Four frontier models launched in the same eight days, July 8โ€“16, 2026, and Artificial Analysis measured all four on a completely different test: the general-purpose Intelligence Index. Cost per Intelligence Index task: GPT-5.6 Luna at 21 cents, Muse Spark 1.1 at 26 cents, Grok 4.5 at 31 cents, GPT-5.6 Sol at $1.04. Same eight-day window, a different kind of test entirely, and still a 5x spread in what it actually costs to get one answer, independent of what any of the four charge per million tokens.

Intelligence Index: cost per task, four July 2026 launches

GPT-5.6 Luna, score 51$0.21/task
Muse Spark 1.1, score 51$0.26/task
Grok 4.5, score 54$0.31/task
GPT-5.6 Sol, score 59$1.04/task

Source: artificialanalysis.ai โ€” four frontier launches in eight days (Jul 17, 2026).

Anthropic's own numbers, pointing in opposite directions

Which is what makes Anthropic's own move so telling. Sonnet 5 launched June 30, 2026 as, in Anthropic's framing, the cheaper way to run agents. Two months later, its rate card is going up 50%, while its stablemate Opus 5, shipped in between on July 24, just proved the same lab can ship a model that's both smarter and cheaper per finished task than the one it replaced. Sticker price and real cost, moving in opposite directions, at the same company, in the same quarter.

Sonnet 5 sticker price vs. Opus 5 real cost per task, same quarter

Sonnet 5 sticker price change (Sep 1, 2026)+50%
Opus 5 vs. Fable 5, cost per task (max)-20%

Sources: platform.claude.com/docs/about-claude/pricing; artificialanalysis.ai.

The formula underneath both charts is the same one: dollars per task equals turns to completion, times tokens per turn, times dollars per token. A rate-card change moves only the last term. Everything this post just walked through, the 800x AA-Briefcase spread, the Kimi K3 trap, the 5x Intelligence Index spread, comes from the first two terms moving independently of the third.

Auditing cost-per-task yourself

None of this requires a benchmark lab. Every number above traces back to a pricing page or a published benchmark run; the only thing missing from most teams' own model selection is multiplying it out.

# cost_per_task_audit.py โ€” turns x tokens/turn x $/token, not the rate card alone
MODELS = [
    # name, input_$_per_M, output_$_per_M, avg_turns, avg_output_tok_per_turn
    ("Kimi K3",        3.0, 15.0, 83,  1_446),   # 120,000 output tok / 83 turns
    ("Opus 5 (high)",  5.0, 25.0, None, None),   # cost/task measured directly: $10.41
    ("Sonnet 5 (Sep)", 3.0, 15.0, None, None),   # sticker only โ€” turns/task not yet published
]

MEASURED_COST_PER_TASK = {
    "Kimi K3": 10.57,
    "Opus 5 (high)": 10.41,
}


def implied_output_tokens(turns, tok_per_turn):
    if turns is None or tok_per_turn is None:
        return None
    return turns * tok_per_turn


for name, input_rate, output_rate, turns, tok_per_turn in MODELS:
    out_tok = implied_output_tokens(turns, tok_per_turn)
    measured = MEASURED_COST_PER_TASK.get(name)
    print({
        "model": name,
        "rate_card": f"${input_rate}/${output_rate} per M tok",
        "implied_output_tokens_per_task": out_tok,
        "measured_cost_per_task": measured,
    })
$ python3 cost_per_task_audit.py
{'model': 'Kimi K3', 'rate_card': '$3.0/$15.0 per M tok', 'implied_output_tokens_per_task': 120018, 'measured_cost_per_task': 10.57}
{'model': 'Opus 5 (high)', 'rate_card': '$5.0/$25.0 per M tok', 'implied_output_tokens_per_task': None, 'measured_cost_per_task': 10.41}
{'model': 'Sonnet 5 (Sep)', 'rate_card': '$3.0/$15.0 per M tok', 'implied_output_tokens_per_task': None, 'measured_cost_per_task': None}

The point of running this isn't the specific numbers, it's the missing column. Sonnet 5's September rate card is public today; its AA-Briefcase-style turns-per-task and cost-per-task are not, because that data only exists once someone actually runs the benchmark. Kimi K3 and Opus 5 (high) sit on comparable measured cost per task, $10.57 and $10.41, despite a very different per-token rate card ($3/$15 vs $5/$25), which is exactly the inversion this whole post is about. Until a model has a published cost-per-task number, its pricing page is a guess about your bill, not a fact about it.

If you want to build your own always-on version of this

Re-running this audit every time a lab ships a new model or moves a price, on top of everything else a platform team already tracks, is exactly the kind of job that shouldnโ€™t depend on someone remembering to re-check a pricing page. A BrainClaw agent, running 24/7 on its own isolated VM, can pull fresh AA-Briefcase and Intelligence Index data on a schedule, re-run the audit above against your own workload, and flag the moment a modelโ€™s real cost per task drifts past what its rate card implied.

And once you know which model actually wins on cost per finished task, not cost per token, MegaBrain gives you one API across 500+ models with transparent, at-cost pricing, so switching from Sonnet 5 to Opus 5, or away from a model that looks cheap on paper, is a one-line config change, not a re-integration.

MegaBrain Gateway

500+ models. One API. No markup.

Use in Claude Code, Cline, Cursor, or any coding agent.

Try MegaBrain free โ†’

Newsletter

Stay in the loop

Get the latest model comparisons and guides โ€” no spam, unsubscribe anytime.