AI EconomicsAnthropicClaudeLLM PricingAI Benchmarks

Anthropic Cut a Price 75%. Its Own Benchmark Partner Says You Pay 20% More.

On September 1, 2026, Anthropic shipped Claude Fable 5.1 and cut its cache-read price from $1.00 to $0.25 per million tokens — a 75% cut — with a specific, numeric promise attached: an estimated 25% lower cost for typical workloads, up to 45% lower for highly agentic ones. Artificial Analysis, the lab Anthropic gave pre-release access to for this exact launch, published its own independent cost measurement the same day. At max effort, the setting you reach for when accuracy can't slip, Fable 5.1 costs $3.76 per Intelligence Index task. Fable 5, its predecessor, cost $3.14. That is not 25% less. That is 20% more.

2026-09-02·13 min read

TL;DR

  • 🏆 Record score— Claude Fable 5.1 (max effort) scores 66 on Artificial Analysis's Intelligence Index, the highest of 192 models the lab has ever measured, ahead of Claude Opus 5 (63), Claude Fable 5 (62), GPT-5.6 Sol (61) and Grok 4.6 (61).
  • 💰 The claim— Anthropic cut cache-read pricing 75% ($1.00 to $0.25 per 1M tokens) and estimates Fable 5.1 costs 25% less than Fable 5 for typical workloads, up to ~45% less for highly agentic work.
  • 📊 The measurement— Artificial Analysis's own benchmark, run at max effort on the same task set, puts Fable 5.1 at $3.76 per Intelligence Index task versus Fable 5's $3.14 — 20% more, not less.
  • 🔧 Why— the cache cut saves ~$1.40/task (max-effort cost without it would be ~$5.16), but Fable 5.1 burns ~1.7x the output tokens of Fable 5 on the identical task at the identical effort setting. The token increase outruns the discount.
  • 🎚️ The reconciliation— both numbers are real, measured at different points on a 5-level effort dial that spans 11x in token usage (13.1M to 143.7M tokens). One notch down from max (xhigh), Fable 5.1 scores 65 at $2.72/task — 13% cheaper than Fable 5 ever was.
  • 🤥 A second, unrelated gap— on Artificial Analysis's hallucination test (AA-Omniscience), Fable 5.1 attempts more questions and gets more right, but also guesses more often when it's wrong (72.6% vs. Fable 5's 63.6%). The two effects cancel: the hallucination-adjusted score lands flat.
  • ⚙️ The default that matters— Claude Code defaults to High effort, not Max; claude.ai and Claude Cowork default to Medium. The +20% figure only bites at the one setting almost nobody ships with by default.

Two primary sources, same launch day, two different numbers

Most “we fact-checked the price cut” posts pit a vendor's marketing copy against Reddit sentiment. This one doesn't need to: Anthropic and Artificial Analysis both published on September 1, 2026, about the same model, and their numbers point in opposite directions at the same effort setting. Anthropic's own launch page for Claude Fable 5.1 states plainly that the model “will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token,” rising to “approximately 45%” for highly agentic work, driven by a cut to cache-read pricing. Artificial Analysis, which Anthropic gave pre-release access to specifically to benchmark this launch, published its own independent cost-per-task number the same day: $3.76 at max effort, versus $3.14 for Fable 5 at the same setting. Neither figure is wrong. They're measuring different things, and the gap between them is the actual story.

What actually shipped on September 1

Claude Fable 5.1 is the higher-capability half of a two-model launch (Claude Mythos 5.1 is the trusted-access sibling, gated for cybersecurity and life-sciences work). Artificial Analysis supported Anthropic with pre-release evaluation, using Anthropic's “default” server-side safety fallback — a detail worth flagging for rigor: about 4% of output tokens across the Intelligence Index run were quietly routed to Claude Opus 4.8 or Opus 5, not generated by Fable 5.1 itself. Small slice, but it means the headline score isn't 100% pure Fable 5.1 output.

With that caveat noted, the intelligence gain is real and broad, not a single cherry-picked benchmark:

Model (effort)AA Intelligence IndexNotes
Claude Fable 5.1 (max)66Highest of 192 models AA has ever measured
Claude Opus 5 (max)63Anthropic’s own flagship reasoning model
Claude Fable 5 (max)62Prior generation
GPT-5.6 Sol (max)61OpenAI
Grok 4.6 (high)61SpaceXAI

Fable 5.1 also sets record scores on several individual evaluations inside the index: 91.4% on Terminal-Bench v2.1 and 62.0% on SciCode, both the highest Artificial Analysis has measured on those benchmarks. On Humanity's Last Exam it scores 59.1%, up from 55.5% for Fable 5. On the agentic knowledge-work benchmarks, GDPval-AA v2 (1,853 Elo) and AA-Briefcase (1,694 Elo), it leads Claude Opus 5 (1,824 and 1,685 respectively), though Artificial Analysis notes the GDPval-AA v2 gap is within the confidence interval — effectively tied, not a clean win. On AA-Briefcase specifically, Fable 5.1 pulls ahead on analytical quality (2,025 Elo vs. Opus 5's 1,980) but falls behind on presentation (1,495 vs. 1,572), a split that matters if your use case cares about polish over correctness.

Anthropic's promise, in its own words

The pricing claim is specific enough to check, which is rarer than it should be. From Anthropic's launch page: “Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token. This is because we're reducing our pricing on cache reads… For highly agentic work, the savings will often be much larger — up to approximately 45%.” The mechanism behind it: cache-read pricing drops from $1.00 to $0.25 per million tokens, a 75% cut, while standard input ($10/M), output ($50/M), and cache-write ($12.50/M) pricing stay unchanged from Fable 5.

The cache-read price cut Anthropic's savings claim is built on

Fable 5 cache read$1.00 / 1M tokens
Fable 5.1 cache read$0.25 / 1M tokens

Source: Anthropic, “Introducing Claude Fable 5.1 and Claude Mythos 5.1,” anthropic.com, Sept 2026

For a workload that's mostly cache hits — the norm for an agent re-reading a long system prompt or codebase turn after turn — a 75% cut to that one line item is a real, large saving. The claim isn't fabricated. It's also not the whole cost equation, and Anthropic's own footnote gives away where the rest of it hides: Fable 5.1 defaults to High effort in Claude Code and Medium in Claude Cowork and on claude.ai — never Max. The 25-45% estimate is built around those real-world defaults, not the setting Artificial Analysis benchmarked.

The number that breaks the promise

Artificial Analysis measured cost per Intelligence Index task directly, using input, cache hit, cache write, reasoning, and answer token prices weighted by each evaluation's share of the index — the same methodology it uses for every model on its leaderboard, not a bespoke calculation for this launch. At max effort, the setting that produces the record 66 score, the numbers don't match Anthropic's framing at all:

Model (max effort)Cost per Intelligence Index taskvs. Fable 5
Claude Fable 5 (max)$3.14baseline
Claude Fable 5.1 (max)$3.76+20%

Twenty percent more, not twenty-five percent less. The two claims aren't reconcilable by rounding error — they're a 45-point swing between what the vendor advertised and what the independent benchmark measured, on the same launch, at the effort level that produces the headline capability number.

Where the extra money actually goes

The mechanism isn't hidden; Artificial Analysis publishes the token counts alongside the cost. The cache-read cut is doing real work — it saves roughly $1.40 per task at max effort. Strip it back out and max-effort Fable 5.1 would cost approximately $5.16, not $3.76. But Fable 5.1 also generates roughly 1.7x the output tokens Fable 5 did on the identical task at the identical effort setting (Artificial Analysis's own verbosity metric shows 140M output tokens for the full Intelligence Index run, well above the 192-model median of 71M). Output tokens price at $50/M, ten times the discounted cache-read rate, so a 1.7x increase in generation volume overwhelms a 75% cut to a much cheaper line item.

The cache cut saves ~$1.40/task. 1.7x the output tokens costs more than that.

Fable 5 (max)$3.14
Fable 5.1, without the cache cut (est.)$5.16
Fable 5.1, with the cache cut (actual)$3.76

Source: Artificial Analysis, Sept 1, 2026 (artificialanalysis.ai/models/claude-fable-5-1)

You can verify the arithmetic yourself; it's just two line items pulling against each other:

# Reproducing Artificial Analysis's reported max-effort cost delta.
# Figures are AA's published numbers for the Sept 1, 2026 launch;
# this script only does the arithmetic, it doesn't re-derive AA's
# own weighted-average methodology.

cache_price_old = 1.00   # USD per 1M cached input tokens, Fable 5
cache_price_new = 0.25   # USD per 1M cached input tokens, Fable 5.1
output_price = 50.00     # USD per 1M output tokens (unchanged)

fable5_cost = 3.14                 # AA-reported, max effort
fable51_cost_actual = 3.76         # AA-reported, max effort
token_multiplier = 1.7             # AA-reported output-token increase

cache_savings = fable51_cost_actual * (1 - cache_price_new / cache_price_old) * 0.485
# 0.485 approximates the cache-read share of task cost at max effort,
# back-solved so this matches AA's reported ~$1.40 saving -- shown for
# transparency, not claimed as AA's exact internal weighting.
implied_without_cut = fable51_cost_actual + 1.40

delta_pct = (fable51_cost_actual / fable5_cost - 1) * 100

print(f"Fable 5 (max):                    ${fable5_cost:.2f}/task")
print(f"Fable 5.1 (max), without cut est.: ${implied_without_cut:.2f}/task")
print(f"Fable 5.1 (max), with cut actual:  ${fable51_cost_actual:.2f}/task")
print(f"Change vs. Fable 5:                {delta_pct:+.1f}%")

# Fable 5 (max):                    $3.14/task
# Fable 5.1 (max), without cut est.: $5.16/task
# Fable 5.1 (max), with cut actual:  $3.76/task
# Change vs. Fable 5:                +19.7%

The 0.485 cache-share constant above is a transparency device, not a disclosed Anthropic or Artificial Analysis figure — it's chosen so the arithmetic reproduces AA's reported ~$1.40 saving and ~$5.16 without-cut estimate. The inputs that matter (the $3.14, $3.76, and 1.7x figures) are AA's own reported numbers, not our estimate.

The effort dial nobody puts in the headline

Here's where the story turns a second time, and it's the part almost no coverage of this launch mentioned. Claude Fable 5.1 ships five effort settings — low, medium, high, xhigh, max — and they aren't a rounding error. Output token usage spans 11x across that range: 13.1 million tokens at low effort, up to 143.7 million at max. The Intelligence Index score moves from 58 to 66 across the same range. Drop one notch from max, to xhigh, and the picture inverts: score 65 (barely below the record), at $2.72 per task — 13% cheaper than Fable 5 ever was at its own max effort.

ConfigurationIntelligence IndexCost per taskOutput tokens (full run)
Fable 5.1, low effort58not disclosed13.1M
Fable 5, max effort62$3.14baseline
Fable 5.1, xhigh effort65$2.72between low and max
Fable 5.1, max effort66$3.76143.7M

So which claim is “true”? Both, at different points on the same dial. Anthropic's 25-45% savings estimate is built around real-world default effort levels — and Anthropic says so directly: Fable 5.1 defaults to High effort in Claude Code, Medium in Claude Cowork and on claude.ai. Artificial Analysis's +20% figure is a max-effort number, the setting that also produces the record-breaking 66 score everyone quoted in the headline. Reporting the score from one setting and assuming the vendor's savings claim from a different one is how you end up with two numbers that look like a contradiction instead of two honest measurements of the same dial.

The hallucination trade nobody wrote up

There's a second, unrelated gap buried in the same launch data. On Artificial Analysis's AA-Omniscience evaluation, which scores knowledge reliability by rewarding correct answers, penalizing confident wrong ones, and not penalizing refusals, Fable 5.1 attempts 93.4% of questions — more than Claude Opus 5's 87.8% — and posts the highest raw accuracy Artificial Analysis has measured on the benchmark, 67.2%, versus 65.4% for Fable 5.

That sounds like an unambiguous win until you look at what happens on the questions it still gets wrong. Of those, Fable 5.1 still attempts an answer 72.6% of the time, versus 63.6% for Fable 5 — a meaningfully higher guess rate on misses. More right answers, and more confidently wrong ones. The two effects roughly cancel: on the hallucination-adjusted AA-Omniscience Index, Fable 5.1's score lands flat against Fable 5, despite the higher raw accuracy number that would look better in a headline on its own.

More right answers. Also more confident wrong ones.

Attempt rate: Opus 587.8%
Attempt rate: Fable 5.193.4%
Accuracy: Fable 565.4%
Accuracy: Fable 5.1 (best-ever)67.2%
Guess rate when wrong: Fable 563.6%
Guess rate when wrong: Fable 5.172.6%

Source: Artificial Analysis, Sept 1, 2026. Net effect: AA-Omniscience Index score lands flat vs. Fable 5.

None of this means Fable 5.1 is a worse model — the accuracy gain is real and the highest AA has measured on this benchmark. It means the honest summary of the launch has three separate trade-offs bundled into it (intelligence up, cost dependent on effort level, hallucination-adjusted reliability flat), and most coverage collapsed all three into a single “new record score, cheaper too” headline.

What this means if you're actually running agents, not chatting

The gap between Anthropic's claim and Artificial Analysis's measurement isn't a gotcha so much as a reminder of a fact that's easy to forget when the only thing most people read is the discount percentage: reasoning-effort settings are a real lever on cost and quality, and the setting that matters for your bill is whichever one your workload actually runs at, not whichever one produced the number in the press release.

A chat session that runs a handful of turns a day at Medium effort really is likely to see something close to Anthropic's 25% figure — the cache-read cut dominates when each turn's output is modest. An agent running continuously, unattended, pushed to Max effort because the task genuinely can't tolerate a dropped edge case, is closer to Artificial Analysis's number: more tokens generated per task, a bigger share of the bill riding on the ten-times-more-expensive output-token price, and a discount on one input line item that doesn't scale with it. Neither Anthropic nor Artificial Analysis is wrong. The mistake is reading either number as if it applies to every workload.

If you're building or running agents that actually operate this way — long-running, high-effort, unattended — the only way to know your real number is to measure your own workload's effort-level mix, not extrapolate from either press release. MegaBrainroutes every call across 500+ models at zero markup and breaks out the real per-model, per-call cost instead of a single blended number at the end of the month, so an effort-level or pricing change on any model you route through shows up as a number you can actually see, not a percentage you have to trust. And because the whole point of this story is that agents don't stop generating tokens when a person logs off, mhermes, MegaBrain's always-on agent runtime, is built to run continuously on its own isolated cloud VM and surface exactly this kind of cost-per-task data as it happens, not after the invoice arrives.

Sign up at getmegabrain.com to see what your own agent traffic actually costs, effort level by effort level, instead of taking either side's percentage on faith.

MegaBrain Gateway

500+ models. One API. No markup.

Use in Claude Code, Cline, Cursor, or any coding agent.

Try MegaBrain free →

Newsletter

Stay in the loop

Get the latest model comparisons and guides — no spam, unsubscribe anytime.