AI BenchmarksMetaAI EconomicsLLM PricingArtificial Analysis

Meta's Best AI Model Isn't For Sale. The One You Can Buy Ties Mid-Tier.

On September 2, 2026, Meta shipped Muse Spark 1.3 — its fourth release in five months — and every headline led with the same number: 62 on the Artificial Analysis Intelligence Index, third overall behind only Claude Fable 5.1 and Claude Opus 5. That number belongs to a variant called “max,” which is limited to Meta's own partners and has no public price. The variant anyone can actually buy today, “xhigh,” scores 61 — and ties Grok 4.6 and GPT-5.6 Sol, not the frontier leaders the coverage implied.

2026-09-03·12 min read

TL;DR

  • 🚀 Cadence— Meta shipped 4 Muse Spark releases in 5 months: the original (April), 1.1 (July), 1.2 (August 5), and 1.3 (September 2). The gap between releases shrank from 92 days to 27 to 28.
  • 🔒 Two products, one name— Muse Spark 1.3 shipped as “xhigh” (score 61, publicly available now) and “max” (score 62, limited to Meta partners, no public price).
  • 📊 What 61 actually ties— Grok 4.6 (high) and GPT-5.6 Sol (max), not Claude Opus 5 (63) or Claude Fable 5.1 (66), the model on top of the same leaderboard.
  • 💰 Price— $0.55 per Intelligence Index task, 42% cheaper than GPT-5.6 Sol's $0.95, and roughly a quarter of Claude Fable 5.1's $3.69 for 5 fewer points.
  • Speed— 235 output tokens/second, second-fastest on the chart behind only Gemini 3.8 Flash (305 tok/s).
  • 🧮 The wrinkle— Artificial Analysis recalibrated its Intelligence Index twice this year. The original Muse Spark scored 52 at launch and scores 43 on today's index. Same model, two numbers, depending which month you read about it.

The number in every headline this week

Artificial Analysis put it plainly the day Muse Spark 1.3 shipped: the “max” variant scores 62 on the Intelligence Index, placing it behind only Claude Fable 5.1 (66) and Claude Opus 5 (63) — third out of every model the lab tracks. Most of the coverage that followed repeated that framing more or less verbatim: Meta, back near the frontier. What most of it didn't repeat, at least not in the headline, is the line right after it — “max,” the variant that scores 62, is in limited preview for Meta's partners. It has no public price, no public API access, and no way for anyone outside that partner list to independently confirm the number.

The variant that actually shipped to the public this week is called “xhigh.” It scores 61 — one point lower. That one point sounds like rounding error until you look at what it costs to close: it's the difference between a model that ties Claude Fable 5.1 near the top of the leaderboard and a model that ties Grok 4.6 and GPT-5.6 Sol, comfortably mid-to-upper tier but nowhere near frontier-leading.

Four releases, five months, an accelerating clock

Whatever you make of the two-tier framing, the release cadence itself is real and worth sitting with. Meta shipped the original Muse Spark in April 2026, 1.1 in July, 1.2 on August 5, and 1.3 on September 2 — four releases in five months, with the gap between each one shrinking: 92 days between the original and 1.1, then 27 days to 1.2, then 28 days to 1.3. That's a lab that went from a quarterly cadence to something close to monthly inside a single product cycle.

Days Between Muse Spark Releases

Original -> 1.192 days
1.1 -> 1.227 days
1.2 -> 1.328 days

Source: officechai.com, Muse Spark 1.3 coverage citing Artificial Analysis, Sep 2, 2026.

Score the whole family on the same, current version of Artificial Analysis's Intelligence Index — more on why that qualifier matters shortly — and the climb tracks the cadence: 43, then 51, then 57, then 61. The jump from 1.2 to 1.3, four points, is described as the largest single jump the model family has posted between any two releases.

Intelligence Index by Release (Recalibrated, Comparable Scale)

Original (April)43
1.1 (July)51
1.2 (August)57
1.3, xhigh (September)61

Source: officechai.com, citing Artificial Analysis Intelligence Index v4.1.1.

One model, two products

Here's the split that the “62, third place” headline glosses over. Muse Spark 1.3 did not ship as a single model with a single score. It shipped as two effort-tier variants that behave, for a customer's purposes, like two different products:

VariantIntelligence IndexAvailabilityPrice
xhigh61Public today (Muse Code, Meta Model API)$1.25 / $4.25 per 1M tokens
max62Limited to Meta partnersUndisclosed

The number that made every launch-day headline is attached to the row you can't order. The number attached to the row you can order — the one that will actually show up in your bill if you build on this model today — is one point lower, and that one point is the difference between beating Claude Opus 5 and simply matching Grok 4.6.

What 61 actually ties

Precision matters here, so here's the exact cluster. Muse Spark 1.3 (xhigh) scores 61 on the Intelligence Index. So do Grok 4.6 (high) and GPT-5.6 Sol (max). Claude Opus 5 sits two points above that cluster at 63. Claude Fable 5.1 — the model currently on top of this exact leaderboard, and the subject of its own pricing reconciliation we covered here yesterday — sits five points above at 66.

Who Scores What Muse Spark 1.3 (xhigh) Scores

Muse Spark 1.3 (xhigh)61
Grok 4.6 (high)61
GPT-5.6 Sol (max)61
Claude Opus 5 (max)63
Claude Fable 5.1 (max)66

Source: officechai.com and implicator.ai, both citing Artificial Analysis, Sep 2-3, 2026.

None of that makes Muse Spark 1.3 a bad model — tying two genuinely strong, widely-used frontier-adjacent models is a real result for a lab that was largely absent from this conversation a year ago. It just isn't the “third place, behind only the top two” result that ran in most of the coverage, because that result belongs to a different, unavailable product.

The price is the actual story

Where Muse Spark 1.3 (xhigh) earns its keep is cost. Artificial Analysis's weighted-average cost per Intelligence Index task puts it at $0.55. GPT-5.6 Sol, the model it ties on score, costs $0.95 per task — 42% more for identical intelligence. Claude Fable 5.1, five points ahead on the leaderboard, costs $3.69 per task: almost seven times the price for one grade's worth of additional intelligence.

Cost per Intelligence Index Task

Muse Spark 1.3 (xhigh)$0.55
Gemini 3.8 Flash$0.58
GPT-5.6 Sol (max)$0.95
Claude Fable 5.1 (max)$3.69

Source: implicator.ai and officechai.com, citing Artificial Analysis cost-per-task methodology.

Under the headline per-token price sits a second tier most of the launch coverage skipped entirely. Meta's “Contributor” pricing runs $0.10 per million input tokens and $0.20 per million output — roughly a twelfth of the standard rate, explicitly aimed at developers experimenting with early agentic workloads:

# Muse Spark 1.3 (xhigh) pricing, per Meta's published rate card
standard = {"input_per_m": 1.25, "output_per_m": 4.25, "cached_input_per_m": 0.15}
contributor = {"input_per_m": 0.10, "output_per_m": 0.20}

ratio_input = standard["input_per_m"] / contributor["input_per_m"]
ratio_output = standard["output_per_m"] / contributor["output_per_m"]
print(f"Contributor tier: {ratio_input:.1f}x cheaper on input, {ratio_output:.1f}x cheaper on output")
# Contributor tier: 12.5x cheaper on input, 21.3x cheaper on output

That's a deliberate below-cost-of-the-flagship on-ramp — a company buying usage data and developer habit formation at a price that undercuts its own standard tier by an order of magnitude, priced specifically for the “early agentic work” segment where switching costs compound fastest.

It's also fast

Muse Spark 1.3 (xhigh) generates 235 output tokens per second — second on Artificial Analysis's entire speed chart, behind only Gemini 3.8 Flash's 305 tokens per second (itself a $0.58-per-task model, so the two occupy adjacent price-performance territory). For a model that ties mid-to-upper-tier competitors at a fraction of their price and runs near the top of the speed chart, that's a genuinely strong, shippable product on its own terms — which is exactly why burying it under an unavailable sibling's score is such an unforced error in how this launch got covered.

Even the history doesn't hold still

One more wrinkle, and it's a data-integrity one rather than a marketing one — and one we already flagged in detail back in July, when the original Muse Spark's own score moved from 52 to 43 with zero code changes, purely from Artificial Analysis rebaselining its index. That same recalibration is still live in every number this piece uses: the Index has since moved to v4.1.1, and every historical Muse Spark score has been restated on the current scale to stay comparable.

Read an article about Muse Spark's April launch and one about its September launch side by side, and you'll see two different Intelligence Index numbers for the identical model, nine points apart, with no note that the ruler itself moved — the exact trap our July piece traced back to a rebaselined eval. Every progression chart in this piece uses the current, recalibrated index applied consistently across all four releases, specifically to avoid repeating it.

What actually shipped this week

Strip out the partner-only variant and the version confusion, and what Meta actually put in developers' hands on September 2 is real and, on its own merits, good: a publicly available model that ties upper-mid-tier frontier competitors, undercuts them by double digits on price, and outruns nearly everything else on speed. That's a legitimate product story. It just isn't the “near-frontier, third place” story that ran in most of the coverage, because that story's headline number belongs to a model nobody outside Meta can currently test, price, or buy.

Meta is shipping faster than almost anyone else in the industry right now. What it hasn't shipped yet is a public price for the variant that would actually justify the frontier framing. Until “max” has one, 62 is a claim and 61 is the product.

If you're actually routing traffic across these models

The practical takeaway for anyone building rather than reading headlines: a model's launch-day score and its available score are frequently not the same number, and the gap between them can be the difference between “ties the leaderboard leader” and “ties the mid-tier cluster.” That gap is invisible if you're reading a press release once at launch and hard-coding a routing decision around it. MegaBrainroutes every call across 500+ models at zero markup and surfaces the real, current per-model cost and performance instead of a launch-day number frozen in a blog post from three weeks ago — so when a “max” variant finally gets a public price, or a “xhigh” gets quietly repriced, that shows up as a number you can see and route around, not a headline you have to remember to re-check.

Sign up at getmegabrain.com to route across Muse Spark 1.3 and 500+ other models with real-time, per-call cost visibility — not a number from launch day.

MegaBrain Gateway

500+ models. One API. No markup.

Use in Claude Code, Cline, Cursor, or any coding agent.

Try MegaBrain free →

Newsletter

Stay in the loop

Get the latest model comparisons and guides — no spam, unsubscribe anytime.