Thinking Machines Shipped a Model That Loses 5 of 6 Benchmarks. It Is Also the Only One That Does Not Lie.
On July 15, 2026, Thinking Machines Lab released Inkling, a 975-billion-parameter open-weights model, and published its own benchmark table showing Inkling losing to GLM-5.2, Kimi K2.6, and DeepSeek V4 on 5 of 6 head-to-head exam benchmarks, by as much as 19 points. Buried in a second, independent Artificial Analysis report is a different number: on AA-Omniscience, the benchmark that scores whether a model knows what it does not know instead of how many exam questions it gets right, Inkling is the only model in the entire comparison with a positive score. Here is the full data set, and why the industry markets the wrong one.
TL;DR
- 🎯 The release— Thinking Machines Lab shipped Inkling, a 975B-parameter (41B active) open-weights model, on July 15, 2026, calling it “the leading open-weights release from a U.S. lab.”
- 📉 The exam scores— Head-to-head against GLM-5.2, Kimi K2.6, and DeepSeek V4 on 6 benchmarks, Inkling loses 5. Down 18.9 points on Terminal-Bench 2.1, 10.4 on HLE, 3.9 on GPQA Diamond, 3.0 on SWE-bench Verified, 5.5 on MMMU Pro. It wins AIME 2026 by 0.7 of a point.
- 🇺🇸 The footnote in “leading”— Artificial Analysis's Intelligence Index ranks Inkling #1, but only against 3 other U.S. open models (Nemotron 3 Ultra, Gemma 4, gpt-oss-120b). Every model beating it on the exams is a different comparison set.
- 🔁 The turn— On GDPval-AA (real professional deliverables), Inkling scores 1238 Elo against Kimi K2.6's 1190 and DeepSeek V4 Flash's 1189. On τ³-Banking (simulated real banking workflows), Inkling wins too: 24% vs. 21% and 23%.
- 💸 Token efficiency— Inkling burns 25K output tokens per task vs. 37K–43K for the models beating it on exams, up to 42% cheaper for the same finished work on a metered API.
- 🏪 Distribution— Inkling ships day-1 on 6 inference providers. GPT-5.6 Sol, the model beating it on aggregate benchmarks, ships through exactly 1.
- 🕵️ The payoff— AA-Omniscience, which scores accuracy minus hallucination rate, gives Inkling +2. Nemotron 3 Ultra scores -1. DeepSeek V4 Pro (max) scores -10. Inkling is the only positive number in the set.
The chart Thinking Machines published about itself
Model launch posts are marketing documents, and marketing documents do not usually contain a chart where your own flagship loses. Thinking Machines Lab's launch post for Inkling does exactly that: a benchmark table, published by the lab itself, showing Inkling scoring 63.8% on Terminal-Bench 2.1 against GLM-5.2's 82.7%, an 18.9-point gap, on the same day, same benchmark, same launch post.
That is not the only place it loses. Line Inkling up against the three labs actually shipping frontier open-weights models right now, GLM (Zhipu), Kimi (Moonshot), and DeepSeek, on 6 head-to-head exam benchmarks, and it loses 5 of 6.
Inkling vs. best Chinese open rival, 6 head-to-head exam benchmarks
Source: Thinking Machines Lab, 'Inkling: Our open-weights model,' effort=0.99, July 15, 2026.
Not a rounding error. Eighteen points down on the benchmark closest to real terminal use. Ten points down on pure factual recall. The one benchmark it wins, AIME 2026, it wins by seven tenths of a point, inside anyone's margin of error.
What Inkling actually is
Inkling is a 975-billion-parameter mixture-of-experts model with 41 billion active parameters, pretrained on 45 trillion tokens of text, images, audio, and video, with a context window stretching to 1 million tokens in the open-weights release (256K on Thinking Machines' own Tinker API). It is multimodal natively, not bolted on, and it is priced on Tinker at $1.87 input / $4.68 output per million tokensat 64K context, roughly double that at 256K context, with a 50%-off launch discount. Thinking Machines' own description is unusually blunt for a launch post: Inkling “is not the strongest overall model available today,” but is positioned as “a good open-weights base for customization,” leaning on multimodality, efficient thinking, and Tinker fine-tuning support rather than leaderboard dominance.
The footnote hiding inside the word “leading”
So where does “leading” come from, if Inkling loses 5 of 6 exams to Chinese open models? Artificial Analysis's own write-up ranks Inkling #1 on their composite Intelligence Index, comfortably ahead of Nvidia's Nemotron 3 Ultra, Google's Gemma 4 31B, and OpenAI's gpt-oss-120b.
Artificial Analysis Intelligence Index — U.S. open-weights models only
Source: Artificial Analysis, 'Thinking Machines has released Inkling, the new leading U.S. open weights model,' July 15, 2026.
Every single model that just beat Inkling on the exam benchmarks above, GLM-5.2, Kimi K2.6, DeepSeek V4, is absent from this leaderboard, because it is scoped to U.S. labs only. Both facts are true at the same time: Inkling is the leading open-weights model from a U.S. lab, and Inkling loses most of its head-to-head exams against the open models that are not from a U.S. lab. Neither Thinking Machines nor Artificial Analysis hid this; the comparison set is stated plainly in both reports. It is just a footnote that a launch headline, and most of the coverage that followed it, did not read out loud.
When the test stops being an exam and starts being a job
Here is where the story turns. GDPval-AA is Artificial Analysis's agentic benchmark, and it does not ask trivia questions. It scores agents on real professional deliverables, the kind of multi-step work you would actually hand to an autonomous system rather than a chat window.
GDPval-AA v2 — agentic Elo on real professional deliverables
Source: Artificial Analysis, GDPval-AA v2, July 15, 2026.
The model that just lost 5 exam benchmarks in a row is suddenly 48 Elo points clear of both models that beat it on the exams. The same pattern holds on τ³-Banking, a benchmark built to simulate an agent handling real banking workflows under real operational constraints, not a multiple-choice test:
τ³-Banking — simulated real banking workflow success rate
Source: Artificial Analysis, July 15, 2026.
On the tests that look like an actual job instead of a test written to be studied for, it is not close, and Inkling is the one ahead.
Fewer tokens, same job, six providers instead of one
Inkling also does that job cheaper. Artificial Analysis measured output tokens burned per task: Inkling uses 25,000, against 37,000–43,000 for the three models that beat it on raw exam scores.
Output tokens burned per task (lower is cheaper on a metered API)
Source: Artificial Analysis, July 15, 2026.
That is up to a 42% smaller billfor the same finished work, on a per-token API. It is also available from more places to buy it from: Inkling shipped day one on 6 inference providers (Tinker, Together AI, Fireworks, Modal, Databricks, Baseten). GPT-5.6 Sol, the OpenAI model that outscores Inkling on Artificial Analysis's aggregate Intelligence Index at $1.04 per task, ships through exactly 1 provider: OpenAI's own API. One of these models you can price-shop across six vendors on day one. The other, you cannot shop at all.
The test nobody markets
Artificial Analysis also runs a factuality benchmark called AA-Omniscience, part of the broader Intelligence Index suite. It does not just check whether a model gets a question right. It scores accuracy and non-hallucination as separate components, so a model that says “I don't know” scores better than one that confidently states a wrong answer, and a model that only gets a fact right when it is genuinely unsure loses points for guessing.
AA-Omniscience — net factuality score (accuracy, less hallucination)
Source: Artificial Analysis, AA-Omniscience, July 2026.
Positive means the model, more often than not, knows what it does not know. Negative means it is, net, confidently wrong more often than it is honestly uncertain. Out of every model in this entire comparison, across five labs, exactly one posts a positive number. It is the model that just lost 5 benchmarks in a row.
The industry ships a new leaderboard almost every week: intelligence indices, coding agent indices, terminal benchmarks, agentic Elo. It has shipped one honesty leaderboard, period, and it is roughly a year old, runs on a single research site, and gets zero mentions in most launch posts, including the one for the model that actually wins it.
Why benchmark marketing selects for the wrong number
None of this is a scandal about Thinking Machines, or Artificial Analysis, or any single lab. Both organizations published their full methodology and their full comparison set in the open; anyone triangulating the two reports gets to the same conclusion this piece does. The actual story is a selection effect. Exam-style benchmarks (HLE, GPQA, AIME, Terminal-Bench) are cheap to run, cheap to publish, and easy to put in a launch-day chart with a single headline number. A benchmark like AA-Omniscience requires scoring confidence calibration across a large question set, is harder to compress into one number a press release can lead with, and answers a question, “will this model quietly lie to me when I am not watching,” that matters enormously for anyone running an unattended agent and close to zero for anyone reading a leaderboard tweet.
Easiest to publish is not the same axis as hardest to fake. Every model this month gets marketed off the same handful of exam questions the entire industry has already memorized the format of. The benchmark that predicts whether an autonomous agent confidently invents an answer at 2 a.m., when nobody is reviewing its output before it reaches a customer, is the one nearly everyone skips.
If you want to route on honesty, not just raw intelligence
Once you accept that a model's raw Intelligence Index and its Omniscience score can point in opposite directions, the practical question for anyone deploying agents is not “which model is smartest,” it is “which of these numbers should gate my routing decision.” A simple honesty filter, applied before you even get to cost or latency, looks like this:
# route_policy.py — drop hallucination-prone models before ranking on cost/speed
CANDIDATES = [
{"model": "inkling", "aa_omniscience": 2},
{"model": "nemotron-3-ultra", "aa_omniscience": -1},
{"model": "deepseek-v4-pro-max", "aa_omniscience": -10},
]
def unattended_agent_safe(candidates, min_omniscience=0):
"""A model with a negative net Omniscience score is, on balance,
confidently wrong more often than it is honestly uncertain."""
return [c for c in candidates if c["aa_omniscience"] >= min_omniscience]
safe = unattended_agent_safe(CANDIDATES)
print(f"{len(CANDIDATES)} candidates -> {len(safe)} pass the honesty filter")
for c in CANDIDATES:
mark = "✓" if c in safe else "⚠ FIX"
print(f" {mark} {c['model']}: Omniscience {c['aa_omniscience']:+d}")$ python3 route_policy.py
3 candidates -> 1 pass the honesty filter
✓ inkling: Omniscience +2
⚠ FIX nemotron-3-ultra: Omniscience -1
⚠ FIX deepseek-v4-pro-max: Omniscience -10That is a toy filter, not a production ranking, and it is deliberately simple: it only proves the point that raw benchmark rank and honesty rank are separate axes, and that most routing logic today only checks the first one. A model choice made off an Intelligence Index alone would rank DeepSeek V4 Pro above Inkling in some deployments and hand an unattended agent to the model with the worst hallucination score of the three.
This is exactly the kind of check that should not be a one-time decision made from a launch-day chart. Benchmarks get revised, models get updated point releases, and a hallucination score measured in July is not guaranteed to hold in September. A mhermes agent is a natural fit for keeping this filter live: it runs on its own isolated VM with shell and network access, so it can pull the latest Artificial Analysis Intelligence Index and Omniscience scores on a schedule, re-run the honesty filter against whatever models your routing config currently allows, and flag the moment a model you are already using in production drifts into negative territory.
And once honesty and cost have both narrowed the field, MegaBrain gives you one API across 500+ models with transparent, at-cost pricing and automatic routing, so acting on that shortlist is a config change to your provider list, not a re-integration of your whole agent stack.
What the data actually says
| Benchmark | What it measures | Inkling result |
|---|---|---|
| 6 exam benchmarks (HLE, GPQA, SWE-bench, Terminal-Bench, MMMU, AIME) | Static knowledge & reasoning | Loses 5 of 6, by up to 18.9 pts |
| AA Intelligence Index (U.S. open models only) | Composite exam score, scoped comparison | #1 of 4 (41 vs. 38/29/24) |
| GDPval-AA v2 | Real professional deliverables, agentic Elo | 1238, ahead of Kimi K2.6 (1190) and DeepSeek V4 Flash (1189) |
| τ³-Banking | Simulated real banking workflows | 24%, ahead of DeepSeek (23%) and Kimi (21%) |
| Output tokens / task | Cost efficiency on a metered API | 25K vs. 37K–43K for exam-winning rivals |
| Day-1 inference providers | Market distribution / price-shoppability | 6 providers vs. GPT-5.6 Sol’s 1 |
| AA-Omniscience (net) | Accuracy minus hallucination rate | +2, the only positive score in the comparison |
Sources for every figure above: Thinking Machines Lab's Inkling launch post (thinkingmachines.ai/news/introducing-inkling, July 15, 2026), MarkTechPost's benchmark writeup of the same release, and Artificial Analysis's independent report, “Thinking Machines has released Inkling, the new leading U.S. open weights model” (artificialanalysis.ai, July 15, 2026).
MegaBrain Gateway
500+ models. One API. No markup.
Use in Claude Code, Cline, Cursor, or any coding agent.
Newsletter
Stay in the loop
Get the latest model comparisons and guides — no spam, unsubscribe anytime.