← MegaBrain BioScience Blog
Feature · August 1, 2026 · 9 min read

An AI Beat Clinicians 2.56-to-1 in a 13,917-Person Trial. Five Days Later, a Benchmark Found the Same Model Class Scores 47 Points Lower on Real Clinical Work.

An AI beat clinicians at diagnosis, 2.56-to-1, across 13,917 real patients. Five days later, a different benchmark found the same class of model scores 44.8% on real clinical work, not the roughly 92% it gets on a medical licensing exam. Both findings are true. Neither is wrong. Here's the 47-point reason not to trust either claim by itself, and what it means for every "AI beats doctors" headline you'll read this year.

2.56-to-1
Odds SymptomAI's differential diagnosis beat independent clinicians on the same transcript (p < 0.001)
44.8% vs. ~92%
BRIDGE's real-world clinical task score for LLMs, vs. their score on medical licensing exams
87 / 59 / 9
Tasks, real-world data sources, and languages behind that 44.8% — 95 models tested

The trial where the AI won

"SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment" (arXiv 2605.04012, Breda, Sunshine, McDuff et al., submitted May 5, 2026) is built on Gemini Flash 2.0 and deployed inside the Fitbit app from June 2025 through April 2026. 13,917 consented participants were randomized across five interview strategies — Base, Fixed Canonical, Flexible Canonical, Dynamic Live, and Dynamic Final — describing real symptoms in their own words, with up to 30 days of prior wearable data available as context. 1,228 of those conversations came with a clinician-confirmed diagnosis attached, and a clinical evaluation panel spent over 250 hours annotating a 517-conversation subset, comparing SymptomAI's differential diagnosis against one written by an independent clinician reviewing the identical transcript, blind to which agent produced which list.

The result, as reported in the paper and Google Research's own summary: SymptomAI's differential diagnoses were preferred over the independent clinician's in 53.3% of head-to-head comparisons, and were rated significantly more accurate overall, at an odds ratio of 2.56 (p < 0.001). The advantage wasn't uniform — it was largest specifically in the cases where the reviewing clinician reported the least confidence in their own differential, exactly the cases where a second opinion is most useful. A separate 1,509-conversation panel drawn from the general US population, not just wearable-device users, was run to check the finding wasn't an artifact of who wears a Fitbit. The paper is explicit that these are research-only AI-derived assessments, not clinical diagnoses, and the deployment ran inside a research arm of a consumer product, not a hospital.

The same week, a benchmark says the exam score was never the right test

"Toward a test of medical AI superintelligence" (Chen, Goh et al., Nature Medicine, DOI 10.1038/s41591-026-04539-8, published July 27, 2026) is a framework paper out of Stanford's Division of Computational Medicine and Clinical Excellence Research Center, with co-authors from Mass General Brigham, Beth Israel Deaconess Medical Center, MIT, the Oxford Internet Institute, and Harvard Medical School. Its central argument: claims that a model has reached medical "superintelligence" are close to unfalsifiable with the benchmarks the field actually uses, because those benchmarks are exam questions, and exam questions aren't clinical practice.

The paper leans on BRIDGE (Wu, Gu, Zhou et al., Nature Biomedical Engineering, DOI 10.1038/s41551-026-01719-2), a benchmark built from 87 clinical tasks drawn from 59 real-world data sources across 9 languages and 14 specialties — discharge summaries, triage notes, lab-order reasoning, coding from messy free text, the actual paperwork and judgment calls that make up clinical work, rather than a multiple-choice exam. Run 95 contemporary LLMs against it, and the aggregate score is 44.8%. Run the same models against standardized medical licensing exams, and they routinely score around 92%. A 47-percentage-point gap, on language models that, on paper, "passed the boards" a while ago.

Neither paper is arguing the other is fraudulent or wrong. BRIDGE and the framework paper are making a narrower, more useful point: an exam score has never been a valid proxy for real clinical task performance, and the size of the gap is now measured, not assumed.

Two real findings, not a contradiction

DimensionSymptomAI (arXiv 2605.04012)BRIDGE (Nature Biomedical Engineering)
InputA patient describing symptoms in their own words, in a live conversational interview, with up to 30 days of prior wearable data for contextRaw real-world clinical documents — notes, orders, results — pulled from 59 real clinical data sources across 9 languages and 14 specialties
TaskConduct one structured interview, then rank a differential diagnosis87 distinct clinical-practice tasks (extraction, summarization, coding, reasoning over messy documents), each graded against task-specific ground truth
ComparatorIndependent clinicians, blind, reviewing the identical transcriptNone — 95 LLMs graded directly against ground truth, no human baseline in the loop
Reported resultPreferred over the clinician-written differential in 53.3% of head-to-head cases; odds ratio 2.56 (p < 0.001)44.8% aggregate score, versus roughly 92% for the same class of models on standardized medical licensing exams

Sources: Breda, Sunshine, McDuff et al., "SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment," arXiv 2605.04012 (May 5, 2026) and the Google Research blog (July 22, 2026); Wu, Gu, Zhou et al., "BRIDGE: benchmarking large language models for understanding real-world clinical practice texts," Nature Biomedical Engineering, DOI 10.1038/s41551-026-01719-2; Chen, Goh et al., "Toward a test of medical AI superintelligence," Nature Medicine, DOI 10.1038/s41591-026-04539-8 (July 27, 2026).

Put the two studies side by side and the apparent tension dissolves into a difference in what's actually being tested. SymptomAI's task is narrow and self-contained: one conversation, complete information volunteered by the patient in real time, judged against a single human comparator working from the same transcript. That's close to the best-case setup for a language model — bounded context, a clear input-to-output mapping, and a comparator who has no more information than the model does. BRIDGE's tasks are the opposite: fragmented real documents, incomplete records, 14 specialties' worth of domain-specific reasoning, and no equivalent human baseline working under the same constraints — just the model against ground truth, cold.

That's not a knock on SymptomAI's result. An odds ratio of 2.56 across nearly 14,000 real participants, with a clinician panel blind to which output came from which source, is a genuinely well-controlled trial, and the finding that AI helps most exactly where clinicians are least confident is a specific, useful, falsifiable claim. It's a knock on treating that result as evidence the model is broadly "better than doctors." It's better than doctors at one particular, well-defined task, under one particular set of conditions. BRIDGE is the measurement of how far that narrow win travels once the task stops being narrow — and the answer, this week, is: not very far. 44.8% is not a passing grade on anyone's licensing board.

What this means for reproducible, local-first science

Neither paper names or tests any commercial research product, MegaBrain BioScience included — this isn't a claim about who's ahead in that market. What the two papers add up to, read together, is a specific and reusable rule for reading the next "AI beats doctors" headline: ask what task was actually tested, what the comparator was, and whether the win generalizes past that exact setup, before generalizing the claim. A 2.56-to-1 odds ratio on one interview format is real and worth taking seriously. It is not the same claim as "44.8% on real clinical work," and conflating the two is how a genuinely well-run trial gets misquoted into a superintelligence claim it never made.

That's also, mechanically, a reproducibility problem: you can only tell these two claims apart if you can see the exact task definition, the exact comparator, and the exact scoring rule behind each number — not just the topline percentage. A workbench that exports a reproducibility record for every run, and pairs it with an independent review step, doesn't make a model better at real-world clinical reasoning. It makes it possible to check, before you cite a number, exactly which of these two kinds of task it was actually measuring.

Try MegaBrain BioScience

A research workbench that runs on your machine and exports a reproducibility record for every result, so a narrow win and a general claim don't get quoted as the same thing.