← MegaBrain BioScience Blog
Feature · September 3, 2026 · 8 min read

11.1 Points With a Certified AI Diagnostic Copilot. 11.7 Points With Nothing At All. Same Randomized Trial.

Prof. Valmed is the first LLM-based clinical decision support system in Europe certified as a medical device. In a randomized trial across seven hospitals, the 42 physicians given access to it improved their top-1 diagnostic accuracy by 11.1 percentage points. The 40 physicians given nothing but conventional resources — UpToDate, AMBOSS, textbooks — improved by 11.7, a gap with a p-value of 0.979. What did separate the two groups: when the certified tool's top suggestion was wrong, physicians followed it into the wrong diagnosis anyway, 95% of the time.

11.1 vs 11.7
Percentage-point gain in top-1 diagnostic accuracy: certified AI copilot vs. no AI at all
94s vs 206s
Case-processing time with vs. without the AI copilot
95%
Of the time physicians followed the copilot's top suggestion into a wrong diagnosis, when it was wrong

The first certified LLM medical device in Europe

Prof. Valmed is a subscription clinical decision support platform built on retrieval-augmented generation: it grounds its outputs in curated medical literature rather than relying only on a pretrained model's memory, a design meant to reduce hallucination. In April 2025 it became, per its manufacturer and VDE's own announcement, the first AI-based clinical decision support tool certified as a Class IIb medical device in the EU — a regulatory bar that general-purpose chatbots like ChatGPT do not clear, and that has restricted their use in formal clinical workflows. That certification is what made the ALLIANCE trial (NCT07166692, first posted September 10, 2025) the first randomized controlled trial anywhere to test a medically certified LLM-based CDS system against real physicians, not just against a benchmark. Eighty-two physicians — 21 rheumatologists and 61 non-rheumatologists, recruited from seven hospitals in Germany and Norway — were randomized 1:1. Each diagnosed three rheumatology vignettes (Cogan syndrome, dermatomyositis, familial Mediterranean fever) twice: once unassisted, once with whatever resources their arm allowed. Outcome assessment was blinded; three board-certified rheumatologists rated every diagnosis, with substantial inter-rater agreement (Cohen's κ=0.799).

The number the trial was designed to move, didn't move

Before assistance, the two arms started from the same place: 22.2% top-1 accuracy in the intervention group, 23.3% in control — both close to stand-alone Prof. Valmed's own accuracy on the same vignettes (33.3%). After assistance, both arms climbed to roughly the same place too: 33.3% with the AI copilot, 35.0% without it. The primary, pre-registered analysis — a mixed-effects model adjusted for age, gender and specialty — found no group-by-time interaction (adjusted OR 0.99, 95% CI 0.45 to 2.19, p=0.979). Top-3 accuracy, diagnostic reasoning score, and post-assistance confidence told the same story: numerically favoring the AI arm on two of three, significant on none. A second look at the same case, it turns out, helps almost as much as a certified AI copilot does — and in this trial, numerically more.

OutcomeIntervention (Prof. Valmed)Control (conventional resources)Between-group effect
Top-1 accuracy, before → after22.2% → 33.3%23.3% → 35.0%adjusted OR 0.99 (0.45–2.19), p=0.979
Case-processing time, after94 s206 s−112 s (−141 to −83), p<0.001
Overconfidence index, before → after16.3% → 23.9%13.7% → 19.7%both arms rose; not formally contrasted
Perceived support quality (1–5)4.33.4+0.93 (0.47–1.39), p<0.001

Source: Kremer, Schlicker, Hasnaj, et al., "Certified large language model-based diagnostic decision support in rheumatology: the ALLIANCE multicentre randomised controlled trial," medRxiv 10.64898/2026.08.29.26361715, posted September 2, 2026. Trial registration NCT07166692. n=82 physicians (42 intervention, 40 control), 7 hospitals in Germany and Norway, 3 rheumatology vignettes per physician. Preprint, not yet peer-reviewed.

What actually moved: speed, not judgment

The trial was not a wash. Case-processing time collapsed with AI assistance — 94 seconds versus 206, an adjusted mean difference of −112 seconds (95% CI −141 to −83, p<0.001) — while both arms started from an identical ~100 seconds before assistance. Physicians using Prof. Valmed also rated information timeliness and diagnostic support quality significantly higher. Read alongside the accuracy result, the honest summary the authors themselves reach is that this class of tool "may be most useful for improving efficiency… rather than increasing top-1 diagnostic accuracy" — a real, measurable win, just not the one a diagnostic-support system is built and certified to deliver.

Calibration and over-reliance: the numbers the headline accuracy figure hides

The trial's most useful finding sits in its exploratory analyses. Diagnostic calibration — mean confidence minus mean observed accuracy — showed overconfidence in every single condition tested, and stand-alone Prof. Valmed was the most overconfident of all of them: 65% mean confidence against 33.3% mean accuracy, a calibration gap of 31.5 percentage points. Confidence rose in both study arms after assistance, but so did overconfidence: from 16.3% to 23.9% in the intervention group, 13.7% to 19.7% in control. And when the authors isolated what physicians did with the copilot's top suggestion specifically, the pattern sharpens further: AI over-reliance — the share of incorrect AI suggestions physicians accepted anyway — was 0.95. Under-reliance, rejecting a correct AI suggestion, was 0.10. Put plainly: when Prof. Valmed was right, physicians used it correctly. When it was wrong, they followed it into the wrong answer 95% of the time. User acceptance survey data from the same 42 physicians shows why that is not surprising — 95% found the tool easy to use and 83% would use it again, but only 64% said they trusted its answers, and just 36% agreed that a mistake made while using it could be corrected quickly and easily. Trust was lower than usability. Reliance was high anyway.

The honest limits

This is a preprint, not yet peer-reviewed, and the authors are direct about what it cannot show. Three vignettes per physician, built entirely from case-history text with no exam findings, labs, or imaging, is a narrow slice of what a rheumatology diagnosis actually involves. Every session ran remotely under a study coordinator's live supervision, with standardized, coordinator-mediated prompts into Prof. Valmed — more controlled than how a physician would query the tool alone at a workstation. Recruitment was voluntary convenience sampling, and the sample already skewed toward LLM familiarity: 89% of enrolled physicians had used an LLM for medical purposes before the trial started, and 83% said they'd welcome certified AI diagnostic support going in — a population primed to find the tool acceptable, which the acceptance numbers bear out, independent of whether it made them more accurate. The calibration and over-reliance findings are exploratory, run on the same 82-physician sample as everything else, and the authors label them accordingly. None of that erases the primary result: a properly randomized, pre-registered, blinded-assessment trial of the only CE-certified LLM diagnostic copilot in Europe found it did not move the outcome it exists to move.

What this means for reproducible, local-first science

A CE mark answers a safety and quality-management question, not an efficacy question — the certification says the tool was built and documented to a standard, not that using it changes what a physician gets right. ALLIANCE is the first trial to test that gap directly for an LLM-based CDS system, and the gap turned out to run the wrong way: the arm with no AI access improved numerically more than the arm with a certified one. That is not an argument against building these tools, and the authors don't make it — faster case turnaround and higher perceived support quality are real outcomes clinics can act on today. It is an argument for treating "certified" and "validated by a randomized trial against the outcome it's marketed for" as two different claims, checked separately, with the second one requiring exactly the kind of exportable, re-runnable trial record — the vignettes, the blinded ratings, the calibration analysis, the raw accept/reject decisions behind that 95% over-reliance figure — that lets the next hospital deciding whether to buy a certified copilot check the number itself, instead of the certificate.

Try MegaBrain BioScience

A research workbench that runs on your machine and exports a reproducibility record for every result, so a claim like a certified tool's trial outcome can be checked, not just cited.