← MegaBrain BioScience Blog
BioSignal #28 · Field notes · August 26, 2026 · 4 min read

A Clinical AI Was Right 99% of the Time. Clinician Adoption Fell From 68% to 30% in Four Weeks Anyway.

This week's sharpest number isn't about whether a clinical AI works. It's about what happens after it does. A DECIDE-AI Stage 1 trial of an emergency-department decision-support tool found experts rated 99 of 100 sampled outputs appropriate. Clinician adoption fell from 68% to 30% over the same four weeks anyway. Being right bought the tool no trust. Below that: a liver-malignancy AI validated across 22,251 real-world patients that caught 51 lesions overlooked on first read, a third FDA-cleared cardiac device from Tempus, a model-guided search for a more efficient compact genome editor, a single-cell benchmark where two scoring methods disagree on which model is best, and an AI co-data-scientist whose biomarker picks held up under a 12-clinician review.

68% → 30%
Clinician adoption of a 99%-accurate ED decision-support AI, over 4 weeks
51 lesions
Overlooked liver lesions an AI-radiologist collaboration caught across 22,251 patients
56.9% vs 18.8%
Manuscript content clinician reviewers kept from an AI co-scientist's report vs. baseline systems

HealthTech: right 99% of the time, trusted less every week

A DECIDE-AI Stage 1 evaluation (Nature Medicine, s41591-026-04601-5, published August 19, 2026) tracked "SHAKED," a multi-LLM decision-support tool for emergency-department triage, across 2 units and 1,138 patients over 4 weeks. Expert review rated 99 of 100 sampled outputs appropriate. Clinician adoption still fell from 68% to 30% over the trial period. Emergency-department length of stay was unchanged (4.9 hours in both arms), and a 9.4-minute reduction in consultation time didn't reach statistical significance. The tool wasn't wrong. Clinicians stopped using it anyway — and DECIDE-AI's reporting structure is exactly what surfaced that gap instead of letting an accuracy number stand in for it.

MedTech: an AI-radiologist catch rate, and a third FDA clearance

"LION" (Nature Medicine, s41591-026-04589-y, published August 19, 2026) trained a liver-malignancy detection model on 6,443 patients, then validated it on a 22,251-patient real-world multicenter cohort (AUC 0.952) and ran a 10,333-patient single-arm trial. Overall AUC across the program reached 0.975 (95% CI 0.971–0.979), holding up in harder subgroups: 0.971 with steatosis, 0.924 with cirrhosis. In the single-arm trial, AI-assisted review caught 51 lesions overlooked on the initial read — 15 of them malignant — triggering 37 amended reports and 22 multidisciplinary-team escalations. Separately, Tempus received its third FDA 510(k) clearance for a cardiovascular AI product on August 24, 2026: Tempus ECG-PH, which flags signs of pulmonary hypertension from a standard 12-lead resting ECG in symptomatic patients 40 and older with no known PH history. It's explicitly not a stand-alone diagnostic and isn't cleared for serial monitoring or paced rhythms.

Biotech: a model-guided search beats the previous best compact editor by 2.6x

A Nature Biotechnology paper (s41587-026-03272-4, published August 24, 2026) used a model-guided optimization pipeline, EvoMax, to search variants of a compact Fanzor nuclease. The resulting editor, FanzMAX v3-hLa, reaches up to 97% editing efficiency at its best locus and roughly 33% mean efficiency across 19 endogenous genomic loci — more than 2.6× the mean efficiency of the previous best compact nucleases tested (enNlovFz2 and enCnCas12f1). Off-target rates aren't reported in the abstract; a compact nuclease with this kind of editing-efficiency jump is worth reproducing directly rather than taking the topline number on faith.

Single-cell: two ways to score a perturbation model, two different winners

"scDrugPerturb-Bench" (bioRxiv 10.64898/2026.08.19.745729, posted August 22, 2026) scores 12 single-cell perturbation-prediction models plus 3 baselines against 423 annotated drug-response cases (717 key genes) drawn from 181 datasets covering 2.5 million cells. Scored the standard way, by expression-similarity to the ground truth, versus scored by the paper's own Mechanism Fidelity Score — which checks whether a model recovers the correct key-gene direction, effect size, and pathway polarity rather than just an overall expression pattern — the two metrics select different "best" model configurations across the 10 train/test splits tested. The metric you pick decides which model wins.

Clinical & translational AI: an AI co-scientist's picks survive expert review

Google Research's "CoDaS" (arXiv 2604.14615), a multi-agent AI co-data-scientist, mined wearable-sensor data across three cohorts (9,279 participant-observations) to prioritize candidate biomarkers, running each candidate through its own replication, stability, and leakage checks before handing it to a clinician. Feeding CoDaS-derived features into a prediction model lifted cross-validated R² by 0.040 for depression and 0.021 for insulin resistance. In a 12-clinician review totaling roughly 25 active hours, clinician confidence tracked CoDaS's own confidence tiers (ρ = 0.67, p = 0.005), reviewers kept 56.9% of CoDaS's report content versus 18.8–30.4% for the baseline systems it was compared against, and it ranked first in 9 of 13 head-to-head sessions.

What this means for reproducible, local-first science

Every system this week clears its own accuracy bar. That's not the number that ended up mattering. SHAKED was correct 99% of the time and lost most of its users anyway; LION's 0.975 AUC only became useful once a radiologist could see the 51 specific cases it caught and act on them; CoDaS earned its 56.9% retention rate one clinician-reviewed report at a time, not from its R² improvement alone. None of that is a capability gap — it's a trust-and-verification gap, and it closes the same way every time: someone has to be able to check the specific claim, on their own terms, before they'll act on it. That's the same problem a local, reproducibility-first workbench is built around.

Try MegaBrain BioScience

A research workbench that runs on your machine and exports a reproducibility record for every result.