← MegaBrain BioScience Blog
BioSignal #25 · Field notes · August 21, 2026 · 6 min read

A Single-Cell AI Model Passed the Simulation Benchmark. Against Real CRISPR Data, It Called the Direction Right 40.9% of the Time. Here's Everything Else That Shipped.

CellOracle was one of only two out of eight tested methods that reliably detected a known regulatory signal in a systematic single-cell perturbation benchmark. Checked against real CRISPR interference knockdowns, it predicted the correct direction of the effect 40.9% of the time — statistically indistinguishable from a coin flip. Full breakdown below. Four more results shipped this week: a 511-antibody, 29-organization blinded AI antibody-design competition where every submitted model but one performed worse than random clone picking at ranking affinity; a 70-million-parameter single-cell model that matched or beat billion-parameter rivals by training on proteomics data instead of more RNA; a Harvard benchmark showing Claude Opus 5 outranks 49 of 95 published protein-variant predictors while still trailing the strongest specialist tool; and a 32-person randomized trial where AI-assisted ultrasound cut scan time 9% without measurably reducing overall sonographer workload.

40.9%
CellOracle's direction accuracy against real CRISPR ground truth — chance level
511 / 29
AI-designed antibodies submitted, by organizations, in the AIntibody blinded competition
70M vs. 3B
Parameters: a proteomics-trained single-cell model vs. the RNA-only model it matched

Single-cell: the lead result — a benchmark winner that loses to chance

Full breakdown in today's feature. Short version: a systematic benchmark of eight single-cell perturbation-prediction methods across four datasets found only two, CellOracle and DDIM, that reliably detect a known transcription-factor-to-pathway signal. Checked against real CRISPRi knockdown data in K562 cells, CellOracle's predicted direction of effect matched the true direction 40.9% of the time, not statistically different from chance. A second finding from the same paper: two other methods, DDIM and scTenifoldKnk, produced gene rankings on the identical task that are anti-correlated, rho = −0.811 — method choice alone can reverse a study's biological conclusion.

Cheminformatics & drug discovery: AI beat random at one task out of three

"A blinded, prospective benchmark of in silico antibody discovery anchored to experimental affinity and developability" (Erasmus, Bedinger, et al.; Nature Biotechnology, DOI 10.1038/s41587-026-03238-6, published August 19, 2026) is the second round of the AIntibody challenge, an experimentally validated, CASP-style blind competition for computational antibody design. It tested 511 AI-designed or predicted antibodies from 29 organizations across three tasks, validated with diverse wet-lab assays.

AIntibody taskResult across 511 submitted antibodies, 29 organizations
Affinity maturation (from phase-1 sequencing outputs)Modeled effectively
Affinity ranking within HCDR3 clustersWorse than random clone picking, except for one model
De novo CDR design of out-of-library proteinsHighly variable; many submissions failed to beat standard selection

Source: Erasmus, Bedinger, et al., "A blinded, prospective benchmark of in silico antibody discovery anchored to experimental affinity and developability," Nature Biotechnology, DOI 10.1038/s41587-026-03238-6, published August 19, 2026.

Several groups produced developable antibodies with affinities under 100 pM, but the paper's own framing is that these were exceptions that did not transfer across tasks: affinity maturation from real sequencing data was modeled effectively, but predicting which antibody in an existing cluster binds best — the task closest to how discovery teams actually triage candidates — was, for every model but one, worse than picking a clone at random.

Single-cell: proteomics data beat 40x more parameters

"Single-cell foundation models benefit from cross-modal training: adding proteomics data beats parameter scaling" (Burq, Stepec, Kim, Cimermancic; Tesorai, Inc.; bioRxiv, DOI 10.64898/2026.08.14.744845, posted August 14, 2026) fine-tuned a 70-million-parameter Tahoe-x1 single-cell model for one epoch on 48,843 proteomic samples drawn from 440 mass-spectrometry studies. That cross-modal model matched or exceeded the original 1-billion- and 3-billion-parameter RNA-only Tahoe-x1 models across most of the benchmark suite, and it transferred better to a held-out protein-perturbation benchmark, where scaling the RNA-only model further provided no comparable benefit. Targeted proteomics curation, not more parameters, was the lever that moved this benchmark.

Proteomics: a fair look at how far general-purpose LLMs actually get

"PG-LLM: Benchmarking General-Purpose Language Models for Protein Variant Ranking" (Arora, Chen, Du, Marks, Church; Harvard Medical School; bioRxiv, DOI 10.64898/2026.07.27.741045, posted July 27, revised August 19, 2026) built 276 variant-prioritization tasks — 217 from ProteinGym plus 59 held out from studies first published after January 2026, to guard against training-data contamination — and tested 13 general-purpose LLMs against 95 published, purpose-built protein predictors on the same variants. Claude Opus 5 (Max) was the strongest LLM tested, at a Spearman correlation of 0.406, edging out GPT 5.6 Sol (Max) at 0.402. Opus 5 beat 49 of the 95 published predictors, including 41 of 46 sequence-only methods, and nearly matched ESM2-650M (0.411) — but stayed well behind the leading specialist predictor, VenusREM, at 0.523. Ranking performance scaled with test-time compute across the GPT, Claude, and Gemini families tested, but the gains tapered before closing the gap to the best specialist tools.

MedTech: AI ultrasound saved time, not effort across the board

"Quantifying Human-AI Workflow in Abdominal Ultrasound: A Prospective Randomised Crossover Study" (Hsiao, Clifford, Lin, et al.; Queensland University of Technology; medRxiv, DOI 10.64898/2026.08.17.26360254, posted August 17, 2026) had 32 healthy adults each undergo one manual and one AI-assisted (ACUSON Sequoia AI Abdomen Release 3.5) upper-abdominal ultrasound exam, in randomized order, with a depth-camera system tracking sonographer hand movement. AI assistance cut scan time by 52.4 seconds (about 9%, p = 0.001), keystrokes by about 28% (p < 0.001), and hand travel by about 39% (p < 0.001) — though the time saving was concentrated in one of the two sonographers. Overall self-reported workload (NASA-TLX) did not differ significantly between conditions, but mental demand and effort subscales both dropped, with no compensating increase elsewhere. Sonographers manually modified 48 of 184 AI-generated measurements, a reminder that the tool assisted the exam rather than replaced the sonographer's judgment.

What this means for reproducible, local-first science

The pattern across this week's strongest result and its lead item's digest companions is the same one: a number that looks like a verdict is usually a number that measured one specific thing, and the thing it measured is worth checking before you build on it. A perturbation model that passes a correlational benchmark has not been shown to understand causation. An antibody-design model that produces one sub-100-pM hit has not been shown to beat random selection at scale. A 70-million-parameter model beating a 3-billion-parameter one says more about what data was added than what scale is worth. And a general-purpose LLM beating half of 95 specialist predictors is real progress that is still, honestly, a loss to the best specialist tool. None of that is a reason to distrust the results — it is a reason to keep the condition each number was measured under attached to the number itself.

Try MegaBrain BioScience

A research workbench that runs on your machine and exports a reproducibility record for every result.