65% of the Data Powering AI Virtual-Cell Models Is Statistically Unreliable. The Models Trained Without It Won Anyway. Here's Everything Else That Shipped.
A new ETH Zurich analysis of 7,170 single-cell perturbations across 29 public datasets found that 65% are statistically unreliable — noise sitting inside the exact ground truth every perturbation-prediction model in the field is trained and graded against. Filter it out, and training on only the reliable slice still matches or beats training on everything. Full breakdown below. One more item shipped this week: eight days after that audit posted, Arc Institute opened its second annual Virtual Cell Challenge — a $100,000, zero-shot global contest scored on exactly this kind of perturbation-response prediction.
Single-cell: the lead result — two-thirds of the ground truth is noise
Full breakdown in today's feature. Short version: "Reliable single-cell perturbations explain and improve model performance" (Wang, Kuipers, Hugi, Platt, Beerenwinkel; ETH Zurich, SIB Swiss Institute of Bioinformatics, University of Basel; bioRxiv, DOI 10.64898/2026.08.11.744177, posted August 12, 2026) classified 7,170 perturbations from 29 public perturb-seq datasets as unreliable (65%), shared (11%), or specific (24%) — only the last category carries the kind of signal these models are supposed to learn a causal relationship from. Applying those labels to published benchmarks changes which method wins. The counterintuitive part: training with reliable perturbations alone matches or outperforms full-data performance, using just 55% of all training perturbations. The paper also shows a 28-cell pilot experiment can predict, before a full screen runs, how many cells it will need to be reliable — a direct cost lever for the next perturb-seq atlas.
Biotech: a $100,000 contest opens on the same kind of data
On August 20, 2026, Arc Institute announced the second edition of its Virtual Cell Challenge, registration at virtualcellchallenge.org. The task is zero-shot: predict CRISPRi knockdown responses in six cell lines entrants have never seen perturbed, with no training set provided for those lines and an evaluation set the organizers describe as significantly larger than last year's. The grand prize is $100,000, sponsored by NVIDIA, 10x Genomics, and Ultima Genomics. Nothing here says anything about the reliability composition of this year's held-out cell lines specifically — that is not published, and it would be unfair to assume it mirrors the 65% figure above. But the ETH paper's method is now sitting there, published, as a way to check.
What this means for reproducible, local-first science
The useful habit here is not distrust of any particular leaderboard — it is treating "this dataset passed standard preprocessing" and "this dataset is trustworthy ground truth" as two different claims, and having a cheap, routine way to check the gap between them before a result gets built on top of it. A 28-cell pilot that flags an unreliable perturbation before the full screen runs is exactly that kind of check: useful only if it is fast enough to run locally, on data you already have, before committing budget to the real experiment.
Try MegaBrain BioScience
A research workbench that runs on your machine and exports a reproducibility record for every result.