← MegaBrain BioScience Blog
BioSignal #20 · Field notes · August 13, 2026 · 5 min read

An Audit of 32 AI Models Found a 50.7% Rate of Working Toxin Designs. Refusing More Often Didn't Make a Model Safer.

A new audit tested 32 large language models against 631 curated toxin-design prompts and scored each one's Functional Harmfulness Rate: whether its output was biologically plausible enough to actually be dangerous, not just whether the model complied. Across the field, that rate came out to 50.7%, and it tracked each model's raw biological-generation capability, not how often it refused. A model that says no more often can still be the one most likely to produce a working design when it doesn't. Also this week: a mechanism-based AI model doubled a cystic fibrosis gene-correction rate while testing 8 candidate designs against a rival's 16, and a Yale-designed synthetic cell-surface protein matched or beat nature's own best performer.

50.7%
Functional Harmfulness Rate across 32 LLMs on toxin-design prompts, uncorrelated with refusal rate
22% vs. 11%
CFTR F508del correction: OptiPrime vs. conventional design, from just 8 candidates
7 / 120
AI-designed cell-surface proteins that matched or beat the best natural one, of those validated

Biotech & biosecurity: refusal rate is measuring the wrong thing

"A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models" (Quan, Hao, Fang, Geng, Zhou, Chen, Wang, Hong, Dai, Yang, Ji; arXiv:2608.02684, posted August 3, revised August 5; accepted to COLM 2026) audited 32 LLMs against 631 curated toxin-design prompts spanning 7 functional categories and scored each model's Functional Harmfulness Rate (FHR): whether the model's output was biologically plausible enough to be actually dangerous, not just whether it complied. The FHR came out to 50.7% across the field, and the paper's central finding is that this tracked each model's underlying biological-generation capability, not how often it refused. A model that says no often can still be the one most likely to produce a working design when it doesn't, which means refusal rate, the metric most safety reporting leans on, is measuring compliance, not risk. The authors' own mitigation, a classifier called BioSafe-Guard, cut predicted functional risk without cutting benign scientific utility, which is the actual fix: a better filter, not a blunter refusal policy.

Biotech: a design model that beats the field by naming what it beat

"Mechanistic machine learning for prediction of prime editing outcomes" (Hsu, Chen, Li, Hemez, Gao et al., David R. Liu lab, Nature Biotechnology, DOI 10.1038/s41587-026-03261-7, published August 12) introduces OptiPrime, which models prime editing as the multi-step biochemical process it actually is (nicking, mismatch repair, flap resolution) instead of treating pegRNA design as a black-box sequence-to-efficiency regression the way DeepPrime and PRIDICT do. On the CFTR p.F508del correction task, the single most common cystic-fibrosis mutation, screening just 8 model-ranked candidates reached 22% correction, against 11% from the field's standard epegRNA+PE2 optimization workflow and under 1% from the top 16 candidates competing black-box models proposed. In a Kif1a-mutant mouse model of a neurological disorder, OptiPrime-guided design reached 40% bulk correction in brain-cortex tissue in a 4-week optimization cycle; the Broad Institute's own writeup says that level of in vivo efficiency "normally take[s] months or longer to achieve." The model and a public webserver are open at optipri.me.

Biotech: an AI-designed surface protein beats nature's best, 7 times out of 120

"Discovery and design of potent cell surface display elements" (Fang, Saskin, Lee, Zou, Xin, Huang, Pan, Dong, Abiri, Feng, Sahni, Yi, Peng, Chen; Sidi Chen lab, Yale, Nature Biotechnology, DOI 10.1038/s41587-026-03144-x, published August 12) built DeepSCan, a deep-learning suite trained on measured surface expression across more than 570 chimeric antigens, to design new cell-surface display (CSD) modules for mRNA antigen display, the anchor that determines whether an antigen actually shows up on a cell's surface for something like a CAR-T or vaccine construct. Of 3,700 computationally designed CSDs, about 120 were experimentally validated, and 7 matched or exceeded the translocation strength of the single best naturally occurring CSD the team could find, a 5.8% hit rate among validated designs, and the first time a generative model has out-designed nature on this specific property rather than just approaching it.

What this means for reproducible, local-first science

All three results this week turn on the same question: what does the number actually show, once you look past the headline? SPIKE-Bench's answer is uncomfortable, the safety metric the field reports doesn't predict the risk it claims to measure, and a model looks safer by refusing than by actually being safer. OptiPrime and DeepSCan show what the opposite discipline looks like applied to capability claims: OptiPrime reports the competing black-box baseline's under-1% result right next to its own 22%, and DeepSCan reports 7 winners out of 120 validated attempts, not just "we beat nature." A number is only as trustworthy as the denominator and the baseline published next to it, whether the claim is a capability or a safety guarantee.

Try MegaBrain BioScience

A research workbench that runs on your machine and exports a reproducibility record for every result.