← MegaBrain BioScience Blog
BioSignal #29 · Field notes · August 31, 2026 · 5 min read

Claude Designed Protein Binders for 15 Targets. Two Independent Labs Confirmed 14 Worked, at Up to 3.5x the Field's Usual Hit Rate.

This week's sharpest number belongs to a rival, not a benchmark leaderboard. Anthropic ran Claude through an autonomous protein-binder design campaign against 15 targets, then sent every design out for blind wet-lab synthesis and testing at two independent labs. Fourteen of fifteen targets produced at least one working binder, and 22–35% of individual designs bound — against a field-typical hit rate of 10–15%. Below that: a cryo-EM trick that fixed a resolution number that looked fine on paper and wasn't usable in practice, a pathology model whose own press coverage couldn't agree on its accuracy until someone read the primary text, a longer-context RNA foundation model, a citywide rapid-genome-sequencing program that changed treatment for over half its patients, and a retrieval method that ties spatial transcriptomics predictions to actual biology instead of just pixels.

14 of 15
Protein targets where at least one Claude-designed binder bound, in blind independent wet-lab testing
22–35%
Per-design hit rate for Claude's binders, vs. a field-typical 10–15%
0.939 vs 0.871
AUC for a 3D-context pathology model vs. an isolated-2D-slice baseline, prostate biopsies

Cheminformatics & drug discovery: a rival's AI protein designs, checked by someone else's lab

Anthropic's own account of the work, " How Claude is accelerating protein design and analytical chemistry," published August 18, 2026, describes an agentic design loop where Claude proposed binder sequences against 15 protein targets. The designs weren't graded by Anthropic's own model — they were synthesized and tested blind by two outside labs, Adaptyv Bio and Twist Bioscience. Fourteen of the fifteen targets came back with at least one confirmed binder, and 22–35% of individual designs bound successfully, well above the 10–15% hit rate the post cites as typical for the field. This is a first-party vendor write-up, not a peer-reviewed paper, and the per-target breakdown behind that 22–35% range isn't published in full. It still clears our bar for a reason most competitor claims don't: the grading happened outside the company making the claim.

Structural biology: a resolution number that looked fine and wasn't

A Nature Methods paper (s41592-026-03184-w) encapsulates target proteins in MS2-derived protein shells — "nanocrates" — before cryo-EM sample prep, to stop the protein from sticking to the air-water interface, the step responsible for most orientation bias and particle damage in cryo-EM. Conventionally prepped thyroglobulin reported a 2.29 Å global resolution — a good-looking number that was unusable in practice, because severe preferred-orientation artifacts left the map too anisotropic to interpret. Encapsulated in a nanocrate, the same protein resolved to 2.94 Å: a nominally worse number that was actually a usable, interpretable structure. Apoferritin reached 2.16 Å and 7,8-dihydroneopterin aldolase reached 2.80 Å with the same method. The headline resolution figure isn't the number that decides whether a structure is real.

MedTech: a pathology model that survived a primary-source recheck

"CARP3D" (Nature Biomedical Engineering, s41551-026-01760-1) uses attention over 2.5D context — a few adjacent slices at a time, rather than one isolated 2D slice or a full 3D volume — to triage which regions of a volumetric biopsy are highest-risk. On prostate-cancer risk stratification (112 biopsies, 54 patients, 121 annotated levels), CARP3D reached an AUC of 0.939 versus 0.871 for an isolated-2D-slice baseline (p<0.005). On Barrett's esophagus dysplasia and cancer screening (95 specimens, 24 patients, 334 annotated levels), it reached 0.921 versus 0.895 (p<0.005). We flagged this paper in an earlier sweep because secondary coverage reported at least three different, partly contradictory AUC pairs for it. The numbers above are quoted directly from the paper's own PMC-hosted text, not from any press summary — which is the only way we'd trust them enough to print here.

Genomics: a longer memory for RNA, and a citywide sequencing program that changed care

"RIBOSPAN" (arXiv 2608.22849) is a 1.61-billion-parameter RNA foundation model natively pretrained on sequences up to 10,240 nucleotides — long enough to see a full mRNA transcript instead of a fragment. Against existing encoder-only RNA models, it improves zero-shot mutation-fitness prediction (RNAGym AUROC 0.6280 vs. AIDO.RNA-CDS's 0.6024; Spearman correlation 0.2524 vs. RNA-FM's 0.1674) and RNA-biotype classification accuracy (0.8989 vs. RNA-FM's 0.8650, on 89,955 sequences). Separately, a Nature Medicine paper (s41591-026-04598-x) reports on "Little Falcon," a centralized, citywide rapid trio whole-genome-sequencing program for critically ill children across Dubai's NICUs and PICUs, serving patients referred from 18 countries. Across 100 sequenced patients, diagnostic yield reached 53% overall (95% CI 43.3–62.5%) and 80% in consanguineous families, with a median turnaround of 3.4 days. Management changed for 53% of patients, and 12% carried more than one molecular finding driving their presentation.

Single-cell: teaching a retrieval model what "biologically similar" means

"BioKERN" (arXiv 2608.24823) adds a biology-aware kernel — combining transcriptomic similarity with spatial proximity — as explicit training supervision for retrieving biologically matched neighborhoods between histology images and spatial transcriptomics data. Against the BLEEP baseline, it lifts multi-scale retrieval accuracy from 0.4960 to 0.6716 on a mouse brain Visium dataset (+35.4% relative) and from 0.0312 to 0.0408 on a human liver dataset (+30.8% relative). The paper's ablations attribute 63–91% of that gain to the kernel regularization itself, not to any added model capacity — the model got better mostly by being told what biological similarity means, not by getting bigger.

What this means for reproducible, local-first science

Two results this week only hold up because someone went back to the primary text. CARP3D had three conflicting AUC pairs floating around before its own paper settled the question. Thyroglobulin's conventional cryo-EM prep reported a resolution number that was better than the fix that replaced it — and worse in every way that mattered, because the map underneath it wasn't usable. Even the strongest result this week, Claude's protein-binder hit rate, only clears our bar because the grading happened outside the company that built the model. None of that is an argument against any of these results. It's an argument for being able to check them yourself, on your own machine, before you build on top of them — which is the whole premise a local, reproducibility-first workbench is built around.

Try MegaBrain BioScience

A research workbench that runs on your machine and exports a reproducibility record for every result.