DeepMind Is Aiming AlphaFold at the Virus Family Behind Ebola. A Weapons Lab Is a Partner.
Google DeepMind and Isomorphic Labs published their joint bioresilience approach this week, naming Lawrence Livermore National Laboratory among 15+ partners built over the past 12 months. AlphaFold 3 is pointed at a pan-filovirus broad-spectrum antibody design effort, aimed at the virus family that includes Ebola and Marburg. Two more releases this week put hard numbers on how far AI research agents still are from doing that kind of work unsupervised, and how much a narrow, purpose-built agent can outrun a generalist one when it's scoped to a single job.
DeepMind and Isomorphic Labs: AlphaFold, pointed at outbreak response
"Our approach to bioresilience" (Google DeepMind, July 16) lays out a Prevent/Detect/Respond framework built around three concrete AI-for-science pieces. AlphaEvolve is being used to optimize the algorithms behind metagenomic sequencing analysis, aimed at faster, cheaper pathogen surveillance so new outbreaks get caught sooner. AlphaFold's protein-structure predictions are being used to accelerate vaccine and countermeasure design. And Isomorphic Labs has stood up a dedicated rapid-response unit that can deploy its AlphaFold 3-powered drug design engine during a novel outbreak — with Lawrence Livermore, the UK AI Security Institute, CEPI, and the Francis Crick Institute named as partners, and a specific pan-filovirus antibody-design collaboration with Lawrence Livermore called out by name. The post reports 15+ partnerships with government and biosecurity bodies built over the last year, though it doesn't publish results from the antibody program itself yet — this is a capacity announcement, not a results paper.
Meanwhile: grading an agent against the paper it's copying
ResearchClawBench(arXiv 2606.07591, revised to v5 on July 3) is a 40-task, 10-domain benchmark that grades AI research agents on whether they can reproduce a real published paper's actual findings, not an LLM-judged approximation of them. Across seven autonomous agents and seventeen standalone LLMs, the best score belongs to Claude Code at 21.5 out of 100; the best standalone model, Claude Opus 4.7, scores 20.7. The authors' own conclusion: current systems "remain far from reliable re-discovery." We go deep on what that gap means — and where the real capability is actually showing up — in today's feature.
And the counterexample: narrow beats broad, by a landslide
DeepEvidence, published in Nature Machine Intelligencethis month, is a two-agent orchestrator built for one job: searching and synthesizing evidence across biomedical knowledge graphs. Tested against general-purpose frontier models on the same benchmarks, it isn't close: 40.0% versus 3.3% for both Claude Sonnet 4.5 and GPT-5 on HLE-Medicine, and 80.0% versus 32.0% for the peer-reviewed Biomni agent on LabBench-LitQA2. A model built for the specific task in front of it beat every generalist alternative tested, by as much as 12x.
What this means for reproducible, local-first science
Read together, this week's releases argue for the same thing from two directions. DeepMind's bioresilience program is explicitly betting on domain-specific AI tools — AlphaFold for structure, AlphaEvolve for sequencing algorithms — rather than a single generalist system doing everything. DeepEvidence proves that bet out in benchmark numbers. And ResearchClawBench shows what's still missing when a general-purpose agent tries to do end-to-end science without that kind of purpose-built scaffolding: real reproduction, checked against a real paper, still tops out around 21%.
Try MegaBrain BioScience
A research workbench that runs on your machine and checks its own work.