← MegaBrain BioScience Blog
Feature · July 28, 2026

Every AI Science Agent Is Also Trying to Do Astrophysics. We're Dropping Everything That Isn't Life Sciences.

MegaBrain Science is now MegaBrain BioScience. The app, the skills, the reviewer checks, and this blog all narrow to one field: biotech, healthtech, medtech, genomics, single-cell, proteomics, structural biology, cheminformatics, and the clinic. Here is the number that decided it.

58.0%
Best generalist score on AstaBench
21.5 / 100
Best score when graded against a real paper
12x
Margin a narrow biomedical agent won by in its own lane

The evidence was in our own archive

Two weeks ago we published The Specialization Dividend. Claude Opus 4.7 tops AstaBench, the leaderboard the field uses to rank research agents, at 58.0%. Grade the same model class against a benchmark that checks the work against a real published paper and the best score anywhere is 21.5 out of 100. A narrow biomedical agent, built for one job, beat frontier generalists by up to 12x on the same exam.

We wrote that up as a finding about the field. It is also a finding about us. We were building a general research workbench — bio, chem, physics, materials, ML — and citing evidence that general research workbenches lose to narrow ones on the tasks that count. That is not a position you get to hold for long.

What actually changes

The product. Every skill, connector, renderer, and eval in the app now points at the life sciences. The database surface is bio-first — UniProt, PDB, AlphaFold DB, Ensembl, ClinVar, gnomAD, ChEMBL, PubChem, GEO, Human Cell Atlas, PRIDE, cBioPortal, ClinicalTrials.gov, PubMed, bioRxiv, medRxiv. The kernel still runs on your machine, because in this field the data is patient cohorts and unpublished screens, and local compute is the condition for using the tool at all.

The reviewer. A domain-blind reviewer checks whether the code ran. A bio reviewer checks whether a batch effect is being read as biology, whether ambient RNA got annotated as a cell type, whether a gene symbol was silently coerced to a date, and whether the reference build matches the annotation. Those checks are the whole point of specializing, and they only exist if you pick a field.

This blog. Nine beats, listed below. Every post has to land in one of them, with a primary source and the numbers attached. The selection rules are written down in the repo so they bind future posts rather than drifting.

The nine beats

Biotech

Discovery platforms, lab automation, synthetic biology, bioprocess. The companies actually running wet labs.

HealthTech

Clinical decision support, EHR-scale modeling, patient-facing systems, and the evaluation of all three.

MedTech

Devices, imaging, diagnostics, and the regulatory path an AI-enabled instrument has to clear.

Genomics

Variant calling and interpretation, GWAS, population-scale cohorts, sequence models.

Single-cell

scRNA and multiome atlases, integration and batch correction, cell-type annotation, perturbation screens.

Proteomics

Mass-spec pipelines, PTM discovery, protein quantification, and the ML that now sits on top of them.

Structural biology

Folding and co-folding, docking, cryo-EM, and the growing gap between predicted and experimental structures.

Cheminformatics & drug discovery

Generative chemistry, ADMET prediction, retrosynthesis, and whether any of it survives a wet-lab check.

Clinical & translational AI

Trial design and matching, real-world evidence, evidence synthesis, and translational research operations.

What we are giving up

This is a real cost, so it is worth naming. We have covered AI in mathematics, materials discovery, climate modeling, quantum computing, and general ML-research agents — some of our best-read posts are in exactly those lanes. Going forward those stories only run when they touch the life sciences directly. A folding-model result runs. A kagome superconductor does not.

General AI-agent methodology — reward hacking, memory budgets, citation verification, multi-agent scaling — stays, but only when the evidence is drawn from a bio task or a bio benchmark. The method matters here because the failure lands on a drug program, not because the method is interesting on its own.

Nothing already published moves or disappears. Every old URL redirects: the archive stays exactly where search engines and your bookmarks left it.

Why we think this wins

Claude Science, Google Co-Scientist, and Microsoft Discovery are all excellent, and all general. Their skills budget is split across every field a scientist might work in. Ours is not split at all. In a field where the difference between a real result and a plausible-looking one is a batch effect nobody checked for, that concentration is the product.

Try MegaBrain BioScience

A workbench built only for the life sciences — running on your machine, checking its own work.