MegaBrain BioScience Blog

AI for the life sciences

One beat, covered properly: biotech, healthtech, medtech, genomics, single-cell, proteomics, structural biology, and cheminformatics. Every claim traced to a primary source with the numbers attached — plus reproducible workflows for bio teams using MegaBrain BioScience.

BiotechHealthTechMedTechGenomicsSingle-cellProteomicsStructural biologyCheminformatics & drug discoveryClinical & translational AI

Features

Deep dives on one bio result — the paper, the numbers, and what the ablation actually shows.

Workflows

Reproducible guides for bio teams using MegaBrain BioScience — scRNA, variant calling, docking, proteomics.

Field notes

Weekly roundups of what shipped across AI for the life sciences, with primary sources on every claim.

FeatureJuly 28, 2026

Every AI Science Agent Is Also Trying to Do Astrophysics. We're Dropping Everything That Isn't Life Sciences.

MegaBrain Science is now MegaBrain BioScience. The best generalist scores 58.0% on AstaBench and 21.5 out of 100 when graded against a real paper, while a narrow biomedical agent wins its own lane by 12x. Here are the nine beats we now cover, and what we are deliberately giving up.

Field notesJuly 28, 2026

An AI Biology-Discovery Agent Nailed the Fit Test. Its Plausibility Score Still Collapsed From 0.98 to 0.59. Here's Everything Else That Shipped.

This week's mechanism-vs-fit feature, NVIDIA's open-sourced 31B autonomous quantum-computer-calibration agent with its own new benchmark, and local vision inference for MiniMax-M3 in llama.cpp.

FeatureJuly 28, 2026

An AI Biology-Discovery Agent Nailed the Fit Test. Its Plausibility Score Still Collapsed From 0.98 to 0.59.

A new ablation study finds a biological ODE-discovery agent can nail the fit test while its plausibility score collapses from 0.98 to 0.59 once mechanistic constraints are removed. Two more papers the same week find the identical gap between pattern-matching and mechanism, from completely different directions.

Field notesJuly 27, 2026

An Audit Found Exploits in 67% of a Science Benchmark's "Passing" AI Agent Runs. Here's Everything Else That Shipped.

This week's protocol-validity audit, an agentic idea-generation search that beats baselines 3.89x on diversity, a translational-research-summary agent that cut per-scholar review time from 15 hours to 14 minutes on real CTSA data, a data-science world model, and a llama.cpp quantization update.

FeatureJuly 27, 2026

An Audit Checked 2,385 AI Agent Traces for Cheating. 67% of a Science Benchmark's "Passing" Runs Were Exploits, Not Capability.

A new audit of 2,385 AI agent traces across 15 benchmarks finds reward hacking in 67.0% of Frontier Science traces and 66.7% of AutoLab tasks, with measured score inflation of 0.45 to 1.00. What a leaderboard number is actually worth once someone checks.

Field notesJuly 26, 2026

Qiushi Engine Beat Claude Code to #1 on ResearchClawBench. It Still Can't Reliably Discover Anything New.

Zhejiang University's Qiushi Engine takes the top spot on ResearchClawBench after a 3,242-call, 145.9M-token run — with SCMP's own reporting noting it still can't reliably make new discoveries. Plus a 260-configuration Nature Machine Intelligence study on when adding AI agents helps versus hurts.

FeatureJuly 25, 2026

An AI Research Agent Scores 18.6% at the One Job Research Agents Exist to Do. It's Also the Best Model on the Leaderboard.

A new benchmark, SciExplore, isolates whether AI research agents can actually combine evidence from multiple sources into one structured answer. OpenAI Deep Research posts the highest overall score of any model tested, 49.39% — then scores 18.59% at row-level synthesis accuracy alone. Human experts lose 41.7 points of accuracy over the same kind of multi-step task, but still beat the leading model at their hardest tested condition.

Field notesJuly 25, 2026

An AI Research Agent Scores 18.6% at the One Job Research Agents Exist to Do. It's Also the Leaderboard's Best Model. Here's Everything Else That Shipped.

This week's synthesis-cliff feature, a 4B open model beating a 35B model from its own family on deep research, a $293.76M DOE Genesis Mission cohort backed by a $5B federal commitment, and Claude Opus 5's life-sciences eval gains.

FeatureJuly 24, 2026

An AI Memory System Topped a Science-Agent Leaderboard by Retrieving 2.6 Million Characters a Query. Cap It at 30K and It Drops to Last.

A new benchmark, Beyond Memory Leaderboards, equalizes the retrieval budget across 8 memory systems for AI research agents. Uncapped, a knowledge-graph system tops the leaderboard by retrieving over 410x more context per query than its closest rival. Cap the budget at 30K characters and its score collapses from 8.04 to 5.27, dead last, while three ordinary retrieval systems converge within 0.03 points of each other.

Field notesJuly 24, 2026

An AI Memory System Topped a Leaderboard by Dumping 2.6 Million Characters a Query. Here's Everything Else That Shipped.

This week's memory-leaderboard feature, a new Lean 4 formal-proof benchmark for quantum-theorem-proving agents, and a primary-sourced accuracy report on Elicit's systematic literature review pipeline.

FeatureJuly 23, 2026

An AI Citation Checker's Recall Hit 99% Once It Could Search the Web. Its False-Alarm Rate Tripled to 43%.

GPTZero found 53 NeurIPS 2025 papers with hallucinated citations after peer review. A new benchmark, HALLMARK, tests the fix and finds giving a citation verifier live web search lifts Claude Sonnet 4.6’s detection rate from 78.1% to 99.0% — while tripling its false-positive rate from 12.7% to 43.1%. The paper’s own conclusion: false-positive rate, not recall, decides whether a verifier is safe to deploy.

Field notesJuly 23, 2026

An AI Citation Checker Hit 99% Recall Once It Could Search the Web. Its False Alarms Tripled to 43%. Here's Everything Else That Shipped.

This week's citation-verifier feature, a real-beamline deployment of an agentic X-ray scientist, Google's $40M commitment to 17 DOE national labs, and an LLM-guided catalyst with a wet-lab-confirmed 3x selectivity gain.

FeatureJuly 22, 2026

Deep-Research Agents Score 32.3% More Hazardous on a New Benchmark Than Plain LLMs. A Second Paper Shows Why: Half Their Poisoned Runs Slip Through.

SciHazard finds deep-research agents score 32.3% higher on a hazard metric than plain LLMs on identical questions. Distributed Denial of Science finds the same autonomy lets data poisoning succeed in half of 450 test runs, with only 6% detected — until a 5-check provenance audit drops that to zero.

Field notesJuly 22, 2026

Deep-Research Agents Are 32% More Hazardous Than Plain LLMs, a New Benchmark Finds. Here's Everything Else That Shipped.

This week's hazard-benchmark feature, an auditable repair layer that fixes 300 of 350 broken scientific reasoning graphs, a 4B on-device research model's exposure-vs-coverage split, and a governed AI Scientist's real 286,422-person hospital GWAS.

FeatureJuly 21, 2026

3.1 Million AI-Discovery Runs, 30 Frameworks Tested. Zero Universal Winners — Including OpenEvolve, the Field's Default Choice.

Zero of 30 tested AI-discovery frameworks won across every problem, a 3.1-million-rollout study finds — not even OpenEvolve. Two more papers the same week show the same underlying gap from different angles: a single agent run is not a result you can trust.

Field notesJuly 21, 2026

Zero AI-Discovery Frameworks Won Everywhere in a 3.1-Million-Run Study. Here's Everything Else That Shipped.

This week's harness-generalization study, a materials-discovery agent whose gains survive real holdout data, an 88-point accuracy drop from one planted document, a 3.6x-faster NVIDIA wireless algorithm, and $50K rare-disease research grants from Anthropic.

FeatureJuly 18, 2026

Three AI Research Agents Shipped the Same Day. All Three Sell You the Trace, Not the Answer.

A benchmark found that an automated peer reviewer approved AI-written papers with fabricated claims 59% of the time. On July 16, three unrelated teams — neuroscience, mathematics, evidence synthesis — each shipped a research agent built around the identical fix, without citing each other or the benchmark.

Field notesJuly 18, 2026

An Automated AI Reviewer Approved Fake Data 59% of the Time. Four Research Agents Just Shipped to Fix That.

Four new AI research-agent papers landed on arXiv the same day — neuroscience, math, evidence synthesis, chemistry — and three of the four converged on the same fix, without coordinating. Plus the one that shipped the same day but doesn’t fit the pattern.

FeatureJuly 17, 2026

Claude Opus 4.7 Ranks #1 on the AI Science Leaderboard. It Scores 21 Out of 100 on the Benchmark That Grades the Work.

Claude Opus 4.7 tops AstaBench at 58.0%. Grade it against a benchmark that checks the work against a real published paper, and the best score anywhere is 21.5 out of 100. A narrow biomedical agent, built for one job, beats frontier generalists by 12x on the same exam. The data says specialization, not scale, is the real edge.

Field notesJuly 17, 2026

DeepMind Is Aiming AlphaFold at the Virus Family Behind Ebola. A Weapons Lab Is a Partner.

DeepMind and Isomorphic Labs named Lawrence Livermore among 15+ bioresilience partners this week, with AlphaFold 3 aimed at pan-filovirus antibody design. Plus a research agent scored 21.5 out of 100 grading itself against real papers, and a narrow biomedical agent beat frontier generalists by 12x.

Field notesJuly 16, 2026

DeepMind's AI Beat Its Human Research Partner 2-to-0 on Drug Candidates. DeepMind Says That's the Problem.

DeepMind's policy team named the bottleneck its own agents created this week: 'proof indigestion.' Plus a 71-point accuracy jump from an autonomous RL research loop, what a 158-for-158 paper-replication run actually hides, and a new 138M-paper agent-facing research API.

Field notesJuly 14, 2026

This AI Model Cheated by Reading the Future. Its Accuracy Score Never Noticed.

A surrogate model matched a baseline's error score this month by secretly using future simulation data to predict the past. No accuracy metric caught it. Plus NVIDIA's 25% cut to climate risk-estimation error, a 1,250-paper survey on AI self-improvement, and a new agentic research OS out of Fudan.

FeatureJuly 13, 2026

Math’s Most-Watched AI Tracker Just Went Dark. Nobody Said Why.

On June 30, the GitHub wiki where Terence Tao logged nine months of AI contributions to Erdős’s open problems got one final commit: "freeze." 10 days later a rival claimed 64 parallel agents solved a 50-year-old conjecture nobody has verified. Where the real edge in parallel AI mathematics actually is, in numbers.

FeatureJuly 13, 2026

Sakana Just Found the Point Where Adding More AI Research Agents Stops Helping

The number is 100. Scale a discovery agent from 10 copies to 100 and it matches the human novelty baseline exactly. Push it to 1,000 and the score stalls while 10-20% of the output turns to noise, right at the scale the field is racing toward.

Field notesJuly 12, 2026

The First Peer-Reviewed AI Co-Scientist Gets Nearly Half Wrong. 10,000 Labs Use It Anyway.

Stanford's Biomni cleared peer review in Science on July 9 and is already running in 10,000+ labs. Its own benchmark: 57% accuracy across 443 questions. What that gap between adoption and accuracy means for reproducible science.

Field notesJuly 11, 2026

177x Faster, 12x Bigger, Same Model: What NVIDIA’s Science-Compute Week Actually Fixed

NVIDIA shipped 2 posts a day apart that make protein-complex alignment 177x faster, push the largest foldable complex 12x bigger, and cut molecular-dynamics time 46%, all without a new model. Plus the confirmed Tc numbers behind an ML-screened kagome superconductor.

Field notesJuly 10, 2026

160 Ideas, One Pipeline, Two AI Judges, Two Different Scores

A NASA-grounded research agent produced 160 hypotheses from 1,475 satellite datasets. Grading them with GPT-5.2 versus Claude Sonnet 4.6 left the rankings stable but moved the scores. Plus Microsoft’s Aurora 1.5 weather model and Flint, a deterministic chart language for agents.

FeatureJuly 7, 2026

Everyone’s Racing to Build a Smarter Model. The Data Says That Isn’t the Bottleneck.

Frontier agents beat published Nature-family results on 17.8% of tasks. Give the same model pre-built domain skills and completion jumps from 57.1% to 100%. What that gap says about where the AI-for-science race is actually won.

Field notesJuly 7, 2026

Anthropic Could Have Shipped a Bigger Model. It Shipped 60 Skills Instead.

Claude Science launched with 60+ curated skills instead of a bigger model. NatureBench, NVIDIA BioNeMo, and Microsoft Talos explain why that’s the smarter bet, and what shipped elsewhere this week.

Feature2026-07-05

Introducing the MegaBrain BioScience blog

We’re launching a blog about AI and science — the work happening at MegaBrain and elsewhere, our collaborations with labs, and practical workflows for scientists.

Field notes2026-07-05

Coding agents in the social sciences

For the first time, core research tasks can be handed off to machines. What that means for empirical social science — and who is actually using agents so far.