← MegaBrain BioScience Blog
Science Signal #11 · Field notes · July 23, 2026 · 5 min read

An AI Citation Checker Hit 99% Recall Once It Could Search the Web. Its False Alarms Tripled to 43%. Here's Everything Else That Shipped.

53. That's how many NeurIPS 2025 papers GPTZero found with hallucinated citations after peer review missed them. A new benchmark tested the fix: give an AI citation verifier live web search, and its detection rate jumps from 78.1% to 99.0% on the identical base model — while its false-positive rate triples from 12.7% to 43.1% in the same upgrade. We go deep on why the benchmark's authors say false-positive rate, not detection rate, is what decides whether a verifier ships, in today's feature. Here's everything else that landed this week.

< 5°
orientation mismatch for an AI agent aligning a real synchrotron beamline
$40M
Google commitment in AI tools and credits to 17 DOE national labs
3x
acetate-selectivity gain from a wet-lab-synthesized, LLM-guided catalyst

A citation checker's recall vs. false-alarm trade-off, in one line

HALLMARK(arXiv 2607.18360) benchmarks four classes of AI citation verifier on 2,526 real BibTeX entries spanning 14 hallucination types. Its headline finding: adding tool access consistently buys recall and costs false-positive rate, and the paper argues false-positive rate is what determines whether a verifier is safe to actually deploy — not the recall number that usually leads the release note. Full breakdown, including the per-model table, in today's feature.

An agentic X-ray scientist goes from simulator to a real beamline

A paper we've been tracking since it first surfaced without published numbers is now out in Nature Machine Intelligence: An agentic artificially intelligent X-ray scientist (Chen et al., published July 21). An LLM-driven agent trained entirely on a virtual diffractometer transferred its procedure to a real instrument — SLAC's SSRL beamline BL17-2, aligning a Co₃Sn₂S₂ single crystal at 16.0 keV — and consistently achieved orientation mismatches under 5° across most runs, locating a reference reflection and building the orientation matrix without a human in the loop. The authors are careful about the limits of the result: they describe it explicitly as a "low degrees of freedom" problem with a "well-defined goal," and flag that demands on simulator fidelity climb sharply as a task's degrees of freedom grow.

Google commits $40M in AI tools to 17 DOE national labs

Announced at the DOE Genesis Mission Summit, Google is giving all 17 Department of Energy national laboratories accelerated access to AlphaEvolve, AlphaFold 3, AlphaGenome, WeatherNext, and AlphaEarth Foundations, plus a year of Gemini for Government seats and tokens for tens of thousands of users — a $40 million commitment in AI tokens and cloud credits. The concrete example in Google's own post: at the National Laboratory of the Rockies, materials data scientist Steven Spurgeon used the tools to cut autonomous electron-microscope calibration time from over 90 minutes to about 13 minutes, and the manual steps to focus an image from as many as 50 down to two.

An LLM-guided catalyst gets a real wet-lab result, not just a benchmark score

Most AI-for-science papers this year report a benchmark number. A new one reports a synthesized material: this catalyst-discovery paper(arXiv 2607.08003, Choudhury et al.) constrains a frontier language model to reason strictly over explicit reaction-network graphs rather than free-form chemistry text, and applies it to CO2 electroreduction. The framework identified ketene desorption and hydroxide capture as the acetate-forming pathway, isolated three physical control levers — local alkalinity, controlled iron incorporation, and restricted interfacial proton-donor access — and used them to guide the prospective synthesis of a copper-iron oxide catalyst. Result, measured in the lab and not simulated: a threefold increase in acetate selectivity over matched copper-rich baseline catalysts.

What this means for reproducible, local-first science

Two of this week's three field-notes items share something the AI-for-science field still produces too rarely: a result checked against reality instead of a held-out test set. The X-ray agent had to work on an actual beamline with actual detector noise; the catalyst had to get synthesized and actually favor acetate over side products in a real reaction. Neither team's numbers came from a leaderboard. That's the same bar a reproducibility record exists to enforce at the scale of a single result — not whether an agent's output looks right, but whether it survives contact with the real instrument or the real reaction.

Try MegaBrain BioScience

A research workbench that runs on your machine and checks its own work.