An AI Memory System Topped a Leaderboard by Dumping 2.6 Million Characters a Query. Here's Everything Else That Shipped.
2,552,000. That's how many characters per query a knowledge-graph memory system retrieved to take the number one spot on a new scientific-memory leaderboard for AI research agents: over 410 times what its closest rival used. Cap every system to the same retrieval budget and that top score collapses to dead last. We go deep on the full budget-controlled test, and why every vendor claiming a "memory" edge should have to show their retrieval budget first, in today's feature. Here's everything else that landed this week.
A memory leaderboard's #1 spot turns out to be a context-volume artifact
Beyond Memory Leaderboards(arXiv 2607.16848) tests 8 memory and retrieval systems for AI research agents across two new benchmarks, PAIM and PTr. Uncapped, a knowledge-graph system retrieves 2.55–2.6 million characters per query to top the leaderboard. Cap the budget at 30,000 characters and the same system drops from a score of 8.04 to 5.27 — last place — while three architecturally ordinary retrieval systems converge within 0.03 points of each other. Full breakdown, including the per-system table, in today's feature.
A new formal-proof benchmark for quantum-theorem agents
Benchmarking Agents for Proving Theorems in Quantum Algorithms and Quantum Information (arXiv 2607.21533, submitted July 23) introduces two new Lean 4 formal-proof benchmarks — 36 quantum-algorithm theorem-completion tasks and 40 quantum-information tasks — and tests four frontier models against them: GPT-5.5, Kimi K3, DeepSeek V4-Pro, and MiniMax M3. Difficulty-weighted scores land at 60.4/100 on the quantum-algorithm tasks and 59.6/100 on quantum information. The one intervention that moves the needle most: giving an agent a searchable lemma library to draw on during proof search lifts scores by up to 15.9 points across model-benchmark pairs. Machine-checked Lean 4 proofs are about as strict a ground truth as an automated-math claim can get, which makes this a cleaner read on real agent capability than a benchmark graded by another LLM.
A primary-sourced accuracy report on an AI literature-review pipeline
Elicit published a self-run evaluation of its systematic literature review capabilities on July 23, testing against 994 real Cochrane systematic reviews spanning 38,493 study records across 12 MeSH subject areas. Reported numbers: 95.0% search recall using only a review's title as the query, 96.89% sensitivity / 92.54% specificity on abstract screening, 99.5% paper-level recall on full-text screening, and 95.6% correct extraction on methods, participants, and interventions fields. These are vendor-reported rather than independently audited, but they're primary-sourced and specific enough to be worth tracking — and specific enough that an independent replication would be straightforward to run.
What this means for reproducible, local-first science
All three items this week share a pattern: a number that looks strong in isolation gets more useful, not less, once you can see the conditions it was measured under — a retrieval budget, a machine-checked proof system, an independently reproducible test corpus. The AI-for-science field has plenty of leaderboards. What it's short on is leaderboards that disclose the constraints behind the score.
Try MegaBrain BioScience
A research workbench that runs on your machine and checks its own work.