Physics & materials: learned surrogates
Foundation models are replacing expensive first-principles simulations and inverse-designing new matter.
Chapter 5 of the AI Index 2026 traces AI across the sciences — from physics, chemistry, and astronomy to biology and Earth science, and into autonomous research agents. In 2025 AI moved beyond improving single pipeline steps toward replacing entire workflows, yet rigorous benchmarks still expose large gaps between plausible output and reliable science. The numbers:
AI's role in science falls into three coexisting layers — predictive models over data, assistants that help scientists work, and autonomous systems that generate discoveries. In 2025 the most visible action shifted to the second and third.
In the Web of Science database, AI-related publications in the natural sciences reached approximately 80,150 in 2025, up from 63,547 in 2024 — a one-year increase of roughly 26%. Physical sciences and life sciences followed similar trajectories, reaching about 33,000 and 29,000 publications respectively, each growing 27%–28% year over year, while Earth science — the smallest category at about 20,460 — grew roughly 23%.
Section 5.2 tracks the datasets, benchmarks, foundation models, and agents of three scientific groupings. A consistent finding: most scientific AI models come from academic and government institutions collaborating across countries — the opposite of the industry-dominated, general-purpose AI landscape.
Frontier models outperform human chemists on average but cannot reproduce published research. On ChemBench (2,700+ chemistry questions), the best models beat human-expert averages while still stumbling on basic tasks. On ReplicationBench, frontier models score below 20% on paper-scale replication in astrophysics. On UnivEarth, LLM agents answer Earth-observation questions with just 33% accuracy and their code fails 58% of the time, while on BixBench, frontier models reach only ~17% on real-world bioinformatics analysis.
How AI is reshaping each scientific domain in 2025. Tap any card for the full trend and its numbers.
Foundation models are replacing expensive first-principles simulations and inverse-designing new matter.
AION-1, AstroVisBench, and a 100TB dataset mark a field-wide shift toward AI infrastructure.
Frontier models top human chemists on ChemBench's 2,700+ questions — but stumble on full research.
111M-parameter MSAPairformer and 200M GPN-Star beat far larger models; OpenGenome2 holds 9.3T base pairs.
Aardvark Weather ran an end-to-end pipeline; FourCastNet 3 forecasts 60 days in under 4 minutes.
From silver to gold at the IMO in one year — but Erdős conjectures stay well beyond reach.
On PaperArena the best agent hits 38.8% vs. 83.5% for PhD experts; AstaBench's best scores ~0.53.
Headline findings from Chapter 5 · Science.
In molecular biology, smaller models are outperforming larger ones — an 111-million-parameter model beat methods built on tens of billions.
AI publications in the natural sciences reached about 80,150 in 2025, up roughly 26% in a single year.
Frontier models outperform human chemists on average — yet score below 20% on replicating published astrophysics papers.
Astronomy released its first foundation model in 2025: AION-1, trained on over 200 million celestial objects from five surveys.
On end-to-end research, the best AI agents score about half of PhD experts — 38.8% versus 83.5% on PaperArena.
Chapter 5 (sections 5.1–5.3) with every figure and citation is free from Stanford HAI. Or head back to the 15 takeaways and nine-chapter overview.
Open Chapter 5 · Science →