AI Index Report 2026

AI is racing to replace whole scientific workflows — reliability is still the gap

Chapter 5 of the AI Index 2026 traces AI across the sciences — from physics, chemistry, and astronomy to biology and Earth science, and into autonomous research agents. In 2025 AI moved beyond improving single pipeline steps toward replacing entire workflows, yet rigorous benchmarks still expose large gaps between plausible output and reliable science. The numbers:

80,150AI publications in the natural sciences in 2025
26% growth in AI-for-science publications, 2024→2025
20% ceiling for frontier models on astrophysics paper replication
39% best agent on PaperArena (PhD baseline 83.5%)
200million celestial objects training AION-1, astronomy's first foundation model
4minutes for FourCastNet 3 to run a 60-day global forecast

5.1 — AI for science: from single steps to whole workflows

AI's role in science falls into three coexisting layers — predictive models over data, assistants that help scientists work, and autonomous systems that generate discoveries. In 2025 the most visible action shifted to the second and third.

In the Web of Science database, AI-related publications in the natural sciences reached approximately 80,150 in 2025, up from 63,547 in 2024 — a one-year increase of roughly 26%. Physical sciences and life sciences followed similar trajectories, reaching about 33,000 and 29,000 publications respectively, each growing 27%–28% year over year, while Earth science — the smallest category at about 20,460 — grew roughly 23%.

AI is becoming routine across disciplines

  • As a share of total output, AI-related work remains a single-digit fraction of each field but is climbing fast. By 2025, Earth science had the highest AI penetration at 8.8%, followed by natural sciences overall at 6.8%, life sciences at 6.5%, and physical sciences at 5.8%. In 2010, all four categories sat below 1%.
  • The clearest breakthroughs cluster in domains with strong existing data infrastructure — structural biology, physics, chemistry, and materials science — rather than in fields built on the most sophisticated mathematical or physics-based models.
  • These developments do not automatically translate into scientific progress. Experimental validation remains expensive and slow: AI can propose novel candidate molecules at scale, but clinical trials to determine whether they work remain a costly, multiyear process. The gap between what AI can propose and what scientists can feasibly test recurs across every domain.

AI publications in the natural sciences, 2025

AI-related publications by field in 2025 (a single paper can fall in more than one domain, so the natural-sciences total is de-duplicated and not the simple sum). Unit: publications.

AI publications in the natural sciences, 2025Natural sciences (total): 8015080150Natural sciences (total)Physical sciences: 3305033050Physical sciencesLife sciences: 2891028910Life sciencesEarth science: 2046020460Earth science

5.2 — Across the domains: capable, but not yet reliable

Section 5.2 tracks the datasets, benchmarks, foundation models, and agents of three scientific groupings. A consistent finding: most scientific AI models come from academic and government institutions collaborating across countries — the opposite of the industry-dominated, general-purpose AI landscape.

Physics, astronomy, chemistry, materials — and a small-model surprise in biology

  • In molecular biology, smaller models are outperforming larger ones. MSAPairformer, a 111-million-parameter protein language model, surpassed previous leading methods on the ProteinGym benchmark; and GPN-Star, a 200-million-parameter genomics model, outperformed Evo 2, a model with 40 billion parameters.
  • Virtual-cell models emerged as a new 2025 frontier: Evo 2 (Arc Institute), STATE, and DeepMind's AlphaGenome aim to predict cellular responses to drugs and genetic perturbations without wet-lab experiments — though current systems still require experimental validation. Evo 2 trained on OpenGenome2, a corpus of 9.3 trillion DNA base pairs from across all domains of life.
  • Astronomy released its first foundation model, first visualization benchmark, and a 100TB training dataset in 2025. AION-1, trained on over 200 million celestial objects from five major surveys, is the field's first foundation model; AstroVisBench introduced the first benchmark for LLM scientific computing and visualization.
  • An AI system ran a full weather-forecasting pipeline end-to-end for the first time in 2025: Aardvark Weather replaced the traditional numerical pipeline with a single ML system, and multiple AI weather models reached operational deployment. FourCastNet 3 generates a 60-day global forecast in under 4 minutes — 8 to 60 times faster than prior approaches.

Strong on subtasks, shaky on full research

Frontier models outperform human chemists on average but cannot reproduce published research. On ChemBench (2,700+ chemistry questions), the best models beat human-expert averages while still stumbling on basic tasks. On ReplicationBench, frontier models score below 20% on paper-scale replication in astrophysics. On UnivEarth, LLM agents answer Earth-observation questions with just 33% accuracy and their code fails 58% of the time, while on BixBench, frontier models reach only ~17% on real-world bioinformatics analysis.

Frontier agents fall well short of expert-level science

Selected 2025 benchmark scores for frontier models and agents on full scientific tasks, versus the PhD-expert baseline on PaperArena. Unit: % accuracy.

Frontier agents fall well short of expert-level sciencePhD experts (PaperArena baseline): 8383PhD experts (PaperArena baseline)Best agent · PaperArena: 3939Best agent · PaperArenaLLM agents · UnivEarth (Earth obs.): 3333LLM agents · UnivEarth (Earth obs.)Frontier · ReplicationBench (astrophysics): 2020Frontier · ReplicationBench (astrophysics)Frontier · BixBench (bioinformatics): 1717Frontier · BixBench (bioinformatics)

AI's share of scientific output is climbing fast

AI-related work as a share of total publications by field in 2025 (all four sat below 1% in 2010). Unit: % of total output (rounded).

AI's share of scientific output is climbing fastEarth science: 99Earth scienceNatural sciences (overall): 77Natural sciences (overall)Life sciences: 66Life sciencesPhysical sciences: 66Physical sciences

Six fields, six frontiers

How AI is reshaping each scientific domain in 2025. Tap any card for the full trend and its numbers.

Physics & materials: learned surrogates

Foundation models are replacing expensive first-principles simulations and inverse-designing new matter.

physics

Astronomy: its first foundation model

AION-1, AstroVisBench, and a 100TB dataset mark a field-wide shift toward AI infrastructure.

astronomy

Chemistry: beats experts, fails replication

Frontier models top human chemists on ChemBench's 2,700+ questions — but stumble on full research.

chemistry

Life sciences: small models, big data

111M-parameter MSAPairformer and 200M GPN-Star beat far larger models; OpenGenome2 holds 9.3T base pairs.

biology

Earth science: weather AI goes operational

Aardvark Weather ran an end-to-end pipeline; FourCastNet 3 forecasts 60 days in under 4 minutes.

earth

Mathematics: gold medals, open problems intact

From silver to gold at the IMO in one year — but Erdős conjectures stay well beyond reach.

math

5.3 — Research agents: half of expert level

On PaperArena the best agent hits 38.8% vs. 83.5% for PhD experts; AstaBench's best scores ~0.53.

agents

The chapter in five lines

Headline findings from Chapter 5 · Science.

In molecular biology, smaller models are outperforming larger ones — an 111-million-parameter model beat methods built on tens of billions.
— Chapter 5 · Science
AI publications in the natural sciences reached about 80,150 in 2025, up roughly 26% in a single year.
— Chapter 5 · Science
Frontier models outperform human chemists on average — yet score below 20% on replicating published astrophysics papers.
— Chapter 5 · Science
Astronomy released its first foundation model in 2025: AION-1, trained on over 200 million celestial objects from five surveys.
— Chapter 5 · Science
On end-to-end research, the best AI agents score about half of PhD experts — 38.8% versus 83.5% on PaperArena.
— Chapter 5 · Science

Read the full Science chapter

Chapter 5 (sections 5.1–5.3) with every figure and citation is free from Stanford HAI. Or head back to the 15 takeaways and nine-chapter overview.

Open Chapter 5 · Science →