AI Index Report 2026

AI is rewriting medicine — but the evidence is still catching up

Chapter 6 of the AI Index 2026 traces AI across four layers of medicine — from molecular biology models, through the clinic, to patients and ethics. Capability is racing ahead; rigorous validation is not. The numbers:

258FDA AI medical devices authorized in 2025
1,357cumulative FDA AI/ML devices (passed 1,000 in 2024)
83% less note-writing time (Sharp HealthCare)
71% growth in protein-AI research, 2024→2025
536prospective imaging-AI trials in 2025 (417 in 2024)
92% of health searches triggering an AI Overview

6.1 — Molecular biology: smaller models are winning

AI models for biology span the central dogma — gene sequence → protein structure → therapeutic design. In 2025 the field pivoted from scaling size to efficiency and specialization.

AI-driven protein research grew about 71% between 2024 and 2025. Protein–drug interactions led, rising from 49.9% to 54.4% of papers, while protein-structure prediction's share fell from 28.7% to 23.9%. The headline trend, though, is that smaller, specialized models are now beating much larger ones.

Smaller beats bigger

  • MSAPairformer, a 111-million-parameter protein language model, surpassed previous state-of-the-art methods on the ProteinGym benchmark — at a fraction of the training and parameter budget. 2024's 98-billion-parameter ESM3 gave way to efficiency-first successors.
  • GPN-Star, a 200-million-parameter genomics model, outperformed Evo 2 (40 billion parameters) on multiple variant-effect prediction tasks — a model nearly 200× larger.
  • After AlphaFold 3, cofolding models converged on a similar parameter scale rather than continuing to grow. AlphaFold 3's FoldBench performance has yet to be significantly surpassed, even by larger models released since.

Virtual cells, and a data bottleneck

  • Virtual-cell models emerged as a 2025 frontier: Evo 2 (Arc Institute), STATE, and DeepMind's AlphaGenome aim to predict cellular responses to drugs and genetic perturbations without wet-lab experiments — though current systems still require experimental validation.
  • Development is now bottlenecked on data, not architecture. Training sets expanded from hundreds of thousands to tens of millions of entries via distilled AI-predicted structures and combined experimental sources; Meta's OMol25 alone holds 100M+ quantum-mechanical calculations.
  • Biomni, a general-purpose biomedical AI agent from Stanford, mapped 25 subfields with 150 specialized tools, 105 software packages, and 59 databases — pairing digital reasoning with physical lab validation.

Multimodal biomedical AI is exploding

Publications on multimodal foundation models for biomedical discovery, by year. Vision–language models (image + text) and vision–omics models (imaging + genomics) led the surge.

Multimodal biomedical AI is exploding2021: 2220212022: 161620222023: 17117120232024: 31431420242025: 4624622025

6.2 — Inside the clinic

Where molecular models meet patients. Tap any card for the full trend and its numbers.

The imaging data gap

Medical-imaging training data is still ~100× smaller than nonmedical AI's.

imaging

LLM clinical reasoning

On management reasoning, o1-preview scored 86% vs. 34% for physicians using conventional resources.

reasoning

AI agents enter the clinic

Microsoft's Diagnostic Orchestrator + o3 hit 85.5% on hard cases vs. 20% for unaided physicians.

agents

FDA device authorizations surge

258 AI devices authorized in 2025; cumulative passed 1,000 in 2024 — yet only 2.4% are backed by RCTs.

regulation

Ambient AI notes — the widest adoption

The broadest-adopted clinical AI; Northwestern reported a 112% ROI and 11.3 more patients/month.

deployment

AI sepsis alerts — deployment that saves lives

COMPOSER at UC San Diego cut sepsis mortality 17% across 6,217 admissions — an estimated 50 lives/year.

deployment

Evidence gaps and real risks

A review of 500+ clinical AI studies found nearly half used exam-style questions, not real patient data.

evidence

FDA-authorized AI devices, by specialty

Cumulative authorized AI/ML devices. Radiology dominates at 1,039 of 1,357 (76.6%), but cardiology, neurology, and others have accelerated since 2020 — AI is spreading beyond imaging.

FDA-authorized AI devices, by specialtyRadiology: 10391039RadiologyCardiovascular: 130130CardiovascularNeurology: 6161Neurology

Prospective imaging-AI trials are rising

Prospective trials validating medical-imaging AI grew 28.5% year over year — a sign the field is starting to test in the real world, not just on benchmarks.

Prospective imaging-AI trials are rising2024: 41741720242025: 5365362025

6.3 — What patients actually want

As AI reaches patients through clinics and consumer platforms, research on how they perceive it grew tenfold (9 → 102 papers, 2020–2025).

AI Overviews already top health searches

Google's AI-generated 'AI Overviews' now appear atop most health-related search results — on average 84%–92% of health queries triggered one across five query types. Symptom and common-health questions were the most likely to trigger an overview (92%), followed by treatment-related queries (90%) and condition-based ones (84%–88%).

Assistance, not autonomy

  • Patients endorse AI in assistive roles rather than autonomous decision-making, especially in high-stakes clinical contexts.
  • Preserving the human relationship is a consistent theme — patients name the potential loss of empathic care as a primary concern.
  • Provider endorsement is a key determinant of patient acceptance, and transparency and disclosure of AI use are prioritized across populations.

6.4 — The ethics conversation is growing, but lopsided

A bibliometric analysis of PubMed Central (Jan 2021 – Dec 2025) tracked ethical disclosure in medical-AI papers.

Of all medical-AI publications in 2025, 43.4% discussed ethics topics — up from 37.1% in 2024 (medical-AI-and-ethics papers rose from 1,114 to 2,378). The volume is climbing, but the focus is uneven.

Governance-heavy, biosecurity-light

Ethical discussion skews toward governance. In 2025, only 14 publications discussed biosecurity, with even fewer directly addressing the ethical implications of misuse or dual use — a notable blind spot as biological models grow more capable.

Global health is the exception

Global health departs from the governance-dominated pattern: among 2025 global-health publications, 51.8% (100 of 193) also mentioned ethics, and societal concerns — equity, justice, and accessibility — ranked highest, surpassing both governance and algorithmic concerns.

The chapter in five lines

Headline findings from Chapter 6 · Medicine.

In molecular biology, smaller models are outperforming larger ones.
— Chapter 6 · Medicine
Ambient AI documentation saw the broadest adoption of any clinical AI category in 2025 — one health system reported a 112% return on investment.
— Chapter 6 · Medicine
The FDA authorized 258 AI medical devices in 2025 — yet only 2.4% of devices with clinical studies were backed by randomized-trial data.
— Chapter 6 · Medicine
A multi-agent AI system scored 85.5% on complex published cases, versus 20% for unaided physicians.
— Chapter 6 · Medicine
Nearly half of 500+ reviewed clinical AI studies used exam-style questions rather than real patient data.
— Chapter 6 · Medicine

Read the full Medicine chapter

Chapter 6 (sections 6.1–6.4) with every figure and citation is free from Stanford HAI. Or head back to the 15 takeaways and nine-chapter overview.

Open Chapter 6 · Medicine →