The imaging data gap
Medical-imaging training data is still ~100× smaller than nonmedical AI's.
Chapter 6 of the AI Index 2026 traces AI across four layers of medicine — from molecular biology models, through the clinic, to patients and ethics. Capability is racing ahead; rigorous validation is not. The numbers:
AI models for biology span the central dogma — gene sequence → protein structure → therapeutic design. In 2025 the field pivoted from scaling size to efficiency and specialization.
AI-driven protein research grew about 71% between 2024 and 2025. Protein–drug interactions led, rising from 49.9% to 54.4% of papers, while protein-structure prediction's share fell from 28.7% to 23.9%. The headline trend, though, is that smaller, specialized models are now beating much larger ones.
Where molecular models meet patients. Tap any card for the full trend and its numbers.
Medical-imaging training data is still ~100× smaller than nonmedical AI's.
On management reasoning, o1-preview scored 86% vs. 34% for physicians using conventional resources.
Microsoft's Diagnostic Orchestrator + o3 hit 85.5% on hard cases vs. 20% for unaided physicians.
258 AI devices authorized in 2025; cumulative passed 1,000 in 2024 — yet only 2.4% are backed by RCTs.
The broadest-adopted clinical AI; Northwestern reported a 112% ROI and 11.3 more patients/month.
COMPOSER at UC San Diego cut sepsis mortality 17% across 6,217 admissions — an estimated 50 lives/year.
A review of 500+ clinical AI studies found nearly half used exam-style questions, not real patient data.
As AI reaches patients through clinics and consumer platforms, research on how they perceive it grew tenfold (9 → 102 papers, 2020–2025).
Google's AI-generated 'AI Overviews' now appear atop most health-related search results — on average 84%–92% of health queries triggered one across five query types. Symptom and common-health questions were the most likely to trigger an overview (92%), followed by treatment-related queries (90%) and condition-based ones (84%–88%).
A bibliometric analysis of PubMed Central (Jan 2021 – Dec 2025) tracked ethical disclosure in medical-AI papers.
Of all medical-AI publications in 2025, 43.4% discussed ethics topics — up from 37.1% in 2024 (medical-AI-and-ethics papers rose from 1,114 to 2,378). The volume is climbing, but the focus is uneven.
Ethical discussion skews toward governance. In 2025, only 14 publications discussed biosecurity, with even fewer directly addressing the ethical implications of misuse or dual use — a notable blind spot as biological models grow more capable.
Global health departs from the governance-dominated pattern: among 2025 global-health publications, 51.8% (100 of 193) also mentioned ethics, and societal concerns — equity, justice, and accessibility — ranked highest, surpassing both governance and algorithmic concerns.
Headline findings from Chapter 6 · Medicine.
In molecular biology, smaller models are outperforming larger ones.
Ambient AI documentation saw the broadest adoption of any clinical AI category in 2025 — one health system reported a 112% return on investment.
The FDA authorized 258 AI medical devices in 2025 — yet only 2.4% of devices with clinical studies were backed by randomized-trial data.
A multi-agent AI system scored 85.5% on complex published cases, versus 20% for unaided physicians.
Nearly half of 500+ reviewed clinical AI studies used exam-style questions rather than real patient data.
Chapter 6 (sections 6.1–6.4) with every figure and citation is free from Stanford HAI. Or head back to the 15 takeaways and nine-chapter overview.
Open Chapter 6 · Medicine →