AI Index Report 2025

The year AI stopped being a lab curiosity in medicine

Chapter 5 of the AI Index 2025 — written with RAISE Health, a Stanford Medicine and HAI collaboration — tracks AI moving from benchmark scores into hospitals, protein databases, and the Nobel Prize citations. Clinical knowledge benchmarks are nearing saturation, FDA authorizations have gone vertical, and the ethics literature is racing to keep up. The numbers:

223FDA-authorized AI-enabled medical devices in 2023 (6 in 2015)
96% — o1's new state of the art on the MedQA clinical benchmark
1,031medical AI ethics publications in 2024 (288 in 2020)
537clinical trials mentioning AI in 2024 (5 in 2014)
98billion parameters in ESM3, trained on 2.78 billion protein sequences
2Nobel Prizes awarded in 2024 for AI-driven breakthroughs

5.1 & 5.2 — Biology got a new instrument

2022 and 2023 were the early stages of AI-driven scientific breakthroughs. 2024 was the year the tools got big enough, and open enough, to change how protein science is actually done.

The headline release was EvolutionaryScale's ESM3, trained on 2.78 billion protein sequences with 98 billion parameters. Its most striking result was esmGFP, a novel artificial green fluorescent protein that the company estimates would have taken nature roughly 500 million years of evolution to produce — generated through human-led chain-of-thought prompting. Larger ESM3 models solved twice as many atomic-coordination tasks as smaller ones, and the model was open-sourced.

Models kept getting bigger

The trend line is unambiguous. ProGen (2020) had 1.2 billion parameters and ProtBert 0.42 billion; ProGen2 and ProtT5 arrived in 2022 at 6.4 billion and 1.2 billion; ESM2 reached 15 billion in 2023; ESM3 hit 98 billion in 2024. Each jump in scale has come with better protein prediction accuracy — ESM C, released in 2024, posted improved structure-prediction results in the CASP15 challenge.

Four results worth remembering

  • AlphaProteo (Google DeepMind) designed the first protein binders for several targets including VEGF-A, a protein linked to cancer and diabetes. Its binders hold together roughly 10 times more strongly than existing state-of-the-art designs, with some estimated to be up to 300 times more effective; for the viral protein BHRF1, 88% of its designed binders bound successfully in the wet lab.
  • AlphaFold 3 extended the series beyond protein structures to interactions with DNA, RNA, ligands, and antibodies. On protein-ligand docking it reached 80.5% of predictions under 2 Å RMSD — and 93.2% when the binding pocket was specified in advance, against 59.7% for the Vina baseline.
  • A Stanford virtual AI lab — a principal-investigator model, a scientific critic, and three specialist agents in immunology, computational biology, and machine learning — autonomously designed 92 nanobodies targeting SARS-CoV-2, with over 90% successfully binding in validation.
  • Google's Connectomics team reconstructed one cubic millimeter of human brain at synaptic resolution: over 5,000 slices at 30 nanometers each, capturing around 57,000 cells and 150 million synapses.

The databases underneath

None of this works without public data. Since 2021 the major protein science databases have grown sharply: UniProt by 31%, the Protein Data Bank by 23%, and the AlphaFold Database by 585%. The infrastructure is old — PDB dates to 1971 and was the first open-access digital resource in the biological sciences, Pfam to 1995, STRING to 2000, UniProt to 2002 — but AI-generated entries are now what drives the growth. Within biological-sciences publishing in 2024, function prediction was the most researched AI-driven protein topic at 8.4% of papers, followed by structure prediction (7.6%), protein-drug interactions (3.0%), and synthetic protein design (0.7%).

Microscopy is following the same curve. Light-based microscopy foundation models doubled from four to eight in 2024, and while no electron or fluorescence microscopy foundation models were released in 2023, four of each appeared in 2024.

5.6 — Foundation models arrive in the rest of science

Dozens of scientific foundation models shipped in 2024, some fine-tuned language models, some trained from scratch on weather or atomistic data. A year of notable releases, in order.

  1. Feb 6, 2024

    CrystalLLM — materials science

    Researchers fine-tuned LLaMA-2 70B on text-encoded atomistic data to generate stable materials, achieving nearly double the metastability rate of a leading diffusion model (49% vs. 28%) while keeping the structures physically plausible. The approach supports unconditional generation, structure infilling, and text-guided design.

  2. Feb 14, 2024

    LlaSMol — chemistry

    To fix LLMs' poor performance on chemistry, researchers built SMolInstruct, a dataset of over 3 million samples across 14 tasks, and fine-tuned a family of models on it. The Mistral-based LlaSMol beat GPT-4 and Claude 3 Opus by a wide margin while tuning just 0.58% of parameters.

  3. Apr 23, 2024

    ORBIT — Earth science

    Oak Ridge National Lab released ORBIT, a 113-billion-parameter vision transformer and the largest AI model ever built for climate science — 1,000 times larger than prior models. Trained with a novel parallelism technique and tested on the Frontier supercomputer, it sustained up to 1.6 exaFLOPS.

  4. May 20, 2024

    Aurora — Earth system forecasting

    Aurora is a large-scale foundation model trained on over a million hours of Earth system data, delivering state-of-the-art forecasts for air quality, ocean waves, cyclone tracks, and high-resolution weather. It outperforms traditional systems at a fraction of the computational cost and can be fine-tuned across domains with minimal resources.

  5. Jul 22, 2024

    NeuralGCM — weather and climate

    A hybrid model combining a differentiable, physics-based solver with machine learning components. It matches or exceeds leading ML and physics-based models on short- and medium-term forecasts, tracks climate metrics accurately over decades, and captures phenomena such as tropical cyclones — with massive computational savings.

  6. Aug 18, 2024

    PhysBERT — physics

    Physics texts are notoriously hard for NLP because of specialized language and complex concepts. PhysBERT, the first physics-specific text-embedding model, was trained on 1.2 million arXiv papers and fine-tuned with supervised data, outperforming general-purpose models on physics tasks such as information retrieval.

  7. Sep 16, 2024

    FireSat — wildfire detection

    Google's satellite-based wildfire detection system uses AI to identify fires as small as 5 by 5 meters within 20 minutes of ignition, by analyzing real-time imagery and environmental data. Developed with Earth Fire Alliance and Muon Space, it improves disaster response and advances global wildfire research.

  8. Oct 2024

    Two Nobel Prizes for AI-driven research

    Google DeepMind's Demis Hassabis and John Jumper won the Nobel Prize in Chemistry for their pioneering work on protein folding with AlphaFold. John Hopfield and Geoffrey Hinton received the Nobel Prize in Physics for their foundational contributions to neural networks. It was the first year AI-related breakthroughs took top honors in two separate sciences.

  9. Dec 4, 2024

    GenCast — 15-day weather forecasts

    Google DeepMind's diffusion-based weather model delivers highly accurate 15-day forecasts, outperforming traditional systems such as the ENS on nearly all metrics, and generating forecasts in minutes rather than hours. Applications span disaster response, renewable energy, and agriculture.

  10. Dec 9, 2024

    AlphaQubit and Willow — quantum computing

    Google DeepMind and Google Quantum AI released AlphaQubit, an AI-based decoder with state-of-the-art quantum error detection. Soon after came Willow, the first quantum chip to achieve exponential error suppression below the surface code threshold. Willow completed a benchmark task in under five minutes that would take the fastest supercomputer over 10 septillion years — longer than the age of the known universe.

FDA-authorized AI-enabled medical devices, by year

The FDA authorized its first AI-enabled medical device in 1995, and for two decades annual approvals stayed in the single digits. Then the curve went vertical: 6 in 2015, 223 in 2023.

FDA-authorized AI-enabled medical devices, by year2015: 6620152017: 262620172019: 808020192021: 12912920212022: 16016020222023: 2232232023

5.3–5.5 — Better than doctors, worse at teamwork

On clinical knowledge benchmarks AI has essentially caught up. On the harder questions — whether handing a doctor an LLM makes the doctor better, and whether the data and ethics infrastructure can keep up — the 2024 evidence is uncomfortable.

MedQA, introduced in 2020, draws over 60,000 clinical questions from professional medical board exams. OpenAI's o1 set a new state of the art at 96.0%, a 5.8 percentage point gain over the best score posted in 2023 and 28.4 points above where the benchmark stood in late 2022. Like other general-knowledge benchmarks, MedQA now looks close to saturation — the report treats that as a signal to build harder evaluations, not as proof of clinical competence.

Two randomized trials, one awkward finding

  • In a 2024 single-blind randomized trial with 50 US-licensed physicians working through complex clinical vignettes, doctors given GPT-4 alongside conventional resources scored 76% — barely above the 74% of doctors using conventional resources alone. GPT-4 by itself scored 92%, a 16-percentage-point margin over unassisted physicians, and there was no time saving in either direction.
  • A 2024–25 randomized controlled trial with 92 physicians looked at management reasoning rather than diagnosis: treatment decisions, risk-benefit trade-offs, patient preferences. Here GPT-4-assisted physicians did beat the control group, by about 6.5 percentage points — but GPT-4 alone still performed on par with the assisted physicians, and the assisted doctors spent slightly longer per case.
  • The pattern across both is the same: excellent standalone model performance does not automatically transfer to human-AI teams. Closing that gap is a workflow, training, and interface problem, not a model-capability one.

Where AI is actually deployed: paperwork

The clearest real-world win of the year was ambient AI scribes — systems that listen to the physician-patient encounter and draft the note. Kaiser Permanente Northern California launched one in late 2023 and thousands of clinicians adopted it before the pilot ended. A Stanford study of a fully integrated, automated scribe found 55% average uptake among physicians, roughly 30 seconds saved per note and about 20 minutes less EHR time per day, with self-reported burden down 35% and burnout down 26%. Investment in ambient scribe technology reportedly reached almost $300 million in 2024.

Stanford Health Care also fully implemented an AI peripheral arterial disease screening model under its FURM (Fair, Useful, Reliable, Measurable) framework. It is expected to reach roughly 1,400 patients a year and operates without external funding — one of only two of six evaluated use cases to make it all the way through.

The data bottleneck nobody solved

More than 80% of FDA-cleared machine learning software targets medical image analysis, yet the training sets are small. The Cancer Genome Atlas — one of the most comprehensive public histopathology collections — holds 11,125 patient samples across 32 cancer types, and histopathology models are often trained on fewer than 1,000 samples when genomic or proteomic labels are required. The geography is narrow too: most US cohorts used to train clinical machine learning algorithms came from California (22), Massachusetts (15), and New York (14), with most states contributing none.

The scale gap against general-purpose AI is stark. GatorTron, a large clinical LLM built to extract patient information from unstructured records, was trained on 82 billion tokens; Llama 3 was trained on 15 trillion — nearly 182 times more. On the imaging side, RadImageNet contains 16 million image-equivalent tokens against roughly 6 billion for DALL-E, about 375 times more. MIMIC-CXR (377,000 images) and CheXpert Plus (around 226,000 radiographs) are important resources but still small beside ImageNet's roughly 14 million images.

5.5 — And then the ethics field got funded

The AI Index meta-reviewed thousands of medical ethics studies for this chapter. NIH grants for medical AI ethics projects went from 25 in fiscal 2023 to 337 in fiscal 2024 — after just 2, 3, and 7 grants in the three preceding years. Total funding followed the same shape, jumping from $16.3 million in 2023 to $276 million in 2024, an almost 17-fold increase in a single year. Whatever else 2024 was, it was the year the ethics of medical AI stopped being an unfunded side conversation.

What the literature is actually about

  • Bias and privacy are the two dominant concerns in 2024, followed by equity, transparency, and trust. The ordering has flipped since 2020, when privacy was the more prominent topic and bias trailed it.
  • Tool coverage is lopsided. OpenAI's GPT series was discussed in 86 medical AI ethics publications in 2024, up from 42 in 2023 — an order of magnitude more attention than Google, Meta, Anthropic, Mistral, Cohere, or xAI models combined received.
  • Evaluation interest exploded alongside it. A PubMed search for 'large language model' returns 1,566 papers since 2019, of which 1,210 were published in 2024 alone against 353 in 2023. An early-2024 systematic review of NLP performance on healthcare tasks found over 500 papers, concentrated on medical knowledge (419) and diagnosis (178).

Synthetic data as the privacy workaround

Because the real data is scarce and sensitive, 2024 research leaned hard on synthetic data. Studies validated it for privacy-preserving clinical risk prediction — modeling lung cancer risk in ever-smokers from the UK Biobank with ADSGAN, PATEGAN, and DPGAN, where the synthetic distributions closely matched the real ones. A Nature study used a generative image model to optimize drug formulation in silico, correctly predicting the percolation threshold of microcrystalline cellulose in oral tablets. And fine-tuned classifiers augmented with synthetic data outperformed ChatGPT-family models at extracting social determinants of health from clinical notes, while showing less sensitivity to race, ethnicity, and gender descriptors.

Deployment still runs through a handful of vendors. EHR adoption reached roughly 90% for any system and 80% for certified systems as of 2021, and a 2023 American Hospital Association survey found predictive-model use concentrated in Epic, Cerner, and Meditech networks — with Epic, Cerner, and CPSI hospitals mostly running vendor-developed models. Whether AI-enabled health IT reaches rural and underserved settings, which face broadband and infrastructure limits, remains an open question.

Medical AI ethics publications, 2020–24

Attention to ethical issues in medical AI has risen every year for five years, quadrupling from 288 publications in 2020 to 1,031 in 2024. In 2024 bias and privacy were the most cited concerns, followed by equity — a reversal from 2020, when privacy outranked bias.

Medical AI ethics publications, 2020–242020: 28828820202021: 39739720212022: 52352320222023: 67467420232024: 103110312024

Six questions the chapter answers

The findings that are easy to misread, with the numbers attached.

Does an AI scoring 96% on a medical exam mean it can practice medicine?
No. MedQA is built from professional medical board exam questions, and o1's 96.0% is a genuine record — up 28.4 points since late 2022. But the report explicitly flags the benchmark as approaching saturation, meaning it can no longer discriminate between good and excellent models. Researchers at UC Santa Cruz, Edinburgh, and the NIH argue that a single QA benchmark misses the complexity of real clinical work, and tested five models across 19 medical datasets covering concept recognition, summarization, decision support, and medical calculations instead. Hallucinations and inconsistent multilingual performance persist.
If GPT-4 beats doctors, why not just use it?
Because beating doctors on vignettes is not the same as working inside a clinic. In the 2024 diagnostic trial GPT-4 alone scored 92% against 74% for unassisted physicians — but physicians handed the same model only reached 76%, and case completion times did not change. The bottleneck is integration: workflow design, user training, and interface, not raw capability. The 2024–25 management-reasoning trial found a real 6.5-point gain from assistance, but the assisted doctors were slower, which the researchers attributed to deeper reflection rather than friction.
What does 'AI-enabled medical device' actually cover?
Mostly imaging. More than 80% of FDA-cleared machine learning software targets the analysis of medical images, and the field is still dominated by two-dimensional data — chest X-rays, histopathology slides, fundus photography — where CNNs and transformers work well. Extension into 3D modalities like CT, MRI, and 3D histopathology is under way but limited by data: UK Biobank's roughly 100,000 MRI scans and TCIA's roughly 50,000 studies are among the largest public 3D sets, and no public 3D histopathology datasets exist at all.
Are medical foundation models real, or marketing?
They shipped, in volume, and they are specialized. 2024 brought general-purpose multimodal models like Med-Gemini alongside discipline-specific ones: EchoCLIP for echocardiology, VisionFM for ophthalmology, CheXagent and Merlin for radiology, and an unusually dense cluster in pathology — CHIEF, Prov-GigaPath, PathChat, TITAN, Virchow, and UNI all within the year. Whether they generalize outside their training institutions is the open question; domain generalization and cross-modal adaptability are listed as the standing challenges for every modeling approach in the chapter.
How much of medicine's AI research is happening outside the US?
A growing share. Clinical trials mentioning AI rose from 5 in 2014 to 537 in 2024, and in 2024 China led with 105 trials, ahead of the United States at 97 and Italy at 42. The trajectory since 2020 is steep across the board — 249 trials that year, 349 in 2021, 396 in 2022, 448 in 2023 — with the COVID-19 pandemic having accelerated adoption in triage, resource allocation, and outcome prediction before the field expanded into chronic disease management, procedural support, and medication safety.
What is the single biggest unsolved problem?
Data — its size, its diversity, and who it represents. Publicly available histopathology cohorts rarely exceed 10,000 patient samples; longitudinal imaging is badly underrepresented, with ADNI's roughly 2,000 participants over 15-plus years being an exemplar rather than a norm; and acquisition variability across instruments, staining techniques, and institutions introduces batch effects that small training sets amplify. The chapter's proposed remedies are privacy-preserving data sharing such as federated learning, synthetic data generation, and better annotation strategies — none of which is solved.

The chapter in five lines

Headline findings from Chapter 5 · Science and Medicine.

The FDA authorized its first AI-enabled medical device in 1995. By 2015 only six had been approved; by 2023 the number had reached 223.
— Chapter 5 · Science and Medicine
GPT-4 alone outperformed doctors — both with and without AI — in diagnosing complex clinical cases.
— Chapter 5 · Science and Medicine
o1 set a new state of the art of 96.0% on MedQA — 28.4 percentage points above where the benchmark stood in late 2022, and close enough to saturation to need replacing.
— Chapter 5 · Science and Medicine
Since 2021 the AlphaFold Database has grown 585%, UniProt 31%, and the Protein Data Bank 23%.
— Chapter 5 · Science and Medicine
In 2024 AI-driven research took two Nobel Prizes: Chemistry for AlphaFold's protein folding, Physics for the foundations of neural networks.
— Chapter 5 · Science and Medicine

Read the full Science and Medicine chapter

Chapter 5 (sections 5.1–5.6) with every figure and citation is free from Stanford HAI. Or head back to the report highlights and eight-chapter overview.

Open the AI Index Report 2025 →