The year AI stopped being a lab curiosity in medicine
Chapter 5 of the AI Index 2025 — written with RAISE Health, a Stanford Medicine and HAI collaboration — tracks AI moving from benchmark scores into hospitals, protein databases, and the Nobel Prize citations. Clinical knowledge benchmarks are nearing saturation, FDA authorizations have gone vertical, and the ethics literature is racing to keep up. The numbers:
5.1 & 5.2 — Biology got a new instrument
2022 and 2023 were the early stages of AI-driven scientific breakthroughs. 2024 was the year the tools got big enough, and open enough, to change how protein science is actually done.
The headline release was EvolutionaryScale's ESM3, trained on 2.78 billion protein sequences with 98 billion parameters. Its most striking result was esmGFP, a novel artificial green fluorescent protein that the company estimates would have taken nature roughly 500 million years of evolution to produce — generated through human-led chain-of-thought prompting. Larger ESM3 models solved twice as many atomic-coordination tasks as smaller ones, and the model was open-sourced.
Models kept getting bigger
The trend line is unambiguous. ProGen (2020) had 1.2 billion parameters and ProtBert 0.42 billion; ProGen2 and ProtT5 arrived in 2022 at 6.4 billion and 1.2 billion; ESM2 reached 15 billion in 2023; ESM3 hit 98 billion in 2024. Each jump in scale has come with better protein prediction accuracy — ESM C, released in 2024, posted improved structure-prediction results in the CASP15 challenge.
Four results worth remembering
- AlphaProteo (Google DeepMind) designed the first protein binders for several targets including VEGF-A, a protein linked to cancer and diabetes. Its binders hold together roughly 10 times more strongly than existing state-of-the-art designs, with some estimated to be up to 300 times more effective; for the viral protein BHRF1, 88% of its designed binders bound successfully in the wet lab.
- AlphaFold 3 extended the series beyond protein structures to interactions with DNA, RNA, ligands, and antibodies. On protein-ligand docking it reached 80.5% of predictions under 2 Å RMSD — and 93.2% when the binding pocket was specified in advance, against 59.7% for the Vina baseline.
- A Stanford virtual AI lab — a principal-investigator model, a scientific critic, and three specialist agents in immunology, computational biology, and machine learning — autonomously designed 92 nanobodies targeting SARS-CoV-2, with over 90% successfully binding in validation.
- Google's Connectomics team reconstructed one cubic millimeter of human brain at synaptic resolution: over 5,000 slices at 30 nanometers each, capturing around 57,000 cells and 150 million synapses.
The databases underneath
None of this works without public data. Since 2021 the major protein science databases have grown sharply: UniProt by 31%, the Protein Data Bank by 23%, and the AlphaFold Database by 585%. The infrastructure is old — PDB dates to 1971 and was the first open-access digital resource in the biological sciences, Pfam to 1995, STRING to 2000, UniProt to 2002 — but AI-generated entries are now what drives the growth. Within biological-sciences publishing in 2024, function prediction was the most researched AI-driven protein topic at 8.4% of papers, followed by structure prediction (7.6%), protein-drug interactions (3.0%), and synthetic protein design (0.7%).
Microscopy is following the same curve. Light-based microscopy foundation models doubled from four to eight in 2024, and while no electron or fluorescence microscopy foundation models were released in 2023, four of each appeared in 2024.
5.6 — Foundation models arrive in the rest of science
Dozens of scientific foundation models shipped in 2024, some fine-tuned language models, some trained from scratch on weather or atomistic data. A year of notable releases, in order.
- Feb 6, 2024
CrystalLLM — materials science
Researchers fine-tuned LLaMA-2 70B on text-encoded atomistic data to generate stable materials, achieving nearly double the metastability rate of a leading diffusion model (49% vs. 28%) while keeping the structures physically plausible. The approach supports unconditional generation, structure infilling, and text-guided design.
- Feb 14, 2024
LlaSMol — chemistry
To fix LLMs' poor performance on chemistry, researchers built SMolInstruct, a dataset of over 3 million samples across 14 tasks, and fine-tuned a family of models on it. The Mistral-based LlaSMol beat GPT-4 and Claude 3 Opus by a wide margin while tuning just 0.58% of parameters.
- Apr 23, 2024
ORBIT — Earth science
Oak Ridge National Lab released ORBIT, a 113-billion-parameter vision transformer and the largest AI model ever built for climate science — 1,000 times larger than prior models. Trained with a novel parallelism technique and tested on the Frontier supercomputer, it sustained up to 1.6 exaFLOPS.
- May 20, 2024
Aurora — Earth system forecasting
Aurora is a large-scale foundation model trained on over a million hours of Earth system data, delivering state-of-the-art forecasts for air quality, ocean waves, cyclone tracks, and high-resolution weather. It outperforms traditional systems at a fraction of the computational cost and can be fine-tuned across domains with minimal resources.
- Jul 22, 2024
NeuralGCM — weather and climate
A hybrid model combining a differentiable, physics-based solver with machine learning components. It matches or exceeds leading ML and physics-based models on short- and medium-term forecasts, tracks climate metrics accurately over decades, and captures phenomena such as tropical cyclones — with massive computational savings.
- Aug 18, 2024
PhysBERT — physics
Physics texts are notoriously hard for NLP because of specialized language and complex concepts. PhysBERT, the first physics-specific text-embedding model, was trained on 1.2 million arXiv papers and fine-tuned with supervised data, outperforming general-purpose models on physics tasks such as information retrieval.
- Sep 16, 2024
FireSat — wildfire detection
Google's satellite-based wildfire detection system uses AI to identify fires as small as 5 by 5 meters within 20 minutes of ignition, by analyzing real-time imagery and environmental data. Developed with Earth Fire Alliance and Muon Space, it improves disaster response and advances global wildfire research.
- Oct 2024
Two Nobel Prizes for AI-driven research
Google DeepMind's Demis Hassabis and John Jumper won the Nobel Prize in Chemistry for their pioneering work on protein folding with AlphaFold. John Hopfield and Geoffrey Hinton received the Nobel Prize in Physics for their foundational contributions to neural networks. It was the first year AI-related breakthroughs took top honors in two separate sciences.
- Dec 4, 2024
GenCast — 15-day weather forecasts
Google DeepMind's diffusion-based weather model delivers highly accurate 15-day forecasts, outperforming traditional systems such as the ENS on nearly all metrics, and generating forecasts in minutes rather than hours. Applications span disaster response, renewable energy, and agriculture.
- Dec 9, 2024
AlphaQubit and Willow — quantum computing
Google DeepMind and Google Quantum AI released AlphaQubit, an AI-based decoder with state-of-the-art quantum error detection. Soon after came Willow, the first quantum chip to achieve exponential error suppression below the surface code threshold. Willow completed a benchmark task in under five minutes that would take the fastest supercomputer over 10 septillion years — longer than the age of the known universe.
5.3–5.5 — Better than doctors, worse at teamwork
On clinical knowledge benchmarks AI has essentially caught up. On the harder questions — whether handing a doctor an LLM makes the doctor better, and whether the data and ethics infrastructure can keep up — the 2024 evidence is uncomfortable.
MedQA, introduced in 2020, draws over 60,000 clinical questions from professional medical board exams. OpenAI's o1 set a new state of the art at 96.0%, a 5.8 percentage point gain over the best score posted in 2023 and 28.4 points above where the benchmark stood in late 2022. Like other general-knowledge benchmarks, MedQA now looks close to saturation — the report treats that as a signal to build harder evaluations, not as proof of clinical competence.
Two randomized trials, one awkward finding
- In a 2024 single-blind randomized trial with 50 US-licensed physicians working through complex clinical vignettes, doctors given GPT-4 alongside conventional resources scored 76% — barely above the 74% of doctors using conventional resources alone. GPT-4 by itself scored 92%, a 16-percentage-point margin over unassisted physicians, and there was no time saving in either direction.
- A 2024–25 randomized controlled trial with 92 physicians looked at management reasoning rather than diagnosis: treatment decisions, risk-benefit trade-offs, patient preferences. Here GPT-4-assisted physicians did beat the control group, by about 6.5 percentage points — but GPT-4 alone still performed on par with the assisted physicians, and the assisted doctors spent slightly longer per case.
- The pattern across both is the same: excellent standalone model performance does not automatically transfer to human-AI teams. Closing that gap is a workflow, training, and interface problem, not a model-capability one.
Where AI is actually deployed: paperwork
The clearest real-world win of the year was ambient AI scribes — systems that listen to the physician-patient encounter and draft the note. Kaiser Permanente Northern California launched one in late 2023 and thousands of clinicians adopted it before the pilot ended. A Stanford study of a fully integrated, automated scribe found 55% average uptake among physicians, roughly 30 seconds saved per note and about 20 minutes less EHR time per day, with self-reported burden down 35% and burnout down 26%. Investment in ambient scribe technology reportedly reached almost $300 million in 2024.
Stanford Health Care also fully implemented an AI peripheral arterial disease screening model under its FURM (Fair, Useful, Reliable, Measurable) framework. It is expected to reach roughly 1,400 patients a year and operates without external funding — one of only two of six evaluated use cases to make it all the way through.
The data bottleneck nobody solved
More than 80% of FDA-cleared machine learning software targets medical image analysis, yet the training sets are small. The Cancer Genome Atlas — one of the most comprehensive public histopathology collections — holds 11,125 patient samples across 32 cancer types, and histopathology models are often trained on fewer than 1,000 samples when genomic or proteomic labels are required. The geography is narrow too: most US cohorts used to train clinical machine learning algorithms came from California (22), Massachusetts (15), and New York (14), with most states contributing none.
The scale gap against general-purpose AI is stark. GatorTron, a large clinical LLM built to extract patient information from unstructured records, was trained on 82 billion tokens; Llama 3 was trained on 15 trillion — nearly 182 times more. On the imaging side, RadImageNet contains 16 million image-equivalent tokens against roughly 6 billion for DALL-E, about 375 times more. MIMIC-CXR (377,000 images) and CheXpert Plus (around 226,000 radiographs) are important resources but still small beside ImageNet's roughly 14 million images.
5.5 — And then the ethics field got funded
The AI Index meta-reviewed thousands of medical ethics studies for this chapter. NIH grants for medical AI ethics projects went from 25 in fiscal 2023 to 337 in fiscal 2024 — after just 2, 3, and 7 grants in the three preceding years. Total funding followed the same shape, jumping from $16.3 million in 2023 to $276 million in 2024, an almost 17-fold increase in a single year. Whatever else 2024 was, it was the year the ethics of medical AI stopped being an unfunded side conversation.
What the literature is actually about
- Bias and privacy are the two dominant concerns in 2024, followed by equity, transparency, and trust. The ordering has flipped since 2020, when privacy was the more prominent topic and bias trailed it.
- Tool coverage is lopsided. OpenAI's GPT series was discussed in 86 medical AI ethics publications in 2024, up from 42 in 2023 — an order of magnitude more attention than Google, Meta, Anthropic, Mistral, Cohere, or xAI models combined received.
- Evaluation interest exploded alongside it. A PubMed search for 'large language model' returns 1,566 papers since 2019, of which 1,210 were published in 2024 alone against 353 in 2023. An early-2024 systematic review of NLP performance on healthcare tasks found over 500 papers, concentrated on medical knowledge (419) and diagnosis (178).
Synthetic data as the privacy workaround
Because the real data is scarce and sensitive, 2024 research leaned hard on synthetic data. Studies validated it for privacy-preserving clinical risk prediction — modeling lung cancer risk in ever-smokers from the UK Biobank with ADSGAN, PATEGAN, and DPGAN, where the synthetic distributions closely matched the real ones. A Nature study used a generative image model to optimize drug formulation in silico, correctly predicting the percolation threshold of microcrystalline cellulose in oral tablets. And fine-tuned classifiers augmented with synthetic data outperformed ChatGPT-family models at extracting social determinants of health from clinical notes, while showing less sensitivity to race, ethnicity, and gender descriptors.
Deployment still runs through a handful of vendors. EHR adoption reached roughly 90% for any system and 80% for certified systems as of 2021, and a 2023 American Hospital Association survey found predictive-model use concentrated in Epic, Cerner, and Meditech networks — with Epic, Cerner, and CPSI hospitals mostly running vendor-developed models. Whether AI-enabled health IT reaches rural and underserved settings, which face broadband and infrastructure limits, remains an open question.
Six questions the chapter answers
The findings that are easy to misread, with the numbers attached.
Does an AI scoring 96% on a medical exam mean it can practice medicine?
If GPT-4 beats doctors, why not just use it?
What does 'AI-enabled medical device' actually cover?
Are medical foundation models real, or marketing?
How much of medicine's AI research is happening outside the US?
What is the single biggest unsolved problem?
The chapter in five lines
Headline findings from Chapter 5 · Science and Medicine.
The FDA authorized its first AI-enabled medical device in 1995. By 2015 only six had been approved; by 2023 the number had reached 223.
GPT-4 alone outperformed doctors — both with and without AI — in diagnosing complex clinical cases.
o1 set a new state of the art of 96.0% on MedQA — 28.4 percentage points above where the benchmark stood in late 2022, and close enough to saturation to need replacing.
Since 2021 the AlphaFold Database has grown 585%, UniProt 31%, and the Protein Data Bank 23%.
In 2024 AI-driven research took two Nobel Prizes: Chemistry for AlphaFold's protein folding, Physics for the foundations of neural networks.
Read the full Science and Medicine chapter
Chapter 5 (sections 5.1–5.6) with every figure and citation is free from Stanford HAI. Or head back to the report highlights and eight-chapter overview.
Open the AI Index Report 2025 →