AI Index Report 2024

2023 was the year foundation models doubled — and the frontier priced universities out

Chapter 1 of the AI Index 2024 traces where AI knowledge is made: papers, patents, notable models, foundation models, conferences and open-source code. Publication and patent data still run a year behind, so those sections describe 2022; the model data is current through 2023. The through-line is the same everywhere — output keeps rising, and the frontier keeps drifting toward whoever can pay for it. The numbers:

149foundation models released in 2023 (72 in 2022)
65.7% of 2023 foundation models released open-source (44.4% in 2022)
61notable ML models from US institutions in 2023 (EU 21, China 15)
191million dollars of compute to train Gemini Ultra (GPT-4: 78)
62.7% rise in granted AI patents worldwide from 2021 to 2022
12.2million GitHub stars for AI projects in 2023 (4.0 million in 2022)

1.1 — AI publishing nearly tripled in a decade, then grew 1.1% in a year

Between 2010 and 2022 the total number of AI publications nearly tripled, from roughly 88,000 to more than 240,000. But the increase over the last year of available data was a modest 1.1% — while conference papers alone jumped 30.2%.

The headline number hides two very different curves. Journals carry the overwhelming bulk of AI research — roughly 232,700 journal articles in 2022 against about 41,200 conference papers — but conferences are the faster-moving half. Since 2015 conference publications have grown 2.6 times and journal publications 2.4 times, and in the last year conference output rose 30.2% against 4.5% for journals. Conference papers climbed from 22,727 in 2020 to 31,629 in 2021 and 41,174 in 2022, more than doubling since 2010.

What the papers are about

  • Machine learning dominates the field mix, with roughly 72,200 publications in 2022 — an almost sevenfold increase since 2015.
  • The next largest fields are far behind: computer vision (21,309 publications), pattern recognition (19,841) and process management (12,052).
  • Below those sit computer networks, control theory, algorithms, linguistics and mathematical optimization, each between roughly 6,800 and 10,400 publications.
  • This edition draws its publication data from CSET, whose methodology changed since the AI Index last used it, so the totals differ slightly from earlier editions.

Who writes them

  • The academic sector produced 81.1% of AI publications in 2022, keeping the position it has held across every region for the past decade.
  • Industry accounted for 7.9%, government 7.0%, nonprofits 2.6% and everything else 1.5%.
  • Industry involvement is highest in the United States, where 14.1% of AI publications come from companies, ahead of the European Union plus the United Kingdom (9.5%) and China (7.4%).
  • China is the most academic of the three: 81.8% of its AI publications come from the education sector, against 75.5% in the United States and 75.6% in the EU plus the UK.
  • Government-affiliated research runs the other way — 10.1% of Chinese AI publications and 9.3% of EU-plus-UK publications, against 5.6% in the United States.

1.2 — AI patents are exploding, and three in five of them are Chinese

Granted AI patents worldwide rose 62.7% between 2021 and 2022 alone, reaching 62,264, and have grown more than 31-fold since 2010. China accounts for 61.1% of them; the US share has fallen from 54.1% in 2010 to 20.9%.

Growth, and a rising rejection rate

  • Across the whole 2010–14 period, granted AI patents grew 56.1%. Between 2021 and 2022 alone they grew 62.7% — the entire early decade of growth, compressed into one year.
  • Getting a patent granted has become much harder. In 2015, 42.2% of all filed AI patents were not granted; by 2022 that figure had risen to 67.4%.
  • In 2022 there were 128,952 ungranted AI patents against 62,264 granted — more than double.
  • The gap shows up in every major filing region. China recorded roughly 80,500 ungranted filings against 35,300 granted, the United States about 15,100 against 12,100, and the EU plus the UK about 2,170 against 1,170.

Where the patents come from

  • As of 2022, 75.2% of the world’s granted AI patents originated in East Asia and the Pacific, with North America next at 21.2% and Europe and Central Asia at 2.3%. Until 2011 North America led.
  • By geographic area: China 61.1%, the United States 20.9%, the EU plus the UK 2.0%, India 0.2%, and 15.7% for the rest of the world.
  • Per capita the ranking inverts. In 2022 South Korea led with 10.26 granted AI patents per 100,000 inhabitants, followed by Luxembourg (8.73) and the United States (4.23), with Japan (2.53) and China (2.51) close behind.
  • Measured by growth from 2012 to 2022, Singapore moved fastest at +5,366%, then South Korea (+3,801%) and China (+3,569%). The United States grew 1,299% over the same decade.

Three in five granted AI patents in the world are Chinese

Granted AI patents as a share of the world total, 2022 (%). The US share peaked at 54.1% in 2010 and has fallen every year since.

Three in five granted AI patents in the world are ChineseChina: 61.1361.13ChinaUnited States: 20.920.9United StatesRest of world: 15.7115.71Rest of worldEU & UK: 2.032.03EU & UKIndia: 0.230.23India

1.3 — Industry builds the frontier now, and 149 foundation models arrived in a single year

Two counts tell the story. Epoch AI recorded 51 notable machine learning models from industry in 2023 against 15 from academia. Stanford’s Ecosystem Graphs recorded 149 foundation models — more than double the 72 of 2022, and nearly 38 times the count of 2019.

Notable models: industry leads, collaboration hits a high

  • Academia led model releases until 2014. In 2023 industry produced 51 notable machine learning models and academia 15 — but 21 came out of industry-academia collaborations, a new high, so the gap narrowed slightly on last year.
  • Creating a cutting-edge model now demands data, compute and money on a scale universities do not have, which is why the split looks the way it does.
  • By country, the United States produced 61 notable models in 2023, ahead of China (15), France (8), Germany (5), and Canada, Israel and the United Kingdom (4 each).
  • Taken together, the European Union and the United Kingdom produced 25 — the first time since 2019 that they have collectively passed China.
  • Compute explains the corporate tilt. AlexNet needed an estimated 470 petaFLOP to train in 2012, the original Transformer around 7,400 in 2017, and Gemini Ultra 50 billion in 2023.

Foundation models: more of them, and more of them open

  • Of the 149 foundation models released in 2023, 98 were open, 23 limited access and 28 no access. In share terms, 18.8% came with no access at all and 15.4% with limited access.
  • The open share has risen fast: 33.3% in 2021, 44.4% in 2022, 65.7% in 2023.
  • Industry produced 72.5% of 2023’s foundation models — 108 of them — against 28 from academia, 9 from industry-academia collaborations and 4 from government.
  • Google released the most in 2023 (18), followed by Meta (11), Microsoft (9) and OpenAI (7). UC Berkeley, with three, was the most prolific academic institution.
  • Cumulatively since 2019 the leaders are Google (40), OpenAI (20), Meta (19), Microsoft (18) and DeepMind (15). Tsinghua University (7) is the top non-Western institution and Stanford (5) the leading American academic one.
  • By country in 2023: the United States 109, China 20, the United Kingdom 8, the United Arab Emirates 4. Cumulatively since 2019 the totals are 182, 30 and 21 for the US, China and the UK.

The United States produced four times as many notable models as China

Number of notable machine learning models in 2023, attributed to the countries of the researchers’ affiliated institutions. A model with authors in several countries can be counted more than once.

The United States produced four times as many notable models as ChinaUnited States: 6161United StatesChina: 1515ChinaFrance: 88FranceGermany: 55GermanyCanada: 44Canada

From $930 to $191 million in six years

This edition is the first to work with Epoch AI on hard training-cost estimates, derived from training duration and the type, quantity and utilization of the hardware, priced at cloud compute rental rates and adjusted for inflation. Read down the list and it becomes obvious why universities have effectively dropped out of frontier model building.

  1. 2012 · AlexNet

    470 petaFLOP, and the GPU era begins

    AlexNet, one of the papers that popularized the now standard practice of using GPUs to train AI models, required an estimated 470 petaFLOP. Gemini Ultra, eleven years later, required more than 100 million times that.

  2. 2017 · Transformer

    $930 to train the architecture behind every modern LLM

    The original Transformer introduced the architecture that underpins virtually every modern LLM. It required around 7,400 petaFLOP and cost roughly $930 to train in inflation-adjusted terms — an amount a graduate student could put on a credit card.

  3. 2019 · RoBERTa Large

    $160,018, and state of the art on comprehension

    RoBERTa Large posted state-of-the-art results on canonical comprehension benchmarks such as SQuAD and GLUE. It cost around $160,000 to train — roughly 170 times the Transformer, two years later.

  4. 2020 · GPT-3 175B

    $4.3 million, and the first seven-figure model on the list

    GPT-3 175B (davinci) is estimated at $4,324,883 to train. It is the point on the curve where model training stops being a research budget line and starts being a capital decision.

  5. 2021 · Megatron-Turing NLG

    $6.4 million for 530 billion parameters

    Megatron-Turing NLG 530B is estimated at $6,405,653. Costs were still climbing gradually at this stage — the steep part of the curve was one year away.

  6. 2022 · PaLM (540B)

    $12.4 million — and a tenfold spread inside one year

    PaLM (540B) is estimated at $12,389,056, while LaMDA, released the same year, comes in at $1,319,586. By 2022 the answer to what a model costs to train depended almost entirely on which model you meant.

  7. 2023 · GPT-4 and Gemini Ultra

    $78 million and $191 million, in the same year

    AI Index estimates put GPT-4’s training compute at $78,352,034 and Gemini Ultra’s at $191,400,000; OpenAI CEO Sam Altman has said GPT-4’s training cost was over $100 million. Llama 2 70B, released the same year, is estimated at $3,931,897 — the frontier and the merely capable are now two different price brackets. The chapter notes that this escalation has effectively excluded universities from building leading-edge foundation models, and that policy responses such as the US National AI Research Resource exist specifically to hand nonindustry actors the compute they lack.

Running out of data — and why synthetic data is not a clean substitute

A dedicated highlight inside section 1.3. Epoch AI projected when each stock of training data gets exhausted, using both historical growth in training set sizes and a compute-adjusted method; separate 2023 studies tested what happens when models are fed their own output instead. Tap a card for the detail.

High-quality language data: about now

Epoch’s historical projection puts exhaustion at 2024.5, with a 90% confidence interval of 2023.5 to 2025.7.

dataprojections

Low-quality language data: the 2030s

Historical projection 2032.4; the compute-based projection pushes it out to 2040.5.

dataprojections

Image data: late 2030s to mid-2040s

Historical projection 2046; here it is the compute-based method that is the pessimistic one, at 2038.8.

dataprojections

Model collapse: the tails vanish

A team of British and Canadian researchers found that models trained predominantly on synthetic data lose the ability to remember the true data distribution.

syntheticresearch

Model Autophagy Disorder

A 2023 imaging study named the same failure MAD, after mad cow disease, and measured it three ways.

syntheticresearch

1.4–1.5 — Conferences filled back up, and open source went vertical

Attendance at the AI conferences the Index tracks rose 6.7% to roughly 63,300 in 2023, recovering after the return to in-person formats. On GitHub the change was of a different order: AI projects rose 59.3% in a single year, and the stars they collected more than tripled.

Conferences

  • Total attendance across the tracked conferences reached roughly 63,300 in 2023 — a 6.7% rise on the year, and about 50,000 more attendees than in 2015. Part of that growth is new conferences rather than bigger ones.
  • NeurIPS remains the largest, drawing approximately 16,380 participants in 2023, followed by CVPR (about 8,340), ICML (7,920), ICCV (7,330) and ICRA (6,600).
  • The direction was not uniform: NeurIPS, ICML, ICCV and AAAI each grew year over year, while CVPR, ICRA, ICLR and IROS each slipped slightly.
  • Among the smaller venues, IJCAI drew about 1,990 attendees, AAMAS 970 and FAccT 830.
  • These figures deserve caution — several of the recent years ran virtual or hybrid formats, and organizers report that counting virtual attendance accurately is difficult.

Open-source AI software

  • GitHub AI projects grew from 845 in 2011 to approximately 1.8 million in 2023, including a sharp 59.3% rise in the last year alone.
  • US-based developers accounted for 22.9% of GitHub AI projects in 2023, with India second at 19.0% and the European Union plus the United Kingdom at 17.9%. China was at 3.0%, and the rest of the world at 37.1%.
  • The US share has been declining steadily since 2016 — not because American output fell, but because everyone else’s rose faster.
  • New stars awarded to AI projects more than tripled, from 4.0 million in 2022 to 12.2 million in 2023. The most starred repositories are libraries such as TensorFlow, OpenCV, Keras and PyTorch.
  • Cumulatively, US-based projects hold about 10.45 million stars, ahead of the EU plus the UK (4.53 million), China (2.12 million) and India (1.92 million). Every major region sampled gained stars year over year.

The chapter in five lines

Headline findings from Chapter 1 · Research and Development.

In 2023, industry produced 51 notable machine learning models, while academia contributed only 15. There were also 21 notable models resulting from industry-academia collaborations in 2023, a new high.
— Chapter 1 · Research and Development
In 2023, a total of 149 foundation models were released, more than double the amount released in 2022. Of these newly released models, 65.7% were open-source, compared to only 44.4% in 2022 and 33.3% in 2021.
— Chapter 1 · Research and Development
OpenAI’s GPT-4 used an estimated $78 million worth of compute to train, while Google’s Gemini Ultra cost $191 million for compute.
— Chapter 1 · Research and Development
In 2022, China led global AI patent origins with 61.1%, significantly outpacing the United States, which accounted for 20.9%. Since 2010, the U.S. share of AI patents has decreased from 54.1%.
— Chapter 1 · Research and Development
Since 2011, the number of AI-related projects on GitHub has seen a consistent increase, growing from 845 in 2011 to approximately 1.8 million in 2023, with a sharp 59.3% rise in the last year alone.
— Chapter 1 · Research and Development

Read Chapter 1 in full

Chapter 1 (sections 1.1–1.5) — publications, patents, frontier AI research, conferences and open-source software — with every figure, footnote and citation, is free from Stanford HAI.

Open the AI Index 2024 →