High-quality language data: about now
Epoch’s historical projection puts exhaustion at 2024.5, with a 90% confidence interval of 2023.5 to 2025.7.
Chapter 1 of the AI Index 2024 traces where AI knowledge is made: papers, patents, notable models, foundation models, conferences and open-source code. Publication and patent data still run a year behind, so those sections describe 2022; the model data is current through 2023. The through-line is the same everywhere — output keeps rising, and the frontier keeps drifting toward whoever can pay for it. The numbers:
Between 2010 and 2022 the total number of AI publications nearly tripled, from roughly 88,000 to more than 240,000. But the increase over the last year of available data was a modest 1.1% — while conference papers alone jumped 30.2%.
The headline number hides two very different curves. Journals carry the overwhelming bulk of AI research — roughly 232,700 journal articles in 2022 against about 41,200 conference papers — but conferences are the faster-moving half. Since 2015 conference publications have grown 2.6 times and journal publications 2.4 times, and in the last year conference output rose 30.2% against 4.5% for journals. Conference papers climbed from 22,727 in 2020 to 31,629 in 2021 and 41,174 in 2022, more than doubling since 2010.
Granted AI patents worldwide rose 62.7% between 2021 and 2022 alone, reaching 62,264, and have grown more than 31-fold since 2010. China accounts for 61.1% of them; the US share has fallen from 54.1% in 2010 to 20.9%.
Two counts tell the story. Epoch AI recorded 51 notable machine learning models from industry in 2023 against 15 from academia. Stanford’s Ecosystem Graphs recorded 149 foundation models — more than double the 72 of 2022, and nearly 38 times the count of 2019.
This edition is the first to work with Epoch AI on hard training-cost estimates, derived from training duration and the type, quantity and utilization of the hardware, priced at cloud compute rental rates and adjusted for inflation. Read down the list and it becomes obvious why universities have effectively dropped out of frontier model building.
AlexNet, one of the papers that popularized the now standard practice of using GPUs to train AI models, required an estimated 470 petaFLOP. Gemini Ultra, eleven years later, required more than 100 million times that.
The original Transformer introduced the architecture that underpins virtually every modern LLM. It required around 7,400 petaFLOP and cost roughly $930 to train in inflation-adjusted terms — an amount a graduate student could put on a credit card.
RoBERTa Large posted state-of-the-art results on canonical comprehension benchmarks such as SQuAD and GLUE. It cost around $160,000 to train — roughly 170 times the Transformer, two years later.
GPT-3 175B (davinci) is estimated at $4,324,883 to train. It is the point on the curve where model training stops being a research budget line and starts being a capital decision.
Megatron-Turing NLG 530B is estimated at $6,405,653. Costs were still climbing gradually at this stage — the steep part of the curve was one year away.
PaLM (540B) is estimated at $12,389,056, while LaMDA, released the same year, comes in at $1,319,586. By 2022 the answer to what a model costs to train depended almost entirely on which model you meant.
AI Index estimates put GPT-4’s training compute at $78,352,034 and Gemini Ultra’s at $191,400,000; OpenAI CEO Sam Altman has said GPT-4’s training cost was over $100 million. Llama 2 70B, released the same year, is estimated at $3,931,897 — the frontier and the merely capable are now two different price brackets. The chapter notes that this escalation has effectively excluded universities from building leading-edge foundation models, and that policy responses such as the US National AI Research Resource exist specifically to hand nonindustry actors the compute they lack.
A dedicated highlight inside section 1.3. Epoch AI projected when each stock of training data gets exhausted, using both historical growth in training set sizes and a compute-adjusted method; separate 2023 studies tested what happens when models are fed their own output instead. Tap a card for the detail.
Epoch’s historical projection puts exhaustion at 2024.5, with a 90% confidence interval of 2023.5 to 2025.7.
Historical projection 2032.4; the compute-based projection pushes it out to 2040.5.
Historical projection 2046; here it is the compute-based method that is the pessimistic one, at 2038.8.
A team of British and Canadian researchers found that models trained predominantly on synthetic data lose the ability to remember the true data distribution.
A 2023 imaging study named the same failure MAD, after mad cow disease, and measured it three ways.
Attendance at the AI conferences the Index tracks rose 6.7% to roughly 63,300 in 2023, recovering after the return to in-person formats. On GitHub the change was of a different order: AI projects rose 59.3% in a single year, and the stars they collected more than tripled.
Headline findings from Chapter 1 · Research and Development.
In 2023, industry produced 51 notable machine learning models, while academia contributed only 15. There were also 21 notable models resulting from industry-academia collaborations in 2023, a new high.
In 2023, a total of 149 foundation models were released, more than double the amount released in 2022. Of these newly released models, 65.7% were open-source, compared to only 44.4% in 2022 and 33.3% in 2021.
OpenAI’s GPT-4 used an estimated $78 million worth of compute to train, while Google’s Gemini Ultra cost $191 million for compute.
In 2022, China led global AI patent origins with 61.1%, significantly outpacing the United States, which accounted for 20.9%. Since 2010, the U.S. share of AI patents has decreased from 54.1%.
Since 2011, the number of AI-related projects on GitHub has seen a consistent increase, growing from 845 in 2011 to approximately 1.8 million in 2023, with a sharp 59.3% rise in the last year alone.
Chapter 1 (sections 1.1–1.5) — publications, patents, frontier AI research, conferences and open-source software — with every figure, footnote and citation, is free from Stanford HAI.
Open the AI Index 2024 →