AI Index Report 2024

AI finished the tests we had — so 2023 went and built harder ones

Chapter 2 of the AI Index 2024 measures where AI stood at the end of 2023. On the old benchmarks the story is saturation: ImageNet, SuperGLUE and VQA improved so little that this edition drops them. On the new ones built in 2023 — coding, expert reasoning, agents, hallucination — the scores are low enough to leave years of headroom. The numbers:

90.04% top MMLU score in 2023 (Gemini Ultra), the first above the 89.8% human baseline
84.3% of competition-level MATH problems solved, against a 90% human baseline
59.4% best score on MMMU, the new expert-level multimodal reasoning test
41% best score on GPQA, where PhD experts in the domain reach 65%
4.8% of SWE-bench real GitHub issues solved by the best model, Claude 2
24.2% median performance advantage of closed models over open ones

2.1 — The measuring stick broke before the models did

AI cleared the human baseline on one benchmark after another, until the benchmarks stopped saying anything. The defining move of 2023 was not a single model but a wholesale replacement of the tests.

Where AI has passed us, and where it has not

The chapter tracks nine benchmarks against human baselines. AI crossed the human line on image classification in 2015, on basic-level reading comprehension in 2017, on visual reasoning in 2020 and on natural language inference in 2021. As of 2023 the categories where AI still fails to exceed human ability are the more complex cognitive ones: visual commonsense reasoning and advanced, competition-level mathematical problem-solving. Planning belongs on that list too.

The benchmarks are wearing out

  • Fifteen benchmarks that appeared in the 2023 AI Index are dropped from this edition, among them ImageNet, SuperGLUE, VQA, aNLI, SST-5, STL-10, VoxCeleb and the Kinetics family.
  • Their improvement since 2022 explains why. ImageNet moved 1.54%, Kvasir-SEG 1.90%, the Cityscapes Challenge 0.23%. For the other eleven, the report records no improvement at all.
  • In their place the 2024 edition tracks 18 benchmarks, 13 of which were introduced in 2023 — deliberately weighted toward coding, advanced reasoning and agentic behavior, areas underrepresented in previous editions.
  • The 2023 cohort: SWE-bench for coding, HEIM for image generation, MMMU and GPQA for general reasoning, MoCa for moral reasoning, BigToM for causal reasoning, PlanBench for planning, AgentBench and MLAgentBench for agents, HaluEval for hallucinations, EditVal for image editing, VisIT-Bench for image instruction-following, and the Chatbot Arena Leaderboard for human preference.

Multimodal stopped being a special case

AI systems used to be narrow — language models that read text well and handled images badly, or the reverse. In 2023 that split closed: Google’s Gemini and OpenAI’s GPT-4 handle images and text together, and in some cases audio as well. The steering committee’s pick of the year’s notable releases reads like a map of that shift — Claude and GPT-4 on the same day in March, Stable Diffusion v2 and Segment Anything in the spring, Llama 2 in July, DALL-E 3 and SynthID in August, Mistral 7B in September, Ernie 4.0 in October, then a November crowd of GPT-4 Turbo with a 128K context window, Whisper v3 and Claude 2.1 with 200K, and finally Gemini on December 6 and Midjourney v6 on December 21.

The scoreboard at the end of 2023

Best reported score on each benchmark (%). MMLU is finished — 90.04% is past its 89.8% human baseline. MATH is close to its 90% baseline. The two benchmarks written in 2023 sit far lower, and SWE-bench barely registers.

The scoreboard at the end of 2023MMLU: 90.0490.04MMLUMATH: 84.384.3MATHMMMU: 59.459.4MMMUGPQA: 4141GPQASWE-bench: 4.84.8SWE-bench

2.2 — Fluent, preferred, and still making things up

Language is where AI looks most finished and is least trustworthy. The comprehension benchmarks are topping out, human voting has become a serious measure, and the factuality numbers are the ones worth staring at.

Understanding

  • Stanford’s HELM scores models across ten scenarios and reports a mean win rate. As of January 2024 GPT-4 leads the aggregate leaderboard at 0.96, ahead of GPT-4 Turbo at 0.83 and Palmyra X V3 (72B) at 0.82.
  • No single model owns every task, though. Yi (34B) tops NarrativeQA at 0.78, Llama 2 (70B) tops closed-book NaturalQuestions at 0.46, PaLM-2 (Bison) tops the open-book version at 0.81, and Palmyra X V3 (72B) tops WMT 2014 translation at 0.26.
  • MMLU covers 57 subjects across the humanities, STEM and the social sciences. Gemini Ultra holds the top score at 90.0% as of January 2024 — a 14.8 percentage-point improvement since 2022 and 57.6 points since MMLU was created in 2019, and the first score to pass the benchmark’s 89.8% human baseline.

Human preference became a metric

Launched in 2023, the Chatbot Arena Leaderboard lets anyone query two anonymous models and vote for the better answer. By early 2024 it had gathered over 200,000 votes, and users ranked OpenAI’s GPT-4 Turbo as the most preferred model, with an Elo rating of 1,252 for the best closed model against 1,149 for the best open one. The AI Index frames this as a shift in kind: with generative models producing high-quality text and images, benchmarking is moving away from computerized rankings such as ImageNet or SQuAD and toward human evaluation. Public feeling about AI is becoming part of how progress gets measured.

Factuality and hallucination

  • TruthfulQA asks roughly 800 questions across 38 categories, many written around misconceptions that humans get wrong too. GPT-4 (RLHF) posts the highest score so far at 0.6, nearly three times the score of a GPT-2-based model tested in 2021.
  • HaluEval, new in 2023, holds over 35,000 hallucinated and normal samples. Its headline finding: ChatGPT fabricates unverifiable information in roughly 19.5% of its responses, across topics from language to climate to technology.
  • Spotting a hallucination is its own hard problem. On HaluEval’s classification task ChatGPT reaches 62.59% on question answering and 58.53% on summarization; Claude 2 manages 69.78% on question answering but 57.75% on summarization; Llama 2 falls to 43.99% on dialogue and 20.46% on the general category.
  • The report is blunt about why this matters — LLMs are being deployed in law and medicine, and real hallucinations have already surfaced in court cases.

2.6 — The gap that did not close

Reasoning is where the 2023 numbers are honest. On expert multimodal questions, on graduate-level science, on planning and on visual commonsense, the best systems are still short of people — sometimes by a lot.

General reasoning: MMMU and GPQA

MMMU, built in 2023 by researchers in the United States and Canada, asks about 11,500 college-level questions across six disciplines, with charts, maps, tables and chemical structures in the question formats. As of January 2024 Gemini Ultra leads every subject category with an overall 59.4%, and the subject-level table shows how far that still is from a medium-level human expert: humanities and social sciences 78.3 against 85, health and medicine 67.3 against 78.8, business 59.3 against 86, science 54.7 against 84.7, art and design 51.4 against 84.2, technology and engineering 47.1 against 79.1. GPQA is harsher still — 448 multiple-choice questions written by subject-matter experts in biology, physics and chemistry, deliberately designed so that Google searching does not help. PhD-level experts score 65% inside their own domains and nonexperts around 34%. The best model, GPT-4, reaches 41.0% on the main set.

Mathematics

  • GSM8K holds roughly 8,000 grade-school word problems that require multistep arithmetic. A GPT-4 variant, GPT-4 Code Interpreter, scores 97% — a 4.4% improvement on the previous year’s state of the art and 30.4% above 2022, when the benchmark was first introduced.
  • MATH is the harder cousin: 12,500 competition-level problems released by UC Berkeley researchers in 2021, on which systems solved just 6.9% at launch. In 2023 a GPT-4-based model solved 84.3%. Impressive, and still under the 90% human baseline — which is why competition-level mathematics stays on the list of things people do better.
  • HELM’s own sub-leaderboard puts GPT-4 Turbo (1106 preview) on top for MATH with a chain-of-thought equivalence score of 0.86, and GPT-4 (0613) on top for GSM8K at 0.93.

Planning, vision and moral judgment

  • PlanBench, from Arizona State University, tested GPT-4 and I-GPT-3 on 600 Blocksworld problems with one-shot learning. GPT-4 generated correct plans 34.3% of the time and cost-optimal plans 33%; I-GPT-3 managed 6.8% and 5.8%. Checking a plan is easier than making one — GPT-4 verified plans correctly 58.6% of the time, I-GPT-3 12%.
  • On Visual Commonsense Reasoning, where a system must both answer a question about an image and pick the right rationale, the top Q->AR score reached 81.60 in 2023 against a human baseline of 85. AI performance rose 7.93% between 2022 and 2023 without closing the gap.
  • MoCa, a Stanford dataset of human stories with moral elements, found no model perfectly matching human moral systems, but newer and larger models like GPT-4 and Claude aligning more closely than smaller ones such as GPT-3. GPT-4 showed the greatest agreement of all models surveyed.
  • BigToM tests theory of mind with 25 controls and 5,000 model-generated evaluations. GPT-4 came out on top, closely matching human accuracy on forward belief and backward belief and slightly surpassing humans on forward action — nearing, but not surpassing, human levels overall.

The class of 2023

Thirteen of the eighteen benchmarks tracked in this edition were introduced in 2023. Here are six of them, and the scores that show how much room they leave.

SWE-bench · coding

2,294 software engineering problems taken from real GitHub issues in popular Python repositories. The best model solves 4.8%.

coding

MMMU · general reasoning

About 11,500 college-level questions across six disciplines, with charts, maps and chemical structures. Gemini Ultra leads at 59.4%.

reasoning

GPQA · Google-proof questions

448 graduate-level questions in biology, physics and chemistry that cannot be answered by searching. GPT-4 reaches 41.0%.

reasoning

HaluEval · hallucination

Over 35,000 hallucinated and normal samples. ChatGPT fabricates unverifiable information in roughly 19.5% of its responses.

factuality

AgentBench · agent behavior

Eight interactive settings including web browsing, online shopping, household management and card games. GPT-4 scores 4.01 overall.

agents

HEIM · image generation

Twelve aspects of text-to-image models, rated by human evaluators. No single model excels on all of them.

image

In this edition, closed models win everything

Percentage score difference between the best closed and best open model, collected in early January 2024. Across ten benchmarks the median advantage is 24.2%, and the widest gap is off this chart entirely: AgentBench, at 317.7%.

In this edition, closed models win everythingHumanEval: 54.8254.82HumanEvalMATH: 39.5739.57MATHMMLU: 27.5427.54MMLUSWE-bench: 20.9120.91SWE-benchGSM8K: 3.973.97GSM8K

2.3–2.9 — Code, pixels, agents and bodies

Away from the language leaderboards, 2023 looks less like saturation and more like a field still finding its footing — with one exception, where the numbers went vertical.

Coding: solved, then unsolved again

HumanEval was introduced by OpenAI researchers in 2021 with 164 handwritten programming problems. A GPT-4 variant, AgentCoder, now leads it at 96.3% — an 11.2 percentage-point increase over the 2022 high, and 64.1 points of improvement since 2021. That is the vertical line. Then SWE-bench arrived in October 2023 with 2,294 problems drawn from real GitHub issues, requiring changes coordinated across a codebase rather than inside a single function, and the best model in the world managed 4.8%. Same year, same models, two orders of magnitude apart — which is roughly the distance between finishing an exercise and doing the job.

Images and video

  • On HEIM’s human-rated image-text alignment, DALL-E 2 (3.5B) leads at 0.94. On quality, aesthetics and originality the Stable Diffusion-based Dreamlike Photoreal v2.0 (1B) leads at 0.92, 0.87 and 0.98 — a one-billion-parameter model beating far larger ones on the aspects people actually look at.
  • VisIT-Bench measures whether a vision-language model can follow a written instruction about an image, across 592 instructions in about 70 categories such as plot analysis, art knowledge and location understanding. As of January 2024 GPT-4V leads with an Elo of 1,349, marginally above the benchmark’s human reference score of 1,338.
  • Meta’s Segment Anything outperforms leading segmentation methods such as RITM on 16 of 23 datasets, and was then used alongside human annotators to build SA-1B: over 1 billion segmentation masks across 11 million images. Better models make better data, which makes better models.
  • Video generation is measured on UCF101, an action-recognition dataset of 101 categories. The year’s top model, W.A.L.T-XL, posted an FVD16 score of 36 — more than halving the previous year’s state of the art, where lower is better.

Agents

  • Voyager, a GPT-4-based Minecraft agent from Nvidia, Caltech, UT Austin, Stanford and UW Madison, collects 3.3 times more unique items than prior systems, travels 2.3 times further, and reaches key tech-tree milestones 15.3 times faster. It matters because it keeps learning in an open-ended world, which AlphaZero-style systems never had to do.
  • MLAgentBench asks whether agents can run scientific experiments — acting as computer science research assistants. Performance splits violently by task: some agents clear 80% on ogbn-arxiv, improving a baseline paper classification model, while every agent tested scored 0% on BabyLM, training a small language model. GPT-4 was consistently the best of them.
  • On AgentBench’s eight environments, GPT-4’s overall 4.01 is the ceiling, and the team attributes the failures to limited long-term reasoning, decision-making and instruction-following — not to a missing tool.

Robots that can be talked to

The chapter’s most quietly consequential finding is that language modeling improved robotics. Google’s PaLM-E, scaled up to 562 billion parameters and trained on visual language alongside robotics data, beats earlier methods such as SayCan and PaLI on embodied visual question answering and planning, and detects its own failures far more reliably — 0.91 for PaLM-E-12B against CLIP-FT’s 0.65 and zero-shot PaLI’s 0.73, which is what closed-loop planning depends on. DeepMind’s RT-2 trains on tokenized robot trajectories plus visual-language data: on tasks involving objects it has never seen, an RT-2/PaLM-E variant hits an 80% success rate against MOO’s 53%, and beats the previous year’s RT-1 by 43 percentage points. Beyond doing more, these systems can ask questions — a real step toward robots that interact with the world rather than execute in it.

Five things the scores do not tell you

Section 2.11 turns the microscope on the models themselves, and 2.13 on what running them costs. Both are less flattering than the leaderboards.

Can these models actually plan?
Not reliably. PlanBench, proposed by a group at Arizona State University, uses problems from the automated planning community, including those from the International Planning Competition. On 600 problems in the Blocksworld domain — a hand stacking blocks, one at a time, onto the table or onto a clear block — GPT-4 generated correct plans 34.3% of the time and cost-optimal plans 33% of the time under one-shot learning. I-GPT-3 managed 6.8% and 5.8%. Verifying a plan is easier than producing one: GPT-4 got plan verification right 58.6% of the time, I-GPT-3 12%. Planning sits alongside competition-level mathematics and visual commonsense reasoning on the short list of things humans still do better.
Do models get better when they check their own work?
They get worse. Self-correction sounds like the obvious answer to hallucination and faulty reasoning, and intrinsic self-correction — where the model fixes itself without external guidance — is the appealing version. Researchers from DeepMind and the University of Illinois at Urbana-Champaign tested GPT-4 on three reasoning benchmarks and found performance declined across all of them. On GSM8K, standard prompting scored 95.50%, one round of self-correction 91.50%, two rounds 89.00%. On CommonSenseQA: 82.00%, then 79.50%, then 80.00%. On HotpotQA: 49.00%, then 49.00%, then 43.00%. Each round also costs more calls — three, then five, instead of one.
Are emergent abilities real?
Often they are an artifact of how you measure. Many papers have argued that LLMs unpredictably display new capabilities at larger scales, which raised the worry that bigger models might develop surprising and uncontrollable abilities. Stanford researchers argue the emergence is frequently a property of the benchmark rather than the model: with nonlinear or discontinuous metrics such as multiple-choice grading, emergent abilities look obvious, and with linear or continuous metrics they largely vanish. Across BIG-bench, they observed emergent abilities on only five of 39 benchmarks. That is a direct challenge to a prevailing belief in AI safety and alignment research.
Does a model stay the same after it ships?
No, and it can go backwards. Closed models such as GPT-4, Claude 2 and Gemini are updated by their developers in response to new data and user feedback, and there is little public research on what that does to them. A study from Stanford and Berkeley compared the March 2023 and June 2023 versions of GPT-3.5 and GPT-4 and found the June GPT-4 was 42 percentage points worse at generating code, 16 points worse at answering sensitive questions and 33 points worse on certain mathematical tasks. Its ability to follow instructions had also diminished, which the researchers suggest may explain the broader declines. Every leaderboard number in this chapter is a snapshot of a moving object.
What does training one of these cost the planet?
More than most developers will say. Training GPT-3 (175B) released a reported 502 tonnes of CO2 equivalent and consumed 1,287 MWh; Gopher (280B) 352 tonnes and 1,066 MWh; Meta’s Llama 2 70B about 291 tonnes on 400 MWh — nearly 291 times the emissions of one traveler on a round-trip flight from New York to San Francisco, and roughly 16 times an average American’s annual footprint. Smaller models cost far less: Starcoder (15.5B) 16.68 tonnes, Luminous Base (13B) 3.17. The bigger problem is silence. OpenAI, Google, Anthropic and Mistral do not report training emissions, though Meta does, and inference is barely reported at all — even though per-query emissions are small, total inference impact can surpass training once a model is queried millions of times a day.

The chapter, in its own words

Five lines from Chapter 2 that carry the year’s argument.

AI has surpassed human performance on several benchmarks, including some in image classification, visual reasoning, and English understanding. Yet it trails behind on more complex tasks like competition-level mathematics, visual commonsense reasoning and planning.
— Chapter 2 · Chapter Highlights
AI models have reached performance saturation on established benchmarks such as ImageNet, SQuAD, and SuperGLUE, prompting researchers to develop more challenging ones.
— Chapter 2 · Chapter Highlights
With generative models producing high-quality text, images, and more, benchmarking has slowly started shifting toward incorporating human evaluations like the Chatbot Arena Leaderboard rather than computerized rankings like ImageNet or SQuAD.
— Chapter 2 · Chapter Highlights
On 10 select AI benchmarks, closed models outperformed open ones, with a median performance advantage of 24.2%. Differences in the performance of closed and open models carry important implications for AI policy debates.
— Chapter 2 · Chapter Highlights
With AI models growing in size and becoming more widely used, it has never been more critical for the AI research community to diligently monitor and mitigate the environmental effects of AI systems.
— Chapter 2 · 2.13 Environmental Impact of AI Systems

Read Chapter 2 in full

Chapter 2 (sections 2.1–2.13) — language, coding, image and video, reasoning, audio, agents, robotics, reinforcement learning, properties of LLMs, techniques for improvement and environmental impact — with every figure and citation is free from Stanford HAI.

Open the AI Index 2024 →