AI Index Report 2025

AI cleared the benchmarks built to stop it — in a single year

Chapter 2 of the AI Index 2025 tracks language, coding, math, reasoning, vision, speech, agents and robotics. Benchmarks designed in 2023 to be hard for years were largely solved by the end of 2024, the open-weight gap nearly vanished, and Chinese models drew almost level with American ones. The numbers:

71.7% of SWE-bench Verified coding problems solved in early 2025 (4.4% in 2023)
48.9percentage-point jump on GPQA Diamond in a single year
1.7% gap left between the best closed and open-weight models (8.0% in Jan 2024)
142-fold shrink in the smallest model clearing 60% on MMLU (540B → 3.8B params)
8.8% top score on Humanity's Last Exam, the hardest new academic benchmark
150,000paid Waymo robotaxi rides a week across four US cities

2.1 — Four gaps closed at once

The story of 2024 is convergence. The gap between AI and human baselines, between closed and open weights, between American and Chinese models, and between the first- and tenth-ranked systems all narrowed sharply in twelve months.

AI versus humans

There are very few task categories left where human ability clearly surpasses AI, and in those that remain the gap is shrinking fast. On MATH, a competition-level mathematics benchmark, state-of-the-art systems now sit 7.9 percentage points ahead of human performance. On MMMU — a multidisciplinary, expert-level test — the best 2024 model, o1, scored 78.2%, just 4.4 points below the human benchmark of 82.6%; at the end of 2023 Google Gemini managed 59.4%. On GPQA Diamond, o3's 87.7% was the first score to exceed the 81.3% accuracy of expert human validators.

Closed versus open weights

  • In early January 2024 the leading closed-weight model outperformed the top open-weight model by 8.0% on the Chatbot Arena Leaderboard. By February 2025 that gap had narrowed to 1.7%.
  • The same collapse shows up on static benchmarks. In late 2023 closed models led open ones on MMLU by 15.9 points; by the end of 2024 the difference was 0.1 percentage point.
  • The catch-up was driven largely by Meta's summer release of Llama 3.1, followed by other strong open-weight models including DeepSeek's V3.
  • Openness remains contested: advocates point to reduced market concentration, better security scrutiny and transparency; critics warn about disinformation and bioweapon risks. Note too that open weights is not open source — training code and data are usually withheld.

The United States versus China

In January 2024 the top US model outperformed the best Chinese model by 9.3% on the LMSYS Chatbot Arena. By February 2025 that lead was 1.7%. On static benchmarks the shift is starker: at the end of 2023 the gaps on MMLU, MMMU, MATH and HumanEval were 17.5, 13.5, 24.3 and 31.6 percentage points; by the end of 2024 they were 0.3, 8.1, 1.6 and 3.7. DeepSeek's R1 launch drew attention for a second reason — the company reported achieving its results with a fraction of the hardware normally required, which moved US stock markets and raised questions about the effectiveness of semiconductor export controls.

The frontier versus itself

  • The Elo gap between the top and tenth-ranked model on the Chatbot Arena Leaderboard fell from 11.9% to 5.4% in a year, and the difference between the top two models shrank from 4.9% in 2023 to 0.7% in 2024.
  • 2024 was a breakthrough year for small models. In 2022 the smallest model scoring above 60% on MMLU was PaLM at 540 billion parameters; by 2024 Microsoft's Phi-3 Mini did it with 3.8 billion — a 142-fold reduction in two years.
  • New reasoning paradigms arrived: OpenAI's o1 and o3 iterate over their own outputs at inference time. o1 scored 74.4% on an International Mathematical Olympiad qualifying exam, against GPT-4o's 9.3%.
  • That reasoning is not free. o1 costs $15 per million input tokens and $60 per million output tokens, versus $2.50 and $10 for GPT-4o, and takes roughly 40 times longer to produce its first token (29.7 seconds against 0.72).

Where AI actually stands on the hardest tests

Best reported score on each benchmark as of early 2025 (%). The 2023-vintage tests are largely solved; the 2024-vintage ones are not — the top system answers 8.8% of Humanity's Last Exam and 2% of FrontierMath.

Where AI actually stands on the hardest testsGPQA Diamond: 87.787.7GPQA DiamondMMMU: 78.278.2MMMUSWE-bench Verified: 71.771.7SWE-bench VerifiedBigCodeBench (hard): 35.535.5BigCodeBench (hard)Humanity's Last Exam: 8.88.8Humanity's Last ExamFrontierMath: 22FrontierMath

The year in launches

A selection of the model and capability releases that shaped 2024, drawn from the chapter's own timeline.

  1. Feb 15, 2024

    Gemini 1.5 Pro · Google

    Google's new flagship LLM. By early 2025 the Chatbot Arena had gathered over a million votes, and users ranked one of Google's Gemini models as the community's most preferred system.

  2. Mar 4, 2024

    Claude 3 · Anthropic

    Anthropic's new LLM family. Its successor, Claude 3.5 Sonnet, would go on to post a perfect 100% on HumanEval under HPT prompting, 97.72% on GSM8K, and the highest mean safety score on Stanford's HELM Safety suite at 0.977.

  3. May 13, 2024

    GPT-4o · OpenAI

    A natively multimodal model reasoning across text, audio and images, priced at $2.50 per million input tokens and $10 per million output tokens, with a 0.72-second time to first token.

  4. Jul 23, 2024

    Llama 3.1 405B · Meta

    Meta's largest model to date, and the single biggest reason the open-weight gap collapsed. Training took roughly 90 days, drew 25.3 million watts, and cost an estimated $170 million.

  5. Sep 12, 2024

    o1-preview · OpenAI

    The first model in the o series, designed to reason step by step rather than answer autoregressively. The full o1 followed on December 5, alongside ChatGPT Pro at $200 a month. Against GPT-4o, o1 gained 2.8 points on MMLU, 34.5 on MATH, 26.7 on GPQA Diamond and 65.1 on AIME 2024.

  6. Oct 22, 2024

    Computer Use · Anthropic

    A computer control capability that lets a model operate a desktop directly — one of the year's clearest steps from chatbot toward agent.

  7. Dec 12, 2024

    Sora · OpenAI

    Previewed in February and publicly accessible in December, Sora generates 20-second videos at resolutions up to 1080p. It arrived in a crowded year: Stable Video 3D in March, Meta's Movie Gen in October (16-second 1080p clips with sound), and Google's Veo 2, whose output was consistently preferred over Movie Gen, Kling v1.5 and Sora Turbo in user comparisons.

  8. Dec 20 & 27, 2024

    o3 (beta) · OpenAI, then DeepSeek-V3

    o3 posted 75.7% on ARC-AGI — 87.5% when given a compute budget above the benchmark's $10,000 limit — plus 87.7% on GPQA Diamond and 71.7% on SWE-bench Verified. A week later DeepSeek released V3, an open-source model whose performance rivaled the frontier at a reported fraction of the training cost.

2.5–2.7 — Code and math fell first, reasoning is following

The benchmarks that defined AI coding and mathematics two years ago are now saturated or solved. The harder replacements built in 2024 show how much room is left.

Coding

  • HumanEval, introduced by OpenAI researchers in 2021 with 164 handwritten problems, is finished: Claude 3.5 Sonnet under HPT prompting scores 100%.
  • SWE-bench, built from real GitHub issues in Python repositories, was meant to be much harder — it requires coordinating changes across files. The best model at the end of 2023 solved 4.4% of problems. By early 2025, o3 solved 71.7% of the Verified set.
  • BigCodeBench, released in 2024, is the current hard case: 1,140 tasks requiring calls across 139 libraries and seven domains. The best model, o1, averages 35.5 on the hard subset — well below the human standard of 97%.
  • In the Chatbot Arena coding filter, Gemini-Exp-1206 leads with an arena score of 1,369, just ahead of o1 at 1,361. DeepSeek-V3 leads Chinese models at 1,317, trailing the top by 3.8%.

Mathematics

  • GSM8K is nearing saturation: Claude 3.5 Sonnet with HPT prompting scores 97.72%, up from a 91.00% high in 2023, and several Mistral, Meta and Qwen models cluster around 96%. On MATH, the best system solves 97.9% of problems.
  • FrontierMath, introduced by Epoch AI, is the answer to that saturation — original problems vetted by expert mathematicians that can take hours, days or collaborative effort to solve. At release, the best of six leading LLMs, Gemini 1.5 Pro, solved just 2.0%.
  • DeepMind's AlphaProof and AlphaGeometry 2 solved four of six problems at the 2024 International Mathematical Olympiad, a silver-medal-equivalent performance. On IMO-AG-30 geometry problems the systems solved 25, against an IMO silver medalist's average of 22.9.
  • In the Chatbot Arena math filter — over 181 models and more than 340,000 public votes — the top model is an OpenAI o1 variant released in December 2024, breaking the Gemini lead seen in the general and coding arenas.

Reasoning

MMMU, launched in 2023 with about 11,500 college-level questions across six disciplines, went from a 59.4% state of the art to o1's 78.2% in a year. GPQA Diamond — 448 expert-written questions that cannot be answered by web search — went from GPT-4's 38.8% to o3's 87.7%, a 48.9-point jump and the first score above the 81.3% expert human baseline. ARC-AGI, designed to resist memorization, is the most dramatic case: 20% when first run in 2020, still only 33% four years later, then 75.7% from o3 — and 87.5% when given a compute budget beyond the benchmark's $10,000 limit. Researchers attribute the earlier stagnation to an overemphasis on scaling, which improved task-specific skill without improving generalization.

The US–China performance gap at the end of 2024

Percentage-point lead of the top US model over the top Chinese model. A year earlier the same four gaps were 17.5, 13.5, 24.3 and 31.6 points — multimodal reasoning is the only one still meaningfully open.

The US–China performance gap at the end of 2024MMMU: 8.18.1MMMUHumanEval: 3.73.7HumanEvalMATH: 1.61.6MATHMMLU: 0.30.3MMLU

2.9 — Out of the browser and onto the road

New this year, the chapter expands its coverage of robotics and self-driving cars. Autonomous taxis are now a commercial service in four US cities and sixteen Chinese ones, with safety data that is starting to look convincing.

Robotaxis at commercial scale

  • As of January 2025, Waymo operates in Phoenix, San Francisco, Los Angeles and Austin, providing about 150,000 paid rides a week and covering over a million miles. Rider-only miles through September 2024 were 20.823 million in Phoenix, 10.209 million in San Francisco, 1.947 million in Los Angeles and 124,000 in Austin. The company plans to test in ten more cities, deliberately including snowy locations such as upstate New York and Truckee, California.
  • Baidu's Apollo Go reported 988,000 rides across China in Q3 2024, a 20% year-over-year increase, operating 400 robotaxis in October 2024 with plans for 1,000 by the end of 2025. Its RT6 robotaxi, with a battery-swapping system, costs about $30,000.
  • Pony.AI has pledged to grow its fleet from 200 to at least 1,000 vehicles, with 2,000 to 3,000 expected by the end of 2026. China is testing more driverless cars than any other country, rolling them out across 16 cities, and has prioritized national regulations to govern deployment.
  • Tesla unveiled the Cybercab in October 2024 — a two-passenger vehicle with no steering wheel or pedals, slated for 2026 production at under $30,000 — alongside the 20-passenger Robovan. Cruise, by contrast, had its license suspended in 2023 after a series of safety incidents.

Are they safer than us?

The emerging evidence says yes. Compared with the estimated rate for human drivers over the same distance, Waymo vehicles recorded 1.42 fewer airbag deployments, 3.16 fewer crashes with reported injuries and 3.65 fewer police-reported crashes per million miles. A separate study with the reinsurer Swiss Re — benchmarked against a dataset of over 500,000 claims and 200 billion miles of driving — found an 88% reduction in property damage claims and a 92% reduction in bodily injury claims. In absolute terms, across 25.3 million miles Waymo vehicles drew nine property damage claims and two bodily injury claims, where human drivers would have been expected to incur 78 and 26. Waymo also outperformed the latest-generation human-driven vehicles fitted with modern safety features.

Robots that learn

  • 2024 was a notable year for humanoids. Figure AI's Figure 02 handles a 44-pound payload, runs for up to five hours on a charge, and is integrated with OpenAI for speech-to-speech reasoning — it can explain what it is doing. Tesla's Optimus and Boston Dynamics' Atlas continued to develop alongside it.
  • DeepMind's AutoRT autonomously generates training data for robots and has produced a dataset of 77,000 robotic trials spanning 6,650 unique tasks. SARA-RT improves the efficiency of transformer-based robotic models; ALOHA and DemoStart tackle dexterous manipulation with far less data.
  • Foundation models arrived in robotics: Nvidia's GROOT development suite pairs humanoid models with simulation frameworks and the Thor robotics computer, following RT-2, PaLM-E and Open-X Embodiment.
  • New benchmarks matched the ambition. nuPlan offers 1,282 hours of driving scenarios with closed-loop evaluation; Bench2Drive provides over 2 million annotated frames from more than 10,000 clips plus 220 evaluation routes; OpenAD is the first real-world open-world benchmark for 3D object detection in driving.

What AI still gets wrong

The chapter is candid about the limits — and about how much we can trust the scores in the first place.

Can these models actually plan?
Better than before, but not reliably. On PlanBench — 600 block-stacking problems from the automated planning community — o1 scored 97.8% on the Blocksworld zero-shot evaluation, far ahead of Llama 3.1 405B (62.6%) and GPT-4o (35.5%). But on Mystery Blocksworld, where answers are syntactically obfuscated, o1 managed 52.8% against Llama 3.1 405B's 0.8% and GPT-4's 0%. And on instances requiring at least 20 steps, o1 solves just 23.6%. These systems still cannot reliably solve problems for which provably correct answers exist through logical reasoning — a real limit on their suitability for high-stakes settings where precision matters.
Are AI agents ready to be deployed?
Not yet, but the shape of their advantage is becoming clear. On VisualAgentBench, which tests embodied, GUI and visual design agents, the best model — GPT-4o — reaches an overall success rate of just 36.2%, and most proprietary models average around 20%; the authors conclude current models are far from ready for direct deployment. RE-Bench found the trade-off is about time: with a two-hour budget the best AI systems score four times higher than human experts, at eight hours humans slightly pull ahead, and at 32 hours humans outscore AI two to one. In specific tasks agents already match human expertise — they can write custom Triton kernels faster than any human expert, and at lower cost. On Meta's GAIA benchmark the top system reached 65.1%, roughly 30 percentage points above the previous best.
Can we trust the benchmark scores at all?
Treat them carefully. BetterBench researchers systematically analyzed 24 prominent benchmarks and found systemic deficiencies: 14 failed to report statistical significance, 17 lacked scripts for replicating results, and most had inadequate documentation. MMLU adhered poorly to quality standards, while GPQA performed significantly better. Contamination is a second problem — a Scale study found significant contamination in many LLMs' GSM8K performance, prompting periodically updated benchmarks like LiveBench. Third, developers report their own scores, sometimes using nonstandard prompting: Google reported an MMLU score for Gemini Ultra using a chain-of-thought technique other developers did not use, and third-party researchers have found publicly reported scores differing from independent evaluations by as much as five percentage points.
Has AI passed the Turing test?
Effectively, and the more interesting point is that the question has stopped mattering. Recent evidence suggests LLMs have advanced far enough that people struggle to distinguish the best-performing models from a human in text conversation. The AI Index frames this less as a milestone reached than as a milestone that lost its meaning: the test remains a historical and cultural landmark, but its diminishing relevance is itself the measure of progress. Attention has moved to benchmarks the systems cannot yet touch.
What is left that AI cannot do?
The 2024-vintage benchmarks. Humanity's Last Exam — 2,700 multimodal questions written by leading professors and graduate reviewers, each pre-tested against state-of-the-art LLMs and rejected if a model could already answer it — is answered correctly just 8.8% of the time by o1, though its creators speculate performance could exceed 50% by the end of 2025. FrontierMath sits at 2%. BigCodeBench's hard subset sits at 35.5%, against a human standard of 97%. And SimpleQA, OpenAI's own factuality test, is answered correctly only 42.7% of the time by o1-preview, its best performer.

Read the full Technical Performance chapter

Chapter 2 (sections 2.1–2.9) — language, image and video, speech, coding, math, reasoning, agents, robotics and autonomous vehicles — with every figure and citation is free from Stanford HAI.

Open the AI Index Report 2025 →