AI cleared the benchmarks built to stop it — in a single year
Chapter 2 of the AI Index 2025 tracks language, coding, math, reasoning, vision, speech, agents and robotics. Benchmarks designed in 2023 to be hard for years were largely solved by the end of 2024, the open-weight gap nearly vanished, and Chinese models drew almost level with American ones. The numbers:
2.1 — Four gaps closed at once
The story of 2024 is convergence. The gap between AI and human baselines, between closed and open weights, between American and Chinese models, and between the first- and tenth-ranked systems all narrowed sharply in twelve months.
AI versus humans
There are very few task categories left where human ability clearly surpasses AI, and in those that remain the gap is shrinking fast. On MATH, a competition-level mathematics benchmark, state-of-the-art systems now sit 7.9 percentage points ahead of human performance. On MMMU — a multidisciplinary, expert-level test — the best 2024 model, o1, scored 78.2%, just 4.4 points below the human benchmark of 82.6%; at the end of 2023 Google Gemini managed 59.4%. On GPQA Diamond, o3's 87.7% was the first score to exceed the 81.3% accuracy of expert human validators.
Closed versus open weights
- In early January 2024 the leading closed-weight model outperformed the top open-weight model by 8.0% on the Chatbot Arena Leaderboard. By February 2025 that gap had narrowed to 1.7%.
- The same collapse shows up on static benchmarks. In late 2023 closed models led open ones on MMLU by 15.9 points; by the end of 2024 the difference was 0.1 percentage point.
- The catch-up was driven largely by Meta's summer release of Llama 3.1, followed by other strong open-weight models including DeepSeek's V3.
- Openness remains contested: advocates point to reduced market concentration, better security scrutiny and transparency; critics warn about disinformation and bioweapon risks. Note too that open weights is not open source — training code and data are usually withheld.
The United States versus China
In January 2024 the top US model outperformed the best Chinese model by 9.3% on the LMSYS Chatbot Arena. By February 2025 that lead was 1.7%. On static benchmarks the shift is starker: at the end of 2023 the gaps on MMLU, MMMU, MATH and HumanEval were 17.5, 13.5, 24.3 and 31.6 percentage points; by the end of 2024 they were 0.3, 8.1, 1.6 and 3.7. DeepSeek's R1 launch drew attention for a second reason — the company reported achieving its results with a fraction of the hardware normally required, which moved US stock markets and raised questions about the effectiveness of semiconductor export controls.
The frontier versus itself
- The Elo gap between the top and tenth-ranked model on the Chatbot Arena Leaderboard fell from 11.9% to 5.4% in a year, and the difference between the top two models shrank from 4.9% in 2023 to 0.7% in 2024.
- 2024 was a breakthrough year for small models. In 2022 the smallest model scoring above 60% on MMLU was PaLM at 540 billion parameters; by 2024 Microsoft's Phi-3 Mini did it with 3.8 billion — a 142-fold reduction in two years.
- New reasoning paradigms arrived: OpenAI's o1 and o3 iterate over their own outputs at inference time. o1 scored 74.4% on an International Mathematical Olympiad qualifying exam, against GPT-4o's 9.3%.
- That reasoning is not free. o1 costs $15 per million input tokens and $60 per million output tokens, versus $2.50 and $10 for GPT-4o, and takes roughly 40 times longer to produce its first token (29.7 seconds against 0.72).
The year in launches
A selection of the model and capability releases that shaped 2024, drawn from the chapter's own timeline.
- Feb 15, 2024
Gemini 1.5 Pro · Google
Google's new flagship LLM. By early 2025 the Chatbot Arena had gathered over a million votes, and users ranked one of Google's Gemini models as the community's most preferred system.
- Mar 4, 2024
Claude 3 · Anthropic
Anthropic's new LLM family. Its successor, Claude 3.5 Sonnet, would go on to post a perfect 100% on HumanEval under HPT prompting, 97.72% on GSM8K, and the highest mean safety score on Stanford's HELM Safety suite at 0.977.
- May 13, 2024
GPT-4o · OpenAI
A natively multimodal model reasoning across text, audio and images, priced at $2.50 per million input tokens and $10 per million output tokens, with a 0.72-second time to first token.
- Jul 23, 2024
Llama 3.1 405B · Meta
Meta's largest model to date, and the single biggest reason the open-weight gap collapsed. Training took roughly 90 days, drew 25.3 million watts, and cost an estimated $170 million.
- Sep 12, 2024
o1-preview · OpenAI
The first model in the o series, designed to reason step by step rather than answer autoregressively. The full o1 followed on December 5, alongside ChatGPT Pro at $200 a month. Against GPT-4o, o1 gained 2.8 points on MMLU, 34.5 on MATH, 26.7 on GPQA Diamond and 65.1 on AIME 2024.
- Oct 22, 2024
Computer Use · Anthropic
A computer control capability that lets a model operate a desktop directly — one of the year's clearest steps from chatbot toward agent.
- Dec 12, 2024
Sora · OpenAI
Previewed in February and publicly accessible in December, Sora generates 20-second videos at resolutions up to 1080p. It arrived in a crowded year: Stable Video 3D in March, Meta's Movie Gen in October (16-second 1080p clips with sound), and Google's Veo 2, whose output was consistently preferred over Movie Gen, Kling v1.5 and Sora Turbo in user comparisons.
- Dec 20 & 27, 2024
o3 (beta) · OpenAI, then DeepSeek-V3
o3 posted 75.7% on ARC-AGI — 87.5% when given a compute budget above the benchmark's $10,000 limit — plus 87.7% on GPQA Diamond and 71.7% on SWE-bench Verified. A week later DeepSeek released V3, an open-source model whose performance rivaled the frontier at a reported fraction of the training cost.
2.5–2.7 — Code and math fell first, reasoning is following
The benchmarks that defined AI coding and mathematics two years ago are now saturated or solved. The harder replacements built in 2024 show how much room is left.
Coding
- HumanEval, introduced by OpenAI researchers in 2021 with 164 handwritten problems, is finished: Claude 3.5 Sonnet under HPT prompting scores 100%.
- SWE-bench, built from real GitHub issues in Python repositories, was meant to be much harder — it requires coordinating changes across files. The best model at the end of 2023 solved 4.4% of problems. By early 2025, o3 solved 71.7% of the Verified set.
- BigCodeBench, released in 2024, is the current hard case: 1,140 tasks requiring calls across 139 libraries and seven domains. The best model, o1, averages 35.5 on the hard subset — well below the human standard of 97%.
- In the Chatbot Arena coding filter, Gemini-Exp-1206 leads with an arena score of 1,369, just ahead of o1 at 1,361. DeepSeek-V3 leads Chinese models at 1,317, trailing the top by 3.8%.
Mathematics
- GSM8K is nearing saturation: Claude 3.5 Sonnet with HPT prompting scores 97.72%, up from a 91.00% high in 2023, and several Mistral, Meta and Qwen models cluster around 96%. On MATH, the best system solves 97.9% of problems.
- FrontierMath, introduced by Epoch AI, is the answer to that saturation — original problems vetted by expert mathematicians that can take hours, days or collaborative effort to solve. At release, the best of six leading LLMs, Gemini 1.5 Pro, solved just 2.0%.
- DeepMind's AlphaProof and AlphaGeometry 2 solved four of six problems at the 2024 International Mathematical Olympiad, a silver-medal-equivalent performance. On IMO-AG-30 geometry problems the systems solved 25, against an IMO silver medalist's average of 22.9.
- In the Chatbot Arena math filter — over 181 models and more than 340,000 public votes — the top model is an OpenAI o1 variant released in December 2024, breaking the Gemini lead seen in the general and coding arenas.
Reasoning
MMMU, launched in 2023 with about 11,500 college-level questions across six disciplines, went from a 59.4% state of the art to o1's 78.2% in a year. GPQA Diamond — 448 expert-written questions that cannot be answered by web search — went from GPT-4's 38.8% to o3's 87.7%, a 48.9-point jump and the first score above the 81.3% expert human baseline. ARC-AGI, designed to resist memorization, is the most dramatic case: 20% when first run in 2020, still only 33% four years later, then 75.7% from o3 — and 87.5% when given a compute budget beyond the benchmark's $10,000 limit. Researchers attribute the earlier stagnation to an overemphasis on scaling, which improved task-specific skill without improving generalization.
2.9 — Out of the browser and onto the road
New this year, the chapter expands its coverage of robotics and self-driving cars. Autonomous taxis are now a commercial service in four US cities and sixteen Chinese ones, with safety data that is starting to look convincing.
Robotaxis at commercial scale
- As of January 2025, Waymo operates in Phoenix, San Francisco, Los Angeles and Austin, providing about 150,000 paid rides a week and covering over a million miles. Rider-only miles through September 2024 were 20.823 million in Phoenix, 10.209 million in San Francisco, 1.947 million in Los Angeles and 124,000 in Austin. The company plans to test in ten more cities, deliberately including snowy locations such as upstate New York and Truckee, California.
- Baidu's Apollo Go reported 988,000 rides across China in Q3 2024, a 20% year-over-year increase, operating 400 robotaxis in October 2024 with plans for 1,000 by the end of 2025. Its RT6 robotaxi, with a battery-swapping system, costs about $30,000.
- Pony.AI has pledged to grow its fleet from 200 to at least 1,000 vehicles, with 2,000 to 3,000 expected by the end of 2026. China is testing more driverless cars than any other country, rolling them out across 16 cities, and has prioritized national regulations to govern deployment.
- Tesla unveiled the Cybercab in October 2024 — a two-passenger vehicle with no steering wheel or pedals, slated for 2026 production at under $30,000 — alongside the 20-passenger Robovan. Cruise, by contrast, had its license suspended in 2023 after a series of safety incidents.
Are they safer than us?
The emerging evidence says yes. Compared with the estimated rate for human drivers over the same distance, Waymo vehicles recorded 1.42 fewer airbag deployments, 3.16 fewer crashes with reported injuries and 3.65 fewer police-reported crashes per million miles. A separate study with the reinsurer Swiss Re — benchmarked against a dataset of over 500,000 claims and 200 billion miles of driving — found an 88% reduction in property damage claims and a 92% reduction in bodily injury claims. In absolute terms, across 25.3 million miles Waymo vehicles drew nine property damage claims and two bodily injury claims, where human drivers would have been expected to incur 78 and 26. Waymo also outperformed the latest-generation human-driven vehicles fitted with modern safety features.
Robots that learn
- 2024 was a notable year for humanoids. Figure AI's Figure 02 handles a 44-pound payload, runs for up to five hours on a charge, and is integrated with OpenAI for speech-to-speech reasoning — it can explain what it is doing. Tesla's Optimus and Boston Dynamics' Atlas continued to develop alongside it.
- DeepMind's AutoRT autonomously generates training data for robots and has produced a dataset of 77,000 robotic trials spanning 6,650 unique tasks. SARA-RT improves the efficiency of transformer-based robotic models; ALOHA and DemoStart tackle dexterous manipulation with far less data.
- Foundation models arrived in robotics: Nvidia's GROOT development suite pairs humanoid models with simulation frameworks and the Thor robotics computer, following RT-2, PaLM-E and Open-X Embodiment.
- New benchmarks matched the ambition. nuPlan offers 1,282 hours of driving scenarios with closed-loop evaluation; Bench2Drive provides over 2 million annotated frames from more than 10,000 clips plus 220 evaluation routes; OpenAD is the first real-world open-world benchmark for 3D object detection in driving.
What AI still gets wrong
The chapter is candid about the limits — and about how much we can trust the scores in the first place.
Can these models actually plan?
Are AI agents ready to be deployed?
Can we trust the benchmark scores at all?
Has AI passed the Turing test?
What is left that AI cannot do?
Read the full Technical Performance chapter
Chapter 2 (sections 2.1–2.9) — language, image and video, speech, coding, math, reasoning, agents, robotics and autonomous vehicles — with every figure and citation is free from Stanford HAI.
Open the AI Index Report 2025 →