SWE-bench: from 60% to near-100%
Autonomous software engineering rose from ~60% in 2024 to close to 100% of the human baseline.
Chapter 2 of the AI Index 2026 tracks AI across language, reasoning, coding, math, agents, and robotics. Scores are rising fast, the gap between top models is shrinking, and evaluations are saturating in months. The numbers:
AI improved rapidly in 2025 across language, reasoning, coding, and math — but progress is outpacing the evaluations built to measure it, and a clear pattern emerges: the frontier is converging.
Frontier systems now meet or exceed established human baselines on long-running benchmarks like ImageNet, SuperGLUE, and MMLU, and have reached or approached human levels on harder reasoning tests — PhD-level science (GPQA Diamond), multimodal reasoning (MMMU), and competition math (AIME). On SWE-bench Verified, performance rose from roughly 60% in 2024 to close to 100% of the human baseline in 2025.
Evaluations meant to be hard for years are now saturated in months. A Stanford review of nine widely used benchmarks found invalid-question rates ranging from 2% on MMLU Math to 42% on GSM8K, and separate research suggests Arena leaderboard standing may partly reflect adaptation to the platform rather than general capability. With capability no longer a clear differentiator, competition is shifting toward cost, reliability, and real-world usefulness.
Reasoning surged across science, math, and abstraction — yet the same models stumble on tasks most humans find trivial. Researchers call it jagged intelligence.
Frontier models gained 30 percentage points in a single year on Humanity's Last Exam — a benchmark built to be hard for AI and favorable to human experts — climbing from under 10% to 38.3%. On GPQA Diamond, mean accuracy reached 93%, 12 points above the expert validator baseline of 81.2%. On MMMU, Gemini 3.1 Pro Preview scored 88.2%, within 0.4 points of the best human expert.
Software agents are closing in on humans on structured tasks. Physical robots still falter outside the lab — with autonomous vehicles as the standout exception.
AI agents advanced from answering questions to completing multistep tasks in 2025, though they still fail roughly one in three attempts on structured benchmarks. On OSWorld, accuracy rose from about 12% to 66.3%, within 6 points of human performance; on WebArena it reached 74.3%, within 4 points of the 78.2% human baseline; on Cybench the unguided solve rate hit 93%, up from 15% in 2024.
Self-driving cars reached mass-scale deployment in 2025. Waymo operated roughly 2,500 robotaxis across five U.S. cities, recording around 450,000 weekly trips. In China, Baidu's Apollo Go provided approximately 11 million fully driverless rides — a 175% year-over-year increase, up from 1.5 million in 2022. Deployments so far concentrate in favorable-weather areas with off-site humans available to take over.
Where AI is being pushed into expert work — coding, math, video, time, and country-level competition. Tap any card for the full trend and its numbers.
Autonomous software engineering rose from ~60% in 2024 to close to 100% of the human baseline.
MathArena answer accuracy hit 97%, but rigorous proof-writing still lags far behind humans.
DeepSeek-R1 briefly matched the top U.S. model in Feb 2025; the lead is now just 2.7%.
Veo 3, tested on 18,000+ videos, simulated buoyancy and solved mazes it was never trained on.
On ClockBench, the top model read analog clocks correctly 50.6% of the time vs. 90.1% for humans.
Waymo hit ~450,000 weekly trips; Apollo Go logged ~11M driverless rides, up 175% YoY.
Headline findings from Chapter 2 · Technical Performance.
Frontier models gained 30 percentage points in a single year on Humanity's Last Exam — a benchmark built to be hard for AI and favorable to human experts.
The U.S.–China AI model performance gap has effectively closed; as of March 2026 the top U.S. model leads by just 2.7%.
AI can win a gold medal at the International Mathematical Olympiad but still reads analog clocks correctly only 50.6% of the time, versus 90.1% for humans.
AI agents advanced from answering questions to completing tasks — yet still fail roughly one in three attempts on structured benchmarks.
Robots succeed in only 12% of real household tasks even as they reach 89.4% in controlled simulation, while autonomous vehicles hit mass-scale deployment.
Chapter 2 (sections 2.1–2.7) with every benchmark, figure, and citation is free from Stanford HAI. Or head back to the 15 takeaways and nine-chapter overview.
Open Chapter 2 · Technical Performance →