AI Index Report 2026

AI capability is outpacing the benchmarks built to measure it

Chapter 2 of the AI Index 2026 tracks AI across language, reasoning, coding, math, agents, and robotics. Scores are rising fast, the gap between top models is shrinking, and evaluations are saturating in months. The numbers:

100% of human baseline reached on SWE-bench Verified (from ~60% in 2024)
66% accuracy on OSWorld computer-use agents (from ~12%)
30percentage-point jump on Humanity's Last Exam in one year
25Elo points now separate the top four Arena models
89% robotic-manipulation success on RLBench in simulation
12% success on real household robot tasks (BEHAVIOR-1K)

Reasoning benchmarks are clearing the human bar

Top model accuracy on frontier reasoning benchmarks (%), early 2026. MMMU and GPQA Diamond now meet or beat expert humans; ARC-AGI-2 and Humanity's Last Exam stay deliberately hard.

Reasoning benchmarks are clearing the human barGPQA Diamond (mean): 9393GPQA Diamond (mean)MMMU (top): 8888MMMU (top)ARC-AGI-2 (top): 8585ARC-AGI-2 (top)Humanity's Last Exam: 3838Humanity's Last Exam

2.4 — Gold medals, but it still can't tell time

Reasoning surged across science, math, and abstraction — yet the same models stumble on tasks most humans find trivial. Researchers call it jagged intelligence.

Frontier models gained 30 percentage points in a single year on Humanity's Last Exam — a benchmark built to be hard for AI and favorable to human experts — climbing from under 10% to 38.3%. On GPQA Diamond, mean accuracy reached 93%, 12 points above the expert validator baseline of 81.2%. On MMMU, Gemini 3.1 Pro Preview scored 88.2%, within 0.4 points of the best human expert.

Jagged intelligence: IMO gold vs. analog clocks

  • Gemini Deep Think scored 35 points to win gold at the 2025 International Mathematical Olympiad, working end to end in natural language within the 4.5-hour limit — up from the 28-point silver achieved in 2024.
  • Yet on ClockBench, the top model read analog clocks correctly only 50.6% of the time, versus 90.1% for humans. When models read the time wrong, their median error was about one to three hours, compared with three minutes for people.
  • On FrontierMath Tier 4, accuracy rose from near 0% to 31.3% since 2024 — but the best models still fail roughly two of every three problems at the hardest tier. On MathArena, answer-based accuracy hit 97% while rigorous proof-writing lagged far behind.

AI agents went from answering to doing

Top agent success rate (%) on real-task benchmarks, early 2026 vs. their 2024–25 starting points. Agents improved fast but still fail roughly one in three attempts.

AI agents went from answering to doingCybench (unguided): 9393Cybench (unguided)Terminal-Bench 2.0: 7777Terminal-Bench 2.0GAIA: 7575GAIAWebArena: 7474WebArenaOSWorld: 6666OSWorldMLE-bench: 6464MLE-bench

2.6 & 2.7 — Agents rising, robots stuck at the front door

Software agents are closing in on humans on structured tasks. Physical robots still falter outside the lab — with autonomous vehicles as the standout exception.

AI agents advanced from answering questions to completing multistep tasks in 2025, though they still fail roughly one in three attempts on structured benchmarks. On OSWorld, accuracy rose from about 12% to 66.3%, within 6 points of human performance; on WebArena it reached 74.3%, within 4 points of the 78.2% human baseline; on Cybench the unguided solve rate hit 93%, up from 15% in 2024.

Lab triumph, household failure

  • In controlled simulation, robotic manipulation on RLBench reached an 89.4% success rate (EquAct), up from about 48% in 2022. But in the 2025 BEHAVIOR-1K household challenge, the top team completed just 12.4% of real tasks — a stark measure of how far AI is from the unstructured physical world.
  • On ResponsibleRobotBench, even the best model (GPT-4o, safe score 0.64) completed under a third of tasks safely when real hazards were present.
  • Humanoid hardware proliferated: Figure AI's Figure 02 logged over 1,250 hours at a BMW plant, loading 90,000+ parts across 30,000+ vehicles. Yet most milestones remain framed as future ambitions rather than verified operations.

Autonomous vehicles: the deployment exception

Self-driving cars reached mass-scale deployment in 2025. Waymo operated roughly 2,500 robotaxis across five U.S. cities, recording around 450,000 weekly trips. In China, Baidu's Apollo Go provided approximately 11 million fully driverless rides — a 175% year-over-year increase, up from 1.5 million in 2022. Deployments so far concentrate in favorable-weather areas with off-site humans available to take over.

2.5 — Inside the professional domains

Where AI is being pushed into expert work — coding, math, video, time, and country-level competition. Tap any card for the full trend and its numbers.

SWE-bench: from 60% to near-100%

Autonomous software engineering rose from ~60% in 2024 to close to 100% of the human baseline.

coding

Math: right answers, shaky proofs

MathArena answer accuracy hit 97%, but rigorous proof-writing still lags far behind humans.

mathematics

The U.S.–China gap has closed

DeepSeek-R1 briefly matched the top U.S. model in Feb 2025; the lead is now just 2.7%.

competition

Video models start to learn physics

Veo 3, tested on 18,000+ videos, simulated buoyancy and solved mazes it was never trained on.

video

It can't reliably tell time

On ClockBench, the top model read analog clocks correctly 50.6% of the time vs. 90.1% for humans.

reasoning

Autonomous vehicles go mass-scale

Waymo hit ~450,000 weekly trips; Apollo Go logged ~11M driverless rides, up 175% YoY.

deployment

From near-zero to near-human in two years

Best success rate (%) at each benchmark's earlier start vs. early 2026 — the pace of improvement on once-hard tasks.

From near-zero to near-human in two yearsOSWorld · ~2024: 1212OSWorld · ~2024OSWorld · 2026: 6666OSWorld · 2026Cybench · 2024: 1515Cybench · 2024Cybench · 2026: 9393Cybench · 2026MLE-bench · 2024: 1717MLE-bench · 2024MLE-bench · 2026: 6464MLE-bench · 2026

The chapter in five lines

Headline findings from Chapter 2 · Technical Performance.

Frontier models gained 30 percentage points in a single year on Humanity's Last Exam — a benchmark built to be hard for AI and favorable to human experts.
— Chapter 2 · Technical Performance
The U.S.–China AI model performance gap has effectively closed; as of March 2026 the top U.S. model leads by just 2.7%.
— Chapter 2 · Technical Performance
AI can win a gold medal at the International Mathematical Olympiad but still reads analog clocks correctly only 50.6% of the time, versus 90.1% for humans.
— Chapter 2 · Technical Performance
AI agents advanced from answering questions to completing tasks — yet still fail roughly one in three attempts on structured benchmarks.
— Chapter 2 · Technical Performance
Robots succeed in only 12% of real household tasks even as they reach 89.4% in controlled simulation, while autonomous vehicles hit mass-scale deployment.
— Chapter 2 · Technical Performance

Read the full Technical Performance chapter

Chapter 2 (sections 2.1–2.7) with every benchmark, figure, and citation is free from Stanford HAI. Or head back to the 15 takeaways and nine-chapter overview.

Open Chapter 2 · Technical Performance →