GDI-H3 · Atari-57
Nearly doubled MuZero’s score on the 57-game Atari suite while using a hundredth of the training frames.
Chapter 2 of the AI Index 2022 measures technical progress across vision, language, speech, recommendation, reinforcement learning, hardware and robotics. The story of the year is not one breakthrough but a pattern: 9 of the 10 benchmarks tracked here had a 2021 state of the art that was pretrained on extra data, which quietly hands the advantage to whoever owns the largest datasets. Meanwhile training got radically cheaper and reinforcement learning started to generalize. The numbers:
The chapter’s own first highlight is one word repeated three times. Across the ten benchmarks tracked here, nine of the 2021 state-of-the-art systems were pretrained on data beyond the benchmark’s own training set — a trend the AI Index says implicitly favors private sector actors with access to vast datasets.
That matters because it changes what a leaderboard measures. A score that depends on pretraining is partly a measurement of who could assemble the pretraining corpus, not only of who designed the better architecture. The chapter makes the point mostly by showing the two lines side by side: on almost every benchmark it plots, there is now a “with extra training data” curve sitting above a “without extra training data” one.
Text summarization on arXiv is the exception in the chapter: the leading ROUGE-1 score without extra training data, 47.15, sits slightly above the 46.74 achieved with it. Over the five years since benchmarking on arXiv began, summarization models have improved 47.1% — and the curve is visibly flattening. PubMed tells the same story from the other side: 48.25 with extra data against 47.81 without, a 34.6% improvement since 2017 whose pace has slowed. The 2021 leader there was HAT, a hierarchical attention transformer from Birch AI and the University of Washington.
AI now beats the human baseline on SuperGLUE and both SQuAD sets by 1% to 5%. Put a logical or abductive step in front of the same passage and the ranking flips back.
Image classification, pose estimation and polyp segmentation are all within a point or two of their ceilings. The interesting movement in 2021 was not upward but sideways — into narrower, more clinical, more deployable subtasks.
The back half of the chapter is about deployment economics: how general an agent has become, what it costs to train a system, how long that takes, and what a lab has to pay to put a robot arm on a bench.
New this edition, the AI Index ran its own survey of robotics professors at top-ranked universities and in emerging economies: 101 professors and researchers from more than 40 universities, covering 117 robotic arm purchases between 2016 and 2022. The median price fell 46.2% in five years, from $42,000 per arm in 2017 to roughly $22,600 in 2021 — robotics research is becoming materially more accessible. Asked which AI skills they employ, 67.0% of respondents reported deep learning and 46.0% reinforcement learning.
Six systems that ended the year holding a state of the art — and what each result actually demonstrates.
Nearly doubled MuZero’s score on the 57-game Atari suite while using a hundredth of the training frames.
Scores 91.0 on the benchmark built in 2019 to be hard — 1.2 points above the human score its own designers set.
Segments colonoscopy polyps at 94.20% on CVC-ClinicDB and 92.17% on Kvasir-SEG — the year’s clearest case of AI moving toward the clinic.
79.50% mean average precision on COCO, 23.8 points above 2016 — and a textbook case of the extra-data trend.
The first model to top Kinetics-400, 600 and 700 at once, at 89.1%, 89.6% and 82.20%.
A 2.0% word error rate on the noisy half of LibriSpeech — the clean half saw no new record at all in 2021.
Five lines from Chapter 2 that carry the year’s argument.
As of 2021, 9 state-of-the-art AI systems out of the 10 benchmarks in this report are trained with extra data. This trend implicitly favors private sector actors with access to vast datasets.
Humans performed 9 percentage points better on aNLI in 2019. As of 2021, that gap has shrunk to 1.
The top chess software engine now exceeds Magnus Carlsen’s top ELO score by 24%. However, in the last two years AI systems have also improved by 129% on more general reinforcement learning tasks (Procgen) in which they must operate in novel environments.
The fact that progress on SuperGLUE was achieved so rapidly suggests that researchers will need to develop more complex suites of natural language tasks to challenge the next generation of AI systems.
An AI Index survey shows that the median price of robotic arms has decreased by 46.2% in the past five years — from $42,000 per arm in 2017 to $22,600 in 2021. Robotics research has become more accessible and affordable.
Chapter 2 (sections 2.1–2.8) — computer vision in images and video, language, speech, recommendation, reinforcement learning, hardware and robotics — with every chart, footnote and citation is free from Stanford HAI.
Open the AI Index 2022 →