AI Index Report 2022

The best scores of 2021 all came with the same asterisk — extra training data

Chapter 2 of the AI Index 2022 measures technical progress across vision, language, speech, recommendation, reinforcement learning, hardware and robotics. The story of the year is not one breakthrough but a pattern: 9 of the 10 benchmarks tracked here had a 2021 state of the art that was pretrained on extra data, which quietly hands the advantage to whoever owns the largest datasets. Meanwhile training got radically cheaper and reinforcement learning started to generalize. The numbers:

9of the 10 benchmarks in this chapter whose 2021 state of the art used extra training data
1percentage points humans still lead AI by on aNLI (9 points in 2019)
128.6% improvement on Procgen, a general reinforcement learning benchmark, over its 2019 baseline
24.3% by which the top chess engine now exceeds Magnus Carlsen’s peak 2,882 Elo
4.59dollars to train an ImageNet classifier to 93% accuracy in 2021 ($1,112.6 in 2017)
22,600dollar median price of a robotic arm in 2021 ($42,000 in 2017)

Data, data, data — nine of ten leaderboards now run on it

The chapter’s own first highlight is one word repeated three times. Across the ten benchmarks tracked here, nine of the 2021 state-of-the-art systems were pretrained on data beyond the benchmark’s own training set — a trend the AI Index says implicitly favors private sector actors with access to vast datasets.

That matters because it changes what a leaderboard measures. A score that depends on pretraining is partly a measurement of who could assemble the pretraining corpus, not only of who designed the better architecture. The chapter makes the point mostly by showing the two lines side by side: on almost every benchmark it plots, there is now a “with extra training data” curve sitting above a “without extra training data” one.

What the extra data buys

  • ImageNet top-1 accuracy: 90.88% with extra training data against 87.80% without.
  • COCO object detection: 79.50% mean average precision with extra data against 77.10% without.
  • Cityscapes pixel-level semantic labeling: 86.20% mean IoU with extra data against 84.30% without.
  • WMT 2014 translation: 46.40 BLEU on English-French and 35.14 on English-German with extra data, against 43.95 and 31.26 without.
  • LibriSpeech Test Clean: a 1.4% word error rate with extra data against 1.7% without — for every 100 words heard, the top model transcribes 99 correctly.

The one place it does not help

Text summarization on arXiv is the exception in the chapter: the leading ROUGE-1 score without extra training data, 47.15, sits slightly above the 46.74 achieved with it. Over the five years since benchmarking on arXiv began, summarization models have improved 47.1% — and the curve is visibly flattening. PubMed tells the same story from the other side: 48.25 with extra data against 47.81 without, a 34.6% improvement since 2017 whose pace has slowed. The 2021 leader there was HAT, a hierarchical attention transformer from Birch AI and the University of Washington.

2021’s best scores, on the five benchmarks that publish a human baseline

Top score reported in 2021. The human baselines, in the same order, are 89.8, 91.2, 92.9, 80.8 and 85.0 — AI is 1 to 5 points clear on straightforward reading comprehension, and still behind everywhere the task also demands abduction or commonsense reasoning.

2021’s best scores, on the five benchmarks that publish a human baselineSuperGLUE: 9191SuperGLUESQuAD 1.1: 95.7295.72SQuAD 1.1aNLI: 91.8791.87aNLIVQA: 79.7879.78VQAVCR: 7272VCR

2.3–2.4 — Reading comprehension is finished; reasoning about what was read is not

AI now beats the human baseline on SuperGLUE and both SQuAD sets by 1% to 5%. Put a logical or abductive step in front of the same passage and the ranking flips back.

English language understanding

  • SuperGLUE was released in May 2019 because systems had begun saturating GLUE. It is already topped out: the SS-MoE model scores 91.0 against the 89.8 human score set by the benchmark’s own developers. The chapter’s reading is that a harder suite is needed again, and needed soon.
  • SQuAD — 107,785 question-and-answer pairs drawn from 536 Wikipedia articles — stands at 95.72 F1 on version 1.1 and 93.21 on version 2.0, above human baselines of 91.20 and 89.50. Both are marginal gains on the previous year: 0.4% and 0.2%.
  • ReClor, built from LSAT logical-reasoning questions, is where the plateau breaks. The best model answers 91.82% of the easy set and 69.29% of the hard set — a 22.5-point drop for the same reading task with a reasoning requirement attached.
  • aNLI, the Allen Institute’s 170,000-pair abductive inference benchmark, sits at 91.87% against a human baseline of 92.90%. AI has gained 7.7 points since 2019, and the human lead has fallen from 9 points to roughly 1.
  • The easier language tasks are done. SNLI, around 600,000 labeled sentence pairs, was topped in April 2021 by Facebook AI’s EFL at 93.10%. On SemEval 2014 sentiment — 7,686 restaurant and laptop reviews — the state of the art is 88.64%, up from correct estimates 7 times in 10 in 2016 to 9 in 10 now.

Translation and speech

  • On WMT 2014 the best BLEU scores are 46.40 for English-French and 35.14 for English-German. Since submissions began that is a 23.7% improvement on the French pair and 68.1% on the German one — the harder pair is closing fast, but absolute translation quality remains meaningfully higher on French.
  • Access is widening as well. The number of commercial machine translation services has risen nearly fivefold since 2017, and 2021 added three open-source ones: M2M-100, mBART and OPUS.
  • Speech recognition has effectively plateaued on its main benchmark. No new state of the art was set on LibriSpeech Test Clean in 2021 because the top system was already at a 1.4% word error rate. On the noisier Test Other set, W2V-BERT — an MIT and Google Brain collaboration — posted 2.0%.
  • Speaker recognition improved faster. On VoxCeleb, systems that reported an equal error rate of 7.8% in 2017 now report 0.42%.

2.1–2.2 — Vision is running out of headroom, so research moved sideways

Image classification, pose estimation and polyp segmentation are all within a point or two of their ceilings. The interesting movement in 2021 was not upward but sideways — into narrower, more clinical, more deployable subtasks.

Images

  • ImageNet top-1 accuracy is 90.88%: in late 2021 the leading system makes on average one error every ten attempts, against four in ten in late 2012. Top-5 accuracy is 99.02%, far past the 94.90% human baseline, and Microsoft’s Florence-CoSwin-H reached 99.0% in November 2021. The chapter is blunt about what follows — if a system is right 98 or 99 times in 100, there is only so much higher it can go.
  • Image generation is measured by Fréchet Inception Distance, where zero would mean the generated images are identical to the real ones. The 2021 state of the art is 7.71 on STL-10, from KAIST and the University of Seoul, and 2.10 on the lower-resolution CIFAR-10, from NVIDIA.
  • Deepfake detection kept pace with deepfake generation. Averaged across the four FaceForensics++ datasets, accuracy rose from 69.9% in 2012 to 97.7% in 2021. Celeb-DF — 590 celebrity videos turned into 5,639 deepfakes — is about 20 points harder, with a top AUC of 76.88.
  • Pose estimation is saturating. The best model gets 99.50% of keypoints right on Leeds Sports Poses, out of a possible 100.0%. In three dimensions, average per-joint error on Human3.6M fell from 16 centimeters in 2014 — half a school ruler — to 1.9 centimeters in 2021, less than a paper clip.
  • Face recognition: in 2017 some leading FRVT algorithms had error rates above 50.0% on certain tests; by 2021 none exceeded 3.0%, and the best result across all datasets, on visa photos, errs once in a thousand faces. Masks still cost accuracy — a false non-match rate of 0.014 masked against 0.002 unmasked — but that gap has narrowed since 2019, and the new MLFW dataset of 6,000 masked faces, from the Beijing University of Posts and Telecommunications, puts the penalty at 5 to 16 points. 18 of 24 US government agencies already use some kind of facial recognition technology.
  • Visual reasoning is the exception in this section. On the VQA challenge the top score is 79.78%, still short of the 80.80% human baseline, though up 24.4 points from 55.4% in 2015.

Video

  • One model now tops all three Kinetics activity-recognition datasets: MTV, from Google Research, Michigan State University and Brown University, released in January 2022, at 89.1% on the 400 series, 89.6% on the 600 and 82.20% on the 700. More striking is the convergence — the gap between the easiest and hardest set fell from 27.14 points in 2020 to 7.4 in 2021, meaning the harder dataset is improving faster than the easier one.
  • On COCO object detection, GLIP reaches 79.50% mean average precision, 23.8 points better than in 2016.
  • YOLO, which deliberately trades accuracy for inference speed, reached 72.40% — 28.4 points better than 2017 — and its gap to the outright best detector narrowed from 11.7% to 7.1%. Detectors got faster and better at the same time.
  • Temporal action localization on ActivityNet, which requires finding when an activity happens as well as what it is, stands at 44.67% from HUST-Alibaba: 26.9 points above 2016, with the annual gains shrinking every year.
  • Visual Commonsense Reasoning — 290,000 multiple-choice questions from 110,000 movie scenarios, where a system must pick the answer and the rationale behind it — is the widest human-AI gap in the chapter. The best mark is 72.0 against an 85.0 human baseline, a 63.6% rise since 2018 that is now yielding increasingly marginal improvements.

One narrow medical benchmark, three years of attention

Academic papers testing systems against Kvasir-SEG, a dataset of 1,000 gastrointestinal polyp images segmented by doctors. The chapter reads the jump as evidence that AI research is moving toward work with direct, real-world applications.

One narrow medical benchmark, three years of attentionBefore 2020: 33Before 20202020: 6620202021: 25252021

2.5–2.8 — Broader agents, cheaper training, more affordable arms

The back half of the chapter is about deployment economics: how general an agent has become, what it costs to train a system, how long that takes, and what a lab has to pay to put a robot arm on a bench.

From narrow skill toward general skill

  • On Atari-57, DeepMind’s MuZero set the state of the art in late 2019 by performing 48.3% better than the previous best. In 2021 GDI-H3, from Tsinghua University and ByteDance, surpassed and nearly doubled it — using 200 million training frames against MuZero’s 20 billion. Twice as effective, and a hundred times more efficient.
  • Procgen is the test that matters for generality: 16 procedurally generated environments introduced by OpenAI in 2019 specifically to punish systems that had learned one narrow skill. MuZero posted 0.6 in November 2021, a 128.6% improvement over the 2019 baseline.
  • Chess is the opposite case — a single narrow skill, tracked for decades. The top engine now sits at 3,581 Elo, 24.3% above Magnus Carlsen’s 2,882, the highest human rating ever documented, recorded in 2014.
  • Recommendation barely moved. The best MovieLens 20M result is an nDCG of 0.448, 5.2% better than 2018, and the best Criteo click-through AUC is 0.813, 1.8% above the 2016 leader. The chapter adds a caveat worth keeping: most recommendation research happens inside companies with every incentive to keep it proprietary, so these academic measures may not reflect the real state of the art.

Training got cheap

  • MLPerf image classification training time fell from 6.2 minutes in 2018 to 0.2 minutes — 13.8 seconds — in 2021, roughly a 27-fold improvement. Recommendation, light-weight object detection, image classification and language processing can all now be trained to baseline performance in under a minute.
  • Training an ImageNet classifier to 93% accuracy cost $1,112.6 in 2017 and $4.59 in 2021: a factor of 223 in four years. Measured from 2018, when the MLPerf competitions began, cost is down 63.6% and training times have improved 94.4%.
  • That speed is bought with hardware, and the hardware is concentrating. The maximum number of accelerators used rose roughly sevenfold since 2018, to 4,320, while the mean across all entrants rose 3.5 times, to 337. The top system averaged 1,785 accelerators — and the gap between the top systems and the field was nine times larger at the end of 2021 than it had been in 2018.

Robot arms got cheaper

New this edition, the AI Index ran its own survey of robotics professors at top-ranked universities and in emerging economies: 101 professors and researchers from more than 40 universities, covering 117 robotic arm purchases between 2016 and 2022. The median price fell 46.2% in five years, from $42,000 per arm in 2017 to roughly $22,600 in 2021 — robotics research is becoming materially more accessible. Asked which AI skills they employ, 67.0% of respondents reported deep learning and 46.0% reinforcement learning.

How long the fastest system takes to train, by task

Wall-clock training time in minutes for the top MLPerf system in 2021; object detection here is the heavy-weight variant, the light-weight one finishes in 0.34. Reinforcement learning is the outlier — the one category where no faster time was registered in either 2020 or 2021.

How long the fastest system takes to train, by taskRL: 13.5713.57RLDetection: 3.243.24DetectionSpeech: 2.382.38SpeechSegmentation: 1.261.26SegmentationImage class.: 0.230.23Image class.

Who held a record at the end of 2021

Six systems that ended the year holding a state of the art — and what each result actually demonstrates.

GDI-H3 · Atari-57

Nearly doubled MuZero’s score on the 57-game Atari suite while using a hundredth of the training frames.

reinforcementefficiency

SS-MoE · SuperGLUE

Scores 91.0 on the benchmark built in 2019 to be hard — 1.2 points above the human score its own designers set.

languagebenchmarks

MSRF-Net · medical segmentation

Segments colonoscopy polyps at 94.20% on CVC-ClinicDB and 92.17% on Kvasir-SEG — the year’s clearest case of AI moving toward the clinic.

visionmedicine

GLIP · COCO object detection

79.50% mean average precision on COCO, 23.8 points above 2016 — and a textbook case of the extra-data trend.

visionbenchmarks

MTV · all three Kinetics sets

The first model to top Kinetics-400, 600 and 700 at once, at 89.1%, 89.6% and 82.20%.

video

W2V-BERT · LibriSpeech Test Other

A 2.0% word error rate on the noisy half of LibriSpeech — the clean half saw no new record at all in 2021.

speech

The chapter in its own words

Five lines from Chapter 2 that carry the year’s argument.

As of 2021, 9 state-of-the-art AI systems out of the 10 benchmarks in this report are trained with extra data. This trend implicitly favors private sector actors with access to vast datasets.
— Chapter 2 · Chapter Highlights
Humans performed 9 percentage points better on aNLI in 2019. As of 2021, that gap has shrunk to 1.
— Chapter 2 · Chapter Highlights
The top chess software engine now exceeds Magnus Carlsen’s top ELO score by 24%. However, in the last two years AI systems have also improved by 129% on more general reinforcement learning tasks (Procgen) in which they must operate in novel environments.
— Chapter 2 · Chapter Highlights
The fact that progress on SuperGLUE was achieved so rapidly suggests that researchers will need to develop more complex suites of natural language tasks to challenge the next generation of AI systems.
— Chapter 2 · 2.3 Language
An AI Index survey shows that the median price of robotic arms has decreased by 46.2% in the past five years — from $42,000 per arm in 2017 to $22,600 in 2021. Robotics research has become more accessible and affordable.
— Chapter 2 · Chapter Highlights

Read Chapter 2 in full

Chapter 2 (sections 2.1–2.8) — computer vision in images and video, language, speech, recommendation, reinforcement learning, hardware and robotics — with every chart, footnote and citation is free from Stanford HAI.

Open the AI Index 2022 →