Learned plasma control for fusion
DeepMind trained a reinforcement learning agent to find optimal tokamak management procedures for containing hydrogen plasma.
Chapter 2 of the AI Index 2023 measures technical progress during 2022 across vision, language, speech, reinforcement learning and hardware. The headline is not a leap but a plateau: state-of-the-art results kept arriving, yet on most benchmarks they arrived by a hair. What did leap was generative AI — and the cost of training it. The numbers:
The theme running through the whole chapter is saturation. AI kept setting records in 2022, but on most of the tests the AI Index tracks the record was barely better than last year’s — and the speed at which benchmarks reach that ceiling is still increasing.
Measured as relative change since each benchmark launched, the median improvement is 42.4%. Measured over the last year alone, the median is 4%. For all but seven of the benchmarks in this chapter, the year’s improvement came in under 5%. The AI Index also dropped two long-standing favourites, SQuAD1.1 and SQuAD2.0, from this edition entirely — no new state-of-the-art results had been posted on either.
Researchers answered saturation by building evaluations that are harder to finish. In June 2022, 442 authors across 132 institutions launched BIG-bench (Beyond the Imitation Game), a suite of 204 tasks spanning linguistics, childhood development, math, common-sense reasoning, biology, physics, social bias and software development. In November, Stanford researchers released HELM (Holistic Evaluation of Language Models), an attempt to judge language models against unified standards rather than one score at a time. Google’s Imagen team shipped DrawBench alongside the model itself, precisely because the existing text-to-image benchmark was no longer discriminating.
A selection from the chapter’s own timeline of the most significant AI developments of 2022, as chosen by the AI Index Steering Committee.
An AI system that writes computer programs at a competitive level, ranking within the top 54% of participants in a human programming competition — progress on exactly the kind of complex problem-solving AI had traditionally struggled with.
One of the world’s largest language models at 540 billion parameters. PaLM reinforced the prevailing belief of the moment: that performance improves by simply training on more data.
A text-to-image system that creates realistic art and images from written descriptions. Its public release is the moment the AI Index marks as igniting the generative AI craze.
A reinforcement learning agent capable of robotic manipulation, game playing, image captioning and natural language generation in one system — evidence that AI was getting better at generalisation.
To better challenge increasingly capable large language models, 442 authors across 132 institutions launched the Beyond the Imitation Game benchmark: 204 tasks from linguistics and childhood development to physics and software development.
GitHub made Copilot available as a subscription service for individual developers. It turns natural language prompts into coding suggestions across multiple languages; surveys suggest it makes coders more productive and less frustrated. Similar systems include OpenAI’s Codex and Salesforce’s CodeGen.
An open-source text-to-image diffusion model whose weights anyone can use freely. It is trained on existing human-made images and gives no credit or acknowledgment, leaving open questions about the ethical use of image generators.
A large-scale speech recognition system trained on roughly 700,000 hours of audio. It needed neither supervised pre-training nor unsupervised training with fine-tuning, yet performed strongly — further validation of simply scaling up training data.
Holistic Evaluation of Language Models, a new benchmarking approach that judges language models against more unified standards — evidence of the field’s attempt to build transparency around increasingly powerful systems.
A publicly usable chatbot capable of writing university-level essays. Months after launch it reached 100 million monthly active users, making it the fastest-growing consumer application in history — and capping a year in which generative AI became part of the zeitgeist.
While the classification benchmarks flatlined, generation did not. Text-to-image, text-to-video and speech all shipped systems in 2022 that ordinary users could actually touch.
DALL·E 2, Stable Diffusion, Midjourney, Meta’s Make-A-Scene and Google’s Imagen all arrived within months of each other. The AI Index put the same prompt — “a panda playing a piano on a warm evening in Paris” — to DALL·E 2, Stable Diffusion and Midjourney to compare them side by side. On the MS-COCO 256×256 FID-30K benchmark, where a lower Fréchet Inception Distance is better, Imagen leads at 7.27, ahead of Make-A-Scene at 7.55 and DALL·E 2 at 10.39; for scale, AttnGAN scored 35.49 back in 2017. Google shipped the harder DrawBench benchmark alongside Imagen, because text-to-image models had outgrown the old test.
AI has traditionally been strong at narrow tasks and weak at crossing between them. In 2022 that started to break down. Microsoft’s BEiT-3 posted state-of-the-art results across four vision skills and five vision-language skills at once — on NLVR visual reasoning it reached 92.60 against a previous best of 87.00, a 6.44% improvement, the largest of the nine. Google’s PaLI took the top spot on VQA v2 at 84.30%, above the 80.78% human baseline. Both are single systems doing what used to require several.
The scaling recipe reached speech in 2022. Whisper, trained in a weakly supervised way on 700,000 hours of audio, beat wav2vec 2.0 Large across a wide range of English speech recognition benchmarks and outperformed leading translation models on the X→EN subset of CoVoST 2, scoring 29.1 BLEU against MAESTRO’s 25.2. On Kincaid46 its median word error rate of 8.81% beat every commercial ASR service tested. It was not best at everything: on FLEURS language identification, zero-shot Whisper managed 64.5% against mSLAM-CTC’s 77.7%. Meanwhile on the original VoxCeleb speaker recognition dataset, American researchers posted an equal error rate of 0.1%, a 0.28 percentage point improvement on the previous year’s state of the art.
2022 was the year AI stopped only being the subject of research and started doing some. Six cases from the chapter — including two where AI was used to improve AI. Tap any card for the detail.
DeepMind trained a reinforcement learning agent to find optimal tokamak management procedures for containing hydrogen plasma.
A reinforcement learning system that discovered faster algorithms for multiplying matrices — a problem researchers had been chipping at for 50 years.
Nvidia’s PrefixRL agent designs chip circuits smaller and faster than the EDA tools that used to do the job.
Generative models produced novel antibodies zero-shot — one round of generation, no further optimisation.
A reinforcement learning agent cut cooling energy by about 12.7% over a three-month experiment.
Google researchers used one of their language models to improve the reasoning of that very same model.
Two sections that belong together. Training keeps getting faster per dollar, which is exactly why models keep getting bigger — and why this edition is the first to put a carbon figure on them.
Drawing on Luccioni et al., 2022, the chapter compares four large language models on parameters, data centre power usage effectiveness, grid carbon intensity, power consumption and emissions. GPT-3 (175B parameters, 1,287 MWh, PUE 1.10, grid intensity 429 gCO₂eq/kWh) released the most carbon at 502 tonnes: 1.4 times more than Gopher, 7.2 times more than OPT and 20.1 times more than BLOOM. The comparison is imperfect — accounting methodologies for reporting carbon emissions are not standardised — but the relativities are stark. BLOOM, the cleanest of the four at 25 tonnes, still emitted 25 times as much as flying one passenger round trip from New York to San Francisco, 1.4 times what the average American emits in a year, and consumed enough energy to power the average American home for 41 years.
The chapter is unusually blunt about the gap between impressive demos and reliable capability — and about how much the scores themselves are worth.
Five lines from the Chapter Highlights and the section text of Chapter 2 · Technical Performance.
AI continued to post state-of-the-art results, but year-over-year improvement on many benchmarks continues to be marginal. Moreover, the speed at which benchmark saturation is being reached is increasing.
2022 saw the release of text-to-image models like DALL-E 2 and Stable Diffusion, text-to-video systems like Make-A-Video, and chatbots like ChatGPT. Still, these systems can be prone to hallucination, confidently outputting incoherent or untrue responses, making it hard to rely on them for critical applications.
Compared to humans, the large language models performed much worse, suggesting that while they are capable, they lack human reasoning capabilities.
AI models are starting to rapidly accelerate scientific progress and in 2022 were used to aid hydrogen fusion, improve the efficiency of matrix manipulation, and generate new antibodies.
Nvidia used an AI reinforcement learning agent to improve the design of the chips that power AI systems. Similarly, Google recently used one of its language models, PaLM, to suggest ways to improve the very same model. Self-improving AI learning will accelerate AI progress.
Chapter 2 (sections 2.1–2.9) — the 2022 timeline, computer vision for images and video, language, speech, reinforcement learning, hardware, environment and AI for science — with every figure and citation is free from Stanford HAI.
Open the AI Index 2023 →