AI Index Report 2023

The year the ethics problems stopped being the researchers’ problem

Chapter 3 of the AI Index 2023 covers 2022 — the year the technical barrier to shipping a generative model collapsed and the ethical questions moved out of the lab and onto social media. The chapter’s throughline is uncomfortable: bigger models are more capable on bias benchmarks and more toxic on others, the fixes that work are training recipes rather than parameter counts, and metrics that sound like they measure the same thing keep disagreeing with each other. The numbers:

260AI incidents and controversies logged by AIAAIC in 2021 — 26× the 2012 count
19AI fairness and bias metrics counted in 2022, rising steadily since 2016
37research papers using the Perspective API in 2022, up 106% in one year
381NeurIPS papers on fairness and bias in 2022 (2021: 168)
89% Winogender accuracy for Flan-PaLM 62B — plain PaLM 62B scores 3.5%
62.5% of popular commercial chatbots that are female by default

3.1–3.2 — Reported misuse is up 26-fold, and the counting is still catching up

The AIAAIC Repository is an independent, open dataset of incidents and controversies driven by AI, algorithms and automation. It started in 2019 as a private project about the reputational risks of AI. By 2021 it was recording 260 new incidents a year.

The number of newly reported AI incidents and controversies in 2021 was 26 times greater than in 2012. The report is careful about what that means: the rise is evidence both of AI becoming more intermeshed with the real world and of growing awareness of how it can be misused. As awareness grew, tracking improved too — which implies older incidents are underreported and the curve is steeper on paper than it was in life. The 2022 figures are missing from the chart entirely, because submissions to AIAAIC go through a lengthy vetting process before they are added.

What the 2022 caseload actually looked like

  • A deepfake video circulated on social media and on a Ukrainian news website in March 2022 appeared to show President Volodymyr Zelenskyy directing his army to surrender to Russia.
  • American prisons were reported in February 2022 to be using AI systems to scan inmates’ phone calls. Voice-to-text systems are known to transcribe Black speakers less accurately, and a large proportion of the incarcerated population in the United States is Black.
  • Intel and the education startup Classroom Technologies were reported in April 2022 to be building a system that identifies the emotional state of students on Zoom — raising the prospect of students being needlessly monitored and having their emotions mischaracterized.
  • London’s Metropolitan Police Service was reported in February 2022 to maintain a Gangs Violence Matrix of over one thousand alleged street gang members, ranked for risk by AI tools. Studies concluded it was inaccurate and discriminated against ethnic and racial minorities; in October 2022 the force announced the list would be drastically reduced.
  • Midjourney was logged as an incident in September 2022 on three counts at once: copyright, because it trains on human-generated images without crediting their source; employment, because of the fear it will replace working artists; and privacy, because the parent company may not have had permission to use the millions of images it trained on.

Meanwhile the measuring instruments multiplied

Counting only metrics that have been cited in at least one other work, the number of AI fairness and bias metrics reached 19 in 2022 and has risen steadily since 2016. The chapter draws a distinction worth keeping: a benchmark contains labeled data, does not change over time and usually measures something intrinsic to the model — StereoSet measures how often a model picks a stereotype over a non-stereotype, but not whether it performs worse for one subgroup than another. A diagnostic metric measures the model’s impact on a downstream task, which is where disparate real-world impact actually shows up. Previous work found that intrinsic and extrinsic metrics for contextualized language models may not correlate with each other at all — so the choice of metric quietly decides the answer. New 2022 entries went both ways: VLStereoSet extended StereoSet into the text-to-image setting, while HolisticBias assembled sentence prompts for demographic biases nothing previously covered.

A 62B model that was instruction-tuned beats a 540B model that wasn’t

Winogender accuracy (%), zero-shot evaluation in the generative setting. Winogender measures gender bias around occupations by testing how often a system fills in a sentence with stereotypical pronouns. Instruction-tuned models outperform models several times their size: Flan-PaLM 62B reaches 89.00% while the identically sized PaLM 62B manages 3.50%.

A 62B model that was instruction-tuned beats a 540B model that wasn’tPaLM 62B: 3.53.5PaLM 62BPaLM 540B: 5.645.64PaLM 540BFlan-PaLM 8B: 72.2572.25Flan-PaLM 8BFlan-T5-XXL: 76.9576.95Flan-T5-XXLFlan-PaLM 62B: 8989Flan-PaLM 62B

3.3 — Every fairness gain seems to cost something somewhere else

This is the longest section of the chapter and the least reassuring. Scrutiny is up sharply — papers using Alphabet’s Perspective API to measure toxicity grew 106% in a year, to 37 in 2022 — but the more carefully models are measured, the more the measurements contradict each other.

Bigger models are better at the bias benchmark

On the Winogender task from SuperGLUE, results reported on PaLM support earlier findings that larger models are simply more capable — despite their higher tendency to generate toxic outputs. PaLM 540B reached 73.58% and Gopher 280B 71.40%, against 57.90% for iPET (ALBERT) at 31M parameters and 50.00% for WARP at 223M. None of them are close to the human baseline of 95.90%. Instruction-tuning is the intervention that actually moves the number: in the generative setting, Flan-PaLM 62B hits 89.00% where PaLM 62B scores 3.50%.

And then the trade-offs show up

  • HELM plotted model accuracy against fairness and bias across its scenarios. More accurate models are more fair — but the correlation between accuracy and gender bias is not clear. Worse, a correlation analysis found that models scoring better on fairness metrics exhibit worse gender bias, and that less gender-biased models tend to be more toxic. Fairness and bias are not the same axis, and improving one can move the other the wrong way.
  • BBQ measures how bias surfaces in question answering, along nine axes including socioeconomic status, religion, disability status and age. Where the context is ambiguous, models are more likely to fall back on a stereotype than to answer “Unknown” — and that result is made worse, not better, in models fine-tuned with reinforcement learning. Most models are biased along physical appearance and age; the picture along race and ethnicity is less clear.
  • In machine translation into English, models consistently perform worse when the correct translation uses “she” rather than “he” — a drop of 2% to 9% across the models measured. They also mistranslate gendered pronouns as “it”, a dehumanizing harm. Instruction-tuning, which helps so much on Winogender, has no measurable effect on mistranslation.
  • RealToxicityPrompts used to produce a clean story: larger models trained on web data were more toxic. HELM’s evaluation shows that trend has become unclear, because companies now apply different pre-training data filtration and different post-training mitigations. Models of the same size can differ substantially in toxicity, small models can turn out surprisingly toxic, and the datasets are too large and too closely guarded to analyze comprehensively.

3.4–3.5 — Text-to-image made bias visible, and chatbots made it personal

Text-to-image models took over social media in 2022, turning fairness and bias into something people could see rather than read about. Tap any card for the full finding.

Stable Diffusion’s CEO is a man in a suit

Hugging Face’s Diffusion Bias Explorer pairs adjectives with occupations. “CEO” returns men in suits almost regardless of the adjective attached.

text-to-image

DALL-E 2 crosses its arms

Prompted with “CEO”, OpenAI’s model returned four images of older, serious-looking men in suits — three of the four with arms crossed authoritatively.

text-to-image

Midjourney’s “influential person”

Asked for an “influential person”, it produced four older white men. Asked for “someone who is intelligent”, four elderly white men in eyeglasses.

text-to-image

VLStereoSet: the more capable model is the more stereotyped one

Across six pre-trained vision-language models, CLIP has the highest vision-language relevance score but exhibits more stereotypical bias than the rest.

benchmarks

Training on Instagram made the model fairer — and no more ethical

On the Casual Conversations dataset, Precision@1 for darker-skinned women rose from 58.2% (ImageNet-supervised) to 90.3% (Instagram 10B SEER).

fairness

Chatbots are disproportionately women

Of 100 conversational AI systems analyzed in mid-2022, 37% were female-gendered — but 62.5% of popular commercial systems were female by default.

chatbots

A third of one dialogue dataset made labelers uncomfortable

Human labelers judged only 67% of PersonaChat examples comfortable for a robot to say, and only 56% possible for a machine to say truthfully.

chatbots

ChatGPT and the dirty bomb

Researcher Matt Korda got detailed bomb-building instructions by role-playing as a safety researcher. The prompt stopped working one day after he published.

safety

3.6 — In Chinese AI ethics research, privacy comes first

Number of papers raising each topic of concern, among 328 Chinese-language AI ethics papers published 2011–2020 on the China National Knowledge Infrastructure platform and annotated by researchers at the University of Turku. Further down the list: freedom 49, unemployment 41, legality 39, transparency 37, autonomy 32. Proposed remedies favor structural reform (71 papers) and legislation (69) over technological solutions (39).

3.6 — In Chinese AI ethics research, privacy comes firstPrivacy: 9999PrivacyEquality: 9595EqualityAgency: 8888AgencyResponsibility: 5858ResponsibilitySecurity: 5050Security

3.7 — AI ethics stopped being a workshop topic

The clearest signal in the chapter is where this research is being published. Accepted submissions to FAccT doubled between 2021 and 2022 and are ten times their 2018 level, and at NeurIPS the ethics topics are migrating out of side workshops and into the main track.

ACM FAccT — the Conference on Fairness, Accountability, and Transparency — was one of the first major venues built to bring researchers, practitioners and policymakers together around sociotechnical analysis of algorithms. Academic institutions still dominate it, with 772 accepted submissions by affiliation in 2022 against 71 in 2018. But industry reached 503, more than ever before, and government-affiliated authors went from 53 in 2021 to 181 in 2022 — evidence that AI ethics has become a working concern for policymakers and practitioners, not only researchers.

At NeurIPS, the topics moved into the main track

  • Fairness and bias more than doubled in a year: 168 accepted papers in 2021 became 381 in 2022, with the workshop stream spiking from 118 to 310. NeurIPS began requiring broader impact statements from authors in 2020, signaling that ethical and societal consequences belong early in the research process.
  • Causal effect and counterfactual reasoning grew more quietly, from 76 papers in 2021 to 80 in 2022 — but the composition flipped. Main-track papers went from 53 to 61 while workshop papers fell from 23 to 19; in 2019 the same topic ran 20 main-track papers against 58 in workshops.
  • Interpretability and explainability is the one topic where the total fell, from 41 papers in 2021 to 24 in 2022. The main track still grew by a third, from 18 to 24 — the report notes that workshop declines can simply reflect year-over-year differences in workshop themes.
  • Privacy in AI peaked at 150 papers in 2020, falling to 128 in 2021 and 103 in 2022, while its main-track presence rose steadily from 12 to 15 to 27. The pattern across all four topics is the same: fewer workshops, more main track — the mark of a subject that has stopped being niche.

The conference on fairness has a geography problem

Share of accepted FAccT submissions by region in 2022 (% of world total). Europe and Central Asia climbed from 18.7% of submissions in 2021 to 30.59% in 2022, driven by European government and academic actors — but FAccT is still broadly dominated by North America and the rest of the Western world. The Middle East and North Africa and Latin America and the Caribbean each account for 0.69%; Sub-Saharan Africa for 0.00%.

The conference on fairness has a geography problemNorth America: 63.2463.24North AmericaEurope: 30.5930.59EuropeEast Asia: 4.254.25East AsiaSouth Asia: 0.550.55South AsiaSub-Saharan: 00Sub-Saharan

3.8 — Automated fact-checking is not as far along as it looks

Significant resources went into building AI systems for fact-checking and misinformation. The last section of the chapter asks whether the benchmarks underneath them describe the real world at all.

What is wrong with the fact-checking datasets?
They quietly assume the answer already exists. Researchers at the Technical University of Darmstadt and IBM analyzed 16 existing fact-checking datasets and found that 11 of them rely on evidence “leaked” from fact-checking reports that did not exist at the time the claim surfaced. A system built on that assumption cannot assign a veracity score to a new claim in real time, which is the only moment that matters. Several datasets also contain claims that fail the criterion of sufficient evidence or counterevidence in a trusted knowledge base.
Why is missing counterevidence such a problem?
Because automated systems assume a false claim will have contradictory evidence sitting somewhere, and new claims usually have neither proof nor contradiction. The chapter’s example: “Half a million sharks could be killed to make the COVID-19 vaccine” has no counterevidence to retrieve — but a human fact-checker can trace it back to the false premise that vaccines rely on shark squalene and rule it false. Language models trained on static snapshots of data, without continual updates, lack exactly the real-world context that makes this possible.
Are researchers still using the old benchmarks?
Citations have plateaued. In 2022, FEVER was cited 236 times, LIAR 191 times and Truth of Varying Shades 99 times — flat compared with previous years. The chapter reads this as a possible shift in the landscape of research on natural language tools for fact-checking on static datasets: the field may be moving away from the static-benchmark framing rather than doubling down on it.
Does making the model bigger make it more truthful?
Not on its own. TruthfulQA asks questions from health, law, finance and politics that humans get wrong because of common misconceptions — asked what happens if you smash a mirror, GPT-3 replied that you will have seven years of bad luck. In 2021, experiments on DeepMind’s Gopher suggested accuracy improves with model size. Stanford researchers then evaluated models from 60 million to 530 billion parameters and found that while large models broadly still beat smaller ones, midsize instruction-tuned models perform surprisingly well: Anthropic’s 52B model and BigScience’s 11B T0pp do disproportionately well for their size, and the best model overall, InstructGPT davinci v2 175B, is also instruction-tuned.
So what is the single lesson of this chapter?
That the training recipe now matters more than the parameter count. Instruction-tuning turns a 3.50% Winogender score into 89.00% at the same model size, moves midsize models up the TruthfulQA table, and can make larger models less toxic than the old scaling story predicted. But it is not a general fix: it has no measurable effect on gendered mistranslation, and models fine-tuned with reinforcement learning fall back on stereotypes in ambiguous BBQ contexts more, not less. Fairness and bias remain separate axes that can be pushed in opposite directions at once.

The chapter in five lines

Headline findings from Chapter 3 · Technical AI Ethics.

According to the AIAAIC database, which tracks incidents related to the ethical misuse of AI, the number of AI incidents and controversies has increased 26 times since 2012.
— Chapter 3 · Technical AI Ethics
While large models are still toxic and biased, new evidence suggests that these issues can be somewhat mitigated after training larger models with instruction-tuning.
— Chapter 3 · Technical AI Ethics
Fairer models may not be less biased: language models which perform better on certain fairness benchmarks tend to have worse gender bias.
— Chapter 3 · Technical AI Ethics
Text-to-image models took over social media in 2022, turning the issues of fairness and bias in AI systems visceral through image form.
— Chapter 3 · Technical AI Ethics
Researchers find that 11 of 16 automated fact-checking datasets rely on evidence “leaked” from fact-checking reports which did not exist at the time the claim surfaced.
— Chapter 3 · Technical AI Ethics

Read Chapter 3 in full

Chapter 3 (sections 3.1–3.8) — fairness and bias metrics, AI incidents, NLP bias, conversational AI, text-to-image models, AI ethics in China, FAccT and NeurIPS trends, and factuality — with every figure and citation is free from Stanford HAI.

Open the AI Index 2023 →