AI Index Report 2026

Responsible AI infrastructure is growing — but it can't keep pace with deployment

Chapter 3 of the AI Index 2026 tracks responsible AI (RAI) across its many dimensions — safety, fairness, transparency, privacy, and factuality. Safety benchmarks have expanded and more organizations have adopted RAI policies, yet frontier models rarely report RAI results, transparency declined in 2025, and progress in one dimension often costs another. The numbers:

362AI incidents recorded by AIID in 2025 (233 in 2024)
94% highest hallucination rate across 26 models (lowest 22%)
64% GPT-4o accuracy on first-person false beliefs (98.2% on true)
17% growth in AI-specific governance roles in 2025
11% of organizations with no RAI policy in 2025 (24% in 2024)
40Foundation Model Transparency Index average in 2025 (58 in 2024)

3.2 — Incidents are rising, and models can't tell knowledge from belief

Two signals from assessing RAI: real-world harms keep climbing, and even where evaluation is maturing — factuality — the failure modes are getting stranger.

Documented AI incidents continued to rise. The AI Incident Database (AIID), an open repository for cases where AI systems caused or nearly caused harm, recorded 362 incidents in 2025, up from 233 in 2024 — the annual count had stayed under 100 until 2022. High-profile 2025 cases included xAI's Grok generating antisemitic hate speech after a system update relaxed its safety filters, deepfake romance scams impersonating actor Jin Dong, and AI-cloned phishing websites mimicking a bankrupt retailer.

Reporting on RAI benchmarks stays sparse

Almost all frontier developers report capability benchmarks like MMLU, GPQA, AIME, and SWE-bench Verified — but reporting on RAI benchmarks (BBQ for bias, HarmBench/Cybench/StrongREJECT/WMDP for security, SimpleQA for factuality) remains mostly empty. Only Claude Opus 4.5 reports results on more than two RAI benchmarks, and only GPT-5.2 reports StrongREJECT.

Belief vs. fact: performance collapses

  • On the AA-Omniscience benchmark (6,000 questions, six domains), hallucination rates across 26 models ranged from 22% (Grok 4.20 Beta) to 94% (gpt-oss-20B). The HHEM document-summarization leaderboard, a different scale, showed top-15 models hallucinating 1.8%–5.4% of the time.
  • KaBLE, a new benchmark testing whether models can tell knowledge from belief, evaluated 24 models on 13,000 questions. GPT-4o scored 98.2% on true beliefs but only 64.4% on first-person false beliefs; DeepSeek R1 fell from over 90% to 14.4%.
  • Models handle third-person false beliefs far better than first-person ones: newer models reach 95% on third-person but only 62.6% on first-person false beliefs. When a false statement is framed as something the user themselves believes, performance breaks down.

3.3 — Organizations are formalizing RAI, but gaps slow adoption

Drawing on a second-year AI Index × McKinsey survey of business leaders (excluding China), RAI maturity, governance ownership, and barriers all shifted between 2024 and 2025.

Responsible AI maturity improved across all regions but remains early-stage: the global average rose from 2.0 to 2.3 on a four-point scale, meaning most organizations are still integrating RAI practices rather than running them fully. AI governance ownership shifted toward dedicated AI-specific roles (up from 14% to 17%), while the share of organizations with no RAI policy fell sharply from 24% to 11%.

The top obstacles

  • Knowledge and training gaps are the most-cited obstacle to implementing RAI, rising from 51% to 59% in 2025.
  • Resource or budget constraints (48%) and regulatory uncertainty (41%) remained among the top barriers, with technical limitations climbing from 32% to 38%.
  • For scaling agentic AI specifically, security and risk concerns dominated at 62% — far ahead of technical limitations (38%) and regulatory uncertainty (38%).

Documented AI incidents keep climbing

Annual AI incidents recorded by the AI Incident Database (AIID). Counts stayed under 100 until 2022, then accelerated as AI deployment spread. Unit: number of incidents.

Documented AI incidents keep climbing2024: 23323320242025: 3623622025

Which regulations shape RAI practices (2025)

Share of organizations naming each regulation as an influence on responsible-AI decisions. GDPR leads but slipped from 65% to 60%; ISO/IEC 42001 and NIST AI RMF are new 2025 entries. Unit: % of organizations.

Which regulations shape RAI practices (2025)GDPR: 6060GDPREU AI Act: 4343EU AI ActISO/IEC 42001: 3636ISO/IEC 42001NIST AI RMF: 3333NIST AI RMFOECD AI Principles: 1616OECD AI Principles

RAI research, by subtopic (2025)

Responsible-AI papers accepted at six leading conferences (AAAI, AIES, FAccT, ICML, ICLR, NeurIPS), up 19% overall to 1,521. Security and safety is now the largest and fastest-growing area. Unit: accepted papers.

RAI research, by subtopic (2025)Security & safety: 641641Security & safetyFairness & bias: 462462Fairness & biasTransparency & explainability: 405405Transparency & explainabilityPrivacy & data governance: 248248Privacy & data governance

The dimensions of responsible AI

Chapter 3 tracks RAI across many dimensions, each with its own measurement challenge. Tap any card for the full picture and its numbers.

Factuality & truthfulness

Hallucination rates span 22%–94% across 26 models; GPT-4o drops to 64.4% on first-person false beliefs.

factuality

AI incidents

AIID logged 362 incidents in 2025, up from 233 — hate speech, deepfake scams, and AI-cloned phishing sites.

safety

Governance & accountability

AI-specific governance roles grew 17%; organizations with no RAI policy fell from 24% to 11%.

governance

Fairness & the global language gap

Models perform far better in English; on HELM Arabic, a regional model beat GPT-5.1 and Gemini.

fairness

Transparency

Foundation Model Transparency Index average fell from 58 to 40; IBM led at 95, xAI scored just 14.

transparency

Security & safety

Safety institutes spread to more countries; on HELM Safety, models cluster at 0.90–0.98 but break under jailbreaks.

safety

Tradeoffs across dimensions

Optimizing one RAI dimension can degrade another — differential privacy cut accuracy by up to 33 points.

tradeoffs

The chapter in five lines

Headline findings from Chapter 3 · Responsible AI.

Documented AI incidents continued to rise, with the AI Incident Database recording 362 in 2025, up from 233 in 2024.
— Chapter 3 · Responsible AI
Across a new accuracy benchmark, hallucination rates range from 22% to 94% — and models still struggle to tell knowledge from belief.
— Chapter 3 · Responsible AI
AI-specific governance roles grew 17% in 2025, and the share of businesses with no responsible AI policy fell from 24% to 11%.
— Chapter 3 · Responsible AI
Foundation model transparency declined in 2025 — the FMTI average dropped from 58 to 40 — after improving the previous year.
— Chapter 3 · Responsible AI
Improving one responsible AI dimension can come at the cost of another: gains in privacy cut accuracy by up to 33 points.
— Chapter 3 · Responsible AI

Read the full Responsible AI chapter

Chapter 3 (sections 3.1–3.10) with every figure and citation is free from Stanford HAI — covering incidents, factuality, governance, fairness, transparency, safety, and the trade-offs between them. Or head back to the 15 takeaways and nine-chapter overview.

Open Chapter 3 · Responsible AI →