Factuality & truthfulness
Hallucination rates span 22%–94% across 26 models; GPT-4o drops to 64.4% on first-person false beliefs.
Chapter 3 of the AI Index 2026 tracks responsible AI (RAI) across its many dimensions — safety, fairness, transparency, privacy, and factuality. Safety benchmarks have expanded and more organizations have adopted RAI policies, yet frontier models rarely report RAI results, transparency declined in 2025, and progress in one dimension often costs another. The numbers:
Two signals from assessing RAI: real-world harms keep climbing, and even where evaluation is maturing — factuality — the failure modes are getting stranger.
Documented AI incidents continued to rise. The AI Incident Database (AIID), an open repository for cases where AI systems caused or nearly caused harm, recorded 362 incidents in 2025, up from 233 in 2024 — the annual count had stayed under 100 until 2022. High-profile 2025 cases included xAI's Grok generating antisemitic hate speech after a system update relaxed its safety filters, deepfake romance scams impersonating actor Jin Dong, and AI-cloned phishing websites mimicking a bankrupt retailer.
Almost all frontier developers report capability benchmarks like MMLU, GPQA, AIME, and SWE-bench Verified — but reporting on RAI benchmarks (BBQ for bias, HarmBench/Cybench/StrongREJECT/WMDP for security, SimpleQA for factuality) remains mostly empty. Only Claude Opus 4.5 reports results on more than two RAI benchmarks, and only GPT-5.2 reports StrongREJECT.
Drawing on a second-year AI Index × McKinsey survey of business leaders (excluding China), RAI maturity, governance ownership, and barriers all shifted between 2024 and 2025.
Responsible AI maturity improved across all regions but remains early-stage: the global average rose from 2.0 to 2.3 on a four-point scale, meaning most organizations are still integrating RAI practices rather than running them fully. AI governance ownership shifted toward dedicated AI-specific roles (up from 14% to 17%), while the share of organizations with no RAI policy fell sharply from 24% to 11%.
Chapter 3 tracks RAI across many dimensions, each with its own measurement challenge. Tap any card for the full picture and its numbers.
Hallucination rates span 22%–94% across 26 models; GPT-4o drops to 64.4% on first-person false beliefs.
AIID logged 362 incidents in 2025, up from 233 — hate speech, deepfake scams, and AI-cloned phishing sites.
AI-specific governance roles grew 17%; organizations with no RAI policy fell from 24% to 11%.
Models perform far better in English; on HELM Arabic, a regional model beat GPT-5.1 and Gemini.
Foundation Model Transparency Index average fell from 58 to 40; IBM led at 95, xAI scored just 14.
Safety institutes spread to more countries; on HELM Safety, models cluster at 0.90–0.98 but break under jailbreaks.
Optimizing one RAI dimension can degrade another — differential privacy cut accuracy by up to 33 points.
Headline findings from Chapter 3 · Responsible AI.
Documented AI incidents continued to rise, with the AI Incident Database recording 362 in 2025, up from 233 in 2024.
Across a new accuracy benchmark, hallucination rates range from 22% to 94% — and models still struggle to tell knowledge from belief.
AI-specific governance roles grew 17% in 2025, and the share of businesses with no responsible AI policy fell from 24% to 11%.
Foundation model transparency declined in 2025 — the FMTI average dropped from 58 to 40 — after improving the previous year.
Improving one responsible AI dimension can come at the cost of another: gains in privacy cut accuracy by up to 33 points.
Chapter 3 (sections 3.1–3.10) with every figure and citation is free from Stanford HAI — covering incidents, factuality, governance, fairness, transparency, safety, and the trade-offs between them. Or head back to the 15 takeaways and nine-chapter overview.
Open Chapter 3 · Responsible AI →