Everyone agrees AI should be safe. Almost nobody measures it the same way
Chapter 3 of the AI Index 2025 finds a responsible AI ecosystem that is evolving unevenly. Incidents are at a record high, academic attention is rising fast, and governments are moving — but model developers still have no shared safety benchmark, and organizations recognize far more risks than they mitigate. The numbers:
3.2 — Everyone reports MMLU. Nobody agrees on a safety test
Major developers consistently test flagship models on the same general capability benchmarks — MMLU, GPQA, AIME. There is no equivalent consensus for safety and responsibility, which makes models genuinely hard to compare on the dimensions that matter most to regulators and buyers.
This is not because developers ignore safety — many run extensive evaluations. The problem is that those evaluations are internal, proprietary and non-standardized, so their results cannot be validated by the wider community. External evaluators such as Gryphon, Apollo Research and METR assess only a selection of models, and their findings cannot be broadly checked either.
New benchmarks are starting to fill the gap
- HELM Safety, from Stanford's Center for Research on Foundation Models, tests models across BBQ (social bias), SimpleSafetyTests (self-harm, physical harm, CSAM), HarmBench (red-teamed harassment, chemical weapons, misinformation), AnthropicRedTeam (adversarial conversations) and XSTest (false refusals of benign prompts). The safest model measured is Claude 3.5 Sonnet at 0.977, just ahead of o1 at 0.976; GPT-3.5 Turbo from 2022 scored 0.853.
- AIR-Bench 2024 grounds safety evaluation in actual regulation: a four-tier taxonomy of 314 granular microrisks derived from eight government regulations and 16 corporate policies. Across 22 leading models, refusal rates ranged from 91% for Anthropic's Claude series to 25% for DBRX Instruct — a spread that points to widespread misalignment with rules like the EU AI Act.
- On hallucination, the Hughes Hallucination Evaluation Model leaderboard has GLM-4-9b-Chat and Gemini-2.0-Flash-Exp tied for the lowest rate at 1.3%, followed by o1-mini (1.4%) and GPT-4o (1.5%).
- On factuality, OpenAI's SimpleQA is answered correctly just 42.7% of the time by its best performer, o1-preview. Some models decline rather than guess: the Claude-3 family refrained from responding to 75% of prompts. Among attempts actually made, o1-preview was correct 47.0% of the time and Claude 3.5 Sonnet 44.5%.
Incidents keep climbing
The AI Incident Database recorded 233 AI-related incidents in 2024, a record high and a 56.4% increase over 2023. Because tracking relies on publicly available media reports, the real number is almost certainly higher. The 2024 cases span the range of harms the chapter is built around: a UK shopper wrongfully identified as a shoplifter by a facial recognition system, then publicly accused and banned from stores; a Texas high school student targeted with AI-generated intimate images made from photos taken from her private account; a chatbot recreating the identity of a murdered teenager without her family's knowledge; and a lawsuit against Character.AI over a teenager's death that alleges the product lacked safeguards for users in distress. In 2024 the community also debated how to define a serious incident at all — no consensus was reached, which is itself part of the problem.
3.3 — Inside companies, responsible AI has no home and no consensus
Two surveys — one run with McKinsey across 30-plus countries, one run by Stanford researchers with Accenture across 1,500 large organizations — paint the same picture. Leaders believe in responsible AI. Almost nothing about how to do it is settled.
Who owns it?
- No single department dominates AI governance. The most common answer was information security, covering cyber, fraud and privacy, at 21%, followed by data and analytics at 17%. Notably, 14% of respondents now report dedicated AI governance roles.
- Investment scales with size: 27% of organizations with $10 billion to $30 billion in revenue, and 21% of those above $30 billion, invest $10 million to $25 million a year in operationalizing responsible AI. Smaller organizations allocate fewer dollars but many still report substantial investment as a share of revenue.
- Regulation is the strongest external driver. 65% of organizations report being influenced by GDPR in their responsible AI decisions, 41% by the EU AI Act, and 21% by the OECD AI Principles.
- The main obstacles are practical rather than political: knowledge and training gaps (51%) and resource or budget constraints lead the list, while only 16% cite a lack of executive support.
- Where policies are in place, 42% of organizations report improved business operations such as greater efficiency and lower costs, and 34% report increased customer trust.
What actually goes wrong
The Global State of Responsible AI survey — its second iteration, covering 1,500 organizations with revenues of at least $500 million across 20 countries and 19 industries, fielded in January and February 2025 — asked what incidents organizations had actually experienced. Adversarial attacks and privacy violations topped the list. More striking, 51% of respondents reported unintended decision making and 47% reported model bias, which suggests that many organizations are struggling to anticipate and control how their AI systems behave. Only 8% of organizations in the McKinsey survey reported experiencing an AI-related incident at all, so these figures describe the ones paying attention.
What changed in a year
- Companies have grown markedly more concerned about financial risks (up 38 percentage points), brand and reputational risks (+16), privacy and data-related risks (+15) and reliability risks (+14) between 2024 and 2025.
- Two categories moved the other way: societal risks fell 7 points and socio-environmental risks fell 8 — the risks that are hardest to price are the ones losing attention.
- On almost every question of philosophy, responses split roughly evenly: whether open- or closed-weight models are safer, whether risk mitigation belongs to model providers or to users, whether agents are too risky for large-scale adoption. The industry has no unified strategic direction.
- The one clear exception is a contradiction: 64% of respondents lean toward a safety-first approach, and yet 58% are already exploring minimally supervised agents — a combination that sits uneasily with the current state of responsible AI maturity.
3.5 — The year governance went multilateral
Where 2023 was a year of national AI strategies, 2024 was a year of coordination. Every major international body published a responsible AI framework, and the first cross-border safety network was formalized.
- May 2024 · OECD
Updated AI principles
The OECD refined its framework to reflect the latest developments in AI governance, emphasizing inclusive growth, transparency and explainability, and respect for the rule of law.
- May 2024 · Council of Europe
The first legally binding AI treaty
The Framework Convention on Artificial Intelligence and Human Rights, Democracy, and the Rule of Law was adopted to ensure that activities across the AI life cycle align with human rights, democracy and the rule of law.
- Jun 2024 · European Union
The EU AI Act
The first comprehensive regulatory framework for AI in a major global economy. It categorizes AI systems by risk and regulates them accordingly, placing most obligations on the providers and developers of high-risk systems.
- Jul 2024 · African Union
The Continental AI Strategy
A unified vision for AI development, ethics and governance across the continent, emphasizing ethical, responsible and equitable development of AI within Africa.
- Sep 2024 · United Nations
Governing AI for Humanity
The UN AI Advisory Body updated its report on establishing global AI governance mechanisms, recommending a blueprint for AI-related risks and calling on standards organizations, technology companies, civil society and policymakers to collaborate.
- Oct 2024 · G7 and ASEAN–US
Open markets and shared standards
The G7 Digital Competition Communiqué reaffirmed commitments to fair and open AI markets and coordinated regulatory approaches. Following the 12th ASEAN–United States Summit, leaders issued a statement on promoting safe, secure and trustworthy AI and committed to cooperating on international governance frameworks and standards.
- Nov 2024 · Nine countries and the EU
The International Network of AI Safety Institutes
The first cross-border network of AI safety institutes was established, uniting technical organizations committed to advancing AI safety, helping governments and societies understand the risks of advanced AI systems, and proposing solutions.
- Feb 2025 · Arab League
The Arab Dialogue Circle on AI
Launched at the Arab League headquarters, the dialogue on Artificial Intelligence in the Arab World focuses on innovative applications while placing strong emphasis on ethical challenges.
3.6–3.8 — The data commons is closing, and bias went underground
Two of the chapter's most consequential findings sit far from the safety headlines: the open web is being fenced off from AI training, and models that pass explicit bias tests keep failing implicit ones.
The web is closing its doors
- A longitudinal audit of consent protocols across C4, RefinedWeb and Dolma found a sharp rise in data use restrictions, enforced mainly through updated robots.txt files and terms of service that explicitly prohibit AI training.
- The proportion of tokens in the top C4 web domains under full restriction rose from 10% in 2017 to 48% in 2024 — with 25 percentage points of that increase arriving between 2023 and 2024 alone. In actively maintained domains, restricted tokens jumped from 5%–7% to 20%–33%.
- Enforcement is inconsistent and uneven: OpenAI's crawlers encounter the highest level of restrictions, while smaller developers face fewer barriers. Signaling mechanisms like robots.txt are ineffective, and stated policies often do not match enforced ones.
- A separate audit of over 1,800 widely used text datasets found more than 70% lacked adequate license information, and 50% of licenses were miscategorized — creating legal and ethical risk for developers who may unknowingly violate copyright or data usage policies.
- The consequences run beyond law. Less public data means less data diversity, weaker model alignment and a harder path to further scaling, since so many recent performance gains have come from training on ever-larger datasets.
Bias that survives the benchmark
In 2024, researchers applied two new detection methods — LLM Implicit Bias, which analyzes automatic associations between words and concepts, and LLM Decision Bias, which captures the behaviors those associations produce — to eight notable models including GPT-4 and Claude 3 Sonnet across 21 stereotype categories. The models disproportionately associated negative terms with Black individuals, more often associated women with the humanities than with STEM, and favored men for leadership roles. Crucially, implicit bias increased as models scaled, even though decision bias and rejection rates did not. Bias metrics have improved on standard benchmarks, creating an illusion of neutrality; the underlying associations have not gone away.
Scaling makes it worse in vision too. Evaluating 14 vision-language models trained on LAION-400M and LAION-2B against the Chicago Face Dataset, researchers found that larger datasets improved human classification — reducing the misidentification of people as nonhuman entities — while simultaneously amplifying racial bias. In the larger ViT-L models, Black and Latino men were disproportionately classified as criminals, with classification probabilities rising by up to 69% as the dataset grew from 400 million to 2 billion samples. The authors advocate transparent dataset curation, detailed hyperparameter documentation and open access for independent audits.
Transparency and the research community
- The Foundation Model Transparency Index v1.1 recorded an average score of 58 out of 100 in May 2024, up from 37 in October 2023, with developers improving on 89 of the 100 indicators. Significant opacity remains in data access, copyright status and downstream impact, and open-source developers outperformed closed-source ones on upstream transparency, particularly on data and labor disclosures.
- Responsible AI papers accepted at six leading conferences rose 28.8%, from 992 in 2023 to 1,278 in 2024. Proportionally, FAccT (69.14%) and AIES (63.33%) led, matching their remits; at NeurIPS the share fell from 13.8% to 9.0%, while at ICML it rose from 3.4% to 8.2%.
- Security and safety submissions nearly doubled in a year, from 276 to 521. Fairness and bias papers roughly doubled to 408. Transparency and explainability reached 355, four times the 2019 figure. Privacy and data governance was the exception, falling 14.5%.
- By country, the United States led with 669 responsible AI papers in 2024, ahead of China (268) and Germany (80). Since 2019 the cumulative totals are the US at 3,158, China at 1,100 and the United Kingdom at 485.
3.9–3.10 — Five uncomfortable findings
The security and special-topics sections are where the chapter is least reassuring, and most specific.
How deep does safety training actually go?
Can agents be trusted with real tools?
What happens when agents start talking to each other?
Did AI misinformation decide the 2024 elections?
Is there anything that actually makes models more robust?
The chapter in five lines
Headline findings from Chapter 3 · Responsible AI.
Reported AI-related incidents rose to 233 in 2024 — a record high and a 56.4% increase over 2023 — and because tracking relies on media reports, the true number is almost certainly higher.
Major developers all test their models on MMLU and GPQA. None of them agree on a single responsible AI benchmark, which makes safety claims almost impossible to compare.
The share of tokens fully restricted from AI training in the top C4 web domains rose from 10% in 2017 to 48% in 2024 — 25 percentage points of that arrived in a single year.
In every risk category, fewer organizations mitigate than recognize: intellectual property infringement is relevant to 57% and actively mitigated by 38%.
Models explicitly trained to be unbiased still show implicit bias — and it increases as models scale, even while the standard benchmark scores improve.
Read the full Responsible AI chapter
Chapter 3 (sections 3.1–3.10) — incidents, benchmarks, organizations, academia, policymaking, privacy, fairness, transparency, security and agents — with every figure and citation is free from Stanford HAI.
Open the AI Index Report 2025 →