Back to the news portal
AI Risks & SafetyNew analysis today · source 8 October 2026Research paperResearchSource analysisUnited StatesGlobal AI evaluation

Does one AI safety score measure one thing?

Not in this psychometric audit of HarmBench. Responses from 81 models on 398 items were better explained by separate response processes than by one harmful-refusal trait; a three-dimensional model cut held-out log loss from 0.322 to 0.258. The preprint does not rank developers or prove which system is safer.

By The Impact of AI Editorial DeskReleased 9 October 2026 at 14:04 BST6 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1The audit analysed 81 models from HELM Safety v1.17.0 on 398 HarmBench items: 199 standard, 99 contextual and 100 copyright prompts.
  • 2Across five repeated train-test splits, a three-dimensional model beat a one-dimensional model; headline log loss fell from 0.322 to 0.258 and Brier score from 0.096 to 0.078.
  • 3Developer-linked item flags largely disappeared after models were matched on the relevant subscore; the analysis does not prove developers behave identically or rank overall safety.
Key themesAI safetyBenchmark validityRefusal behaviourPsychometricsModel evaluation

Research topic

Whether HarmBench's aggregate score represents one stable harmful-refusal construct across models, prompt types and developers

The Impact of AI research cover asking whether one AI safety score measures one thing, with a conceptual score dial splitting into three lenses; it states 81 models, 398 items and preprint status.
AI-generated editorial illustration. The dial and three lenses are conceptual psychometric motifs, not a real benchmark dashboard or measured chart.

The direct answer: the headline score mixes distinct behaviours

The preprint finds that HarmBench performance is not well described by one stable harmful-refusal trait. A model's response to a plainly harmful request, a contextual request and a copyright prompt appears to involve separable response processes. Combining them into one pass rate can conceal why systems differ and make a change in benchmark composition look like a change in safety.

A three-dimensional item-response model predicted held-out responses better than a unidimensional model across five repeated splits. In the headline comparison, held-out log loss fell from 0.322 to 0.258 and Brier score from 0.096 to 0.078. A seven-domain model was only slightly better on log loss at 0.255. The message is not that three is a universal number; it is that one aggregate number is too coarse for these data.[1]

The Impact Brief · Free

Follow the evidence in ai risks & safety.

Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.

Choose your topics (optional)

One concise, source-linked briefing. Unsubscribe at any time.

What was analysed

The authors began with HELM Safety version 1.17.0, containing results for 87 models. Six with incomplete coverage were excluded, leaving 81. Four datasets plausibly targeted refusal, but three had mean pass rates from 0.92 to 0.94. Those near-ceiling results provided little variation, so the detailed audit focused on HarmBench.

HarmBench contributed 398 items: 199 standard requests, 99 contextual prompts and 100 copyright prompts, also spanning seven harm domains. The pool included systems from many developers, with 21 OpenAI and 11 Anthropic models in the developer comparison. This is a broad snapshot, not a random sample of deployment; versions from the same developer are also related rather than independent observations.[1][2]

How the psychometric test worked

The audit compared a one-dimensional item-response model with exploratory and confirmatory alternatives. The three-dimensional model represents the three response processes, while a seven-dimensional model follows harm-domain labels. Item-response models separate estimated model ability from item difficulty and discrimination, making them useful for testing whether one latent trait plausibly explains the response pattern.

For each of five repetitions, researchers split model-item responses 80% for training and 20% for testing. They used 20 random starts and 2,000 training epochs, then compared held-out log loss and Brier score. The three-dimensional model won in every split. Removing copyright items did not eliminate the issue: a two-dimensional model still beat one dimension, with log loss of 0.278 versus 0.293.[1]

Why predictive fit matters for a safety score

A score is useful when it supports comparisons beyond the exact items observed. If one trait explained harmful refusal, a model placed higher on it should have consistently higher probability of passing after item difficulty is accounted for. Worse held-out prediction indicates that the single-trait account misses systematic structure. Here, prompt scope explains enough structure to improve predictions repeatedly.

The categories also reflect different policy choices and failure costs. Refusing clear physical harm, navigating a contextual prompt and withholding copyrighted text can involve different objectives. An overall average weights those objectives according to the number and difficulty of items present. A leaderboard can move when those weights change even if underlying behaviours do not.[1]

The developer comparison exposes a matching problem

The paper examines differential item functioning between OpenAI and Anthropic models. When matched on one overall score, the Mantel-Haenszel method flagged 13 items and logistic regression flagged 17 as behaving differently by developer. After matching on the relevant scope-specific score, the counts fell to one and two. The apparent developer effect was therefore largely explained by combining models with different profiles into one dimension.

This does not prove the developers' models are equivalent. The analysis includes unequal numbers of related versions, power varies by item, and absence of a flag is not evidence of identical safety. It demonstrates a measurement risk: conditioning on a misspecified total can manufacture item-level differences that shrink when matching respects the benchmark's multidimensional structure.[1]

Practical meaning, funding and limits

A better leaderboard would publish subscale results, item coverage and uncertainty. If a composite remains, its weights should be explicit and justified for a stated use. Purchasers should not use one benchmark score as a certification threshold unless the weighted construct matches their deployment harms. HarmBench still measures responses in a harness, not incidents, interactive bypasses, operator controls or real-world harm.

The unreviewed preprint was presented at the COLM 2026 AIMS workshop. Authors are affiliated with Carnegie Mellon University, Indiana University and Google; the paper reports Google support and three Google-affiliated authors. It binarises a five-point rubric and uses HELM's two language-model judges rather than reproducing the original HarmBench classifier. Confidence would rise through independent replication across graders, benchmark versions and held-out model families, plus evidence connecting subscales with adversarial tests and deployment outcomes.[1][2]

What this means for people

  • One refusal percentage is not a complete statement of safety.
  • Purchasers need category-specific results that match local harms.
  • Developers should disclose weighting and item-level failures.

Global context

The benchmark is used internationally, but definitions of prohibited harm, acceptable assistance and copyright compliance differ across jurisdictions and cultures. Disaggregated results make those choices more visible; they do not remove the need for locally relevant prompts, languages, laws and deployment evidence.

What the evidence does not yet show

  • Unreviewed preprint and workshop paper.
  • One HELM Safety snapshot, focused on HarmBench after three datasets saturated.
  • Convenience sample with related versions from the same developers.
  • Binarised rubric and HELM language-model judges rather than the original classifier pipeline.
  • Better benchmark fit does not establish real-world safety.

What to watch next

  • Independent replication across versions and graders.
  • Leaderboards exposing subscales and uncertainty.
  • Links between subscales, adversarial tests and incidents.
  • Models accounting for related versions.
  • Preregistered construct validation.

Living evidence record

Impact record IAI-0A0JGOM

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent or research support

Present

Record status

Monitoring

Last checked

9 October 2026

Source trail

2 direct sources across 2 source types.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

Evidence trail

Sources used for this report

Links checked 9 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

AI Risks & Safety

What do the OpenAI firings prove?

OpenAI says three safety researchers were dismissed for violating policies on sensitive information; the researchers say they worked within their mandates and warn that the process could chill outside safety collaboration. The public record establishes a serious governance dispute, but not whose account is correct.

9 min · 3 sources

AI Risks & Safety

Can a few poisoned documents reframe medical AI answers?

Yes, in a peer-reviewed laboratory study across three biomedical corpora and five open-weight language models. Injecting five crafted documents per question often put at least three in the top five retrieved results, and one or two retrieved poisons could be enough in successful attacks.

6 min · 1 source

AI Risks & Safety

Can AI agents learn what is appropriate?

Perhaps, but the new Google-led report is a research agenda, not a tested safeguard. More than 50 contributors propose contextual policy engines, dynamic sandboxes and multi-agent benchmarks without demonstrating that they prevent real privacy or security failures.

8 min · 2 sources

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.