The paper applies a construct validity framework to HELM Safety, a popular AI safety benchmark.
HELM Safety's HarmBench fails to measure a single harmful refusal construct.
Psychometric tests reveal that HarmBench's responses require multiple dimensions and are influenced by model developers.
Big Tech
Christopher M. Stewart · Preston Botter · Natalie Sarabosing · Muye Zhang · Rachel Phinnemore · Shalini Ghosh · +2 more
Carnegie Mellon University · Indiana University · Google
Research Digest··3 min read
The authors audited HELM Safety's HarmBench dataset for construct validity of 'harmful refusal'.
Why this paper
From Google and 2 others
In one line
HarmBench scores collapse distinct harm behaviors, not measuring a single harmful refusal attribute.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§