SEE MATPROOF ON YOUR STACK — BOOK A 30-MINUTE DEMO
AI_SAFETYarxiv_cscr10 Aug 2026

arXiv: Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks

AI_SAFETY. Sourced from arxiv_cscr, summarised by Matproof.

AI Analysis

What changed and what to do.

A new research paper, published on arXiv, challenges the reliability of internal harmfulness scores used to evaluate AI safety. The study demonstrates that these scores, which are often used to rank how dangerous a model's outputs are, can be systematically manipulated. Specifically, the authors found that successful jailbreaks—prompts designed to bypass safety filters—frequently receive low harmfulness scores, meaning the internal measurement system fails to flag them as dangerous. This indicates that relying solely on these scores for safety assurance is fundamentally flawed.

This finding directly impacts any organization deploying or developing large language models, particularly those in regulated sectors like finance, healthcare, and legal services, where compliance with AI safety standards is critical. It also affects cloud providers and AI vendors who use these internal metrics to certify their models for enterprise use. The paper suggests that current evaluation frameworks, which may be used to demonstrate compliance with emerging EU AI Act requirements, could be providing a false sense of security.

Compliance teams should immediately review their AI risk assessment procedures to determine if they depend on internal harmfulness scores. They should treat these scores as a supplementary signal, not a primary safety guarantee. The next step is to implement additional, independent testing methods, such as red-teaming with diverse jailbreak attempts and human review of edge cases. This will help ensure that safety evaluations are robust and that the organization is not unknowingly deploying models with exploitable vulnerabilities.

This summary is AI-generated for orientation purposes. For regulatory action, always consult the original source linked above.

More AI_SAFETY updates

Latest in AI_SAFETY.

Live regulatory monitoring

Never miss a compliance update.

Get weekly digests of DORA, NIS2, GDPR, MaRisk, and ISO 27001 changes — straight to your inbox. Free.

No spam. Weekly digest only. Unsubscribe anytime.

DORANIS2GDPRMaRiskISO 27001

Map this to your controls

Connect regulatory changes to your compliance work.

Matproof maps every regulator update directly to your controls and surfaces the ones that affect your organisation — across 21 frameworks.