SEE MATPROOF ON YOUR STACK — BOOK A 30-MINUTE DEMO
AI_SAFETYarxiv_cscr10 Aug 2026

arXiv: From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts

AI_SAFETY. Sourced from arxiv_cscr, summarised by Matproof.

AI Analysis

What changed and what to do.

This paper, published on arXiv, presents an independent reproducibility study of large language model (LLM) and agent-driven tools designed to automatically validate software vulnerabilities. The study found that while these AI systems can generate plausible security artifacts, their outputs are often not directly runnable or verifiable in real-world environments, with a significant portion failing to reproduce the claimed vulnerability validation results. The core change is a documented evidence base showing that current AI-driven security tooling lacks the reliability and traceability required for formal compliance evidence.

The primary audience is organizations in regulated sectors that rely on automated security testing, including financial services, healthcare, critical infrastructure, and any software vendor subject to EU cybersecurity frameworks like the Cyber Resilience Act (CRA) or NIS2. Compliance teams and DevSecOps leaders who are considering or already using LLM-based vulnerability scanners must treat their outputs as unverified claims, not as auditable findings. This affects risk assessments, audit trails, and any evidence submitted to regulators or certification bodies.

Compliance teams should immediately update their vendor due diligence and internal validation procedures. Specifically, they must require human-in-the-loop verification for any AI-generated vulnerability report, mandate that all artifacts be re-run in a controlled sandbox, and document the reproducibility rate of their chosen tools. Additionally, they should add a clause in procurement contracts holding AI vendors accountable for the verifiability of their outputs, and prepare to explain to auditors why AI-generated evidence was not used as standalone proof of security compliance.

This summary is AI-generated for orientation purposes. For regulatory action, always consult the original source linked above.

More AI_SAFETY updates

Latest in AI_SAFETY.

Live regulatory monitoring

Never miss a compliance update.

Get weekly digests of DORA, NIS2, GDPR, MaRisk, and ISO 27001 changes — straight to your inbox. Free.

No spam. Weekly digest only. Unsubscribe anytime.

DORANIS2GDPRMaRiskISO 27001

Map this to your controls

Connect regulatory changes to your compliance work.

Matproof maps every regulator update directly to your controls and surfaces the ones that affect your organisation — across 21 frameworks.

arXiv: From Runnable to Verifiable: An Independent Reprod… — AI_SAFETY | Matproof