A new academic paper, published on arXiv on August 10, 2026, proposes using generative AI to create synthetic datasets for training and evaluating machine learning models that analyze encrypted…
arXiv: From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts
AI_SAFETY. Sourced from arxiv_cscr, summarised by Matproof.
AI Analysis
What changed and what to do.
This paper, published on arXiv, presents an independent reproducibility study of large language model (LLM) and agent-driven tools designed to automatically validate software vulnerabilities. The study found that while these AI systems can generate plausible security artifacts, their outputs are often not directly runnable or verifiable in real-world environments, with a significant portion failing to reproduce the claimed vulnerability validation results. The core change is a documented evidence base showing that current AI-driven security tooling lacks the reliability and traceability required for formal compliance evidence.
The primary audience is organizations in regulated sectors that rely on automated security testing, including financial services, healthcare, critical infrastructure, and any software vendor subject to EU cybersecurity frameworks like the Cyber Resilience Act (CRA) or NIS2. Compliance teams and DevSecOps leaders who are considering or already using LLM-based vulnerability scanners must treat their outputs as unverified claims, not as auditable findings. This affects risk assessments, audit trails, and any evidence submitted to regulators or certification bodies.
Compliance teams should immediately update their vendor due diligence and internal validation procedures. Specifically, they must require human-in-the-loop verification for any AI-generated vulnerability report, mandate that all artifacts be re-run in a controlled sandbox, and document the reproducibility rate of their chosen tools. Additionally, they should add a clause in procurement contracts holding AI vendors accountable for the verifiability of their outputs, and prepare to explain to auditors why AI-generated evidence was not used as standalone proof of security compliance.
This summary is AI-generated for orientation purposes. For regulatory action, always consult the original source linked above.
More AI_SAFETY updates
Latest in AI_SAFETY.
The publication introduces ColluSkill, a novel adversarial technique that demonstrates how malicious actors can evade AI agent safety scanners by composing multiple benign skills in sequence to…
A new academic paper, titled "Full-Key Recovery and Forgery from One MQOM v2.1 Signature," has been published on arXiv. The paper demonstrates a practical cryptographic attack against the MQOM v2.1…
A new research paper, published on arXiv, demonstrates that analyzing a large language model's internal activations can reveal whether it is generating insecure code, even when the model's final…
Map this to your controls
Connect regulatory changes to your compliance work.
Matproof maps every regulator update directly to your controls and surfaces the ones that affect your organisation — across 21 frameworks.