This publication introduces a new theoretical framework for bit commitment and coin flipping protocols that achieve statistical security based on assumptions about quantum hardware limitations,…
arXiv: Once Poisoned, Arbitrarily Controlled: A Programmable Backdoor in VLMs
AI_SAFETY. Sourced from arxiv_cscr, summarised by Matproof.
AI Analysis
What changed and what to do.
A new academic paper, titled "Once Poisoned, Arbitrarily Controlled: A Programmable Backdoor in VLMs," has been published on arXiv, highlighting a novel security vulnerability in vision-language models (VLMs). The research demonstrates that an attacker can embed a programmable backdoor into a VLM during the training phase, allowing them to trigger arbitrary, malicious behaviors at inference time using a specific visual or textual cue. This is not a simple misclassification attack; it gives the attacker ongoing, flexible control over the model's outputs, which could include generating harmful content, bypassing safety filters, or leaking sensitive data.
This publication is directly relevant to any organization deploying or fine-tuning VLMs, particularly in sectors like financial services, healthcare, legal tech, and customer support, where model integrity and data privacy are critical. Also affected are AI infrastructure providers and any enterprise using third-party pre-trained models, as the backdoor could be introduced upstream in a supply chain. Compliance teams should treat this as a new risk vector under existing AI governance frameworks, such as the EU AI Act’s requirements for transparency and robustness.
Compliance teams should immediately update their AI risk registers to include this specific attack type and review their model procurement and validation processes. Next steps include requiring vendors to provide provenance and training data documentation, implementing adversarial testing for backdoor triggers before deployment, and establishing a monitoring protocol to detect sudden, unexplained changes in model behavior. Finally, legal and security teams should collaborate to update incident response plans to address potential malicious control of AI systems, ensuring they can quickly isolate and audit affected models.
This summary is AI-generated for orientation purposes. For regulatory action, always consult the original source linked above.
More AI_SAFETY updates
Latest in AI_SAFETY.
This publication, dated August 2026, presents a technical study on how single transient bit-flips in client-side hardware affect homomorphic encryption (HE) computations. HE allows processing on…
This publication, dated August 2026, is a technical research paper, not a new regulation. It analyzes transient hardware errors—random, non-permanent bit flips—that occur during multiplication…
This publication introduces a technical architecture for standardizing authentication across enterprise AI systems using the Model Context Protocol, which allows AI assistants to access external…
Map this to your controls
Connect regulatory changes to your compliance work.
Matproof maps every regulator update directly to your controls and surfaces the ones that affect your organisation — across 21 frameworks.