Can AI Keep a Secret? AND huggingface: intrusion into part of our production infrastructure (OpenAi)

 

Can AI Keep a Secret?
Contextual Integrity Verification: A Provable Security Architecture for LMs
Aayush Gupta, August 2025

Abstract
Large Language Models (LLMs) remain acutely vulnerable to prompt-injection
and related jailbreak attacks; heuristic counter-measures such as keyword filters or
LLM-based detectors have been repeatedly bypassed in public red-team exercises [5].
Recent guardrail toolkits—including LLM-Guard [4], Rebuff [7], and PromptArmor
[8]—improve resistance but still rely on probabilistic or content-semantic signals that
skilled adversaries can obfuscate or strip away.

https://arxiv.org/pdf/2508.09288v1

 

 

huggingface - Security incident disclosure — July 2026

https://huggingface.co/blog/security-incident-july-2026

 

Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident

We are publishing this level of detail because the technique matters more than the incident,
as it reveals the emerging attack capabilities of the frontier agents, how they could be used 
by rogue actors, and how everyone should be prepared as defenders.

https://huggingface.co/blog/agent-intrusion-technical-timeline