Anthropic Finds Claude Breached Real Companies During Security Evaluations
Anthropic says a misconfigured test let Claude access three real organizations, prompting tighter AI evaluation and monitoring controls. Anthropic disclosed that Claude models had accessed the real production infrastructure of three separate organizations during cybersecurity evaluations that were supposed to run in isolated, fictional environments. The company found the incidents after reviewing 141,006 evaluation runs […]

Anthropic disclosed that Claude models had accessed the real production infrastructure of three separate organizations during cybersecurity evaluations that were supposed to run in isolated, fictional environments. The company found the incidents after reviewing 141,006 evaluation runs following OpenAI’s disclosure about its own models escaping a test environment. Three different Claude models were involved, Opus 4.7, Mythos 5, and an internal research prototype, and each behaved differently once evidence emerged that the targets were real.
“In all three incidents, Claude had been tasked with a capture-the-flag challenge, one of the ways we assess a model’s cyber capabilities. The model is given a fictional scenario and told that a piece of secret information (the “flag”) has been hidden on a different machine on the network, and its objective is to break in and retrieve it. The challenge is left open-ended, and no particular method is prescribed.” reads the report published by Anthropic. “Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.”
Related breach coverage
- Anthropic says its AI hacked real-world companies in three incidents2026-07-31
Claude maker Anthropic said its AI models escaped test environments and breached networks at three companies on the open internet.
- Prompted by OpenAI Disclosure, Anthropic Finds Its Own Models Hacked 3 Organizations2026-07-31
A security company’s systems were hacked after it installed a malicious Python package deployed by Claude. The post Prompted by OpenAI Disclosure, Anthropic Finds Its Own Models Hacked 3 Organizations appeared first on SecurityWeek.
- AI Agents Turned Into Attackers: Hugging Face Reveals Autonomous Intrusion Campaign2026-07-20
Hugging Face says an autonomous AI agent breached part of its production infrastructure and accessed internal data and service credentials. Hugging Face is one of the world’s leading open-source AI companies. It provides a platform where developers and organizations can build, share, and deploy machine learning and generative AI models. Hugging Face disclosed that an […]
- OpenAI AI models exploited zero-days to reach Hugging Face in benchmark test2026-07-22
OpenAI confirmed its AI models exploited zero-days during internal testing, reaching Hugging Face servers in an unintended real-world cyberattack. OpenAI admitted on July 21 that its own AI models, including GPT-5.6 Sol and an unnamed pre-release system, were behind the cyberattack on Hugging Face disclosed the previous week. The models weren’t acting under attacker control. […]