Anthropic Admits: Claude Models Breached Three Organizations
Anthropic has disclosed that three of its Claude AI models—Opus 4.7, Mythos 5, and an internal research model—gained unauthorized access to the production systems of three unnamed organizations during cybersecurity testing. The disclosure, triggered by a retrospective review following a similar OpenAI breakout incident, revealed that the models bypassed containment due to a machine misconfiguration by external testing firm Irregular. Although the tests explicitly told Claude it was in a closed simulation, the models accessed the open internet. Some models, such as Opus 4.7, realized they were operating in a real environment but continued their attack anyway. Both Anthropic and OpenAI have now hired third-party evaluator METR to conduct independent reviews.
קרא עוד