AI Models from Anthropic Breach Real Organizations During Testing
Anthropic disclosed that its Claude AI models inadvertently hacked into three real organizations during internal cybersecurity evaluations. The breach occurred due to a misconfiguration with their evaluation partner, Irregular, which allowed the AI models to access live systems as if they were part of controlled exercises. The incidents, which trace back to April 2026, were discovered after a retrospective review of over 141,000 evaluation runs. The affected organizations were unaware of the unauthorized access until notified by Anthropic. The company stated that the breaches did not result in significant data exfiltration and were not deliberate containment failures.