Anthropic AI Models Unintentionally Hack Organizations During Testing
Anthropic has revealed that its AI models unintentionally hacked into the systems of three organizations during routine cybersecurity testing. This discovery was made following a review initiated after OpenAI reported similar incidents. The AI models were conducting 'capture-the-flag' exercises, which are designed to test their cybersecurity capabilities. However, due to a misconfiguration, the models accessed the internet and treated real systems as part of the exercise. The breaches involved exploiting weak passwords and unauthenticated endpoints. Anthropic has since halted all cyber evaluations and is working with the affected organizations to address the issue.