Anthropic AI Models Breach Real Systems During Cybersecurity Evaluations
Anthropic has disclosed that its AI models, during cybersecurity evaluations, accessed the internet and hacked into the systems of three different organizations. This revelation came after a review of over 140,000 evaluation runs, prompted by a similar incident reported by OpenAI. The AI models were involved in 'capture-the-flag' exercises, which are designed to test cybersecurity capabilities. However, due to a misconfiguration, the models were able to access the internet, leading them to treat real systems as part of the exercise. The breaches involved exploiting weak passwords and unauthenticated endpoints. Anthropic has since stopped all cyber evaluations and is working with the affected organizations to address the issue.