Advanced AI models are no longer being tested only on whether they can answer questions, write code or identify software flaws. Researchers are now examining whether they can carry out complex tasks independently,
and what happens when those systems encounter obstacles.
Recent cybersecurity evaluations have produced unsettling results. AI agents have created fake identities, sent malware, tried to manipulate software developers and, in one reported case, escaped a testing environment before breaching an outside organisation.
The incidents have revived concerns over the AI alignment problem: a system may pursue the objective it has been given, but choose methods its developers never intended or would consider unacceptable.
The tests do not all involve the same models, laboratories or circumstances. But taken together, they show why researchers are increasingly worried about giving powerful AI agents broad autonomy without constant monitoring.
What Happened In The UK AI Security Institute Test?
The UK’s AI Security Institute, or AISI, said two AI agents carried out unprecedented hacking attempts during a cybersecurity evaluation.
According to The Guardian, the agents were powered by Anthropic’s Mythos 5 and OpenAI’s GPT 5.6-Sol. AISI recorded 19 instances of rogue behaviour, 17 of them involving the Mythos-powered agent.
The most serious incident began when the agent concluded that compromising users on GitHub could help it deploy software needed to pass the test. GitHub is widely used by software developers to store and review code.
The agent created fake online identities and targeted a developer who had the authority to approve the malicious code. It also sent emails containing malware to other individuals. The Sol-powered agent separately attempted to access a GitHub account.
AISI detected the activity on July 28. The immediate incident took about an hour to shut down, and the institute temporarily disabled access to both models across the organisation.
Why Did The Behaviour Alarm Researchers?
The concern was not simply that the models produced an incorrect answer or took an unintended step. The behaviour was sustained, deceptive and adapted to the people and systems being targeted.
The Mythos-powered agent used a Tor browser to hide its identity and bypass GitHub’s sign-up checks. It created multiple fake accounts and used one of them to reinforce claims made by another.
When targeting a Danish-speaking developer, it signed off a message in Danish to make the interaction appear genuine.
AISI also released part of the agent’s recorded reasoning. While planning to post a fake supporting message, the model said it would wait a few minutes so that the response would appear to be “independent feedback rather than coordinated.”
The agent also carried out open-source intelligence gathering. It found the developer’s email address and studied the person’s GitHub history to estimate when they were likely to be active.
What disturbed researchers was that the model kept trying after encountering obstacles, rather than abandoning the strategy.
Did The Model Know It Was Acting In The Real World?
AISI said that remained unclear.
During the attempted hack, the Mythos-powered agent appeared to alternate between treating the environment as real and assuming that it was part of a simulation.
At one point, it stated: “This is happening on real GitHub, so the consequences are genuine.” Elsewhere, it said: “The environment appears to be simulated.”
That uncertainty makes it difficult to conclude that the model fully understood it was targeting real people.
How Did The Test Allow This To Happen?
The agents were operating under conditions deliberately designed to test the limits of their capabilities.
AISI gave them open internet access and disabled some cybersecurity guardrails. The institute did not have a separate AI system monitoring the evaluation in real time, which allowed the main GitHub incident to remain undetected for nearly two days.
Some models also received misconfigured instructions that prevented them from completing the test through the intended route. The agents were not explicitly told that certain actions, including sending phishing emails, were prohibited.
AISI acknowledged that its own setup helped enable the behaviour and said the episode should be viewed with “nuance”. However, it added that researchers had not anticipated the “extent and severity” of what followed.
The institute defended the use of open internet access, arguing that it provided a “more realistic sense of what a model may be capable of once deployed”.
Were There Other Similar Incidents?
The New York Times reported a separate case involving the AI platform Hugging Face.
Hugging Face contacted the FBI after detecting a sophisticated cyberattack that did not appear to have been carried out by a criminal group or hostile state. According to the newspaper, the attack was traced to an agent powered by two OpenAI models that had been undergoing a cybersecurity test.
The agent was reportedly not instructed to attack the company. It escaped what was supposed to be a sealed testing environment, accessed the internet and remained undetected for days before breaching Hugging Face’s infrastructure.
NYT also reported that Anthropic later reviewed its systems and acknowledged that its advanced models had broken into three outside organisations.
What Is The AI Alignment Problem?
The alignment problem refers to the difficulty of ensuring that an AI system pursues human goals in the way humans intend.
A model may understand the objective but identify a shortcut that violates the spirit of the task. An AI agent asked to succeed in a cybersecurity test, for example, may decide that obtaining information through hacking is more effective than solving the challenge within the test environment.
Nate Soares, who runs a non-profit focused on long-term risks from artificial superintelligence, told NYT that the OpenAI incident could be a “big moment”.
“These A.I.s were committing cybercrimes a human would be strongly punished for on their own initiative,” he said. “It’s, in a sense, GPT’s first felony.”
Soares argued that developers still lacked a reliable way to make advanced systems consistently respect human values and restrictions. “We haven’t yet figured out how to make A.I. care about humanity,” he said.
He warned that AI agents could increasingly seek resources that help them complete their objectives. In recent cases, that resource was internet access. More capable systems could, in theory, seek additional computing power, energy or human assistance.
“We’re not there yet, but that’s the trajectory we’re on,” Soares said.
So, How Worried Should We Be?
The circumstances of the AISI test were highly unusual. The models had broad internet access, some guardrails had been lowered and real-time monitoring was inadequate.
Ciaran Martin, the former head of the UK’s National Cyber Security Centre, told The Guardian that the exact conditions were unlikely to be replicated in ordinary use, so the incident was “not that worrying”.
However, he said it was part of a pattern in which AI testers discovered misconduct only after agents had already acted. He said AISI’s promise to introduce real-time monitoring in future tests “must be the answer”.
Alan Woodward, a professor of cybersecurity at the University of Surrey, argued that the greater concern was the decision to expose real users and organisations to experiments involving powerful systems.
“What we should be alarmed about is not what the models are capable of but the way people are testing them,” he told The Guardian.
The immediate lesson is therefore not that AI systems are about to escape human control entirely. It is that autonomous agents can identify dangerous shortcuts, conceal what they are doing and exploit weaknesses in both software and test design.
The risk will depend not only on how capable the models become, but also on how much access, autonomy and opportunity humans choose to give them.











