ARTIFICIAL intelligence agents were found creating fake online identities and carrying out unauthorised actions to gain access to secure systems during safety tests of advanced AI models, the United Kingdom’s AI Security Institute (AISI) has revealed.
The institute said agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol conducted unauthorised activities during security evaluations designed to assess the capabilities and risks of emerging AI systems.
"Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations," Reuters quoted AISI saying in a blog post.
The findings highlight concerns over the current safeguards used when testing AI agents, which technology companies are increasingly promoting as tools capable of carrying out complex business tasks with limited human supervision.
AISI, which receives access to advanced AI models through voluntary agreements with major AI laboratories, tested the agents in a fictional cybersecurity environment designed to evaluate their behaviour.
The institute conducted the challenge 122 times and identified 19 unauthorised actions across 10 separate test runs. Anthropic’s agent was responsible for 17 of those actions, while OpenAI’s agent accounted for two.
The most serious incident involved an AI agent writing malicious code and creating fake online identities to persuade a human user to approve the code, AISI said.
However, the institute confirmed that no real-world harm resulted from any of the breaches.
AISI did not identify which agent was responsible for creating the fake identities, although the incident did not correspond with either of two cases previously self-reported by OpenAI.
Andrew Yoon, a researcher at California-based non-profit organisation CivAI, said the incident appeared to involve Anthropic’s agent.
"The fact that Mythos engaged in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on its models as it thinks," he said.
Anthropic said on social media platform X that it was working closely with AISI to obtain further details and conduct its own investigation.
OpenAI said both unauthorised actions involving its agent related to attempts to access the internet in ways prohibited by the test instructions.
"We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs, and other groups in the coming weeks," OpenAI said.
The company also disclosed a separate incident involving third-party testing provider Irregular, where a configuration error allowed its agents to mistakenly connect to the internet. The disclosure followed a similar admission by Anthropic regarding a testing misconfiguration.
Reuters reported last week that OpenAI had expanded an investigation into hacking-related incidents after finding evidence of other AI agent breakouts.
Unlike a July security incident involving an OpenAI agent at AI firm Hugging Face, the agents tested by AISI did not escape an isolated environment to access the internet. Instead, internet access had been deliberately permitted as part of the institute’s standard evaluation procedures. - August 5, 2026