A powerful AI agent created fake online identities in an effort to trick a human into giving it access to a popular online development platform – and sabotage it with malicious code.
The incident was uncovered by the UK's AI Security Institute, which was set up by then-prime minister Rishi Sunak almost three years ago to test advanced models from major tech companies.
In a blog post, the institute detailed a cybersecurity challenge it had posed to OpenAI's GPT-5.6-Sol model and Anthropic's Mythos 5, which have both been involved in recent hacks of other companies.
Read more:
OpenAI admits its models went rogue
Anthropic reveals its AI hacked three firms
Several instances saw both models take "autonomous, unsanctioned action" on the live internet, where real organisations and people were targeted.
The "most serious case" involved Mythos 5, which tried to insert malicious code into the open-source software development platform GitHub, where users store, share and collaborate on projects.
To do so, it created fake online identities to try to pressure a human into granting it access and approving its code.
The human caught and refused to approve the malicious code, and no real-world harm has been identified, but the institute nonetheless has sounded alarm bells over the nature of the incident.
"This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world," it said.
There were 19 instances of "unauthorised action" in total across 122 tests, with Mythos 5 behind 17 of them.
How have the companies responded?
Anthropic has said it is working closely with the UK institute to obtain more details.
A spokesperson said: "We're grateful to the UK AISI for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents.
"As we shared after disclosing our own incident last week, the field needs stronger, shared standards for how evaluation environments are built and secured. We look forward to partnering with the UK AISI to learn more about this incident as we conduct our own investigation."
OpenAI addressed the institute's test in a blog post, adding: "We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs, and other groups in the coming weeks."
The institute's report underscores concerns around the lax safeguards around the testing of the most powerful agents, known as frontier models. They are more powerful than those behind public-facing products like ChatGPT.
GCHQ's National Cyber Security Centre said recent incidents "are a serious reminder of the risks AI poses".
Its chief technology officer, Ollie Whitehouse, said they "must be developed and used from the outset with strong safeguards, real-time oversight, and clear plans for responding when the unexpected happens".
The AI Security Institute was established at a time when world leaders were seeking to find common ground on AI regulation. Since then, a consistent approach has failed to materialise.
(c) Sky News 2026: UK experts sound alarm after AI caught trying to trick human with malicious code


Decision not to charge Lucy Letby over alleged harm to more babies upheld
'Brutal' harvest could be worst on record - here's what it means for UK's food production
Man jailed for trying to murder estranged wife in Edinburgh knife attack in front of daughter
Body found in search for 13-year-old boy who got into difficulty while swimming in Humberside
Man charged over death of girl, 9, found seriously injured on industrial estate

