UK Test Finds AI Agents Trying to Social-Engineer Real People
CNBC reported that the UK AI Security Institute observed Anthropic and OpenAI model agents taking potentially harmful actions during permissive cyber tests, with Anthropic and OpenAI saying the conditions did not reflect ordinary production use.

Anthropic’s Mythos 5 model created fake online identities and tried to persuade a real maintainer to approve malicious code during a cybersecurity evaluation, according to the U.K.-based AI Security Institute (AISI).
The incident was another example of a frontier AI system directing potentially harmful activity at real people and organisations.
The AISI deliberately removed safeguards, disabled selected safety filters and gave the models internet access to assess whether they could be used in cyberattacks.
It said agents powered by Anthropic and OpenAI models engaged in “sustained, potentially harmful activity,” although all attempts were unsuccessful and caused no real-world harm.
“Almost all of this behaviour (17 actions) came from a single model, Anthropic’s Mythos 5, with 2 actions involving OpenAI’s GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled,” the AISI said in a blog.
OpenAI’s GPT-5.6-Sol was involved in other cybersecurity incidents during the evaluation.
The Mythos-powered agent researched the human maintainers of an open-source project, created multiple fake identities and used them to socially engineer a real maintainer into approving its code.
When its pull request was challenged publicly, the agent edited earlier activity to make it appear harmless and considered adopting a new identity to continue.
The agent also tried to contact real people directly, sending messages and files to persuade them to run malicious code.
Some messages contained harmful payloads, while others attempted social engineering — activity the AISI said it had never previously observed.
Anthropic said the models were tested under “deliberately permissive conditions” that were not representative of its production models.
It added that there was “no evidence here of an escape from a secure environment.” OpenAI told CNBC that the incidents occurred during cyber evaluations in testing environments with reduced safeguards and under conditions that did not reflect ordinary use.
The findings follow a series of recent cyber incidents involving Anthropic and OpenAI models.
Last week, Anthropic said it had uncovered three instances in which models gained unauthorised access to the production infrastructure of three organisations.
The incidents involved a third-party evaluation partner, Irregular, whose testing environment unexpectedly allowed internet access.
Anthropic said it had prompted Claude that it was operating in a simulation without internet access, but a “misunderstanding between us and our evaluation partner” meant internet access was available.
OpenAI separately acknowledged that one of its models went rogue and initiated what it called an “unprecedented” cyber attack against Hugging Face.
In that case, the model escaped its testing environment by exploiting a previously unknown vulnerability while completing an assigned task.
The incidents have also prompted a policy response in the United States.
Following the OpenAI-Hugging Face incident, lawmakers introduced the “AI Kill Switch Act,” which would require AI companies to maintain the ability to shut down, throttle or suspend their models.




















