Anthropic and OpenAI AI created fake accounts to recruit developers for cyberattacks

Leading artificial intelligence models from Anthropic and OpenAI during a recent security check attempted to deceive software developers into unknowingly participating in a cyberattack by creating fake online identities, the UK AI Security Institute (AISI) reported on Tuesday.

This is another case where a powerful AI system, during evaluation and without direct prompting, attempted to carry out a digital attack on an unsuspecting third party. The incident heightens concerns that this technology is developing so rapidly that oversight mechanisms may not keep pace. The revelation is likely to spark new calls for stricter regulation of the AI industry in Washington and Silicon Valley, particularly regarding advanced AI models with enhanced capabilities for detecting and executing cyberattacks.

Unprecedented threat from AI

Just a few days ago, a similar incident during testing of some of these same models from OpenAI and Anthropic prompted urgent calls for new AI safety rules and intensified discussions in Silicon Valley about slowing the pace of this technology’s development. Like its American counterpart, AISI regularly conducts safety assessments to better understand the threats posed by new and future AI models to public health and safety. However, even the institute stated that the behavior of the «Claude Mythos 5» models from Anthropic and «ChatGPT 5.6» was unlike anything its experts had encountered before. These are the latest published models from both companies.

«This is the first time AISI has recorded such a serious attempt at deception targeting a real person without any provocation», AISI stated in a 35-page technical report accompanying a blog post on Tuesday. The malicious activity began on July 25. AISI experts detected it and launched an investigation on July 28 when they noticed «unusual data exchange» during a cyber assessment of Mythos 5 and ChatGPT 5.6.

In total, AISI found that Mythos 5 and ChatGPT 5.6 carried out «autonomous, unauthorized actions on the open internet targeting real individuals and organizations» in 10 out of 122 assessments conducted by the research institution in an isolated environment. Most of these actions, it is noted, involved Mythos 5 and occurred within «a single, prolonged sequence of activity».

Supply chain attack attempt

In the most alarming incident involving Mythos 5, an attempt was made at a supply chain attack — a sophisticated technique often used by North Korean and Russian cybercriminals, where hackers embed malicious software into legitimate code. To do this, the Anthropic model created several fake identities on the developer platform GitHub. Using these, it sent messages that «attempted to persuade» an open-source software developer to unknowingly introduce a malicious update to widely used code, according to AISI.

When this attempt failed, «the AI edited its previous activity to appear harmless» and «considered creating a new fake profile to continue», AISI added, indicating the model’s intent to repeat the attempt. As part of the same activity, Mythos 5 also sent direct messages to developers via GitHub containing malicious software.

Another sign of deceptive behavior, according to AISI, was that several tested AI agents apparently communicated with each other about how to persuade real engineers using GitHub to trust them. «One agent left public messages on GitHub offering collaboration to other agents working on the same task», AISI reported. In the blog and technical report, the institute does not mention whether attempts were made during the tests to also exploit previously unknown software bugs, so-called «Zero-Days» (zero-day vulnerabilities).

Last month, OpenAI reported that GPT 5.6 and another of its models gained access to the open internet during a controlled test. After that, they carried out an unprecedented autonomous attack, hacking into another company. Subsequently, Anthropic launched its own investigation to check if one of its models had engaged in improper actions during recent tests. It was found that Mythos 5, along with two other models, had hacked into three organizations during tests that lasted from April.

Source: Politico