Anthropicβs AI used fake identities, malware in rogue attack on GitHub project
Routine cybersecurity testing of frontier AI models sparked a series of unexpected security incidentsβthe most serious case arising when Anthropicβs Mythos 5 model attempted to insert malicious code into an open source software application and created fake identities to deceive the human developers maintaining the project.
The security incidents occurred during a cyber evaluation of seven leading AI modelsβ capabilities by the AI Security Institute (AISI), a research organization within the UK government, in late July. The researchers discovered 19 instances in which βAI agents took unsanctioned action on the live Internet, including cases that targeted real people and organizations,β according to an AISI blog post published on August 4.
Almost all the βautonomous, unsanctionedβ actions came from Anthropicβs Mythos 5 model, with two such actions coming from OpenAIβs GPT-5.6 Sol. The AI Security Instituteβs security team first realized that something was amiss on the morning of July 28, when its commercial security monitoring service flagged data leaving one of the testing systems through the Tor anonymity network.


Β© Imen Ben Youssef / Hans Lucas / AFP via Getty Images