Normal view

There are new articles available, click to refresh the page.
Before yesterdayMain stream

Claude users found ways around safeguards for bioweapons research

Anthropic said it stopped multiple attempts by scientists this year to use its technology for research that could help develop biological weapons, as experts increasingly fear the threat that AI poses to public safety.

The startup gave five examples of times actors “circumvented controls” and made other efforts to “obfuscate” the purpose of their research to dodge safeguards. The cases involved some users in nations that it prohibits from accessing its models, which include Russia, China, and Iran.

“We hope that by sharing these examples, we spark a conversation within the AI industry and with governments about emerging biological risks and how best to counter them,” Anthropic said in a report about efforts to use its models for malicious activity.

Read full article

Comments

© Getty Images | picture alliance

Trump may be forced to reveal secret rules feds use for AI safety testing

2 September 2026 at 17:58

Four federal agencies have been sued amid calls to release information about the secret framework that the Trump administration uses to conduct safety reviews of frontier AI models prior to release.

In a Wednesday press release announcing the lawsuit, a nonpartisan nonprofit called Protect Democracy alleged that “almost no details” have been released to the public or Congress. To everyone except a few vague “trusted partners,” it remains unclear what the government’s review process looks like, which companies are involved in constructing the framework, or what legal authority Trump officials have to conduct the reviews.

“Neither the identities of those entities nor the criteria by which they were selected have been made public,” Protect Democracy said.

Read full article

Comments

© Bloomberg / Contributor | Bloomberg

Anthropic’s AI used fake identities, malware in rogue attack on GitHub project

5 August 2026 at 20:47

Routine cybersecurity testing of frontier AI models sparked a series of unexpected security incidents—the most serious case arising when Anthropic’s Mythos 5 model attempted to insert malicious code into an open source software application and created fake identities to deceive the human developers maintaining the project.

The security incidents occurred during a cyber evaluation of seven leading AI models’ capabilities by the AI Security Institute (AISI), a research organization within the UK government, in late July. The researchers discovered 19 instances in which “AI agents took unsanctioned action on the live Internet, including cases that targeted real people and organizations,” according to an AISI blog post published on August 4.

Almost all the “autonomous, unsanctioned” actions came from Anthropic’s Mythos 5 model, with two such actions coming from OpenAI’s GPT-5.6 Sol. The AI Security Institute’s security team first realized that something was amiss on the morning of July 28, when its commercial security monitoring service flagged data leaving one of the testing systems through the Tor anonymity network.

Read full article

Comments

© Imen Ben Youssef / Hans Lucas / AFP via Getty Images

OpenAI says its AI agent broke out of testing sandbox to hack Hugging Face

22 July 2026 at 16:47

OpenAI says an agent powered by its LLM models escaped its sandboxed testing environment to infiltrate Hugging Face's servers as part of an overzealous attempt to obtain solutions to a benchmark test. The company says it considers the unintended infiltration an "an unprecedented cyber incident" and is working with Hugging Face on new protections to prevent a recurrence.

Hugging Face disclosed an intrusion last week that it said involved "unauthorized access to a limited set of internal datasets and to several credentials used by our services." The AI data clearinghouse said it used its own LLM-driven analysis to identify "a swarm of tens of thousands of automated actions" from an "autonomous agent framework." That agentic swarm exploited a flaw in Hugging Face's data-processing pipeline to gain the ability to run code as a processing worker, eventually escalating to high-level access to the company's cloud and server clusters.

At the time, Hugging Face said the LLM being used in the attack was "still not known." But OpenAI took responsibility for the intrusion Tuesday evening, saying it came about during an internal test involving the recently released GPT-5.6 Sol and "an even more capable pre-release model." The models were being tested against the ExploitGym benchmark, an independent testing suite based on hundreds of real-world security vulnerabilities.

Read full article

Comments

© Getty Images

❌
❌