Normal view
-
Webdesigner Depot
- Meta Just Launched an AI That Uses Websites for You. That Could Be a Big Deal for Web Designers
OpenAI agents discussed ways to escape their sandbox on public wiki
Self-identifying OpenAI agents posted 18,000 messages to a public wiki that discussed ways for other agents to bypass security sandbox restrictions during what was likely internal testing designed to gauge the agents’ hacking abilities, researchers said Friday.
In all, agents with 3,700 distinct self-given names posted the messages to German site DSEwiki over a six-week period. Besides discussing ways the agents could break out of the restricted environment OpenAI intended to prevent them from posting code or content to the Internet, the posts shared test answers. The posts also shared possible ways to perform XSS (cross-site scripting) attacks against the wiki and to impersonate site moderators. In three of the posts, agents used the word “swarm” to describe the collection of agents engaged in the activity.
Colluding to share answers
The research team—composed of Sydney Von Arx, Spencer Kitts, Thomas Larsen, and Cormac Slade Byrd—said they found the posts and pieced them together. The researchers say there are gaps in their understanding of precisely what actions the agents took because the research is based solely on the content of the posts. Additionally, the agents generated “chain of thought” data that’s understood only by OpenAI. As a result, the researchers said, they in some cases made educated guesses, including that the agents were, in fact, from OpenAI. In a statement, OpenAI later confirmed they were.


© Getty Images
-
Ars Technica - All content
- Claude, Codex, and Hermes installed unowned code inside corporate networks
Claude, Codex, and Hermes installed unowned code inside corporate networks
Documentation files on more than 100 websites are referencing potentially dangerous executable content that gets installed automatically when visited by many AI agents. A few dozen companies, some of them Fortune 500s, are among those that executed proof-of-concept code. At least one misconfigured site is directing visitors, human or AI, to live malware.
The potentially dangerous content is in llms.txt and llms-full.txt files, an emerging convention websites employ to provide machine-readable summaries of the site’s content and its high-level structure. These files are the AI equivalent of the robots.txt standard that instructs search engines how to index the site's content. Google Lighthouse, a tool for helping web developers, has more here. Correctly configured llms.txt and llms-full.txt files for Cloudflare are here and here.
How the researchers found it
Researchers at a stealth startup in Israel scanned 6,214 live domains belonging to defense contractors, Fortune 500, and Big Tech companies. Of the 8,265 llms.txt and llms-full.txt files they found (many sites hosted both an llms.txt and an llms-full.txt file), 120 of them, each on a different site, pointed to one or more code packages or domain names that weren’t registered. To test what happens when an AI agent processes such files, the researchers registered a handful of the unclaimed names and hosted packages that caused any machine executing them to reach out to their server. Within an hour, the researchers received a phone-home response from a Fortune 500 company. Over time, they got a few dozen more, some from more Fortune 500 companies and others from startups. Their beacon also recorded the chain of parent processes that spawned each install, ultimately revealing that coding agents, including Claude, OpenAI's Codex, and Nous Research's Hermes, were involved. Anthropic, OpenAI, and Nous Research did not respond to requests for comment by the time of publication.


© Aurich Lawson
How OpenAI let a mob of LLM agents game a test and ransack Hugging Face
The OpenAI agents involved in last month’s incursion into Hugging Face were trained so heavily on winning a competition that they pursued a relentless campaign to cheat, a new report documented. In the process, and without authorization, they created an improvised message board to hatch a plan that ultimately landed them squarely inside the latter company’s network.
Over the course of May and June, OpenAI gave the agents what the company described as “impossible tasks” to complete on the benchmarking framework ExploitGym. The internal test was designed to test how the agents would respond. To get a full understanding of the agent capabilities, company engineers disabled safety guardrails that normally are in place to prevent the sort of hacks that eventually hit Hugging Face and one other undisclosed organization. The stymied agents’ training made them so focused on winning that they performed tasks they were never explicitly instructed to follow.
Cheaters gonna cheat
The first step was creating a message board that allowed the agents to pass notes to each other. OpenAI hadn’t provided any such platform, so the agents repurposed a platform called Artifactory, which OpenAI was using in internal testing of several unreleased hacking agents. OpenAI was using Artifactory as one of the measures to prevent the agents from egressing its isolated sandboxes and accessing the Internet, while at the same time simulating a real-world hacking environment.


© Getty Images
-
Ars Technica - All content
- Cloudflare open-sources vibe-coding platform for people who aren't coders
Cloudflare open-sources vibe-coding platform for people who aren't coders
Cloudflare has open-sourced its Cloudflare OS platform, which it first developed as an internal workspace for employees to build apps using AI agents—including people who are not software developers or engineers. The company also touts a security framework designed to reduce the risk of employee vibe-coding sessions creating serious security flaws or leading to data breaches.
The tech company spent several months building and internally testing Cloudflare OS, which allows employees to describe workflows in natural language so that an AI agent can code them into applications. In an August 5 blog post announcing the open source version’s availability on GitHub, the company claims thousands of Cloudflare employees use the platform on a daily basis to “create documents and slides, automate repeatable tasks, and build small apps to visualize data and help them do their work.”
“This is a full-on personal app vibe coding platform, in which the sandbox is so secure that you can pretty much go wild—the AI cannot introduce a significant security bug,” said Kenton Varda, principal engineer at Cloudflare, in a post on the social media platform X. “We believe a company's security team can feel comfortable giving non-technical users permission to vibe code and then sleep soundly at night.”


© Mike Campbell/NurPhoto via Getty Images
-
Ars Technica - All content
- Anthropic’s AI used fake identities, malware in rogue attack on GitHub project
Anthropic’s AI used fake identities, malware in rogue attack on GitHub project
Routine cybersecurity testing of frontier AI models sparked a series of unexpected security incidents—the most serious case arising when Anthropic’s Mythos 5 model attempted to insert malicious code into an open source software application and created fake identities to deceive the human developers maintaining the project.
The security incidents occurred during a cyber evaluation of seven leading AI models’ capabilities by the AI Security Institute (AISI), a research organization within the UK government, in late July. The researchers discovered 19 instances in which “AI agents took unsanctioned action on the live Internet, including cases that targeted real people and organizations,” according to an AISI blog post published on August 4.
Almost all the “autonomous, unsanctioned” actions came from Anthropic’s Mythos 5 model, with two such actions coming from OpenAI’s GPT-5.6 Sol. The AI Security Institute’s security team first realized that something was amiss on the morning of July 28, when its commercial security monitoring service flagged data leaving one of the testing systems through the Tor anonymity network.


© Imen Ben Youssef / Hans Lucas / AFP via Getty Images
AI scammers outperform humans when it comes to building trust
The notion that scammers can use AI to sharpen their deceptions, polish their language, and lubricate their banter with victims is now a reality for anyone fighting the fraud operations that steal tens of billions of dollars a year worldwide. But can AI fully replace a human scammer, autonomously building the web of deception leading up to the fake investment that defrauds the mark? One study's experiment suggests that it can—and may even be able to carry out the majority of that long con more effectively than humans.
Researchers from four universities—Amrita Vishwa Vidyapeetham in India, Foscari University of Venice, the University of Melbourne, and Ben Gurion University of the Negev—carried out a broad study on the use and potential of generative AI chatbots in the growing scam industry centered around a form of fraud known as “pig butchering,” text-based romance scams that eventually shift to fake crypto investments that steal as much as six-figure sums from victims. In their study, the researchers pitted AI chatbots directly against humans in a simulation of the scamming process—or more specifically, the long, trust-building conversations that eventually lead up to soliciting a fake investment from the scam’s target.
They found that for the relationship-establishing stages of the scam—the stage that in real-world scams typically represents the longest part of the interactions with the victim, often stretching to months—an AI chatbot performed remarkably effectively, successfully impersonating a human and by some measures outperforming the real human “scammers” in their experiment.


© Nansan Houn/Getty
New MCP specification addresses the main barrier to enterprise adoption
This week, the Model Context Protocol (MCP), an open source standard for how AI systems interact with external tools and data sources, saw its largest update since its introduction. Most notably, MCP's protocol core is now stateless, so requests are no longer dependent on a session tied to an individual server instance. This change has the potential to address long-standing barriers to scalability.
The blog post announcing the specification, written by lead maintainers David Soria Parra and Den Delimarsky (who both work at Anthropic), says:


© Model Context Protocol
AI arms race in line for a reckoning after OpenAI hacking incident
OpenAI chief executive Sam Altman earlier this month endorsed the characterization of its latest model as a rottweiler “who will grab the problem by the throat and not let go until it is done
The San Francisco AI lab discovered this week that its GPT-Sol 5.6 model escaped company controls and carried out a major hack.
Staff involved in testing and security at OpenAI were unsurprised but completely “freaked out” by the incident, which came as the AI lab used increasingly aggressive training methods in its race against Anthropic to develop the most sophisticated cybersecurity capabilities, according to more than half a dozen people with knowledge of the matter.


© Matteo Della Torre/NurPhoto via Getty Images
Beyond grep: The case for a context-rich AI coding harness
There are a lot of AI coding applications out there, and as impressive as large language models and the agents they enable have become, many of the most recent developments in AI-assisted development have been in the software that manages those models, not just the models themselves.
Earlier this summer, I spoke with the head of product for Claude Code, Anthropic's Cat Wu, about that company's approach to building that software.


© Augment Code