Normal view

There are new articles available, click to refresh the page.
Today — 18 September 2026Main stream

LLMs respond differently to harmful prompts when AI watermarking is used

17 September 2026 at 18:33

In response to a new European Union law, AI platforms are implementing new schemes for watermarking the content they generate. Anthropic recently disclosed its future Claude models will use SynthID-Text, an approach Google created and released as open source. It uses a secret key that subtly changes the process a model uses for choosing the next word in a sentence. Whereas a top next word choice might be “cloudy,” the key might change it to “overcast.” Anyone who knows the key can determine if it was generated by the platform using it.

New research shows that SynthID-Text can change not just word selection but also the tools a model invokes and the chances it will adhere to or disregard safety guardrails it has been trained to follow. The threat can become greater in the face of an adversarial prompt, in which an attacker attempts to cause a model to carry out a harmful action, such as revealing a password or other sensitive information. Instructions that normally wouldn’t be followed will, in some cases, be performed once the watermarking is deployed. The finding underscores the need for developers to thoroughly test how their LLMs and agents behave when watermarking is in place.

Changing safety behavior

“As compared to the same models without watermarking, it is definitely going to change their behavior, especially when we place it under adversarial conditions, or we make these models call tools when they’re powering an agent,” Andrea Siposova, an AI security researcher at Lasso Security, told Ars. “Watermarking is made to not be perceptible to a reader, but we know that when we are changing anything about what the model is generating, it is going to cause some tradeoffs, it’s going to show up somewhere.”

Read full article

Comments

© Getty Images

Yesterday — 17 September 2026Main stream

Cloudflare Just Gave AI Training Bots the Middle Finger

16 September 2026 at 17:18
Cloudflare just gave website owners a new weapon against AI crawlers: keep the search traffic, block the AI training. After years of watching bots consume the web’s content, publishers finally have an easier way to tell AI companies where to go.
Before yesterdayMain stream

Exclusive: Paying for frontier AI models buys 4-month head start at 5x the cost

15 September 2026 at 12:00

The performance gap between frontier AI models from US tech companies and the best open-weights models from Chinese companies has closed to just 4.4 months, according to a Mozilla report. That explains why many companies are shifting to the significantly cheaper open models for routine work—and helps reveal a narrow band of workloads where frontier models are worth the cost.

Most organizations should ideally be using open models as the default for the majority of their work, according to the latest State of Open Source AI report from Mozilla, published on September 15 and shared with Ars prior to publication. The report highlights how a leading open model, Moonshot AI’s Kimi K3, achieves a composite AI performance score on the Artificial Analysis Intelligence Index that is just three points behind Anthropic’s Fable 5 closed frontier model, all while costing just 30 percent of the latter.

“[A Closed model] earns its premium in a few places: expert professional work, high-intensity retrieval, and long context,” Raffi Krikorian, chief technology officer at Mozilla, said in an email to Ars. “We see the decision to pay for closed [models] as workload-specific rather than organization-specific.”

Read full article

Comments

© Imen Ben Youssef / Hans Lucas / AFP via Getty Images

AI leaders want to hit the brakes after years of reckless speed

14 September 2026 at 19:06

For years now, the major frontier AI labs have all been acting as if they're in an all-out, winner-take-all race with control of world-changing machine superintelligence (or at least market-changing artificial general intelligence) at the finish line. This weekend, the industry as a whole rapidly started turning away from that posture, urging coordination on slowing down the development of frontier AI that they say could soon be too dangerous and unknowable to control.

Anthropic's Dario Amodei was at the forefront of this change in tone, arguing in a nearly 4,000-word essay this weekend that "we must slow the pace at which we improve the capabilities of AI models" to avoid "a race to the bottom, spurred by commercial incentives, [that] can make [catastrophic] risks more acute."

Within hours, other AI leaders were echoing the same call. OpenAI co-founder and CEO Sam Altman posted his agreement on social media and said similar pacing discussions had been taking place at OpenAI. Alphabet Chief Scientist and Google DeepMind cofounder and chair Demis Hassabis said that Amodei's essay "points towards the right path forward," and renewed his own recent call for an industry-wide standards body. Microsoft CEO Satya Nadella posted that the company "welcome[s] the research, focus, and deliberate pacing needed to get alignment right as the design goal," ahead of the release of a lengthy "humanist AI" code of conduct for its models.

Read full article

Comments

© Getty Images

Claude users found ways around safeguards for bioweapons research

Anthropic said it stopped multiple attempts by scientists this year to use its technology for research that could help develop biological weapons, as experts increasingly fear the threat that AI poses to public safety.

The startup gave five examples of times actors “circumvented controls” and made other efforts to “obfuscate” the purpose of their research to dodge safeguards. The cases involved some users in nations that it prohibits from accessing its models, which include Russia, China, and Iran.

“We hope that by sharing these examples, we spark a conversation within the AI industry and with governments about emerging biological risks and how best to counter them,” Anthropic said in a report about efforts to use its models for malicious activity.

Read full article

Comments

© Getty Images | picture alliance

Anthropic researcher quits with a warning: Self-improving AI could "kill us all"

9 September 2026 at 16:59

When a prominent researcher quits a job at a frontier AI lab these days, it's often to pursue a new startup or protest a new business model. But AI researcher Jacob Coxon is using his departure from Anthropic to publicly warn that frontier AI companies are "gambling with our lives" with systems that they "earnestly believe... could kill us all by the end of the decade."

In a social media thread Tuesday night, Coxon said that this existential risk is inherent not so much in today's models but more in the impending prospect of "self-improving superintelligence" creating "superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources." Others working on these models have either not "internalized the civilizational stakes" or believe that they need to "speedrun" the race to superintelligence to prevent an irresponsible party from getting there first, he wrote.

Lest you think this is just one departing researcher expressing an unpopular opinion, Anthropic Alignment Science lead Evan Hubinger piped in on social media to say that "Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade."

Read full article

Comments

© Getty Images

Trump administration throws toys out of the pram after court rules its ‘woke’ Anthropic blacklist illegal

The Trump administration’s attempt to punish AI company Anthropic after it refused to allow its technology to be used for autonomous warfare and mass surveillance has been ruled illegal by a federal judge, exposing the administration’s claim that the company posed a national security threat as little more than retaliation against a critic.

Anthropic, the company behind Claude AI, was blacklisted by the Trump administration after a dispute over how its technology could be used by the US government and military. The company had refused to remove restrictions on the use of its AI models for lethal autonomous weapons and mass surveillance of Americans.

Rather than accept those limits, the Trump administration responded by taking sweeping measures against the company.

In March, Anthropic sued the administration, arguing that it had been designated a national security supply-chain risk in retaliation for exercising its First Amendment rights. The company said it had the right to express its views to both the public and the government about how its AI services should be used and about AI safety.

It also argued that the government had failed to follow procedures set out by Congress when designating it a supply-chain risk.

The court has now sided with Anthropic, finding that the government illegally retaliated against the company by designating it a supply-chain risk to national security.

The ruling vacated the government’s actions and ordered the administration to rescind its directives against Anthropic.

Trump and Defence Secretary Pete Hegseth ordered federal agencies to permanently stop using Anthropic’s products and barred defence contractors from doing business with the company, including business unrelated to military work.

Judge Rita Lin rejected the administration’s justification for those actions.

“Though the Department of War is undisputedly free to select the AI vendor of its choice, the evidence demonstrates that the broad measures imposed on Anthropic were illegal and baseless,” Lin wrote. “The empty invocation of national security is not a blank check to punish and retaliate against government critics.”

For an administration that has repeatedly presented political disagreement as a threat to national security while portraying its opponents as part of some amorphous “radical left,” the White House’s response to Anthropic’s lawsuit was as predictable as they come.

“President Trump will never allow a radical left, woke company to jeopardise our national security by dictating how the greatest and most powerful military in the world operates,” it said.

And for all the Trump administration’s talk of fighting “woke” corporations, the episode raises the question: who, exactly, is being ideological when a government responds to a company’s concerns about the use of its technology by attempting to destroy its access to government business?

On the evidence of this ruling, it was not Anthropic that overstepped the line.

The post Trump administration throws toys out of the pram after court rules its ‘woke’ Anthropic blacklist illegal appeared first on Left Foot Forward: Leading the UK's progressive debate.

Anthropic’s $2 trillion IPO puts powerful external trustees in spotlight

Anthropic’s prospective public-market investors must reckon with an external group of trustees that control the majority of the AI company’s board, as its planned blockbuster initial public offering forces close scrutiny of its experimental governance structure.

The company’s Long-Term Benefit Trust (LTBT) is a small group of advisers created to safeguard the lab’s mission of developing AI for the long-term benefit of humanity, even as commercial pressures intensify.

The trust holds no equity in Anthropic, but has significant influence, with the San Francisco-based company planning to preserve its role after a stock market debut that could value the Claude maker at as much as $2 trillion.

Read full article

Comments

© Imen Ben Youssef / Hans Lucas / AFP via Getty Images

Four major AI models suffer rare overlapping downtime

3 September 2026 at 18:10

Cloud-based AI models operated by OpenAI, Anthropic, xAI, and Google suffered a rare and overlapping set of significant service interruptions over a period of hours Thursday morning.

Anthropic first reported a "partial outage" related to "elevated errors on requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5" at 9:23 am (all times Eastern). The company reported that it had "identified the cause" of the error roughly 15 minutes later, before reporting that "a fix has been deployed" and the issue was resolved by 12:16 pm. A separate incident report indicated "elevated errors on requests to Claude Sonnet 5" for a brief period just after noon.

OpenAI, meanwhile, reported "elevated errors across ChatGPT and Codex" were resulting in "degraded performance" as of 10:43 am Thursday morning. A mitigation put in place a little more than half an hour later led to the issue being marked as "resolved" by 12:55 pm.

Read full article

Comments

© Getty Images

Trump may be forced to reveal secret rules feds use for AI safety testing

2 September 2026 at 17:58

Four federal agencies have been sued amid calls to release information about the secret framework that the Trump administration uses to conduct safety reviews of frontier AI models prior to release.

In a Wednesday press release announcing the lawsuit, a nonpartisan nonprofit called Protect Democracy alleged that “almost no details” have been released to the public or Congress. To everyone except a few vague “trusted partners,” it remains unclear what the government’s review process looks like, which companies are involved in constructing the framework, or what legal authority Trump officials have to conduct the reviews.

“Neither the identities of those entities nor the criteria by which they were selected have been made public,” Protect Democracy said.

Read full article

Comments

© Bloomberg / Contributor | Bloomberg

“Zlibrary my beloved”: Anthropic staff chats extolling piracy cited in Sony suit

31 August 2026 at 18:10

Some of the world’s leading music publishers think that Anthropic got off too light in a historic settlement where the Claude maker paid authors $1.5 billion after admitting to pirating more than 7 million books to train AI.

“$1.5 billion is obviously not a large enough settlement to deter infringing conduct by a company that has parlayed such mass infringement into a staggering $2-trillion-dollar valuation,” music publishers said in a lawsuit filed Friday.

Music publishers suing Anthropic include Sony, EMI, and Warner Chappell. They alleged that Anthropic’s illegal torrenting also included “thousands upon thousands” of their copyrighted musical compositions.

Read full article

Comments

© Steven Errico | Photographer's Choice RF

Trump blacklisting of "woke" Anthropic deemed illegal by federal judge

28 August 2026 at 18:07

The Trump administration's blacklisting of Anthropic was illegal, a federal judge ruled in an order vacating government directives against the use of the firm's AI technology.

The government illegally retaliated against Anthropic by designating it a supply-chain risk to national security, said yesterday's ruling by Judge Rita Lin in the US District Court for the Northern District of California. The maker of Claude AI technology was barred by the US after it refused to drop restrictions on the use of its products for lethal autonomous warfare and mass surveillance of Americans, the ruling said.

"The undisputed record shows that the challenged actions constituted unlawful retaliation in violation of the First Amendment," Lin wrote in an order that granted key portions of Anthropic's motion for summary judgment.

Read full article

Comments

© Getty Images | picture alliance

Anthropic's new hardware standard lets AI agents control the physical world

27 August 2026 at 22:15

For all the interest in and uptake of agentic AI systems over the past year or so, the world of automated AI has thus far been primarily limited to text, images, code, and other data and actions that take place inside a computer. Anthropic is now aiming to change that somewhat with what it's calling the Model Hardware Standard (MHS), a set of standardized drivers designed to let AI agents easily interface with and control arbitrary devices.

For now, the "research preview" of the MHS effort is being sold mainly as a way to help scientists streamline the arduous process of creating the custom software integrations that are often needed to get disparate components of an experiment working in concert. MHS can provide a common interface and common format for data sharing between these devices, Anthropic says, allowing them to talk to each other across a network "without needing a bespoke 'translator' program in between." The standardized system could reduce weeks or months of exacting experimental setup down to "hours or minutes," Anthropic writes.

An Anthropic graphic illustrating how MHS serves as a "translation" layer between AI agents and multiple types of devices. Credit: Anthropic

In a video posted alongside the announcement, Anthropic Technical Staffer Alek Kemeny says the MHS effort was inspired by observing neuroscientist Arco Bast work through an experiment on memory formation in the brain at the HHMI Janelia Research Campus in Ashburn, Virginia. Kemeny said Bast had worked out an interface to get the rotating laser beams, microscopes, cameras, and myriad other components of the experiment to coordinate through a common interface. "This idea could be used to have AI run any science experiment in the world," Kemeny recalls thinking at the time.

Read full article

Comments

© Anthropic

OpenAI and Anthropic in price war as Chinese AI rivals gain ground

14 August 2026 at 14:27

Leading US AI labs such as OpenAI and Anthropic are releasing cheaper models as they fight to retain cost-conscious customers who are switching to cut-price alternatives from Chinese rivals.

The price war comes as rising AI bills push companies to curb usage and seek cheaper models, helping Chinese developers including Moonshot and DeepSeek make inroads with users from Silicon Valley to Europe.

OpenAI recently said that it was slashing prices for GPT-5.6 Luna, its “fastest and most affordable model”, by 80 percent. Anthropic has launched Claude Opus 5, touting the system’s “frontier intelligence... at half the price” of Fable 5, the company’s most capable model.

Read full article

Comments

© MARTIN LELIEVRE

Anthropic could be worth $2 trillion when it goes public

Anthropic investors expect the AI startup to float at a valuation of $2 trillion or more in October, a dizzying figure that would eclipse SpaceX and make the AI lab’s debut the largest-ever initial public offering.

Half a dozen of the company’s backers told the FT that Anthropic’s rapidly rising revenue would enable it to more than double its current valuation in a planned autumn float.

A listing at that level could unlock billions of dollars in gains for the five-year-old company’s early investors but would also test public markets that are growing more nervous about the AI boom.

Read full article

Comments

© Getty Images | picture alliance

Claude's new Scarlet Letter watermark is invisible—for now

13 August 2026 at 11:10

Anthropic has revealed that it will soon watermark content that is processed (not just generated!) by any of its models. In a support article, Anthropic explained that it was rolling out machine-readable watermarks to comply with the European Union’s AI Act, which requires all AI system providers to watermark AI-generated or manipulated audio, image, text, and video outputs. The law applies to any AI model released after August 2 and provides a grace period until December 2026 for providers to update previously released models.

Anthropic confirmed that moving forward, all new models offered globally—not just in the EU—will mark AI-generated content “from day one.” Text outputs will “carry embedded watermarks,” invisible to the user, and other “generated files will include digitally signed provenance metadata where supported,” Anthropic said.

Notably, Anthropic is deploying a "nuke it from orbit" approach, applying the watermarks to all processed content where supported, even though the EU does not require it for cases where an AI system performs "an assistive function for standard editing" (the guidance's own example is grammar correction), or where it doesn't "substantially alter" the user's text or its meaning.

Read full article

Comments

© Aurich Lawson | American Psycho (Lions Gate Films)

Booksellers suspect AI firms are buying and then destroying rare books

12 August 2026 at 15:19

If you can truly appreciate an old book—and maybe even marvel at how its fragile, yellowing pages contain some of the earliest ways that people tried to make sense of the world around them—then headlines about tech companies that are destroying books to train AI likely torture a tender part of your soul.

It’s indeed depressing to imagine piles of book spines waiting to be fed into wood chippers while torn-out pages are cropped, scanned, and trashed. But that’s the cheapest and easiest way to scan books as fast as possible, and AI companies are in a race to advance their models by training on the kind of engaging, high-quality long-form texts that can only be found in books. So book lovers fear it’s likely that the practice is happening on a grander scale than is currently being reported and that some physical copies of books will be lost forever.

What makes this destruction extra painful, though, is that it doesn’t have to be this way.

Read full article

Comments

© hexvivo | iStock / Getty Images Plus

AI chatbots have failed people in crisis. Can that be fixed?

7 August 2026 at 13:49

This year alone, there have been numerous known instances—often via lawsuits—of AI chatbots (most often, OpenAI’s ChatGPT) that have gone horrifically wrong.

A January lawsuit described the story of a man who took his own life after being allegedly “coached” into suicide. A college student in Georgia sued OpenAI, claiming that ChatGPT “pushed him into psychosis.”

In June, a Canadian family also sued OpenAI and argued that ChatGPT agreed with the young woman’s dismissiveness when it first gave her the option to seek professional mental health advice. ChatGPT allegedly “encouraged” her to end her life, too, and she did so.

Read full article

Comments

© Oscar Wong via Getty

Anthropic will design its own hardware to power Claude

6 August 2026 at 20:03

Anthropic is hiring a "custom silicon team" to design chips on which to run its models, the company has revealed.

Yesterday, Business Insider noticed a job listing for a senior engineer with experience shipping semiconductor designs. (You can see listings for a silicon engineer and a technical program manager, silicon on Anthropic's job board right now.) A spokesperson for Anthropic then confirmed the plans to both Business Insider and TechCrunch.

Read full article

Comments

© Chesnot via Getty Images

❌
❌