Normal view

There are new articles available, click to refresh the page.
Today — 18 September 2026Main stream

Microsoft exec called AI scraping the “largest theft of labor in human history”

17 September 2026 at 20:10

For years, Microsoft and OpenAI have fought to keep certain information out of the public eye in their fight with news organizations that have accused the AI firms of teaming up to violate copyright laws by stealing tons of news content to train AI.

However, now the details that should never have been marked confidential are starting to leak. In a motion for summary judgment that was unsealed Thursday from news plaintiffs led by The New York Times, internal documents are exposed that news groups alleged show exactly how Microsoft and OpenAI viewed the threat to news before unleashing new AI products like ChatGPT and Copilot.

Perhaps most explosively, Microsoft Director of Applied Science Brent Hecht repeatedly warned in documents that scraping news for AI training was “an astonishing theft of unprecedented proportions,” calling it perhaps the “largest theft of labor in human history,” news orgs said. In another document, Hecht contradicted Microsoft and OpenAI’s argument that training AI on news content is fair use, suggesting that the plan to widely scrape news made “a complete mockery of the idea of ‘fair use.’”

Read full article

Comments

© Charles Taylor | iStock / Getty Images Plus

Yesterday — 17 September 2026Main stream

Cloudflare Just Gave AI Training Bots the Middle Finger

16 September 2026 at 17:18
Cloudflare just gave website owners a new weapon against AI crawlers: keep the search traffic, block the AI training. After years of watching bots consume the web’s content, publishers finally have an easier way to tell AI companies where to go.
Before yesterdayMain stream

Flight attendants freaked out that Google is buying tons of Spirit employee data

19 August 2026 at 20:04

Last Friday, Google won an auction to acquire a huge amount of Spirit Airlines data.

The data doesn’t include personal information or customer data, but instead nearly covers the airline's entire employment and workplace record.

To ensure that no individual can be identified in the dataset, Google agreed to use a court-appointed ombudsman to oversee a process to strip any personally identifying information (PII) from the data before it’s transferred to Google. Under the deal, Google agreed to maintain the data in this de-identified form and to never intentionally re-identify the data. And if Google sells access to the data, third parties would supposedly be bound by the same terms.

Read full article

Comments

© Sarah Rice / Stringer | Getty Images News

Hidden Airtag reveals Amazon is trashing rare books to train AI

17 August 2026 at 18:13

For the past year or so, booksellers have suspected that AI firms are buying up huge lots of rare books, then destroying them after scanning them to train AI. But this was hard to prove until now, as 404 Media reports that an Airtag hidden in a rare book shows that at least one tech giant, in the race to advance its frontier models, is behind some of the bulk orders: Amazon.

On Monday, 404 Media revealed that it had connected with a bookseller who agreed to plant an Airtag in a rare book that was part of a bulk order. That Airtag was then tracked to an Amazon AI training facility in Las Vegas that housed a team focused on tearing books from their spines and scanning pages, 404 Media reported. Apparently tone-deaf to the escalating backlash over destructive book scanning, a logo on the door of that team’s warehouse, VGT3, showed a Tyrannosaurus rex preparing to devour a book, 404 Media documented.

Amazon deflects

Amazon declined to comment on 404 Media’s findings, only providing Ars with the same statement it gave to 404 Media, which does not mention AI training specifically.

Read full article

Comments

© fotek | iStock / Getty Images Plus

Booksellers suspect AI firms are buying and then destroying rare books

12 August 2026 at 15:19

If you can truly appreciate an old book—and maybe even marvel at how its fragile, yellowing pages contain some of the earliest ways that people tried to make sense of the world around them—then headlines about tech companies that are destroying books to train AI likely torture a tender part of your soul.

It’s indeed depressing to imagine piles of book spines waiting to be fed into wood chippers while torn-out pages are cropped, scanned, and trashed. But that’s the cheapest and easiest way to scan books as fast as possible, and AI companies are in a race to advance their models by training on the kind of engaging, high-quality long-form texts that can only be found in books. So book lovers fear it’s likely that the practice is happening on a grander scale than is currently being reported and that some physical copies of books will be lost forever.

What makes this destruction extra painful, though, is that it doesn’t have to be this way.

Read full article

Comments

© hexvivo | iStock / Getty Images Plus

Anthropic’s $1.5B copyright settlement approved; only 350 authors opted out

21 July 2026 at 17:33

On Monday, a judge approved a $1.5 billion settlement between Anthropic and authors, ending the largest copyright class-action ever certified and granting the largest copyright settlement ever reached.

Back in May, some authors fought to block the settlement, which was proposed after the court ruled that Anthropic training AI on books was fair use; however, its piracy of works was likely not.

Authors opposing the settlement argued that lawyers’ fees were too high and authors’ payouts were too low. Hoping to avoid accepting the estimated $3,000-per-work payout and file separate lawsuits to seek higher damages, a handful of authors tried to opt out past the deadline.

Read full article

Comments

© yalcinadali | iStock / Getty Images Plus

OpenAI may have made a fatal misstep in copyright fight with news orgs

OpenAI is facing calls for "serious sanctions" after fighting to keep news organizations from snooping through millions of logs to find evidence of users skirting their paywalls by prompting ChatGPT to regurgitate their articles.

This evidence is considered among the most important to both sides, potentially either dooming OpenAI as an infringer or exonerating its chatbot technology as a transformative fair use of news sites' content.

In a sanctions motion Thursday, news organizations suing OpenAI—led by The New York Times—accused the AI firm of repeatedly lying for years to conceal evidence of infringement that could hobble OpenAI's defense.

Read full article

Comments

© Alona Horkova | iStock / Getty Images Plus

Musk’s X poses “serious risk to Americans’ privacy,” advocates warn FTC

Ahead of a July 2 deadline to submit public comments, advocates are warning the Federal Trade Commission that it must keep close watch over Elon Musk’s X and firmly reject a recent bid to end the agency’s ongoing audits of the platform’s data handling.

Last month, the FTC posted a notice explaining that X had argued that an FTC order was no longer necessary due to changes Musk had made to the platform.

The initial order came as a penalty after the FTC found that a coding error had caused then-Twitter to improperly share users’ contact information for ad targeting that had initially been submitted for two-factor authentication. Under the order, X is subjected to costly independent audits, and the FTC has authority to demand documents to ensure compliance with data privacy laws without taking additional legal action.

Read full article

Comments

© CHARLY TRIBALLEAU / Contributor | AFP

❌
❌