Normal view

There are new articles available, click to refresh the page.
Yesterday — 18 September 2026Main stream

Microsoft exec called AI scraping the “largest theft of labor in human history”

17 September 2026 at 20:10

For years, Microsoft and OpenAI have fought to keep certain information out of the public eye in their fight with news organizations that have accused the AI firms of teaming up to violate copyright laws by stealing tons of news content to train AI.

However, now the details that should never have been marked confidential are starting to leak. In a motion for summary judgment that was unsealed Thursday from news plaintiffs led by The New York Times, internal documents are exposed that news groups alleged show exactly how Microsoft and OpenAI viewed the threat to news before unleashing new AI products like ChatGPT and Copilot.

Perhaps most explosively, Microsoft Director of Applied Science Brent Hecht repeatedly warned in documents that scraping news for AI training was “an astonishing theft of unprecedented proportions,” calling it perhaps the “largest theft of labor in human history,” news orgs said. In another document, Hecht contradicted Microsoft and OpenAI’s argument that training AI on news content is fair use, suggesting that the plan to widely scrape news made “a complete mockery of the idea of ‘fair use.’”

Read full article

Comments

© Charles Taylor | iStock / Getty Images Plus

Before yesterdayMain stream

EFF to Courts: Don’t Rewrite Copyright Over AI Hype

The history of technology is rife with copyright panics.  In the 1980s, major rightsholders ran to Congress and the courts, claiming that videotape recorders (VTR) were “to the American film producer and the American public as the Boston strangler is to the woman home alone.” Then, the Supreme Court declined to embrace the hype, noting that the VTR was capable of all kinds of non-infringing uses, like time-shifting and cautioning courts to avoid rewriting copyright law in response to new technologies. We believe that courts now should be similarly wary about the hype surrounding AI.

Hollywood’s hyperbole has echoed that of composer John Phillip Sousa, who claimed in 1906 that the player piano and the gramophone would destroy music composition; portrait artists who feared the camera would replace the paintbrush. None of these things happened. Cameras, for example, sparked a resurgence of portraiture and, by making it possible for more people to create images, led to unexpected developments—like the rise of photojournalism.

New markets, new ideas, and new creators are actually what copyright is supposed to promote, not restrict. Using copyright to lock in existing gatekeepers and massive rightsholders’ profits helps neither the public nor individual artists.

Generative AI has sparked the latest wave of anxiety and with it a massive wave of litigation. In multiple cases around the U.S. and the world, rightsholders are asking courts to do precisely what the Supreme Court warned against: dramatically expand copyright protections based in substantial part on hyperbole and speculation. They should decline to do so.

Copyright owners claim that unless courts abandon 300-year-old copyright principles—and give rightsholders the power to control non-infringing works created by others—an imagined flood of AI-generated works will devastate creative markets. Under this “market dilution” theory, building generative AI tools cannot be fair use because those tools might be encourage the proliferation of competing works.

As EFF has explained to the courts in multiple amicus briefs in Concord Music Group, Inc. v. Anthropic PBC and In re Mosaic LLM Litigation, that’s not how copyright works. In fact, accepting this theory would undermine copyright’s constitutional purpose: promoting the creation of expressive works for the public’s benefit. Because copyright law is designed to encourage others to build freely on existing works, it punishes infringement, not competition. The “market dilution” theory would eviscerate not only the fair use doctrine, but also other limits on copyright that work specifically to prevent rightsholders from unfairly suppressing competition by claiming broad ownership over tropes, genres, styles, and so on. In other words, publishers would wield unchecked veto power over any expression that might conceivably compete with a work they own.

The result? Art doesn’t get created, ideas are never expressed, and we’re all worse off. Copyright shouldn’t be a tool to silence future creative competitors—whether or not they use AI in their work.

And the plaintiffs in these cases get at least two other things wrong. First, research shows that large generative AI models are unlikely to produce infringing works because the more data on which a model is trained, the less any individual training example matters to any particular output.

Second, AI tools aren’t necessarily displacing human creativity. To take a just a few examples:

  • Boston-based artist Nettrice Gaskins uses AI to create Afro-futurist art, including a portrait of Octavia Butler displayed at the San Francisco Airport
  • Indian artists Prateek Arora and Varun Gupta use generative AI to reimagine Western science fiction.
  • Philadelphia-based artist Alex Smith uses generative AI to reimagine Afrofuturism with queer, plus-sized Black superheroes.
  • Ana Miljački, a professor of architecture at MIT, used generative AI to create a “non-liner documentary” film on Yugoslav World War II memorials and the values they embodied.
  • A research-creation project used AI generated visual art to both amplify the voices of activists in the Iran Woman Life Freedom Movement and evaluate AI’s role in sociopolitical advocacy through art.
  • AI company Bronze works with musicians like Disclosure and Jai Paul to create songs that never sound the same when played back twice, challenging audience conceptions of what music could be.

It is not the place of courts to say these people are not artists or that AI cannot augment human creativity in a positive way.

Given this range of experimentation, courts should be reluctant to decide in advance what tools do and do not foster “human creativity.” Like the VTR, large language models are general purpose tools, used by humans to do a broad variety of things far beyond generating lyrics. The effects of this particular technological innovation will doubtless be far-reaching, disruptive, and potentially harmful for some—but distorting copyright law is not the way to address those harms.

EFF Thanks SerpApi For Helping Us Protect Free Speech Online

EFF is grateful for SerpApi’s generous support, helping us fight for your rights to speak and access information online. SerpApi has been giving to EFF every year since 2018, and alongside our 32,000 individual donors, their gift is critical to keeping up the fight.

Whether in the courts, halls of power, or broader policy debates, we appreciate the work this support has made possible over the years. Some examples:

  • We sued the U.S. Department of Homeland Security and Department of State to stop an unconstitutional social media surveillance program to identify and punish individuals who express viewpoints the government disagrees with.
  • We helped develop the Santa Clara Principles, a framework to reign in overbroad content moderation so that all users are treated fairly and offered consistent tools for recourse if their speech is censored by tech companies.
  • In the whitepaper Unfiltered: How YouTube’s Content ID Discourages Fair Use and Dictates What We See Online, we pushed back on YouTube for silencing individual creators in the interest of protecting a small number of giant copyright holders.
  • We stood with whistleblowers and dissidents persecuted for their online speech.
  • We continued the fight to protect Section 230.

We live in an era when lawful speech and the right to access information are being targeted by Big Tech and governments around the world that are hostile to dissent. Free speech online is core to EFF’s mission, and SerpApi’s support will help us continue the fight to protect everyone’s right to free expression.

Enshittification Merch That Actually Fights Enshittification 

Enshittification isn't just a sweary word to describe the accelerating decay of the online platforms, apps, and services that we rely on.  

It's a framework for understanding the structural incentives that make tech companies enemies of their own users over time—the surveillance business model, the erosion of privacy, the monopoly power that eliminates alternatives, the regulatory capture that prevents accountability.  

SUPPORT EFF

GET LimITED EDITION MERCH + FIGHT ENSHITTIFICATION

These are some of EFF's core fights and have been for over 35 years. EFF sues. EFF advocates. EFF codes. And EFF wins. EFF is the most profound and powerful disenshittifying force on the planet Earth, and I’ve been proud to fight alongside them for nearly 25 of those years.  

One of the lessons you learn in battles with very long timelines against very powerful actors is that these battles are deeply serious, and because of that they must also be fun. “Enshittification” took off as a shorthand in part because of the minor license to vulgarity it confers. It's slightly crass for a reason: getting people to engage with the abstract issues of tech policy can be hard at the best of times. No one knows this better than my colleagues at EFF, who consistently surprise me with their ability to make complex, technical concepts concrete, memorable, and sometimes even joyful. 

Words matter, but so do visuals. For the cover of the U.S. edition of my book, Enshittification, designer Devin Washburn of No Ideas studio created an iconic variation of the "pile of poo" emoji, with angry eyebrows and a grawlix-scrawled censor bar over its mouth. It instantly became the symbol of enshittification I’d been looking for. 

A digital illustration of an angry poop emoji holding a black sign reading "&!#%", set against a blue and gray background tiled with oversized "& !#%" characters.

I liked it so much I ordered a couple hundred enamel pins and a couple thousand vinyl stickers and handed them out to people I met on my 33-city book tour. Even when giving them away, I was inundated with requests to buy more of them.  

I've since bought out Devin's rights to the image and released it under a Creative Commons Attribution 4.0 license—free for anyone to use, remix, or build on, including commercially, with attribution. The high-resolution files are on Wikimedia CommonsFlickr, and the Internet Archive (including a PSD with an ink-density adjustment layer). It belongs to the commons now. 

But I made sure EFF had first crack at the design for their “official merch,” and they've done right by it. There are two items available now in the EFF shop, and all proceeds go directly to EFF's work defending digital rights. I’ve spent years admiring EFF’s merch and consistent, creative visual identity, so it fills me with pride to see this more-than-a-mere-poop-emoji in their shop.  

A recognizable visual shorthand is a genuine organizing tool. When someone sees the enshittification emoji, they know what the conversation is about. When you wear the pin or slap the sticker on your laptop, you're signaling that you understand what's happening to the internet, and that you know we can do better.  

You can get a $5 sticker:

An angry poop emoji sticker affixed to the coiled spring mechanism inside a vending machine. The sticker depicts a scowling poop emoji holding a black sign reading "&$!#%".

A hand with black nail polish and a gold ring holds an angry poop emoji sticker against a white door. The sticker shows a scowling poop emoji holding a black sign reading "&$!#%".

Or a $10 pin:

A close-up of an enamel pin on the lapel of a tan jacket. The pin depicts an angry poop emoji holding a black banner reading "&$!#%".

An enamel pin clipped to the nose bridge of black-framed sunglasses resting on a wooden surface. The pin depicts a scowling poop emoji holding a black banner reading "&$!#%". 

 Because the design is CC-licensed, you don't have to buy one. You can make your own merch, your own swag, your own illustrations. I made a lawn flag for my front garden.

A  small white garden flag on a metal stake, planted among cacti and succulents in a sunny yard. The flag depicts an angry poop emoji holding a sign reading "&$!#%". 

But if you do want to buy a sticker or pin, you can do so while supporting the most profound and powerful disenshittifying force on the planet Earth—the Electronic Frontier Foundation.

SUPPORT EFF

GET LimITED EDITION MERCH + FIGHT ENSHITTIFICATION

 

Tell Congress: Just Say No to NO FAKES

The Senate Judiciary Committee is set to consider and vote on the Nurture Originals, Foster Art, and Keep Entertainment Safe Act (NO FAKES). Instead of targeting the real privacy harms posed by AI-generated replicas, this law would create another layer of internet censorship on top of the already existing legal and voluntary takedown systems. Congress should reject NO FAKES.

Take action

Tell Congress to Say No to NO FAKES

As currently written, NO FAKES proposes to tackle the problems of misleading AI-generated replicas by creating a broad property right in someone's look, voice, and general style. However, there are all kinds of First Amendment-protected expression that would be swept under the NO FAKES regime—think about parody, news, criticism.

NO FAKES also does a laughable job of protecting artists from use of their image in misleading ways. It doesn’t create a privacy right, but rather a property right that can easily be signed away—as major studios and record labels are almost certain to require in their contracts with artists. As a result, NO FAKES actually creates a new avenue for the exploitation of artists by companies instead of protection from misleading replicas. 

The bill also makes it trivially easy for protected speech to be censored. It is a supercharged version of the already flawed copyright takedown regime. It would essentially require platforms to institute filters that don't just look for exact matches of copyrighted material, as current filters do, but anything that might be a digital replica. Even though the latest version of this bill adds some forms of redress for bad faith takedowns, those provisions lack the teeth required to deter a malicious actor. 

NO FAKES targets speech, tools, and innovation instead of focusing on the real concern posed by these replicas: privacy. This bill was a bad idea when it was introduced, and got even worse when it was amended last year. Tell Congress to just say no to NO FAKES.

Take action

Tell Congress to Say No to NO FAKES

California’s AB 412 Still Demands Developers Do The Impossible

California lawmakers are again considering A.B. 412, a bill that would require AI developers to identify and disclose copyrighted works used to train generative AI systems.

The problem this year is the same as last year: it’s practically impossible to comply with this law. The bill demands information that often does not exist, and cannot realistically be obtained. 

EFF submitted an opposition letter to the California Senate Privacy Committee explaining why we continue to believe A.B. 412 is simply unworkable. To the extent developers do follow this law, it will have the effect of locking in the power of the largest companies in AI. 

A Burden That Can’t Be Met

A.B. 412 sounds simple: just have AI developers create and keep a list of all the registered copyrighted works they use in AI training. 

That may seem straightforward. In practice, it’s anything but. 

There is no machine-readable “list” of copyrighted works at the U.S. Copyright Office. And many copyright holders can get a copyright without even depositing a publicly viewable sample of the work—for example, software companies may register copyright on proprietary code without revealing it to the public. 

And on the open internet, copyright information is often incomplete, unavailable, or impossible to verify. One image may be registered with the copyright office, while the next is licensed under a free Creative Commons license (like the images that EFF creates), and the next is public domain. A message forum user might post an original story, photograph, or poem without any indication of ownership or registration status. 

The bill effectively asks developers to continuously cross-reference massive batches of online data against a copyright system that simply wasn’t designed to do so. If California passes A.B. 412, its impact will go far beyond the large AI companies we read about in the headlines. 

Not Just Big Tech

Supporters often frame this bill as a way to help creative workers have some leverage against Big Tech, but the bill reaches much further than the big AI companies. 

Its definition of “developer” extends to anyone who makes a generative AI model available to Californians. That includes indie developers tinkering with an existing model, open-source initiatives, nonprofits, and other non-commercial efforts. Recent amendments added exemptions for universities and government entities, which is important, but that still leaves out a vast swathe of non-commercial tech work that’s done by people without full-time jobs in government or academia. 

Large companies will hire compliance teams and lawyers to navigate these requirements. Smaller organizations and independent developers usually can’t. The result will be fewer opportunities for startups and new entrants. Faced with this massive compliance burden, some won’t even try. 

Courts Are Already Deciding These Questions

The bill is premised on the idea that copyright owners currently don’t have good remedies if they’re mistreated by AI companies. That simply isn’t true. And the growing wave of federal court filings in this space prove it. Content companies that want to sue tech companies, large or small, have no problem doing so. Those courts are still working through important questions about fair use and transformative use. Some courts have already concluded that many AI training activities qualify as fair use. Others continue to evaluate the issue.

California lawmakers should not rush to impose new state regulation while those questions remain unresolved. This is why copyright is governed at the federal level: both creators and fair users benefit from a single set of nationwide rules. 

At this point, the bill remains a solution in search of a problem. Rights holders already have powerful tools to protect their interests under existing federal law. What this bill adds isn’t clarity or transparency, but a costly and essentially impossible compliance burden that will discourage small developers and researchers. 

California has been able to support both artistic creativity and tech innovation for decades now.  But A.B. 412 does not strike the right balance. 

If you are a California resident and interested in speaking out about this bill, you can find and contact your representatives through this website

Another Court Rules Copyright Can’t Stop People From Reading and Speaking the Law

Another court has ruled that copyright can’t be used to keep our laws behind a paywall. The U.S. Court of Appeals for the Third Circuit upheld a lower court’s ruling that it is fair use to copy and disseminate building codes that have been incorporated into federal and state law, even though those codes are developed by private parties who claim copyright in them. The court followed the suggestions EFF and others presented in an amicus brief, and joined a growing list of courts that have placed public access to the law over private copyright holders’ desire for control.

UpCodes created a database of building codes—like the National Electrical Code—that includes codes incorporated by reference into law. ASTM, a private organization that coordinated the development of some of those codes, insists that it retains copyright in them even after they have been adopted into law, and therefore has the right to control how the public accesses and shares them. Fortunately, neither the Constitution nor the Copyright Act support that theory. Faced with similar claims, some courts, including the Fifth Circuit Court of Appeals, have held that the codes lose copyright protection when they are incorporated into law. Others, like the D.C. Circuit Court of Appeals in a case EFF defended on behalf of Public.Resource.Org, have held that, whether or not the legal status of the standards changes once they are incorporated into law, making them fully accessible and usable online is a lawful fair use.

In this case, the Third Circuit found that UpCodes’s copying of the codes was a fair use, in a decision closely following the D.C. Circuit’s reasoning. Fair use turns on four factors listed in the Copyright Act, and the court found that all four favored UpCodes to some degree.

On the first factor, the purpose and character of the use, the court found that UpCodes’s use was “transformative” because it had a separate and distinct purpose from ASTM—informing people about the law, rather than just best practices in the building industry. No matter that UpCodes was copying and disseminating entire safety codes verbatim—using the codes for a different purpose was enough. And UpCodes being a commercial venture didn’t change the outcome either, because UpCodes wasn’t charging for access to the codes.

On the second factor, the nature of the copyrighted work, the Third Circuit joined other appeals courts in finding that laws are facts, and stand at “the periphery of copyright’s core protection.” And this included codes that were “indirectly” incorporated—meaning that they were incorporated into other codes that were themselves incorporated into law.

The third factor looks at the amount and substantiality of the material used. The court said that UpCodes could not have accomplished its purpose—providing access to the current binding laws governing building construction—without copying entire codes, so the copying was justified. Importantly, the court noted that UpCodes was justified in copying optional parts of the codes as well as “mandatory” sections because both help people understand what the law is.

Finally, the fourth factor looks at potential harm to the market for the original work, balanced against the public interest in allowing the challenged use. The court rejected an argument frequently raised by copyright holders—that harm can be assumed any time materials are posted to the internet for all to access. Instead, the court held that when a use is transformative, a rightsholder has to bring evidence of harm, and that harm will be balanced against the public benefit. Because “enhanced public access to the law is a clear and significant public benefit,” and ASTM hadn’t shown significant evidence that UpCodes had meaningfully reduced ASTM’s revenues, the fourth factor was at least neutral. It didn’t matter to the court that ASTM offered to provide copies of legally binding standards to the public on request, because “the mere possibility of obtaining a free technical standard does not nullify the public benefits associated with enhanced access to law.”

This is a good result that will expand the public’s access to the laws that bind us—something that’s more important than ever given recent assaults on the rule of law. In the future, we hope that courts will recognize that codes and standards lose copyright when they are incorporated into law, so that people don’t have to spend years and legal fees litigating fair use just to exercise their rights.

❌
❌