Economy & Policy

Publishers Now Have AI Buyers for Content—and Zero Say on Price

Unsealed filings show OpenAI and Microsoft knew what they were taking. Meanwhile a $1B scraper economy pays publishers almost nothing—and buyers still set the price.

By Grace Kim

4 min read

Updated

What's News

  • Microsoft's Brent Hecht warned in a 2023 memo that AI models hoovering up work would be 'the largest theft of labor in human history.'
  • Data brokers like Exa, Parallel and Tavily form a scraper economy estimated at roughly $1 billion, largely bypassing publishers.
  • Meta accounted for 46.3% of AI bot traffic tracked by DataDome in H1, ahead of OpenAI, while also holding publisher deals with USA Today, CNN and Fox News.

Meta generated 46.3% of all AI bot traffic tracked by DataDome in the first half of the year, well ahead of OpenAI—while simultaneously signing content deals with USA Today, CNN, Fox News and People Inc. That is the core paradox of the emerging market for publisher content: the AI industry's biggest data harvester is also a buyer, and it gets to decide when.

The context around that paradox got sharper on September 17, when a brief made public in The New York Times's lawsuit against OpenAI and Microsoft revealed internal documents showing executives at both companies understood exactly what they were taking. Microsoft's Brent Hecht, a director of applied science, warned in a memo in early 2023 that "millions of people around the world will soon consider large models 'hoovering up' all their work to be an astonishing theft," and called it "the largest theft of labor in human history." When a researcher described getting around the Times paywall, OpenAI President Greg Brockman replied, "ah nice." Nick Turley, OpenAI's head of ChatGPT, called chatbots an "existential threat" to publishers.

The government, for its part, has picked a side. The Department of Justice filed a statement of interest in the case arguing that training AI models on publishers' content is fair use—a position consistent with the administration's view that any concession on copyright would help China win the AI race. President Donald Trump's own summary: "China's not doing it."

But training is only part of the economics. The DOJ itself conceded that "an output reconstructing and disseminating an original copyrighted work may not be transformative," a near description of AI search, which is essentially a machine for summarizing current reporting. The unsealed documents also speak to market harm, another fair-use pillar. Turley wrote that OpenAI's products "are largely substitutive, period." Microsoft CEO Satya Nadella testified that using chatbots "has substituted" for visiting original sources—and said in the court documents, "anything that is paywalled should be licensed."

The Times filed its lawsuit in December 2023, when the debate centered on training. The legal ambiguity since then is precisely why a market for training data never emerged: it is hard to justify investing in a payment framework for something that might be free in a few months. Training is where the lawsuits are. Inference—AI answers about real-time content—is where the money is.

Except that money is mostly bypassing the media. A class of data brokers—firms like Exa, Parallel and Tavily—scrape the internet at scale and resell the data to AI companies, ad agencies, investment firms and even other publishers. One estimate, cited in Matthew Scott Goldstein's widely circulated report on the scraper economy, pegged the market at about $1 billion. Brian Morrissey at the Rebooting argues the AI industry stays away from buying directly because most content isn't unique enough—the AI only needs one world-class enchilada recipe. Commoditization drives a race to the bottom, and the bottom sits close to zero.

The pressure on publishers is intensifying. DataDome's report showed "bad" bot traffic grew 124% in a year, more than nine times faster than human traffic, with scraping alone up 185%. Meanwhile, 65.3% of the more than 21,000 popular websites DataDome tested didn't stop a single one of its test bots.

Payment infrastructure, however, is starting to form. Parallel has introduced a way to pay publishers for their contributions to agent tasks, with The Atlantic and Fortune among its first partners. Cloudflare, TollBit and ProRata have all built payment rails for "good" bots, and Cloudflare just shifted to paying publishers when their content shapes an AI answer, not just when it's fetched. Even Google is opening up payments to some publications when their content contributes significantly to answers in AI Overviews, AI Mode and Gemini. A model resembling the YouTube Partner Program is beginning to look attainable.

The open question is who sets the price. Right now it's the buyer. More than a dozen data brokers will sell content at the lowest possible price, exchanges are nascent with wildly inconsistent pricing, and even Google's program pays whatever Google decides. Publishers in the Google pilot described the math to Digiday as "quite black box," and one called the offers "lowball." Jonathan Roberts, People Inc.'s chief innovation officer, described the situation this summer as having 30 Napsters for content, but no Spotify.

The music analogy is instructive. Spotify didn't get artists paid just by creating a reliable exchange; rights holders also had enough collective weight to insist on it. Publishers have the first part forming and almost none of the second. SPUR, a coalition building standards for how AI systems license and track content—the Associated Press recently joined—might be the beginning of that leverage. The government is clearly sitting this out, so any bargaining power will have to come from publishers banding together or shoring up their own defenses.

AI answers are still only as good as the information they can access—something even Nadella acknowledges. But the market hasn't been told what that access costs. Until publishers find leverage, buyers will keep treating content as free for the taking. And right now, they're not wrong.

Original: commoncrawl.org

Share this article:

More from Grace Kim

Grace Kim

Show full bio

Market editor covering industry trends and analytics at Business Bearings.

406 articles

Related articles

« Previous articleNext article »