Selling Data for AI May Be Publishers’ Salvation
Developers of large language models increasingly are paying internet publishers for access to their content. The payments could supplant ad revenue for some publishers.
Tunguz is the founder of Theory Ventures.
In its first earnings report as a public company, Reddit last month said its average revenue per user in the first quarter rose 8% from a year earlier. The biggest reason: new deals in which Reddit is licensing the billions of words and images on its site to developers of large language models. For the year, Reddit said it expected more than $66 million of revenue from such deals, which would amount to about 6% of its projected annual revenue. But it likely will be more: Later in the month, Reddit said it had reached a deal to license its data to OpenAI.
Reddit isn’t alone. Shutterstock and Freepik also have struck agreements with tech firms to license their libraries of content to train LLMs, as have Tumblr and WordPress. News organizations including The Associated Press, Axel Springer and Reuters are on board, and News Corp (owner of the New York Post and the Wall Street Journal), Vox Media and the Atlantic recently finalized deals with OpenAI. Bloomberg recently reported that Adobe is purchasing videos from creators for $3 per minute to train its AI models.