Exclusive: Mercor’s Fast Growth Relies on Biggest AI Companies, Documents Show Save 25% to unlock this story

Sign in
Subscribe

    Data Tools

    • About Pro
    • Enterprise Software Startup Takeover List 2026
    • The Next GPs 2026
    • The Executives Leading the Data Center Race
    • The Next GPs 2025
    • The Rising Stars of AI Research
    • Leaders of the AI Shopping Revolution
    • Enterprise Software Startup Takeover List 2025
    • Org Charts
    • The Information 50 2025
    • Generative AI Takeover List
    • Generative AI Database
    • AI Chip Database
    • AI Data Center Database
    • Tech IPO Tracker
    • Tech Sentiment Tracker
    • Gigafactory Database

    Special Projects

    • The Information 50 Database
    • VC Diversity Index
    • Enterprise Tech Powerlist
  • Org Charts
  • Deep Research
  • Tech
  • Finance
  • Weekend
  • Charts
  • Events
  • TITV
    • Directory

      Search, find and engage with others who are serious about tech and business.

    • Forum

      Follow and be a part of discussions about tech, finance and media.

    • Brand Partnerships

      Premium advertising opportunities for brands

    • Group Subscriptions

      Team access to our exclusive tech news

    • Newsletters

      Journalists who break and shape the news, in your inbox

    • Video

      Catch up on conversations with global leaders in tech, media and finance

    • Partner Content

      Explore our recent partner collaborations

      XFacebookLinkedInThreadsInstagram
    • Help & Support
    • RSS Feed
    • Careers
    Sign in
  • About Pro
  • Enterprise Software Startup Takeover List 2026
  • The Next GPs 2026
  • The Executives Leading the Data Center Race
  • The Next GPs 2025
  • The Rising Stars of AI Research
  • Leaders of the AI Shopping Revolution
  • Enterprise Software Startup Takeover List 2025
  • Org Charts
  • The Information 50 2025
  • Generative AI Takeover List
  • Generative AI Database
  • AI Chip Database
  • AI Data Center Database
  • Tech IPO Tracker
  • Tech Sentiment Tracker
  • Gigafactory Database

SPECIAL PROJECTS

  • The Information 50 Database
  • VC Diversity Index
  • Enterprise Tech Powerlist
Deep Research
TITV
Tech
Finance
Weekend
Charts
Events
Newsletters
  • Directory

    Search, find and engage with others who are serious about tech and business.

  • Forum

    Follow and be a part of discussions about tech, finance and media.

  • Brand Partnerships

    Premium advertising opportunities for brands

  • Group Subscriptions

    Team access to our exclusive tech news

  • Newsletters

    Journalists who break and shape the news, in your inbox

  • Video

    Catch up on conversations with global leaders in tech, media and finance

  • Partner Content

    Explore our recent partner collaborations

Subscribe
  • Sign in
  • Search
  • Opinion
  • Venture Capital
  • Artificial Intelligence
  • Startups
  • Market Research
    XFacebookLinkedInThreadsInstagram
  • Help & Support
  • RSS Feed
  • Careers

In-depth insights in seconds. Ask Deep Research.

AI Agenda

Author Paranoia About AI Finds New Targets

Art generated by Midjourney.
By
Stephanie Palazzolo
[email protected]Profile and archive

AI founders need to watch out. It’s no secret that publishers and writers are sensitive about large-language models using their data for training. But just how sensitive became clear in the recent shuttering of a six-year-old publishing dataset called Prosecraft.

The service, created in 2017 by Benji Smith, founder and CEO at word processor Shaxpir, was intended to be a helpful resource to authors. It ranked titles based on how passive or vivid their language is. For instance, it granted Lewis Carroll’s “Alice’s Adventures in Wonderland” a 83.94% “vividness” score. One issue: it scraped those books from the web without the authors’ permission. Even so, people didn’t seem to pay it much mind until the rise of generative AI, which has raised concerns around LLMs training on copyrighted material. 

While Prosecraft itself has little to do with LLMs, the sheer amount of easily-accessible textual data it offered rang alarm bells with writers like “Little Fires Everywhere” author Celeste Ng. They worried that Prosecraft could be used by LLMs for training purposes. A few worried tweets this weekend from writers who discovered the site later kicked off a firestorm on X (formerly known as Twitter). Smith quickly buckled, taking down the dataset. Prosecraft had ingested 27,000 books up until that point. Smith did not respond to a request for comment.

Recommended