Legal Problems Are Coming for AI Researchers in Academia; OpenAI Cofounder Shares His AGI Timeline
Artificial intelligence developers have a long history of playing fast and loose with intellectual property rights when it comes to getting data to train their models. While data owners are increasingly suing OpenAI, Microsoft, Meta Platforms and other major firms over such alleged violations, the heat may soon be turned on academics who operate the same way.
In one example I found, researchers from Carnegie Mellon University and the University of Waterloo thought they had uncovered a shortcut to one of the thorniest issues facing AI developers: finding large amounts of special training data that's especially good at improving AI's reasoning skills. The researchers scraped math, science and engineering problems from textbook and tutoring sites such as Chegg, Course Hero, Google’s Socratic learning app and Khan Academy. They then used that data to identify and pull similar examples of problems from the Common Crawl, a massive, open-source repository of internet data, to use to improve a large language model’s reasoning skills. (The researchers have released part of the training dataset on Hugging Face.)
The only problem: Those education sites explicitly ban users from scraping their data.