Are We Running Out of Training Data?
I was catching up with the new AI Index Report from Stanford’s Human-Centered Artificial Intelligence Institute when one section caught my eye. That was a forecast that we’ll likely run out of high-quality language data needed for train AI models sometime this year.
That’s a worrying timeline. So far, AI companies have improved their large language models by training them on lots of data combined with increasingly more computing power (a strategy otherwise known as the “scaling laws”). If the supply of high quality data is about to disappear, LLMs may soon hit a quality wall.
Assessing how troublesome this could be is difficult. We know little about the data used to train OpenAI’s GPT-4 or Anthropic’s Claude 3, suggesting these companies are holding their cards close to their chest and consider it their secret sauce. Still, there’s evidence that things aren’t as bad as the forecast might suggest.