Businesses Want Slower AI Models—And That Might Hurt Nvidia
OpenAI isn’t normally an acquirer of other companies but it has announced two purchases in the last week, of search and analytics firm Rockset and collaboration startup Multi. Both highlight how OpenAI is focused on enterprise offerings, and features that big business appreciate such as retrieval augmented generation, despite the attention given to its consumer products like its “flirtatious” ChatGPT.
That makes sense: there’s surely more money in business-focused AI services than in the consumer market, at least in the near term. By the same token, though, businesses have become much more focused on the cost-effectiveness and returns of AI than on finding the fastest, most advanced models.
One sign of that is the rise of “batch processing.” Today, popular consumer AI products like ChatGPT or Perplexity provide users with near-instantaneous responses, otherwise known as “real-time” inference.
However, businesses don’t always need immediate responses, and are often willing to wait hours, days or even weeks for responses—as long as they don’t have to shell out as much money. Founders of cloud and inference providers have told me that there’s growing pressure from business customers for this kind of flexibility—otherwise known as “batch processing.”