Recap: Taking Generative AI Beyond the Demos
Generative artificial intelligence has exploded this past year. We’ve seen incredibly promising demos—but when scaling to enterprise level, companies often encounter an increased level of uncertainty along with difficulty of implementation. How can organizations harness the power of this new technology while scaling it reliably and cost-effectively?
In partnership with Comcast NBCUniversal LIFT Labs, The Information’s Stephanie Palazzolo spoke with two generative AI pioneers for advice on successful scaling:
- Douwe Kiela is CEO and a co-founder of Contextual AI and an adjunct professor at Stanford University
- Luis Ceze is CEO and a co-founder of OctoML and a professor at the University of Washington
The Blockers: Why It’s Hard to Go From Demo to Production
Palazzolo kicked off the discussion by asking Kiela and Ceze about the problems they’ve run into while implementing large language models for enterprise clients.
The first issue, said Kiela, was staleness. “You train these language models up to a certain point,” he said, “but then we don’t know how to always keep them up to date.”
Ceze added the second issue: reliability. “The No. 1 issue is reliability as the application scales. If you’re doing demos, you can do a couple of requests. The infrastructure runs fine. But then as you ramp up beyond one hundred users, that’s when things start breaking.” And from what he’s seen, “These are all things that don’t pop up during the demo phase but become an issue as you move into production.”
Both guests agreed that scaling from demos to production is full of trade-offs. If you can scale, you don’t have reliability. If you have some reliability, you don’t have performance. And if you do have all that, it’s difficult to scale simply because of the overwhelming cost.
The Chip Shortage: Hype or Reality?
Palazzolo then asked Kiela and Ceze if there’s an actual chip shortage—or rather a perceived one as organizations go for top-of-the-line products when an older-generation product will do just as well.
Ceze believes the problem lies in overprovisioning. “You always think that you want the latest and greatest because you’re going to need it. But there’s a lot more availability on the lower-end chips. They’re often perfectly capable of running what you need—as long as you know how to make the most out of that hardware.”
Kiela believes peak chip shortage is actually behind us. Why? “Nvidia made a very conscious decision to make their products more hardware agnostic so that they run on AMD and other GPUs that are coming up. Hyperscalers are quickly building the next generation of accelerators,” he said. “There’s a massive push to alleviate this shortage.”
Where Should Companies Start?
Currently, there’s pressure on almost every company to incorporate generative AI into its offerings. But where should companies begin their journey? Buying off-the-shelf options? Hiring large AI teams? Or consulting with specialized startups like the ones Kiela and Ceze run?
“I usually tell enterprises that it’s too early to put all your eggs in one basket,” Kiela said. “It makes sense to try different sources and see what really sticks.” He added, “Just be very experimental with the things you’re doing and find things that actually work.”
Ceze said while experimentation is critical, the companies that will succeed are the ones that are ready for scale deployment. “It’s been superinteresting to see folks finding that they can have meaningful results in good-enough deployable quality.”
What’s on the Horizon?
Both Ceze and Kiela are highly optimistic about the future of generative AI. Ceze pointed to possibilities in life sciences.“I’m excited about using generative models to design complex molecules like proteins.”
For his part, Kiela is looking forward to AI systems that can understand the human world better, which of course requires human interaction. “If they understand our world much better, then they can also understand us much better—and be even better aligned with what we really want.”
Realizing the Promise of Generative AI on a Large Scale
As chips and GPUs become more available to implement, the cost barriers will fall away. Reliability will increase. Overprovisioning will settle down. And with enterprise-size scaling, we’ll finally be able to go from demo to reality—fulfilling the promise of what generative AI can help humans achieve.