OpenAI’s Impressive Engineering Feat with GPT-4o AKA ‘Her’
OpenAI’s Google search killer may not yet be ready for prime time, and neither is the long-awaited GPT-5 large language model. But the ChatGPT developer on Monday may have still managed to steal some thunder ahead of Google’s annual conference for software developers, I/O, which starts later this morning.
In a 26-minute demonstration, OpenAI COO Mira Murati showed off GPT-4o, a new model (the “o” stands for “omni”) that lets people talk to ChatGPT on their phone, similar to the way they speak to Siri and other voice assistants. Unlike with Apple’s Siri, though, the ChatGPT voice assistant can also analyze and discuss images or videos that it sees, and it can recognize different emotions when users speak to it. These capabilities could allow ChatGPT to take on a wide variety of new applications, from helping people with math homework to translating language in real-time or helping people prepare for interviews.
In other words, OpenAI’s tech can now do what Google’s tech pretended to do six months ago, when that company released large language models that aimed to compete with OpenAI’s. At I/O, you should expect Google to bust out its own AI upgrades to Google Assistant, the company’s rival to Siri.
Jim Fan, a senior research manager at Nvidia, pointed out on X that while GPT-4o isn’t necessarily a technological breakthrough, it represented an impressive engineering feat because it weaves together different technologies OpenAI had separately developed. Others seemed to agree: Soumith Chintala, who leads work on PyTorch at Meta Platforms, wrote in a post on X that OpenAI was “establishing new expectations for AI” but also that GPT-4o “feels attainable.” By contrast, when OpenAI launched GPT-4 more than a year ago, it felt “magically impossible” to replicate, Chintala wrote. (To be fair, Chintala does have an incentive to downplay OpenAI’s releases, given that Meta is a rival.)