The Three Classes of AI Coding Assistants
The model race continues. On Monday evening, Elon Musk’s xAI released a new family of large language models, Grok 3, which includes a baseline model, a smaller and faster version of the baseline model, and two “reasoning” models.
Early reactions seem promising. Andrej Karpathy, one of the cofounders of OpenAI who left the company last year, said on X that the reasoning version of Grok 3 performs around the level of OpenAI’s strongest models, like o1-pro mode, and is slightly better than the latest reasoning models from DeepSeek and Google. (Following the announcement, X bumped the pricing of its most premium tier, which will first get access to the new Grok models, to $40 per month, but it's still a lot cheaper than the $200-per-month ChatGPT Pro subscription.)
One shortcoming Karpathy pointed out, though, is Grok’s tendency to make up information and sources in its responses. For instance, he said the model would sometimes conjure up URLs for web pages that didn’t exist and avoid citing X as a source in its answers. That’s surprising, since Musk seems to think “X is the only place for real, trustworthy news” compared to biased legacy media outlets like ourselves, if you’re to believe his post here!
It’s also interesting how Grok 2 and Grok 3 had very different responses when asked about publications like The Information. When we asked Grok 2 to explain those differences, it noted “changes in training data, societal trends or the specific programming goals for Grok 3 which might emphasize a more direct critique of legacy media.”
Now, onto today’s column…