Why OpenAI Should Worry About Google’s Pretraining Prowess
Google’s breakthrough with Gemini 3, its latest AI model, has lit a fire under OpenAI and its CEO Sam Altman, who told thousands of his researchers to buckle down and prepare for “rough vibes” and “temporary economic headwinds,” Erin and I reported.
What might particularly worry Altman about Google’s latest model is how the tech company got Gemini 3 to be so good in the first place—specifically, through improvements in pretraining. Not only is pretraining an area in which OpenAI has recently struggled with, but many researchers also believe that improving pretraining is key to getting a model to better generalize, or do things that involve subjects or information outside the data on which they were trained.
As a reminder, pretraining refers to the first phase of model training in which researchers expose a model to data from the web and other sources so it can learn connections between them. This is in contrast to the next stage of model training, post-training, in which the pre-trained model is shown more curated data to learn about specific fields like medicine or law or how to better respond to chatbot users.
Getting a model to better generalize is important because someday, we’d hope that AI models could help us make new discoveries that they’ve never seen before in their training data, like conducting AI research or even coming up with a cure for diseases.