AI’s ‘Split-Brain’ Problem
Developing new artificial intelligence models can sometimes seem like a game of whack-a-mole: fixing a model’s bad answers to certain questions can cause the model to give bad answers to other questions.
One version of this problem, according to researchers at OpenAI and elsewhere, is known as a split-brain problem, in which changes to how a question is phrased can lead to enormous variations in the AI-generated answer.
It’s part of a theme we’ve been writing about a lot recently: today’s models don’t develop an understanding of how the world works the way humans do. Some experts argue that this means they don’t generalize, or handle tasks outside the specific material they’ve been trained to recognize.
That could be a big problem, given that investors are giving tens of billions of dollars to labs like OpenAI and Anthropic so they can train models to make new discoveries in fields like medicine and math. (I expect to see this issue discussed a lot at this week’s Neural Information Processing Systems conference, whose main conference starts today.)