Anthropic Beat OpenAI in Test of AI That Performs AI Research
Since the days of Alan Turing, artificial intelligence developers have been captivated by the prospect of AI powerful enough to improve itself. OpenAI, we recently reported, has already developed an internal AI “research assistant” tool to help its researchers work faster, a possible first step in the development of AI that can conduct AI research on its own.
Later this week, researchers at Model Evaluation and Threat Research, a nonprofit group, will publish a first-of-its-kind evaluation of how large language models from OpenAI and Anthropic perform when asked to solve seven AI research problems. An early look at the results show Anthropic did very well compared to OpenAI.
In five of the seven tests METR ran, the latest version of Anthropic’s most advanced model, Claude Sonnet 3.5, outperformed OpenAI’s most advanced model, o1-preview—and Claude won by a wide margin in two of them. O1-preview beat Claude on the other two tests, including a decisive win in one of the tests.