Evaluating Models is Getting Even Harder
As AI gets more advanced, researchers are increasingly telling me that it’s also getting harder to evaluate how good AI actually is at various tasks.
Most obviously, as AI improves, it quickly saturates all the benchmarks we have today and researchers have to rush to come up with even harder evaluations that the models won’t immediately ace.
There are more subtle reasons as well. Researchers at the International Conference on Machine Learning said last week that today’s models have gotten so good, they can work on a task for hours or days. That’s great, but it also means evaluating their performance can take hours or days as well.
We’re quickly approaching a point where models can “work for weeks or indefinitely,” OpenAI researcher Noam Brown said during an ICML panel last week. (You can imagine that a model would want to work for that long in cases like drug discovery, where a model might spend weeks running experiments and analyzing the results of those experiments.) That means evaluating a model could eventually take longer than training that model in the first place, Brown said. Not only is that difficult to do practically, but it could also slow down the process of model development.