
While innumerable benchmarks and tests exist to determine the savvy and capabilities of AI, one perhaps more obscure benchmark appears to be making waves in the AI community. According to a new report, companies like Google, OpenAI, and Anthropic are now making their models play old-school Pokémon to evaluate performance, as reported by the Wall Street Journal.
"The thing that has made Pokémon fun and that has captured the [machine learning] community’s interest is that it’s a lot less constrained than Pong or some of the other games that people have historically done this on. It’s a pretty hard problem for a computer program to be able to do," Anthropic AI lead David Hershey told the outlet.