
When OpenAI dropped ChatGPT-5.5 last week, I couldn’t wait to see how it compared to other models. I started with Claude 4.7 Opus and the results were unexpected especially since the launch was framed as a state-of-the-art leap forward. OpenAI published benchmark scores that placed it ahead of both Claude Opus 4.7 and Google's Gemini 3.1 Pro.
Gemini 3.1 Pro, released two months earlier in February, came with its own bold claims: more than double the ARC-AGI-2 score of its predecessor, "exceptional instruction following.”