
Claude 4.6 Opus launched just days ago, and I immediately pitted it against ChatGPT-5.2 Thinking to see how it compared to OpenAI’s smartest model. Naturally, with Gemini’s recent dominance, I had to see how it compared to Gemini 3 Flash.
I put the two top models head-to-head across nine challenging tests spanning math, logic, coding, creative writing and more — tasks designed to push each model's reasoning, creativity and practical usefulness to the limit.
My prompts aren’t the kind of questions you can answer by regurgitating training data; they require genuine multi-step thinking, context judgment and the ability to follow complex constraints. Here's how Anthropic's most powerful model stacked up against Google's latest.