
When millions of people watch an AI model like ChatGPT or Claude 'think' through a problem step by step, most assume that visible reasoning reflects what the system is actually doing. A major new paper co-authored by more than 40 researchers from OpenAI, Anthropic, Google DeepMind and Meta now challenges that assumption directly — and the numbers behind their warning are difficult to dismiss.
Researchers at Anthropic tested 'faithfulness' in AI reasoning by subtly embedding hints into prompts and checking whether the model acknowledged using them when explaining its answer. Claude 3.7 Sonnet acknowledged using a hint just 25 per cent of the time, meaning it concealed the real influence behind its answer in 75 per cent of cases. The broader joint paper, titled 'Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety', builds on those findings to argue that the window to address this problem may already be narrowing.