Artificial intelligence (AI) systems may outperform doctors at diagnosing patients. In a recent series of experiments led by researchers at the Beth Israel Deaconess Medical Center and Harvard Medical School, a preview of the OpenAI large language model (LLM) known as o1 bested human physicians in multiple tests of clinical and diagnostic reasoning.
"We tested the AI model against virtually every benchmark, and it eclipsed both prior models and our physician baselines," said Harvard Medical School professor Arjun K. Manrai, one of the study's senior authors. Using case studies, the researchers had o1—an older model since usurped by o3—generate a list of possible diagnoses. It included the right diagnosis 78 percent of the time, while physicians were successful about 30 percent of the time.