Patients with rare diseases are increasingly using artificial intelligence tools to generate possible explanations for symptoms that have remained unexplained for years, and in a growing number of cases, those tools have pointed to the right answer.
What has not been established is how often that happens. The accounts circulating are individual successes reported by patients, families, and clinicians. No one has measured a success rate for patients using chatbots on their own, and the published research that does exist gives figures far more modest than the anecdotes suggest.
The problem those patients are trying to solve is real and well documented. Roughly one in ten Americans has one of more than 10,000 known rare diseases, and the time from symptom onset to accurate diagnosis commonly runs five years or longer.
Ten Thousand Conditions and Almost No Clinician Sees Most of Them
The reason rare disease diagnosis takes so long has less to do with physician competence than with arithmetic.
A rare disease in the United States is one affecting fewer than 200,000 people, and most affect far fewer than that. /With more than 10,000 such conditions, no clinician can carry working familiarity with more than a fraction. A physician may encounter a given rare disease once in a career, or never. Pattern recognition, which is how most diagnosis actually works, fails when the pattern has never been seen.
Early symptoms compound the problem because they are usually nonspecific. Fatigue, developmental delay, recurrent infections, unexplained pain and gastrointestinal problems are consistent with hundreds of ordinary conditions before they are consistent with a rare one. The reasonable clinical move at each step is to rule out the common explanation, and each step takes months.
The referral structure then fragments the picture. A patient with multisystem symptoms sees a gastroenterologist, then a neurologist, then a rheumatologist, each evaluating within their own domain. The information needed to recognize the condition often exists across those records but has never been assembled in one place by one person.
The Undiagnosed Diseases Network, a National Institutes of Health program established in 2014, exists specifically for the cases that fail this way. Rizwan Hamid, director of the Potocsnak Center for Undiagnosed and Rare Disorders at Vanderbilt University Medical Center, said in a university statement that for some patients referred to the network, the diagnostic odyssey "lasts for more than 10 years."
Thirteen Percent Is Better Than the Comparison and Still Mostly Wrong
The most useful published benchmark comes from that same center, and it should be read carefully in both directions.
Researchers prompted two large language models to generate differential diagnoses from clinical summaries for 90 previously solved rare disease cases drawn from the network. Published in JAMA Network Open, the analysis found the models correctly identified the diagnosis in 13.3 percent and 10.0 percent of cases, compared with a historical clinical review rate of 5.6 percent for physicians working from the same medical records. The models produced diagnoses judged helpful but incorrect in a further 23.3 percent and 16.7 percent of cases.
Both readings of that result are true. A two- to threefold advantage over physician chart review is a genuine signal and likely reflects a breadth of pattern exposure that no individual clinician can match. And 13 percent means the tool was wrong 87 percent of the time in cases that eventually had answers.
Purpose-built systems are being developed to outperform general chatbots. DeepRare, described in Nature this year, coordinates specialized tools and knowledge sources, reads free-text clinic notes alongside structured symptom codes and genetic testing results, and produces ranked diagnoses with narrative reasoning and a self-checking step meant to reduce overconfident errors. An accompanying Nature commentary notes that with thousands of rare diseases recognized, even experienced clinicians struggle to connect unfamiliar symptom patterns to underlying causes.
Other approaches target different parts of the problem. A weakly supervised algorithm developed by NIH-funded researchers at Harvard Medical School and Boston Children's Hospital uses incomplete electronic health record data to flag patients who may have an undiagnosed rare condition. A separate algorithm published in Genetics in Medicine ranks candidate disease-causing genes by comparing evolutionary patterns across species, ranking the correct gene first in most test cases.
A Plausible List Is Not a Diagnosis
The gap between generating possibilities and reaching a conclusion is where the practical risk sits, and it is worth stating plainly for anyone considering this route.
A chatbot can produce a fluent, confident, internally coherent explanation that is wrong. It can also produce a list of conditions that are technically consistent with the symptoms but vanishingly unlikely, which for an already frightened patient or parent is not neutral information. Becoming convinced of a specific rare diagnosis before testing can distort the clinical conversation rather than improve it.
Performance also depends heavily on how much information the model is given. A separate study of chatbot performance led by Mass General Brigham researchers, also published in JAMA Network Open, found that leading models failed to produce an appropriate list of possible causes more than 80 percent of the time when working from the partial information a patient would typically supply early on, even though accuracy exceeded 90 percent once all clinical details were available. Early in a workup is exactly when patients are most likely to reach for these tools.
Researchers developing these systems say the same thing clinicians do: the tools require human confirmation, prospective validation in diverse real-world clinics, and oversight to ensure they narrow rather than widen disparities.
Used carefully, the tools have uses that do not require them to be right about the diagnosis. Translating dense medical terminology into plain language, organizing a chronological symptom history, summarizing years of scattered records into a single document a specialist can read in five minutes, and generating a list of questions to ask are all things chatbots do reliably. That assembled history is frequently the missing ingredient, and producing it is valuable regardless of whether the AI names the condition.
Families should also know that formal routes exist. The Undiagnosed Diseases Network accepts applications, and some academic centers accept self-referral for undiagnosed cases. Whole genome sequencing has fallen substantially in cost, and specialists increasingly recommend it earlier rather than as a last resort. Anyone bringing AI-generated possibilities to an appointment is better served presenting them as questions than as conclusions.
Prospective clinical validation studies are underway. MedicalDaily will report results as they are published.
Key Questions Answered
How long does rare disease diagnosis usually take? Commonly five years or longer. For the most complex cases referred to the Undiagnosed Diseases Network, clinicians there say the odyssey can last more than 10 years.
How accurate are AI chatbots at this? In a study of 90 previously solved rare disease cases, two models correctly identified the diagnosis in 13.3 percent and 10.0 percent of cases, compared with a historical clinical review rate of 5.6 percent.
Does that mean AI is better than a doctor? No. It means AI outperformed a historical benchmark for physicians reviewing written records without seeing the patient. It does not compare AI to a full clinical evaluation, and both figures are low.
Why is a rare disease so hard to diagnose? More than 10,000 rare conditions exist, so most clinicians never encounter any one of them. Early symptoms are nonspecific, and specialty referral fragments the picture across separate records.
Does it matter how much information the chatbot gets? Considerably. Research found leading models failed to produce an appropriate differential more than 80 percent of the time with partial information, while accuracy exceeded 90 percent once all clinical details were supplied.
What are chatbots reliably useful for here? Translating medical terminology, assembling a chronological symptom history, summarizing scattered records, and preparing questions for an appointment.
Are there formal programs for undiagnosed patients? Yes. The Undiagnosed Diseases Network accepts applications, and some academic centers accept self-referral for undiagnosed cases.