AI Chatbots Misdiagnose 80% of Early Medical Cases

A study by Mass General Brigham published in JAMA Network Open found that AI chatbots, including models from OpenAI and DeepSeek, misdiagnose over 80% of early medical cases due to their inability to accurately navigate the diagnostic process when provided with incomplete patient data. The research tested 21 large language models and revealed that while they achieved correct final diagnoses over 90% of the time with complete information, their performance in generating differential diagnoses was significantly lacking. This highlights ongoing concerns from public health experts regarding the reliability of AI in medical decision-making.
