Harvard researchers find AI models outperformed doctors in emergency room diagnoses
A new study published in Science reveals that large language models achieved more accurate triage diagnoses than attending physicians when given raw electronic medical records, highlighting both the potential and the regulatory gaps in healthcare technology.

A study published in Science by researchers at Harvard Medical School and Beth Israel Deaconess Medical Centre has found that OpenAI's large language models provided more accurate diagnoses than human attending physicians in real emergency room cases. The research, which examined the performance of the o1 and 4o models, challenges current assumptions about human dominance in initial medical triage.
In an experiment involving 76 patients who presented to the Beth Israel emergency room, the o1 model achieved an exact or very close diagnosis in 67 per cent of triage cases. This figure surpassed the performance of two human attending physicians, who recorded accuracy rates of 55 per cent and 50 per cent respectively. The 4o model also performed competitively, matching or exceeding human performance at every diagnostic touchpoint tested.
The assessment relied strictly on text-based information available in electronic medical records, with no pre-processing applied to the data. To ensure objectivity, the diagnoses generated by both the AI and the doctors were evaluated by two other attending physicians who remained blind to the origin of each assessment. The study noted that the performance gap was most pronounced during the initial ER triage phase, where information is scarce and urgency is high.
Lead author Arjun Manrai, who heads an AI lab at Harvard Medical School, stated that the team tested the AI model against virtually every benchmark, noting that it eclipsed both prior models and their physician baselines. However, the researchers emphasised that the study does not claim AI is ready for real-life life-or-death decisions. Instead, they called for prospective trials to evaluate these technologies in real-world patient care settings before such systems could be delegated critical responsibilities.
Adam Rodman, a doctor at Beth Israel and a lead author of the study, highlighted the current regulatory vacuum surrounding these technologies. He noted that there is no formal framework right now for accountability regarding AI diagnoses. Furthermore, Rodman observed that patients still desire human guidance for critical treatment decisions and for navigating the complexities of life-or-death scenarios.
The findings also underscore the limitations of current foundation models when dealing with non-text inputs. While this specific study focused solely on text-based data from electronic records, existing studies suggest that models remain more limited in reasoning over images or audio. Consequently, the researchers caution that the results should not be extrapolated to multimodal data interpretation without further testing.
