Harvard Study in Science: OpenAI's o1 Outdiagnoses ER Physicians at Triage
A peer-reviewed Harvard Medical School study published in Science found OpenAI's o1 correctly diagnosed 67% of ER patients at triage, beating two attending physicians who scored 55% and 50%. The gap was largest with minimal patient data.
A peer-reviewed study published May 3 in Science by Harvard Medical School and Beth Israel Deaconess Medical Center put OpenAI’s o1 model into a real emergency room — and it outperformed attending physicians.
The study enrolled 76 actual ER patients. At the first diagnostic touchpoint, where clinicians have the least data, o1 reached an “exact or very close diagnosis” in 67% of triage cases. Two attending physicians scored 55% and 50% respectively. GPT-4o was also tested and performed worse than o1.
Why Triage Is the Hard Part
The performance gap at triage is the finding that matters. It’s the moment with the most uncertainty: minimal history, no lab results, the patient still describing their symptoms. That’s exactly where misdiagnosis is most dangerous and most common in the real world.
O1’s architecture, which uses extended chain-of-thought reasoning before producing an output, appears to handle differential diagnosis well even under sparse input conditions. The model can weigh multiple competing explanations against partial symptom sets without anchoring too early on a single hypothesis — a failure mode that human clinicians are well-documented to fall into.
What the Study Does Not Claim
The researchers are clear: this was a text-based study. O1 received written symptom descriptions and patient data; it did not examine anyone. There is no physical exam, no non-verbal cues, no follow-up questions that a human physician conducts through actual interaction.
The paper does not argue that AI is ready to replace ER doctors. It argues that AI tools, deployed as a second opinion or triage support layer, could meaningfully reduce diagnostic errors at the point of care where they happen most often.
That framing matters. The realistic near-term deployment is an AI-assisted triage protocol, not an autonomous diagnostic agent. A physician reviews the model’s output, not the other way around.
The Clinical AI Race
This study arrives as every major health system is evaluating AI tools for clinical decision support. Epic has integrated GPT-4o into its clinical workflow product. Google Health’s MedPaLM 2 has been piloted in real hospital settings. Microsoft’s Nuance DAX handles ambient documentation for over 700 health systems.
What makes the Harvard study stand out is rigor: published in Science, peer-reviewed, drawn from real patients rather than curated benchmark datasets. Most clinical AI studies are conducted on historical records or carefully cleaned datasets that don’t reflect the messy conditions of actual emergency medicine.
The 67% vs. 55% gap sounds modest, but across millions of ER visits, that delta translates to hundreds of thousands of earlier correct diagnoses annually — and the downstream cost savings, improved outcomes, and reduced malpractice exposure that follow.
What Comes Next
Harvard’s team is planning a larger follow-up study with expanded patient cohorts and multi-site participation. The key variables they want to isolate: does performance hold when o1 is given access to real-time lab values? Does it degrade when symptoms are entered by nurses rather than physicians?
The more fundamental question — who is liable when an AI-supported diagnosis is wrong — remains entirely unresolved. That legal ambiguity will shape deployment timelines far more than benchmark scores.
For now, the finding stands: in text-based triage, o1 is already better than a median attending physician. That is not a claim anyone would have made confidently 24 months ago.
Related reading
- AI Models OpenAI Makes GPT-5.6 Luna the Free ChatGPT Default — With Unlimited Text Chats
- AI Models OpenAI Clears US Government Review, Sets GPT-5.6 Sol, Terra, and Luna for Public Launch Thursday
- AI Models OpenAI Ships GPT-5.5 Instant as Default ChatGPT Model — 52.5% Fewer Hallucinations on High-Stakes Prompts