What it is
Most medical AI evaluations use static multiple-choice vignettes that do not reflect how clinicians actually work, so this team built the Sequential Diagnosis Benchmark from 304 diagnostically challenging NEJM clinicopathological cases turned into stepwise encounters where the solver must request findings one at a time from a gatekeeper model. Their MAI Diagnostic Orchestrator (MAI-DxO) simulates a panel of physicians that proposes differentials and selects high-value tests. Paired with OpenAI's o3, it reached 80% diagnostic accuracy, roughly four times the 20% average of generalist physicians under the same constraints, and 85.5% when configured for maximum accuracy.
Why it matters
Static benchmarks reward memorized answers, while real diagnosis is iterative, adaptive, and costly. By scoring both accuracy and the expense of the visits and tests ordered, this benchmark captures the actual clinical tradeoff. The orchestrator did not just win on accuracy: it cut diagnostic costs by 20% versus physicians and 70% versus the off-the-shelf o3 model, and its gains held across models from multiple providers rather than depending on one system.
Underlined numbers link to their source. Every metric and quoted figure is listed under Sources and data below.
Filed underclinical AI, diagnostic reasoning, language models, benchmarks, healthcare