Skip to content

EarlyPreprint

An AI Diagnostic Orchestrator Solves Hard Cases at Four Times the Physician Rate

A panel-of-doctors language model reaches 85.5% accuracy on the toughest NEJM cases while cutting testing costs.

ByHarsha Nori, Mayank Daswani, Kelly, Christopher, Scott Lundberg, Marco Túlio Ribeiro, Marc Wilson and 2 others

What it is

Most medical AI evaluations use static multiple-choice vignettes that do not reflect how clinicians actually work, so this team built the Sequential Diagnosis Benchmark from 304 diagnostically challenging NEJM clinicopathological cases turned into stepwise encounters where the solver must request findings one at a time from a gatekeeper model. Their MAI Diagnostic Orchestrator (MAI-DxO) simulates a panel of physicians that proposes differentials and selects high-value tests. Paired with OpenAI's o3, it reached 80% diagnostic accuracy, roughly four times the 20% average of generalist physicians under the same constraints, and 85.5% when configured for maximum accuracy.

Why it matters

Static benchmarks reward memorized answers, while real diagnosis is iterative, adaptive, and costly. By scoring both accuracy and the expense of the visits and tests ordered, this benchmark captures the actual clinical tradeoff. The orchestrator did not just win on accuracy: it cut diagnostic costs by 20% versus physicians and 70% versus the off-the-shelf o3 model, and its gains held across models from multiple providers rather than depending on one system.

Underlined numbers link to their source. Every metric and quoted figure is listed under Sources and data below.

Filed underclinical AI, diagnostic reasoning, language models, benchmarks, healthcare

Scored 2026-09-28How scoring works

The Frontier Score, in full

How this entry's score is built, term by term. Open any pillar to see each sub-metric's raw value, how it maps to 0 to 100, and the record that holds it.
Evidence 69Impact 36Novelty 40
The three pillars as one shape: a result strong on every axis fills the triangle.
How the Frontier Score is calculated for this breakthrough
TermScoreWeightPoints
EvidenceHow established and verifiable the result is.69× 0.2517.4
ImpactHow much the result matters, judged field-relative and by current momentum.36× 0.258.9
NoveltyHow genuinely new the result is, and how fast-moving.40× 0.5020.0
Weighted pillars46.3
Frontier recency15 months since publication× 0.64
IntegrityNot retracted× 1
Frontier Score29
Confidence: Firming up. Gaining evidence: about 17 citations so far, its trajectory is still forming.Gaining evidence: about 17 citations so far, its trajectory is still forming.

Every sub-metric, traced to its source

Evidence69
Evidence sub-metrics
Sub-metricRaw value0 to 100WeightPointsSource
Peer review & venueWhether the work has cleared peer review, and how selective its venue is. Peer review is necessary but not sufficient: before citations accrue, a fresh result in the most selective venues (Nature, Science, Cell, NEJM, the Lancet, PNAS, Physical Review Letters/X) is a stronger evidence signal than one in a legitimate but very high-volume mega-journal, which in turn outranks a preprint that has not been reviewed at all.Maps to 0 to 100: Peer-reviewed article graded by venue prestige: flagship 100, elite-family 90, high-volume mega-journal 78. Review 90, book 70, dataset 60, preprint 45, other 55.Preprint / unreviewed450.3013.5OpenAlex type / primary_location / source
CorroborationIndependent corroboration that the result is real and being taken up, read as the BREADTH of agreement rather than the size of the citation pile. Counts how many of the five independent citation indices (OpenAlex, Crossref, Semantic Scholar, OpenCitations, Europe PMC) report the work at all, plus orthogonal, non-citation lines of corroboration: entry into the encyclopedia, registered clinical trials, released datasets or software, public funding on record, and technical-community discussion. Citation MAGNITUDE is deliberately judged under Impact, not here, so Evidence and Impact measure genuinely different things instead of both rewarding the same citation pile twice.Maps to 0 to 100: Starts at 28 and rises with each of the five independent citation indices that agree the work is cited (+9 each) and each orthogonal non-citation corroboration line (+6 each: Wikipedia, clinical trials, open datasets or software, public funding, community discussion). Capped at 100.3 of 5 citation indices agree, 1 independent non-citation line610.3018.3OpenAlex, Crossref, Semantic Scholar, OpenCitations, Europe PMC, DataCite, NIH RePORTER, Wikipedia, Hacker News count of agreeing indices + orthogonal reach
VerifiabilityHow openly the result can be read, reused, and reproduced. Rewards open access, a permissive reuse license (CC-BY/CC0), a green repository copy anyone can archive, openly minable full text, and released datasets or software a reader can actually run. Cross-checked across OpenAlex, Unpaywall, Europe PMC, and DataCite.Maps to 0 to 100: Closed 45; open access 80+; a permissive CC license 100; a repository copy, open full text, or a released dataset/software artifact each raise the floor.Open access (green)880.2017.6OpenAlex, Unpaywall, Europe PMC, DataCite is_oa / license / has_repository_copy / open datasets
IntegrityRetraction and post-publication correction status. A retracted result is not evidence of anything and collapses the whole Frontier Score; a correction or expression of concern is a softer flag. Read across OpenAlex and Crossref (both update-to and updated-by notices).Maps to 0 to 100: Clean 100; a correction or expression of concern 75; retracted 0. Flagged if OpenAlex OR Crossref records the notice.No retraction or concern on record1000.2020.0OpenAlex, Crossref is_retracted / update-to / updated-by
Impact36
Impact sub-metrics
Sub-metricRaw value0 to 100WeightPointsSource
Field-normalized impactImpact relative to the world average for the same field, year, and type, so a small field and a large field are judged fairly. Triangulates three independent field-normalized measures: OpenAlex's Field-Weighted Citation Impact, its citation percentile within the exact field-and-year cohort, and the NIH iCite Relative Citation Ratio. 1.0x is average; the percentile is the share of same-field, same-year work it out-cites.Maps to 0 to 100: Each ratio (1.0 = field average) maps 1.0 to 50, 3.0 to 75, 9.0 to 90; the field+year percentile maps directly (top 1% -> ~99). The available measures are averaged. Falls back to a saturating citation count when none is available yet.Not yet field-weightedn/aredistributed0.0OpenAlex, NIH iCite fwci / citation_normalized_percentile / relative_citation_ratio
Influential citationsCitations that genuinely build on the work rather than mention it in passing (Semantic Scholar's influential-citation measure). A sharper signal of real impact than a raw count.Maps to 0 to 100: Saturating: 15 influential citations maps to 50. Falls back to a fraction of consensus citations when unavailable.8 influential350.207.0Semantic Scholar influentialCitationCount
Citation velocityCitations received in the trailing twelve months, judged as a rate rather than a lifetime total. Momentum is what a frontier index cares about: a result being taken up fast right now is landing, whether its lifetime pile is large or still small. Where an expected field citation rate is known, the momentum is also read relative to it, so a fast-moving result in a quiet field is not overlooked.Maps to 0 to 100: Saturating: 30 citations in the last twelve months maps to 50. When NIH iCite provides the field's expected rate, this is blended evenly with the ratio of observed to expected momentum.17 in last 12 months360.3512.7OpenAlex, NIH iCite counts_by_year / field_citation_rate
Novelty40
Novelty sub-metrics
Sub-metricRaw value0 to 100WeightPointsSource
Conceptual noveltyHow genuinely new the contribution is, measured at publication from the work's OWN content, not its date: how atypical its combination of research areas is versus all prior work, and whether it builds on a broad, deep base rather than only the newest papers in a fast-moving area. This is what separates a real advance from a recent-but-commodity result, and it needs no forward citations, so it is stable for brand-new work. An honestly noisy proxy (even state-of-the-art novelty measures agree only moderately with expert judgment), so it carries bounded weight and never drives the ranking alone.Maps to 0 to 100: Blends two publication-time signals, each 0-100: an atypical-combination score (how unusual the pairing of the work's research topics is versus the prior literature, by pointwise mutual information) and a foundational-reach score (median age and depth of the work's references). The stronger signal weighs 0.62, the weaker 0.38. When neither is measurable (a work too thin in topics or references), the pillar renormalizes over its recency signals instead.Machine Learning in Healthcare x Clinical Reasoning and Diagnostic Skills (1,187 prior works made this pairing)180.458.1OpenAlex topics + referenced_works (publication-time)
RecencyHow recently the work appeared. The frontier is now, so newer scores higher.Maps to 0 to 100: Exponential decay with a 12-month half-life from the publication date.15 months old420.3012.6OpenAlex publication_date
Attention accelerationWhether attention is rising: citations in the latest full year versus the year before. Accelerating interest signals an active, opening frontier.Maps to 0 to 100: Latest/prior-year ratio: 1x maps to 50, 2x to 75, 4x to 100. Unknown is neutral.Not enough history500.105.0OpenAlex counts_by_year
Edge of reviewPreprints and brand-new work sit ahead of the peer-review process. That earns novelty here (while it is discounted under Evidence), because it is where the frontier forms first.Maps to 0 to 100: Preprint 95, article <6 months 80, article <18 months 65, else 45.Preprint (ahead of review)950.1514.3OpenAlex type / publication_date
ProvenanceAll sources

Sources and data

Every number behind this entry, grouped by the source that holds it: up to five citation indices cross-checked, field-normalized impact, then real-world reach (Wikipedia, community discussion, clinical trials, released data and public funding) where it exists. Each value links to its record.

OpenAlex

Primary record
Total citations
17
Citations, recent 12 months
17
Publication date
2025-06-27

Semantic Scholar

Citations, all versions
105
Influential citations
8

OpenCitations

Citations
12

arXiv

arXiv preprint
v2

Hacker News

Hacker News points
8
Hacker News comments
1

Figures quoted in the write-up

FigureValueSource
Peak diagnostic accuracy (MAI-DxO configured for maximum accuracy on the benchmark)85.5%arXiv (Cornell University)
Accuracy with o3 (MAI-DxO paired with OpenAI's o3 model)80%arXiv (Cornell University)
Physician average (Generalist physicians under the same stepwise constraints, roughly four times lower)20%arXiv (Cornell University)
Benchmark cases (Hard NEJM clinicopathological conference cases recast as stepwise encounters)304arXiv (Cornell University)
Cost reduction vs physicians (Lower diagnostic testing cost than physicians)20%arXiv (Cornell University)
Cost reduction vs raw o3 (Lower cost than off-the-shelf o3 alone)70%arXiv (Cornell University)
AI & ComputingThe whole field

Back to the Frontier Index, or browse every breakthrough by date.

The Frontier Brief

The week's frontier, scored.

Every Monday, the highest-scoring new breakthroughs in the index, each with its score and a one-line read on why it matters. What you get

One email a week. The top new breakthroughs, scored. Unsubscribe anytime.