What it is
AMIE is a large language model tuned for diagnostic dialogue, trained in a self-play simulation that let it practice history-taking across many conditions and specialties. In a randomized, double-blind study it held text consultations with trained patient-actors across 159 case scenarios and was compared head to head with 20 primary care physicians. Specialist reviewers rated AMIE at least as good as the physicians on 30 of 32 clinical axes, and the patient-actors rated it favorably on 25 of 26.
Why it matters
Diagnostic dialogue, not just answering exam questions, is the core clinical skill, and this is the first blinded, controlled evidence a conversational model can match physicians at it. Winning on 30 of 32 specialist-judged axes, including communication and empathy, suggests the gains are not limited to raw accuracy. The self-play training loop points to a way to scale clinical competence without proportionally more human-labeled cases.
Underlined numbers link to their source. Every metric and quoted figure is listed under Sources and data below.
Filed undermedical AI, large language models, diagnosis, clinical trial, doctor-patient dialogue