What it is
DeepSeek-R1 shows that a large language model can be taught to reason through pure reinforcement learning on verifiable tasks, with no human-written reasoning traces to imitate. Under this training, behaviors like self-reflection, verification, and dynamic strategy adaptation emerge on their own. The resulting model outperforms counterparts trained by conventional supervised learning on human demonstrations across mathematics, coding competitions, and STEM problems.
Why it matters
This is the first peer-reviewed demonstration (a Nature cover) that advanced reasoning can be incentivized by RL alone, removing the dependence on costly human-annotated reasoning data. The result landed with unusual force: roughly 770 citations within a year and a field-weighted citation impact near 916, hundreds of times the norm for its field. The emergent reasoning patterns also transfer, guiding and improving smaller models.
Underlined numbers link to their source. Every metric and quoted figure is listed under Sources and data below.
Filed underreinforcement learning, reasoning, open models, LLM, DeepSeek
Watch
A short explainer of this result.