Skip to content

NotablePeer-reviewed

Centaur: a language model that predicts how people behave

One fine-tuned model forecasts human choices across 160 psychology experiments.

ByMarcel Binz, Elif Akata, Matthias Bethge, Franziska Brändle, Fred Callaway, Julian Coda-Forno and 2 others

Helmholtz Munich · University of Tübingen · University of Oxford · Max Planck Institute for Biological Cybernetics and 4 more

What it is

Researchers fine-tuned a large language model on Psych-101, a dataset of trial-by-trial decisions from more than 60,000 participants making over 10 million choices across 160 experiments. The resulting model, Centaur, predicts and simulates human behavior in any task that can be described in natural language, beating existing bespoke cognitive models on held-out participants. It also generalizes to unseen cover stories, altered task structures, and entirely new domains, and its internal representations became more aligned with human neural activity after fine-tuning.

Why it matters

Cognitive science has relied on separate hand-built models for each task, and none transferred well. Centaur is a single system trained on 10 million choices that out-predicts those specialized models and holds up out of distribution, suggesting one computational model can capture behavior across many domains at once. That is a plausible foundation for a more unified, testable theory of how people decide and learn.

Underlined numbers link to their source. Every metric and quoted figure is listed under Sources and data below.

Filed undercognition, foundation model, LLM, behavior, neuroscience

Scored 2026-09-30How scoring works

The Frontier Score, in full

How this entry's score is built, term by term. Open any pillar to see each sub-metric's raw value, how it maps to 0 to 100, and the record that holds it.
Evidence 94Impact 84Novelty 47
The three pillars as one shape: a result strong on every axis fills the triangle.
How the Frontier Score is calculated for this breakthrough
TermScoreWeightPoints
EvidenceHow established and verifiable the result is.94× 0.2523.4
ImpactHow much the result matters, judged field-relative and by current momentum.84× 0.2520.9
NoveltyHow genuinely new the result is, and how fast-moving.47× 0.5023.6
Weighted pillars67.9
Frontier recency15 months since publication× 0.64
IntegrityNot retracted× 1
Frontier Score43
Confidence: Settled. Established evidence: about 108 citations across 5 independent indices.Established evidence: about 108 citations across 5 independent indices.

Every sub-metric, traced to its source

Evidence94
Evidence sub-metrics
Sub-metricRaw value0 to 100WeightPointsSource
Peer review & venueWhether the work has cleared peer review, and how selective its venue is. Peer review is necessary but not sufficient: before citations accrue, a fresh result in the most selective venues (Nature, Science, Cell, NEJM, the Lancet, PNAS, Physical Review Letters/X) is a stronger evidence signal than one in a legitimate but very high-volume mega-journal, which in turn outranks a preprint that has not been reviewed at all.Maps to 0 to 100: Peer-reviewed article graded by venue prestige: flagship 100, elite-family 90, high-volume mega-journal 78. Review 90, book 70, dataset 60, preprint 45, other 55.Peer-reviewed1000.3030.0OpenAlex type / primary_location / source
CorroborationIndependent corroboration that the result is real and being taken up, read as the BREADTH of agreement rather than the size of the citation pile. Counts how many of the five independent citation indices (OpenAlex, Crossref, Semantic Scholar, OpenCitations, Europe PMC) report the work at all, plus orthogonal, non-citation lines of corroboration: entry into the encyclopedia, registered clinical trials, released datasets or software, public funding on record, and technical-community discussion. Citation MAGNITUDE is deliberately judged under Impact, not here, so Evidence and Impact measure genuinely different things instead of both rewarding the same citation pile twice.Maps to 0 to 100: Starts at 28 and rises with each of the five independent citation indices that agree the work is cited (+9 each) and each orthogonal non-citation corroboration line (+6 each: Wikipedia, clinical trials, open datasets or software, public funding, community discussion). Capped at 100.5 of 5 citation indices agree, 1 independent non-citation line790.3023.7OpenAlex, Crossref, Semantic Scholar, OpenCitations, Europe PMC, DataCite, NIH RePORTER, Wikipedia, Hacker News count of agreeing indices + orthogonal reach
VerifiabilityHow openly the result can be read, reused, and reproduced. Rewards open access, a permissive reuse license (CC-BY/CC0), a green repository copy anyone can archive, openly minable full text, and released datasets or software a reader can actually run. Cross-checked across OpenAlex, Unpaywall, Europe PMC, and DataCite.Maps to 0 to 100: Closed 45; open access 80+; a permissive CC license 100; a repository copy, open full text, or a released dataset/software artifact each raise the floor.Open access (hybrid), CC-BY, repository copy1000.2020.0OpenAlex, Unpaywall, Europe PMC, DataCite is_oa / license / has_repository_copy / open datasets
IntegrityRetraction and post-publication correction status. A retracted result is not evidence of anything and collapses the whole Frontier Score; a correction or expression of concern is a softer flag. Read across OpenAlex and Crossref (both update-to and updated-by notices).Maps to 0 to 100: Clean 100; a correction or expression of concern 75; retracted 0. Flagged if OpenAlex OR Crossref records the notice.No retraction or concern on record1000.2020.0OpenAlex, Crossref is_retracted / update-to / updated-by
Impact84
Impact sub-metrics
Sub-metricRaw value0 to 100WeightPointsSource
Field-normalized impactImpact relative to the world average for the same field, year, and type, so a small field and a large field are judged fairly. Triangulates three independent field-normalized measures: OpenAlex's Field-Weighted Citation Impact, its citation percentile within the exact field-and-year cohort, and the NIH iCite Relative Citation Ratio. 1.0x is average; the percentile is the share of same-field, same-year work it out-cites.Maps to 0 to 100: Each ratio (1.0 = field average) maps 1.0 to 50, 3.0 to 75, 9.0 to 90; the field+year percentile maps directly (top 1% -> ~99). The available measures are averaged. Falls back to a saturating citation count when none is available yet.64.7x FWCI, 12.8x RCR, top 0.0% of field-year vs field970.4543.7OpenAlex, NIH iCite fwci / citation_normalized_percentile / relative_citation_ratio
Influential citationsCitations that genuinely build on the work rather than mention it in passing (Semantic Scholar's influential-citation measure). A sharper signal of real impact than a raw count.Maps to 0 to 100: Saturating: 15 influential citations maps to 50. Falls back to a fraction of consensus citations when unavailable.20 influential570.2011.4Semantic Scholar influentialCitationCount
Citation velocityCitations received in the trailing twelve months, judged as a rate rather than a lifetime total. Momentum is what a frontier index cares about: a result being taken up fast right now is landing, whether its lifetime pile is large or still small. Where an expected field citation rate is known, the momentum is also read relative to it, so a fast-moving result in a quiet field is not overlooked.Maps to 0 to 100: Saturating: 30 citations in the last twelve months maps to 50. When NIH iCite provides the field's expected rate, this is blended evenly with the ratio of observed to expected momentum.81 in last 12 months (9.2x field rate)820.3528.6OpenAlex, NIH iCite counts_by_year / field_citation_rate
Novelty47
Novelty sub-metrics
Sub-metricRaw value0 to 100WeightPointsSource
Conceptual noveltyHow genuinely new the contribution is, measured at publication from the work's OWN content, not its date: how atypical its combination of research areas is versus all prior work, and whether it builds on a broad, deep base rather than only the newest papers in a fast-moving area. This is what separates a real advance from a recent-but-commodity result, and it needs no forward citations, so it is stable for brand-new work. An honestly noisy proxy (even state-of-the-art novelty measures agree only moderately with expert judgment), so it carries bounded weight and never drives the ranking alone.Maps to 0 to 100: Blends two publication-time signals, each 0-100: an atypical-combination score (how unusual the pairing of the work's research topics is versus the prior literature, by pointwise mutual information) and a foundational-reach score (median age and depth of the work's references). The stronger signal weighs 0.62, the weaker 0.38. When neither is measurable (a work too thin in topics or references), the pillar renormalizes over its recency signals instead.Action Observation and Synchronization x Functional Brain Connectivity Studies (1,258 prior works made this pairing); builds on refs a median 4.7y deep330.4514.7OpenAlex topics + referenced_works (publication-time)
RecencyHow recently the work appeared. The frontier is now, so newer scores higher.Maps to 0 to 100: Exponential decay with a 12-month half-life from the publication date.15 months old420.3012.7OpenAlex publication_date
Attention accelerationWhether attention is rising: citations in the latest full year versus the year before. Accelerating interest signals an active, opening frontier.Maps to 0 to 100: Latest/prior-year ratio: 1x maps to 50, 2x to 75, 4x to 100. Unknown is neutral.11.7x rising1000.1010.0OpenAlex counts_by_year
Edge of reviewPreprints and brand-new work sit ahead of the peer-review process. That earns novelty here (while it is discounted under Evidence), because it is where the frontier forms first.Maps to 0 to 100: Preprint 95, article <6 months 80, article <18 months 65, else 45.Peer-reviewed record650.159.8OpenAlex type / publication_date
ProvenanceAll sources

Sources and data

Every number behind this entry, grouped by the source that holds it: up to five citation indices cross-checked, field-normalized impact, then real-world reach (Wikipedia, community discussion, clinical trials, released data and public funding) where it exists. Each value links to its record.

OpenAlex

Primary record
Total citations
108
Field-weighted citation impact
64.7x field average
Citation percentile
Top 0.1% of its field-year
Citations, recent 12 months
81
Publication date
2025-07-02

Crossref

Citations
127
Works it builds on
71

Semantic Scholar

Citations, all versions
242
Influential citations
20

OpenCitations

Citations
73

Europe PMC

Citations
44

NIH iCite

Relative Citation Ratio
12.79x field average
Potential to translate
50%

Unpaywall

Reuse license
CC-BY

Figures quoted in the write-up

FigureValueSource
Participants (People whose decisions are in the Psych-101 training set)60,000+Nature
Choices (Trial-by-trial human decisions used for fine-tuning)10,000,000+Nature
Experiments (Distinct psychology experiments spanned by Psych-101)160Nature
NeuroscienceThe whole field

Back to the Frontier Index, or browse every breakthrough by date.

The Frontier Brief

The week's frontier, scored.

Every Monday, the highest-scoring new breakthroughs in the index, each with its score and a one-line read on why it matters. What you get

One email a week. The top new breakthroughs, scored. Unsubscribe anytime.