What it is
Researchers fine-tuned a large language model on Psych-101, a dataset of trial-by-trial decisions from more than 60,000 participants making over 10 million choices across 160 experiments. The resulting model, Centaur, predicts and simulates human behavior in any task that can be described in natural language, beating existing bespoke cognitive models on held-out participants. It also generalizes to unseen cover stories, altered task structures, and entirely new domains, and its internal representations became more aligned with human neural activity after fine-tuning.
Why it matters
Cognitive science has relied on separate hand-built models for each task, and none transferred well. Centaur is a single system trained on 10 million choices that out-predicts those specialized models and holds up out of distribution, suggesting one computational model can capture behavior across many domains at once. That is a plausible foundation for a more unified, testable theory of how people decide and learn.
Underlined numbers link to their source. Every metric and quoted figure is listed under Sources and data below.
Filed undercognition, foundation model, LLM, behavior, neuroscience