Towards a general model of human reward-based learning
Google Deepmind
Recording
Abstract
Traditional work in the study of human reward-based learning involves designing an experimental task---often inspired by Reinforcement Learning (RL) theory---and fits a small set of computational models---often inspired by RL algorithms---to that dataset. For example, researchers often model human behavior on bandit tasks using variants of Q-learning. While this approach has been highly productive, leading to landmark discoveries such as the dopamine reward prediction error hypothesis, it also has limitations. This talk focuses on the lack of generalizability of such models: Even if they closely fit behavior on the original task, models derived from the one-task-one-model paradigm usually predict behavior on other tasks quite poorly. I argue that this lack of generalizability is a fundamental problem for the cognitive sciences: we intuitively expect our models to be robust to superficial task differences, such as variations in the number of choice options, reward probabilities, or the exact kind of non-stationarity. I will propose potential solutions to this problem along two dimensions: the behavioral dataset and the computational model. Regarding computational models, I will introduce work in which we moved beyond the limitations of hand-crafted one-off models by employing flexible, data-driven methods. These methods allowed us to compare classes of models instead of individual model instances, allowing us to cover the space of possible models more exhaustively, and innovate cognitive mechanisms very efficiently. For the behavioral dataset, we move from using single learning tasks to a comprehensive task space that encompasses most existing paradigms in the literature, while closing the gaps between them in a near-continuous fashion. Our results suggest that more general models in conjunction with broader datasets can pave the road toward increasingly general models of human reward-based learning and decision making, and a persistent departure from many aspects of RL theory. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2026-05-20. Recording duration: 00:51:19.
Topics
Related Job Opportunities
Postdoctoral Scientist - Sarvestani Lab
The Sarvestani Lab at Cornell University is recruiting a postdoctoral scientist in systems neuroscience to study how visual and motor systems across the brain and body support perception and…
PhD Studentship: Mitochondrial Metabolism and Novel Therapeutic Strategies for Metabolic Dysfunction-Associated Steatotic Liver Disease (MASLD) (Fixed Term)
Supervisors: Professor Andrew Murray, Department of Physiology, Development and Neuroscience, University of Cambridge Dr Ross Lindsay, Novo Nordisk Funding: Fully funded PhD studentship (Home/UK…
Research Associate (Fixed Term)
We seek a highly motivated Postdoctoral Research Associate to join the laboratory of Professor Kathy Niakan. We are based in the Loke Centre for Trophoblast Research (LCTR), in the Department of…