Reinforcement Learning

Upcoming events

The Institute for Advanced Study's School of Natural Sciences convenes mathematicians and physicists working on Liouville theory for a three-day workshop spanning its spacelike, timelike and related formulations. Andreas Blommaert and Beatrix Muehlmann organize the programme at Rubenstein Commons, and the official event page states that drop-in attendance is welcome.

Recordings

Wed, May 20, 2026 · 11:00 America/New_York

Traditional work in the study of human reward-based learning involves designing an experimental task---often inspired by Reinforcement Learning (RL) theory---and fits a small set of computational models---often inspired by RL algorithms---to that dataset. For example, researchers often model human behavior on bandit tasks using variants of Q-learning. While this approach has been highly productive, leading to landmark discoveries such as the dopamine reward prediction error hypothesis, it also has limitations. This talk focuses on the lack of generalizability of such models: Even if they closely fit behavior on the original task, models derived from the one-task-one-model paradigm usually predict behavior on other tasks quite poorly. I argue that this lack of generalizability is a fundamental problem for the cognitive sciences: we intuitively expect our models to be robust to superficial task differences, such as variations in the number of choice options, reward probabilities, or the exact kind of non-stationarity. I will propose potential solutions to this problem along two dimensions: the behavioral dataset and the computational model. Regarding computational models, I will introduce work in which we moved beyond the limitations of hand-crafted one-off models by employing flexible, data-driven methods. These methods allowed us to compare classes of models instead of individual model instances, allowing us to cover the space of possible models more exhaustively, and innovate cognitive mechanisms very efficiently. For the behavioral dataset, we move from using single learning tasks to a comprehensive task space that encompasses most existing paradigms in the literature, while closing the gaps between them in a near-continuous fashion. Our results suggest that more general models in conjunction with broader datasets can pave the road toward increasingly general models of human reward-based learning and decision making, and a persistent departure from many aspects of RL theory. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2026-05-20. Recording duration: 00:51:19.

human reward-based learningReinforcement Learning+8 moreSeries: van Vreeswijk Theoretical Neuroscience Seminar

What is So Interesting About Reinforcement Learning?

Andrew Barto · University of Massachusetts Amherst

Wed, Oct 29, 2025 · 11:00 America/New_York

This talk aims to answer these questions along four dimensions. First is history. RL was the basis of AI long before the term AI was introduced in 1956. The first machine learning (ML) systems were based on RL even before digital computers existed. Despite notable early successes of ML based on RL, RL essentially disappeared from ML until relatively recently. A second reason for renewed interest in RL is the clarification of some misunderstandings that have been prevalent in the ML community. A third, and most important, reason for this resurgence is that new, or rediscovered, algorithms and connections to well developed mathematical and engineering methods have been worked out. Finally, a fourth reason for the renewed interest in RL is its strong links to animal reward systems, in particular, to the role that dopamine plays in motivation and learning. VVTNS Sixth Season Opening Lecture. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2025-10-29. Recording duration: 00:55:21.

Reinforcement Learningartificial intelligence+8 moreSeries: van Vreeswijk Theoretical Neuroscience Seminar

Neural mechanisms of memory linking and replay: inhibition matters

Tomoki Fukai · Okinawa Institute of Science and Technology

Wed, May 14, 2025 · 11:00 America/New_York

My talk will consist of three subtopics. The brain remembers episodes not in isolation but with their contextual relationships, such as spatial or temporal proximity. This is an essential feature of the brain’s memory, but the underlying mechanism is yet to be explored. Cell assemblies, or engrams, may provide neural representations for such relationships. First, I will show a class of associative memory models that encode and retrieve multiple memory contents linked by an arbitrary graph structure through experience and demonstrate the crucial role of the balance between two inhibitory subnetwork types in the flexible retrieval of relational memories. Secondly, I propose a theoretical framework to generate a cognitive map, i.e., neural representations of relationships between memory items. This framework aims at the predictive function of the hippocampus and is based on successor representations proposed for reinforcement learning. Intriguingly, the model provides a unified account for grid cells in spatial navigation and concept cells in natural language processing. Finally, I will discuss another crucial role of the hippocampal memory system, memory replay, in a spiking neural network model. Unlike the conventional associative memory models that maintain attractor memory states, this model attempts to maximize the capacity of replayed activity patterns. Our model suggests the crucial role of inhibitory plasticity in optimizing spontaneous memory replay. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2025-05-14. Recording duration: 00:46:49.

Memory linkingMemory Replay+8 moreSeries: van Vreeswijk Theoretical Neuroscience Seminar

Wed, Jan 15, 2025 · 11:00 America/New_York

The dynamics of learning in natural and artificial environments is a problem of great interest to both neuroscientists and artificial intelligence experts. However, standard analyses of animal training data either treat behavior as fixed, or track only coarse performance statistics (e.g., accuracy and bias), providing limited insight into the dynamic evolution of behavioral strategies over the course of learning. To overcome these limitations, we propose a dynamic psychophysical model that efficiently tracks trial-to-trial changes in behavior over the course of training. In this talk, I will describe recent work based on a dynamic logistic regression model that captures the time-varying dependencies of behavior on stimuli and other task covariates, which we applied to mouse training data from the International Brain Lab (IBL). Secondly, I will discuss efforts to infer animal learning rules from time-varying behavior in order to characterize how they adjust their policy in response to reward. Finally, I will describe recent work on adaptive optimal training, which combines ideas from reinforcement learning and adaptive experimental design to formulate methods for inferring animal learning rules from behavior, and using these rules to speed up animal training. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2025-01-15. Recording duration: 00:49:39.

dynamic animal behaviorlearning dynamics+8 moreSeries: van Vreeswijk Theoretical Neuroscience Seminar

Open deadlines

No open deadlines listed.

Recent changes

The Institute for Advanced Study's School of Natural Sciences convenes mathematicians and physicists working on Liouville theory for a three-day workshop spanning its spacelike, timelike and related formulations. Andreas Blommaert and Beatrix Muehlmann organize the programme at Rubenstein Commons, and the official event page states that drop-in attendance is welcome.

We use essential cookies to run the site. Optional analytics and public-page session replay help us improve World Wide. Learn more.