Understanding reward-guided learning using large-scale datasets
Understanding the neural mechanisms of reward-guided learning is a long-standing goal of computational neuroscience. Recent methodological innovations enable us to collect ever larger neural and behavioral datasets. This presents opportunities to achieve greater understanding of learning in the brain at scale, as well as methodological challenges. In the first part of the talk, I will discuss our recent insights into the mechanisms by which zebra finch songbirds learn to sing. Dopamine has been long thought to guide reward-based trial-and-error learning by encoding reward prediction errors. However, it is unknown whether the learning of natural behaviours, such as developmental vocal learning, occurs through dopamine-based reinforcement. Longitudinal recordings of dopamine and bird songs reveal that dopamine activity is indeed consistent with encoding a reward prediction error during naturalistic learning. In the second part of the talk, I will talk about recent work we are doing at DeepMind to develop tools for automatically discovering interpretable models of behavior directly from animal choice data. Our method, dubbed CogFunSearch, uses LLMs within an evolutionary search process in order to "discover" novel models in the form of Python programs that excel at accurately predicting animal behavior during reward-guided learning. The discovered programs reveal novel patterns of learning and choice behavior that update our understanding of how the brain solves reinforcement learning problems.
Understanding reward-guided learning using large-scale datasets
Understanding the neural mechanisms of reward-guided learning is a long-standing goal of computational neuroscience. Recent methodological innovations enable us to collect ever larger neural and behavioral datasets. This presents opportunities to achieve greater understanding of learning in the brain at scale, as well as methodological challenges. In the first part of the talk, I will discuss our recent insights into the mechanisms by which zebra finch songbirds learn to sing. Dopamine has been long thought to guide reward-based trial-and-error learning by encoding reward prediction errors. However, it is unknown whether the learning of natural behaviours, such as developmental vocal learning, occurs through dopamine-based reinforcement. Longitudinal recordings of dopamine and bird songs reveal that dopamine activity is indeed consistent with encoding a reward prediction error during naturalistic learning. In the second part of the talk, I will talk about recent work we are doing at DeepMind to develop tools for automatically discovering interpretable models of behavior directly from animal choice data. Our method, dubbed CogFunSearch, uses LLMs within an evolutionary search process in order to "discover" novel models in the form of Python programs that excel at accurately predicting animal behavior during reward-guided learning. The discovered programs reveal novel patterns of learning and choice behavior that update our understanding of how the brain solves reinforcement learning problems.
Thalamocortical dynamics in a complex learned behavior
Performing learned behaviors requires animals to produce precisely timed motor sequences. The underlying neuronal circuits must convert incoming spike trains into precisely timed firing to indicate the onset of crucial sensory cues or to carry out well-coordinated muscle movements. Birdsong is a remarkable example of a complex, learned and precisely timed natural behavior which is controlled by a brainstem-thalamocortical feedback loop. Projection neurons within the zebra finch cortical nucleus HVC (used as a proper name), produce precisely timed, highly reliable and ultra-sparse neural sequences that are thought to underlie song dynamics. However, the origin of short timescale dynamics of the song is debated. One model posits that these dynamics reside in HVC and are mediated through a synaptic chain mechanism. Alternatively, the upstream motor thalamic nucleus Uveaformis (Uva), could drive HVC bursts as part of a brainstem-thalamocortical distributed network. Using focal temperature manipulation we found that the song dynamics reside chiefly in HVC. We then characterized the activity of thalamic nucleus Uva, which provides input to HVC. We developed a lightweight (~1 g) microdrive for juxtacellular recordings and with it performed the very first extracellular single unit recordings in Uva during song. Recordings revealed HVC-projecting Uva neurons contain timing information during the song, but compared to HVC neurons, fire densely in time and are much less reliable. Computational models of Uva-driven HVC neurons estimated that a high degree of synaptic convergence is needed from Uva to HVC to overcome the inconsistency of Uva firing patterns. However, axon terminals of single Uva neurons exhibit low convergence within HVC such that each HVC neuron receives input from 2-7 Uva neurons. These results suggest that thalamus maintains sequential cortical activity during song but does not provide unambiguous timing information. Our observations are consistent with a model in which the brainstem-thalamocortical feedback loop acts at the syllable timescale (~100 ms) and does not support a model in which the brainstem-thalamocortical feedback loop acts at fast timescale (~10 ms) to generate sequences within cortex. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2024-02-14. Recording duration: 00:39:45.
Variability, maintenance and learning in birdsong
The songbird zebra finch is an exemplary model system in which to study trial-and-error learning, as the bird learns its single song gradually through the production of many noisy renditions. It is also a good system in which to study the maintenance of motor skills, as the adult bird actively maintains its song and retains some residual plasticity. Motor learning occurs through the association of timing within the song, represented by sparse firing in nucleus HVC, with motor output, driven by nucleus RA. Here we show through modeling that the small level of observed variability in HVC can result in a network which is more easily able to adapt to change, and is most robust to cell damage or death, than an unperturbed network. In collaboration with Carlos Lois’ lab, we also consider the effect of directly perturbing HVC through viral injection of toxins that affect the firing of projection neurons. Following these perturbations, the song is profoundly affected but is able to almost perfectly recover. We characterize the changes in song acoustics and syntax, and propose models for HVC architecture and plasticity that can account for some of the observed effects. Finally, we suggest a potential role for inputs from nucleus Uva in helping to control timing precision in HVC.