Learning to Express Reward Prediction Error-like Dopaminergic Activity Requires Plastic Representations of Time
Harel Shouval · The University of Texas at Houston
Wed, Jun 14, 2023 · 05:00 UTC
The dominant theoretical framework to account for reinforcement learning in the brain is temporal difference (TD) reinforcement learning. The TD framework predicts that some neuronal elements should represent the reward prediction error (RPE), which means they signal the difference between the expected future rewards and the actual rewards. The prominence of the TD theory arises from the observation that firing properties of dopaminergic neurons in the ventral tegmental area appear similar to those of RPE model-neurons in TD learning. Previous implementations of TD learning assume a fixed temp