Decomposing motivation into value and salience
Humans and other animals approach reward and avoid punishment and pay attention to cues predicting these events. Such motivated behavior thus appears to be guided by value, which directs behavior towards or away from positively or negatively valenced outcomes. Moreover, it is facilitated by (top-down) salience, which enhances attention to behaviorally relevant learned cues predicting the occurrence of valenced outcomes. Using human neuroimaging, we recently separated value (ventral striatum, posterior ventromedial prefrontal cortex) from salience (anterior ventromedial cortex, occipital cortex) in the domain of liquid reward and punishment. Moreover, we investigated potential drivers of learned salience: the probability and uncertainty with which valenced and non-valenced outcomes occur. We find that the brain dissociates valenced from non-valenced probability and uncertainty, which indicates that reinforcement matters for the brain, in addition to information provided by probability and uncertainty alone, regardless of valence. Finally, we assessed learning signals (unsigned prediction errors) that may underpin the acquisition of salience. Particularly the insula appears to be central for this function, encoding a subjective salience prediction error, similarly at the time of positively and negatively valenced outcomes. However, it appears to employ domain-specific time constants, leading to stronger salience signals in the aversive than the appetitive domain at the time of cues. These findings explain why previous research associated the insula with both valence-independent salience processing and with preferential encoding of the aversive domain. More generally, the distinction of value and salience appears to provide a useful framework for capturing the neural basis of motivated behavior.
Reinforcement Learning
The structure of behavior entrained to long intervals
Interpretation of interval timing data generated from animal models is complicated by ostensible motivational effects which arise from the delay-to-reward imposed by interval timing tasks, as well as overlap between timed and non-timed responses. These factors become increasingly prevalent at longer intervals. To address these concerns, two adjustments to long interval timing tasks are proposed. First, subjects should be afforded with reinforced non-timing behaviors concurrent with timing. Second, subjects should initiate the onset of timed stimuli. Under these conditions, interference by extraneous behavior would be detected in the rate of concurrent non- timing behaviors, and changes in motivation would be detected in the rate at which timed stimuli are initiated. In a task with these characteristics, rats initiated a concurrent fixed-interval (FI) random-ratio (RR) schedule of reinforcement. This design facilitated response-initiated timing behavior, even at increasingly long delays. Pre-feeding manipulations revealed an effect on the number of initiated trials, but not on the timing peak function.
Mechanisms of Perceptual Learning
Perceptual learning (PL) is defined as long-term performance improvement on a perceptual task as a result of perceptual experience (Sasaki, Nanez& Watanabe, 2011, Nat Rev Neurosci, 2011). We first found that PL occurs for task-irrelevant and subthreshold features and that pairing task-irrelevant features with rewards is the key to form task-irrelevant PL (TIPL) (Watanabe, Nanez & Sasaki, Nature, 2001; Watanabe et al, 2002, Nature Neuroscience; Seitz & Watanabe, Nature, 2003; Seitz, Kim & Watanabe, 2009, Neuron; Shibata et al, 2011, Science). These results suggest that PL occurs as a result of interactions between reinforcement and bottom-up stimulus signals (Seitz & Watanabe, 2005, TICS). On the other hand, fMRI study results indicate that lateral prefrontal cortex fails to detect and thus to suppress subthreshold task-irrelevant signals. This leads to the paradoxical effect that a signal that is below, but close to, one’s discrimination threshold ends up being stronger than suprathreshold signals (Tsushima, Sasaki & Watanabe, 2006, Science). We confirmed this mechanism with the following results: Task-irrelevant learning occurs only when a presented feature is under and close to the threshold with younger individuals (Tsushima et al, 2009, Current Biol), whereas with older individuals who tend to have less inhibitory control task-irrelevant learning occurs with a feature whose signal is much greater than the threshold (Chang et al, 2014, Current Biol). From all of these results, we conclude that attention and reward play important but different roles in PL. I will further discuss different stages and phases in mechanisms of PL (Seitz et al, 2005, PNAS; Yotsumoto, Watanabe & Sasaki, Neuron, 2008; Yotsumoto et al, Curr Biol, 2009; Watanabe & Sasaki, 2015, Ann Rev Psychol; Shibata et al, 2017, Nat Neurosci; Tamaki et al, 2020, Nat Neurosci).