Vision Science seminars
January 2022
What happens to our ability to perceive multisensory information as we age?
Fiona Newell· Trinity Collge Dublin
Thu, Jan 13 · 16:00 UTC
Our ability to perceive the world around us can be affected by a number of factors including the nature of the external information, prior experience of the environment, and the integrity of the underlying perceptual system. A particular challenge for the brain is to maintain a coherent perception from information encoded by the peripheral sensory organs whose function is affected by typical, developmental changes across the lifespan. Yet, how the brain adapts to the maturation of the senses, as well as experiential changes in the multisensory environment, is poorly understood. Over the past few years, we have used a range of multisensory tasks to investigate the role of ageing on the brain’s ability to merge sensory inputs. In particular, we have embedded an audio-visual task based on the sound-induced flash illusion (SIFI) into a large-scale, longitudinal study of ageing. Our findings support the idea that the temporal binding window (TBW) is modulated by age and reveal important individual differences in this TBW that may have clinical implications. However, our investigations also suggest the TWB is experience-dependent with evidence for both long and short term behavioural plasticity. An overview of these findings, including recent evidence on how multisensory integration may be associated with higher order functions, will be discussed.
December 2021
Wiring Minimization of Deep Neural Networks Reveal Conditions in which Multiple Visuotopic Areas Emerge
Dina Obeid· Harvard University
Wed, Dec 15 · 05:00 UTC
The visual system is characterized by multiple mirrored visuotopic maps, with each repetition corresponding to a different visual area. In this work we explore whether such visuotopic organization can emerge as a result of minimizing the total wire length between neurons connected in a deep hierarchical network. Our results show that networks with purely feedforward connectivity typically result in a single visuotopic map, and in certain cases no visuotopic map emerges. However, when we modify the network by introducing lateral connections, with sufficient lateral connectivity among neurons within layers, multiple visuotopic maps emerge, where some connectivity motifs yield mirrored alternations of visuotopic maps–a signature of biological visual system areas. These results demonstrate that different connectivity profiles have different emergent organizations under the minimum total wire length hypothesis, and highlight that characterizing the large-scale spatial organizing of tuning properties in a biological system might also provide insights into the underlying connectivity.
Spatial Integration in Normal Face Processing and Its Breakdown in Congenital Prosopagnosia
Galia Avidan· Ben Gurion U
Tue, Dec 14 · 16:00 UTC
Molecular recognition and the assembly of feature-selective retinal circuits
Arjun Krishnaswamy· Department of Physiology, McGill University
Tue, Dec 14 · 05:00 UTC
Roles of attention and consciousness in perceptual learning
Kazuhisa Shibata· RIKEN Center for Brain Science
Mon, Dec 13 · 22:00 UTC
Visual perceptual learning (VPL) is defined as improved performance on a visual task due to visual experience. It was once argued that attention to a visual feature is necessary for VPL of the feature to occur. Contrary to this view, a phenomenon called task-irrelevant VPL demonstrated that VPL can occur due to exposure to a feature which is sub-threshold and task-irrelevant, and therefore, unattended. A series of findings based on task-irrelevant VPL has indicated the following two mechanisms. First, attention to a feature facilitates VPL of the feature while inhibiting VPL of unattended and supra-threshold features. Second, reward paired with a feature enables VPL of the feature irrespective of whether the feature is attended or not. However, we recently found an additional twist; VPL of a task-irrelevant and supra-threshold feature embedded in a natural scene is not subject to the inhibition of attention. This new finding suggests a need to revise the current view or add a new mechanism as to how VPL occurs.
Opponent processing in the expanded retinal mosaic of Nymphalid butterflies
Gregor Belušič· University of Ljubljana
Mon, Dec 13 · 15:00 UTC
In many butterflies, the ancestral trichromatic insect colour vision, based on UV-, blue- and green-sensitive photoreceptors, is extended with red-sensitive cells. Physiological evidence for red receptors has been missing in nymphalid butterflies, although some species can discriminate red hues well. In eight species from genera Archaeoprepona, Argynnis, Charaxes, Danaus, Melitaea, Morpho, Heliconius and Speyeria, we found a novel class of green-sensitive photoreceptors that have hyperpolarizing responses to stimulation with red light. These green-positive, red-negative (G+R–) cells are allocated to positions R1/2, normally occupied by UV and blue-sensitive cells. Spectral sensitivity, polarization sensitivity and temporal dynamics suggest that the red opponent units (R–) are the basal photoreceptors R9, interacting with R1/2 in the same ommatidia via direct inhibitory synapses. We found the G+R– cells exclusively in butterflies with red-shining ommatidia, which contain longitudinal screening pigments. The implementation of the red colour channel with R9 is different from pierid and papilionid butterflies, where cells R5–8 are the red receptors. The nymphalid red-green opponent channel and the potential for tetrachromacy seem to have been switched on several times during evolution, balancing between the cost of neural processing and the value of extended colour information.
Decoding sounds in early visual cortex of sighted and blind individuals
Petra Vetter· University of Fribourg, Switzerland
Thu, Dec 9 · 16:00 UTC
Neurovascular signaling pathways in the mammalian retina
Will Grimes· NINDS/NIH
Mon, Dec 6 · 13:00 UTC
As a developmental outpocket of the brain, the retina exhibits features commonly found in most brain areas, including neurovascular interactions. In this presentation I will discuss various pathways that contribute to neurovascular interactions in the mammalian retina and present newly uncovered elements that likely participate in these pathways. Information obtained from retina could improve our understanding of neurovascular coupling pathways throughout the brain.
Spatial alignment supports visual comparisons
Nina Simms· Northwestern University
Thu, Dec 2 · 16:00 UTC
Visual comparisons are ubiquitous, and they can also be an important source for learning (e.g., Gentner et al., 2016; Kok et al., 2013). In science, technology, engineering, and math (STEM), key information is often conveyed through figures, graphs, and diagrams (Mayer, 1993). Comparing within and across visuals is critical for gleaning insight into the underlying concepts, structures, and processes that they represent. This talk addresses how people make visual comparisons and how visual comparisons can be best supported to improve learning. In particular, the talk will present a series of studies exploring the Spatial Alignment Principle (Matlen et al., 2020), derived from Structure-Mapping Theory (Gentner, 1983). Structure-mapping theory proposes that comparisons involve a process of finding correspondences between elements based on structured relationships. The Spatial Alignment Principle suggests that spatially arranging compared figures directly – to support correct correspondences and minimize interference from incorrect correspondences – will facilitate visual comparisons. We find that direct placement can facilitate visual comparison in educationally relevant stimuli, and that it may be especially important when figures are less familiar. We also present complementary evidence illustrating the preponderance of visual comparisons in 7th grade science textbooks.
NMC4 Short Talk: Sensory intermixing of mental imagery and perception
Nadine Dijkstra· Wellcome Centre for Human Neuroimaging
Thu, Dec 2 · 05:15 UTC
Several lines of research have demonstrated that internally generated sensory experience - such as during memory, dreaming and mental imagery - activates similar neural representations as externally triggered perception. This overlap raises a fundamental challenge: how is the brain able to keep apart signals reflecting imagination and reality? In a series of online psychophysics experiments combined with computational modelling, we investigated to what extent imagination and perception are confused when the same content is simultaneously imagined and perceived. We found that simultaneous congruent mental imagery consistently led to an increase in perceptual presence responses, and that congruent perceptual presence responses were in turn associated with a more vivid imagery experience. Our findings can be best explained by a simple signal detection model in which imagined and perceived signals are added together. Perceptual reality monitoring can then easily be implemented by evaluating whether this intermixed signal is strong or vivid enough to pass a ‘reality threshold’. Our model suggests that, in contrast to self-generated sensory changes during movement, our brain does not discount self-generated sensory signals during mental imagery. This has profound implications for our understanding of reality monitoring and perception in general.
The brain represents the external world through the bottleneck of sensory organs. The network of hierarchically organized neurons is thought to recover the causes of sensory inputs to reconstruct the reality in the brain in idiosyncratic ways depending on individuals and their internal states. How can we understand the world model represented in an individual’s brain, or the neuroverse? My lab has been working on brain decoding of visual perception and subjective experiences such as imagery and dreaming using machine learning and deep neural network representations. In this talk, I will outline the progress of brain decoding methods and present how subjective experiences are externalized as images and how they could be shared across individuals via neural code conversion. The prospects of these approaches in basic science and neurotechnology will be discussed.
NMC4 Short Talk: Hypothesis-neutral response-optimized models of higher-order visual cortex reveal strong semantic selectivity
Meenakshi Khosla· Massachusetts Institute of Technology
Wed, Dec 1 · 11:15 UTC
Modeling neural responses to naturalistic stimuli has been instrumental in advancing our understanding of the visual system. Dominant computational modeling efforts in this direction have been deeply rooted in preconceived hypotheses. In contrast, hypothesis-neutral computational methodologies with minimal apriorism which bring neuroscience data directly to bear on the model development process are likely to be much more flexible and effective in modeling and understanding tuning properties throughout the visual system. In this study, we develop a hypothesis-neutral approach and characterize response selectivity in the human visual cortex exhaustively and systematically via response-optimized deep neural network models. First, we leverage the unprecedented scale and quality of the recently released Natural Scenes Dataset to constrain parametrized neural models of higher-order visual systems and achieve novel predictive precision, in some cases, significantly outperforming the predictive success of state-of-the-art task-optimized models. Next, we ask what kinds of functional properties emerge spontaneously in these response-optimized models? We examine trained networks through structural ( feature visualizations) as well as functional analysis (feature verbalizations) by running `virtual' fMRI experiments on large-scale probe datasets. Strikingly, despite no category-level supervision, since the models are solely optimized for brain response prediction from scratch, the units in the networks after optimization act as detectors for semantic concepts like `faces' or `words', thereby providing one of the strongest evidences for categorical selectivity in these visual areas. The observed selectivity in model neurons raises another question: are the category-selective units simply functioning as detectors for their preferred category or are they a by-product of a non-category-specific visual processing mechanism? To investigate this, we create selective deprivations in the visual diet of these response-optimized networks and study semantic selectivity in the resulting `deprived' networks, thereby also shedding light on the role of specific visual experiences in shaping neuronal tuning. Together with this new class of data-driven models and novel model interpretability techniques, our study illustrates that DNN models of visual cortex need not be conceived as obscure models with limited explanatory power, rather as powerful, unifying tools for probing the nature of representations and computations in the brain.
NMC4 Short Talk: Image embeddings informed by natural language improve predictions and understanding of human higher-level visual cortex
Aria Wang· Carnegie Mellon University
Wed, Dec 1 · 11:00 UTC
To better understand human scene understanding, we extracted features from images using CLIP, a neural network model of visual concept trained with supervision from natural language. We then constructed voxelwise encoding models to explain whole brain responses arising from viewing natural images from the Natural Scenes Dataset (NSD) - a large-scale fMRI dataset collected at 7T. Our results reveal that CLIP, as compared to convolution based image classification models such as ResNet or AlexNet, as well as language models such as BERT, gives rise to representations that enable better prediction performance - up to a 0.86 correlation with test data and an r-square of 0.75 - in higher-level visual cortex in humans. Moreover, CLIP representations explain distinctly unique variance in these higher-level visual areas as compared to models trained with only images or text. Control experiments show that the improvement in prediction observed with CLIP is not due to architectural differences (transformer vs. convolution) or to the encoding of image captions per se (vs. single object labels). Together our results indicate that CLIP and, more generally, multimodal models trained jointly on images and text, may serve as better candidate models of representation in human higher-level visual cortex. The bridge between language and vision provided by jointly trained models such as CLIP also opens up new and more semantically-rich ways of interpreting the visual brain.
NMC4 Short Talk: Untangling Contributions of Distinct Features of Images to Object Processing in Inferotemporal Cortex
Hanxiao Lu· Yale University
Wed, Dec 1 · 09:15 UTC
How do humans perceive daily objects of various features and categorize these seemingly intuitive and effortless mental representations? Prior literature focusing on the role of the inferotemporal region (IT) has revealed object category clustering that is consistent with the semantic predefined structure (superordinate, ordinate, subordinate). It has however been debated whether the neural signals in the IT regions are a reflection of such categorical hierarchy [Wen et al.,2018; Bracci et al., 2017]. Visual attributes of images that correlated with semantic and category dimensions may have confounded these prior results. Our study aimed to address this debate by building and comparing models using the DNN AlexNet, to explain the variance in representational dissimilarity matrix (RDM) of neural signals in the IT region. We found that mid and high level perceptual attributes of the DNN model contribute the most to neural RDMs in the IT region. Semantic categories, as in predefined structure, were moderately correlated with mid to high DNN layers (r = [0.24 - 0.36]). Variance partitioning analysis also showed that the IT neural representations were mostly explained by DNN layers, while semantic categorical RDMs brought little additional information. In light of these results, we propose future works should focus more on the specific role IT plays in facilitating the extraction and coding of visual features that lead to the emergence of categorical conceptualizations.
NMC4 Short Talk: Directly interfacing brain and deep networks exposes non-hierarchical visual processing
Nick Sexton (he/him)· University College London
Wed, Dec 1 · 09:00 UTC
A recent approach to understanding the mammalian visual system is to show correspondence between the sequential stages of processing in the ventral stream with layers in a deep convolutional neural network (DCNN), providing evidence that visual information is processed hierarchically, with successive stages containing ever higher-level information. However, correspondence is usually defined as shared variance between brain region and model layer. We propose that task-relevant variance is a stricter test: If a DCNN layer corresponds to a brain region, then substituting the model’s activity with brain activity should successfully drive the model’s object recognition decision. Using this approach on three datasets (human fMRI and macaque neuron firing rates) we found that in contrast to the hierarchical view, all ventral stream regions corresponded best to later model layers. That is, all regions contain high-level information about object category. We hypothesised that this is due to recurrent connections propagating high-level visual information from later regions back to early regions, in contrast to the exclusively feed-forward connectivity of DCNNs. Using task-relevant correspondence with a late DCNN layer akin to a tracer, we used Granger causal modelling to show late-DCNN correspondence in IT drives correspondence in V4. Our analysis suggests, effectively, that no ventral stream region can be appropriately characterised as ‘early’ beyond 70ms after stimulus presentation, challenging hierarchical models. More broadly, we ask what it means for a model component and brain region to correspond: beyond quantifying shared variance, we must consider the functional role in the computation. We also demonstrate that using a DCNN to decode high-level conceptual information from ventral stream produces a general mapping from brain to model activation space, which generalises to novel classes held-out from training data. This suggests future possibilities for brain-machine interface with high-level conceptual information, beyond current designs that interface with the sensorimotor periphery.
November 2021
Spatial summation for motion detection
Joshua Solomon· City, University of London
Tue, Nov 30 · 16:00 UTC
What transcriptomics tells us about retinal development, disease and evolution
Joshua Sanes· Harvard University
Mon, Nov 22 · 14:00 UTC
Classification of neurons, long viewed as a fairly boring enterprise, has emerged as a major bottleneck in analysis of neural circuits. High throughput single cell RNA-seq has provided a new way to improve the situation. We initially applied this method to mouse retina, showing that its five neuronal classes (photoreceptors, three groups of interneurons, and retinal ganglion cells) can be divided into 130 discrete types. We then applied the method to other species including human, macaque, zebrafish and chick. With the atlases in hand, we are now using them to address questions about how retinal cell types diversify, how they differ in their responses to injury and disease, and the extent to which cell classes and types are conserved among vertebrates.
How our senses work both separately and together involves rich computational problems. I will discuss the spatial and representational problems faced by the visual and auditory system, focusing on two issues. 1. How does the brain correct for discrepancies in the visual and auditory spatial reference frames? I will describe our recent discovery of a novel type of otoacoustic emission, the eye movement related eardrum oscillation, or EMREO (Gruters et al, PNAS 2018). 2. How does the brain encode more than one stimulus at a time? I will discuss evidence for neural time-division multiplexing, in which neural activity fluctuates across time to allow representations to encode more than one simultaneous stimulus (Caruso et al, Nat Comm 2018). These findings all emerged from experimentally testing computational models regarding spatial representations and their transformations within and across sensory pathways. Further, they speak to several general problems confronting modern neuroscience such as the hierarchical organization of brain pathways and limits on perceptual/cognitive processing.