Machine Learning seminars
September 2020
Free will, decision-making and machine learning
Siobhan Hall· Stellenbosch University
Wed, Sep 9 · 17:30 UTC
The question of free will has been topical for millennia, especially considering its links to moral responsibility and the ownership of that responsibility. Free will, or volition, is an incredibly complex phenomenon - and cannot easily be reduced to a single empirical paradigm. Roskies (2010) proposes that there are five cognitive aspects to be considered when developing a more complete understanding of volition. These are: intention, initiation, feeling, executive control and decision-making. Decision-making will be the focus of this talk, which steps through aspects of the philosophy of free will; highlights experimental paradigms stemming from the seminal work of Benjamin Libet et al., and proposes machine learning as a promising method in progressing the empirical studies of decision-making and free will.
In the Learning Salon, we will discuss the similarities and differences between biological and machine learning, including individuals with diverse perspectives and backgrounds, so we can all learn from one another.
Fast and deep neuromorphic learning with time-to-first-spike coding
Julian Goeltz· Universität Bern
Tue, Sep 1 · 16:55 UTC
Engineered pattern-recognition systems strive for short time-to-solution and low energy-to-solution characteristics. This represents one of the main driving forces behind the development of neuromorphic devices. For both them and their biological archetypes, this corresponds to using as few spikes as early as possible. The concept of few and early spikes is used as the founding principle in the time-to-first-spike coding scheme. Within this framework, we have developed a spike-timing-based learning algorithm, which we used to train neuronal networks on the mixed-signal neuromorphic platform BrainScaleS-2. We derive, from first principles, error-backpropagation-based learning in networks of leaky integrate-and-fire (LIF) neurons relying only on spike times, for specific configurations of neuronal and synaptic time constants. We explicitly examine applicability to neuromorphic substrates by studying the effects of reduced weight precision and range, as well as of parameter noise. We demonstrate the feasibility of our approach on continuous and discrete data spaces, both in software simulations and on BrainScaleS-2. This narrows the gap between previous models of first-spike-time learning and biological neuronal dynamics and paves the way for fast and energy-efficient neuromorphic applications.
Back-propagation in spiking neural networks
Timothee Masquelier· Centre national de la recherche scientifique, CNRS | Toulouse
Tue, Sep 1 · 14:10 UTC
Back-propagation is a powerful supervised learning algorithm in artificial neural networks, because it solves the credit assignment problem (essentially: what should the hidden layers do?). This algorithm has led to the deep learning revolution. But unfortunately, back-propagation cannot be used directly in spiking neural networks (SNN). Indeed, it requires differentiable activation functions, whereas spikes are all-or-none events which cause discontinuities. Here we present two strategies to overcome this problem. The first one is to use a so-called 'surrogate gradient', that is to approximate the derivative of the threshold function with the derivative of a sigmoid. We will present some applications of this method for time series processing (audio, internet traffic, EEG). The second one concerns a specific class of SNNs, which process static inputs using latency coding with at most one spike per neuron. Using approximations, we derived a latency-based back-propagation rule for this sort of networks, called S4NN, and applied it to image classification.
August 2020
Synthesizing Machine Intelligence in Neuromorphic Computers with Differentiable Programming
Emre Neftci· University of California Irvine
Mon, Aug 31 · 16:55 UTC
The potential of machine learning and deep learning to advance artificial intelligence is driving a quest to build dedicated computers, such as neuromorphic hardware that emulate the biological processes of the brain. While the hardware technologies already exist, their application to real-world tasks is hindered by the lack of suitable programming methods. Advances at the interface of neural computation and machine learning showed that key aspects of deep learning models and tools can be transferred to biologically plausible neural circuits. Building on these advances, I will show that differentiable programming can address many challenges of programming spiking neural networks for solving real-world tasks, and help devise novel continual and local learning algorithms. In turn, these new algorithms pave the road towards systematically synthesizing machine intelligence in neuromorphic hardware without detailed knowledge of the hardware circuits.
E-prop: A biologically inspired paradigm for learning in recurrent networks of spiking neurons
Franz Scherr· Technische Universität Graz
Mon, Aug 31 · 16:10 UTC
Transformative advances in deep learning, such as deep reinforcement learning, usually rely on gradient-based learning methods such as backpropagation through time (BPTT) as a core learning algorithm. However, BPTT is not argued to be biologically plausible, since it requires to a propagate gradients backwards in time and across neurons. Here, we propose e-prop, a novel gradient-based learning method with local and online weight update rules for recurrent neural networks, and in particular recurrent spiking neural networks (RSNNs). As a result, e-prop has the potential to provide a substantial fraction of the power of deep learning to RSNNs. In this presentation, we will motivate e-prop from the perspective of recent insights in neuroscience and show how these have to be combined to form an algorithm for online gradient descent. The mathematical results will be supported by empirical evidence in supervised and reinforcement learning tasks. We will also discuss how limitations that are inherited from gradient-based learning methods, such as sample-efficiency, can be addressed by considering an evolution-like optimization that enhances learning on particular task families. The emerging learning architecture can be used to learn tasks by a single demonstration, hence enabling one-shot learning.
On temporal coding in spiking neural networks with alpha synaptic function
Iulia M. Comsa· Google Research Zürich, Switzerland
Mon, Aug 31 · 14:55 UTC
The timing of individual neuronal spikes is essential for biological brains to make fast responses to sensory stimuli. However, conventional artificial neural networks lack the intrinsic temporal coding ability present in biological networks. We propose a spiking neural network model that encodes information in the relative timing of individual neuron spikes. In classification tasks, the output of the network is indicated by the first neuron to spike in the output layer. This temporal coding scheme allows the supervised training of the network with backpropagation, using locally exact derivatives of the postsynaptic spike times with respect to presynaptic spike times. The network operates using a biologically-plausible alpha synaptic transfer function. Additionally, we use trainable synchronisation pulses that provide bias, add flexibility during training and exploit the decay part of the alpha function. We show that such networks can be trained successfully on noisy Boolean logic tasks and on the MNIST dataset encoded in time. The results show that the spiking neural network outperforms comparable spiking models on MNIST and achieves similar quality to fully connected conventional networks with the same architecture. We also find that the spiking network spontaneously discovers two operating regimes, mirroring the accuracy-speed trade-off observed in human decision-making: a slow regime, where a decision is taken after all hidden neurons have spiked and the accuracy is very high, and a fast regime, where a decision is taken very fast but the accuracy is lower. These results demonstrate the computational power of spiking networks with biological characteristics that encode information in the timing of individual neurons. By studying temporal coding in spiking networks, we aim to create building blocks towards energy-efficient and more complex biologically-inspired neural architectures.
Effective and Efficient Computation with Multiple-timescale Spiking Recurrent Neural Networks
Sander Bohte· Centrum Wiskunde & Informatica, Amsterdam
Mon, Aug 31 · 14:10 UTC
The emergence of brain-inspired neuromorphic computing as a paradigm for edge AI is motivating the search for high-performance and efficient spiking neural networks to run on this hardware. However, compared to classical neural networks in deep learning, current spiking neural networks lack competitive performance in compelling areas. Here, for sequential and streaming tasks, we demonstrate how spiking recurrent neural networks (SRNN) using adaptive spiking neurons are able to achieve state-of-the-art performance compared to other spiking neural networks and almost reach or exceed the performance of classical recurrent neural networks (RNNs) while exhibiting sparse activity. From this, we calculate a 100x energy improvement for our SRNNs over classical RNNs on the harder tasks. We find in particular that adapting the timescales of spiking neurons is crucial for achieving such performance, and we demonstrate the performance for SRNNs for different spiking neuron models.
Student´s Oral Presentation III: Emotional State Classification Using Low-Cost Single-Channel Electroencephalography
Francisco López-Guzmán, Rodrigo Sanz, Montevideo, Uruguay· Universidad de la República, Montevideo, Uruguay
Thu, Aug 20 · 12:30 UTC
Although electroencephalography (EEG) has been used in clinical and research studies for almost a century, recent technological advances have made the equipment and processing tools more accessible outside laboratory settings. These low-cost alternatives can achieve satisfactory results in experiments such as detecting event-related potentials and classifying cognitive states. In our research, we use low-cost single-channel EEG to classify brain activity during the presentation of images of opposite emotional valence from the OASIS database. Emotional classification has already been achieved using research-grade and commercial-grade equipment, but our approach pioneers the use of educational-grade equipment for said task. EEG data is collected with a Backyard Brains SpikerBox, a low-cost and open-source bioamplifier that can record a single-channel electric signal from a pair of electrodes placed on the scalp, and used to train machine learning classifiers.
Machine learning methods applied to dMRI tractography for the study of brain connectivity
Pamela Guevara· Department of Electrical Engineering, Faculty of Engineering, Universidad de Concepción, Chile
Wed, Aug 19 · 13:15 UTC
Tractography datasets, calculated from dMRI, represent the main WM structural connections in the brain. Thanks to advances in image acquisition and processing, the complexity and size of these datasets have constantly increased, also containing a large amount of artifacts. We present some examples of algorithms, most of them based on classical machine learning approaches, to analyze these data and identify common connectivity patterns among subjects.
July 2020
Predicting Patterns of Similarity Among Abstract Semantic Relations
Nick Ichien· UCLA
Thu, Jul 9 · 16:00 UTC
In this talk, I will present some data showing that people’s similarity judgments among word pairs reflect distinctions between abstract semantic relations like contrast, cause-effect, or part-whole. Further, the extent that individual participants’ similarity judgments discriminate between abstract semantic relations was linearly associated with both fluid and crystallized verbal intelligence, albeit more strongly with fluid intelligence. Finally, I will compare three models according to their ability to predict these similarity judgments. All models take as input vector representations of individual word meanings, but they differ in their representation of relations: one model does not represent relations at all, a second model represents relations implicitly, and a third model represents relations explicitly. Across the three models, the third model served as the best predictor of human similarity judgments suggesting the importance of explicit relation representation to fully account for human semantic cognition.
Multi-resolution Multi-task Gaussian Processes: London air pollution
Ollie Hamelijnck· The Alan Turing Institute, London
Thu, Jul 9 · 13:00 UTC
Poor air quality in cities is a significant threat to health and life expectancy, with over 80% of people living in urban areas exposed to air quality levels that exceed World Health Organisation limits. In this session, I present a multi-resolution multi-task framework that handles evidence integration under varying spatio-temporal sampling resolution and noise levels. We have developed both shallow Gaussian Process (GP) mixture models and deep GP constructions that naturally handle this evidence integration, as well as biases in the mean. These models underpin our work at the Alan Turing Institute towards providing spatio-temporal forecasts of air pollution across London. We demonstrate the effectiveness of our framework on both synthetic examples and applications on London air quality. For further information go to: https://www.turing.ac.uk/research/research-projects/london-air-quality. Collaborators: Oliver Hamelijnck, Theodoros Damoulas, Kangrui Wang and Mark Girolami.
Untangling the web of behaviours used to produce spider orb webs
Andrew Gordus· John Hopkins University
Wed, Jul 8 · 06:00 UTC
Many innate behaviours are the result of multiple sensorimotor programs that are dynamically coordinated to produce higher-order behaviours such as courtship or architecture construction. Extendend phenotypes such as architecture are especially useful for ethological study because the structure itself is a physical record of behavioural intent. A particularly elegant and easily quantifiable structure is the spider orb-web. The geometric symmetry and regularity of these webs have long generated interest in their behavioural origin. However, quantitative analyses of this behaviour have been sparse due to the difficulty of recording web-making in real-time. To address this, we have developed a novel assay enabling real-time, high-resolution tracking of limb movements and web structure produced by the hackled orb-weaver Uloborus diversus. With its small brain size of approximately 100,000 neurons, the spider U. diversus offers a tractable model organism for the study of complex behaviours. Using deep learning frameworks for limb tracking, and unsupervised behavioural clustering methods, we have developed an atlas of stereotyped movement motifs and are investigating the behavioural state transitions of which the geometry of the web is an emergent property. In addition to tracking limb movements, we have developed algorithms to track the web’s dynamic graph structure. We aim to model the relationship between the spider’s sensory experience on the web and its motor decisions, thereby identifying the sensory and internal states contributing to this sensorimotor transformation. Parallel efforts in our group are establishing 2-photon in vivo calcium imaging protocols in this spider, eventually facilitating a search for neural correlates underlying the internal and sensory state variables identified by our behavioural models. In addition, we have assembled a genome, and are developing genetic perturbation methods to investigate the genetic underpinnings of orb-weaving behaviour. Together, we aim to understand how complex innate behaviours are coordinated by underlying neuronal and genetic mechanisms.
Learning Theory for Continual and Meta-Learning
Christoph Lampert· Institute of Science and Technology Austria
Thu, Jul 2 · 13:00 UTC
June 2020
High-dimensional geometry of visual cortex
Carsen Stringer_· Janelia Research Campus
Thu, Jun 25 · 17:00 UTC
Interpreting high-dimensional datasets requires new computational and analytical methods. We developed such methods to extract and analyze neural activity from 20,000 neurons recorded simultaneously in awake, behaving mice. The neural activity was not low-dimensional as commonly thought, but instead was high-dimensional and obeyed a power-law scaling across its eigenvalues. We developed a theory that proposes that neural responses to external stimuli maximize information capacity while maintaining a smooth neural code. We then observed power-law eigenvalue scaling in many real-world datasets, and therefore developed a nonlinear manifold embedding algorithm called Rastermap that can capture such high-dimensional structure.
Understanding machine learning via exactly solvable statistical physics models
Lenka Zdeborová· CNRS & CEA Saclay
Wed, Jun 24 · 13:00 UTC
The affinity between statistical physics and machine learning has long history, this is reflected even in the machine learning terminology that is in part adopted from physics. I will describe the main lines of this long-lasting friendship in the context of current theoretical challenges and open questions about deep learning. Theoretical physics often proceeds in terms of solvable synthetic models, I will describe the related line of work on solvable models of simple feed-forward neural networks. I will highlight a path forward to capture the subtle interplay between the structure of the data, the architecture of the network, and the learning algorithm.
Disentangling the roles of dimensionality and cell categories in neural computations
Srdjan Ostojic· École Normale Supérieure
Fri, Jun 19 · 13:00 UTC
The description of neural computations currently relies on two competing views: (i) a classical single-cell view that aims to relate the activity of individual neurons to sensory or behavioural variables, and organize them into functional classes; (ii) a more recent population view that instead characterises computations in terms of collective neural trajectories, and focuses on the dimensionality of these trajectories as animals perform tasks. How the two key concepts of functional cell classes and low-dimensional trajectories interact to shape neural computations is however at present not understood. Here I will address this question by combining machine-learning tools for training recurrent neural networks with reverse-engineering and theoretical analyses of network dynamics.
Thinking Fast and Slow in AlphaZero and the Brain
Wed, Jun 17 · 11:30 UTC · Online
In his bestseller 'Thinking, Fast and Slow', Daniel Kahneman popularized the idea that there are two fundamentally different process of thought: a 'System 1' process that is unconscious and instinctive, and a 'System 2' process that is deliberative and requires conscious attention. There is a growing recognition that machine learning is mostly stuck at the 'System 1' level of cognition, and that moving to 'System 2' methods are key to solving long-standing challenges such as out-of-distribution generalization. In this talk, AlphaZero will be used as a case-study of the power of combining 'System 1' and 'System 2' processes. The similarities and differences between AlphaZero and human learning will be explored, along with drawing lessons for the future of machine learning.
Deep learning for model-based RL
Timothy Lillicrap· Google Deep Mind, University College London
Fri, Jun 12 · 13:00 UTC
Model-based approaches to control and decision making have long held the promise of being more powerful and data efficient than model-free counterparts. However, success with model-based methods has been limited to those cases where a perfect model can be queried. The game of Go was mastered by AlphaGo using a combination of neural networks and the MCTS planning algorithm. But planning required a perfect representation of the game rules. I will describe new algorithms that instead leverage deep neural networks to learn models of the environment which are then used to plan, and update policy and value functions. These new algorithms offer hints about how brains might approach planning and acting in complex environments.
The geometry of abstraction in artificial and biological neural networks
Stefano Fusi· Columbia University
Thu, Jun 11 · 13:00 UTC
The curse of dimensionality plagues models of reinforcement learning and decision-making. The process of abstraction solves this by constructing abstract variables describing features shared by different specific instances, reducing dimensionality and enabling generalization in novel situations. We characterized neural representations in monkeys performing a task where a hidden variable described the temporal statistics of stimulus-response-outcome mappings. Abstraction was defined operationally using the generalization performance of neural decoders across task conditions not used for training. This type of generalization requires a particular geometric format of neural representations. Neural ensembles in dorsolateral pre-frontal cortex, anterior cingulate cortex and hippocampus, and in simulated neural networks, simultaneously represented multiple hidden and explicit variables in a format reflecting abstraction. Task events engaging cognitive operations modulated this format. These findings elucidate how the brain and artificial systems represent abstract variables, variables critical for generalization that in turn confers cognitive flexibility.