Machine Learning seminars
April 2024
Improving Language Understanding by Generative Pre Training
Tue, Apr 23 · 11:00 UTC · Online
Natural language understanding comprises a wide range of diverse tasks such as textual entailment, question answering, semantic similarity assessment, and document classification. Although large unlabeled text corpora are abundant, labeled data for learning these specific tasks is scarce, making it challenging for discriminatively trained models to perform adequately. We demonstrate that large gains on these tasks can be realized by generative pre-training of a language model on a diverse corpus of unlabeled text, followed by discriminative fine-tuning on each specific task. In contrast to previous approaches, we make use of task-aware input transformations during fine-tuning to achieve effective transfer while requiring minimal changes to the model architecture. We demonstrate the effectiveness of our approach on a wide range of benchmarks for natural language understanding. Our general task-agnostic model outperforms discriminatively trained models that use architectures specifically crafted for each task, significantly improving upon the state of the art in 9 out of the 12 tasks studied. For instance, we achieve absolute improvements of 8.9% on commonsense reasoning (Stories Cloze Test), 5.7% on question answering (RACE), and 1.5% on textual entailment (MultiNLI).
Biologically motivated learning dynamics: parallel architectures and nonlinear Hebbian plasticity.
Michael Buice· Allen Institute
Wed, Apr 10 · 15:00 UTC
Learning in biological systems takes place in contexts and with dynamics not often accounted for by simple models. I will describe the learning dynamics of two model systems that incorporate either architectural or dynamic constraints from biological observations. In the first case, inspired by the observed mesoscopic structure of the mouse brain as revealed by the Allen Mouse Brain Connectivity Atlas, as well as multiple examples of parallel pathways in mammalian brains, I present a mathematical analysis of learning dynamics in networks that have parallel computational pathways driven by the same cost function. We use the approximation of deep linear networks with large hidden layer sizes to show that, as the depth of the parallel pathways increases, different features of the training set (defined by the singular values of the input-output correlation) will typically concentrate in one of the pathways. This result is derived analytically and demonstrated with numerical simulation with both linear and non-linear networks. Thus, rather than sharing stimulus and task features across multiple pathways, parallel network architectures learn to produce sharply diversified representations with specialized and specific pathways, a mechanism which may hold important consequences for codes in both biological and artificial systems. In the second case, I discuss learning dynamics in a generalization of Hebbian rules and show that these rules allow a neuron to learn tensor decompositions of higher-order input correlations. Unlike the case of the Oja rule and PCA, the resulting learned representation is not unique but selects amongst the tensor eigenvectors according to initial conditions. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2024-04-10. Recording duration: 00:30:39.
Computational NeuroscienceDynamical Systems+1 moreSeries: van Vreeswijk Theoretical Neuroscience SeminarVideo
March 2024
Large Language Models (LLMs) have recently demonstrated remarkable capabilities in natural language processing tasks and beyond. This success of LLMs has led to a large influx of research contributions in this direction. These works encompass diverse topics such as architectural innovations, better training strategies, context length improvements, fine-tuning, multi-modal LLMs, robotics, datasets, benchmarking, efficiency, and more. With the rapid development of techniques and regular breakthroughs in LLM research, it has become considerably challenging to perceive the bigger picture of the advances in this direction. Considering the rapidly emerging plethora of literature on LLMs, it is imperative that the research community is able to benefit from a concise yet comprehensive overview of the recent developments in this field. This article provides an overview of the existing literature on a broad range of LLM-related concepts. Our self-contained comprehensive overview of LLMs discusses relevant background concepts along with covering the advanced topics at the frontier of research in LLMs. This review article is intended to not only provide a systematic survey but also a quick comprehensive reference for the researchers and practitioners to draw insights from extensive informative summaries of the existing works to advance the LLM research.
January 2024
Discovering learning-induced changes in neural representations from large-scale neural data tensors
N Alex Cayco Gajic· Ecole normale supérieure, Paris
Wed, Jan 24 · 16:00 UTC
Learning induces changes in neural activity over slow timescales. These changes can be summarized by restructuring neural population data into a three-dimensional array or tensor, of size neurons by time points by trials. Classic dimensionality reduction methods often assume that neural representations are constrained to a fixed low-dimensional latent subspace. Consequently, this view does not capture how the latent subspace could evolve over learning, nor how high-dimensional neural activity could emerge over learning. Furthermore, the link between these empirically-observed changes in neural activity as a result of learning and circuit-level changes in recurrent dynamics is unclear. In this talk I will discuss our recent efforts towards developing dimensionality reduction and data-driven modeling methods based on tensors in order to identify how neural representations change over learning. First we introduce a new tensor decomposition, sliceTCA, which is able to disentangle latent variables of multiple covariability classes that are often mixed in neural population data. We demonstrate in three datasets that sliceTCA is able to capture more behaviorally-relevant information in neural data than previous methods. Second, to probe for how circuit-level changes in neural dynamics implement the observed changes in neural activity, we develop a data-driven RNN-based framework in which the recurrent connectivity is constrained to be low tensor rank. We demonstrate that such low tensor rank RNNs (ltrRNNs) are able to capture changes in neural geometry and dynamics in motor cortical data from a motor adaptation task. Together, both sliceTCA and ltrRNN demonstrate the utility of interpretable, tensor-based methods for discovery of learning-induced changes in neural representations directly from data. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2024-01-24. Recording duration: 00:45:54.
Matrix Factorization with Neural Networks
Marc Mézard· Bocconi University, Milano
Wed, Jan 3 · 16:00 UTC
The factorization of a large matrix into the product of two matrices is an important mathematical problem encountered in many tasks, ranging from dictionary learning to machine learning. Statistical physics can provide on the one hand theoretical limits on the possibility of factorizing matrices in the limit of infinite size, and also practical algorithms. While this program has been successful in the case of finite rank matrices, the regime of extensive rank (scaling linearly with the dimension of the matrix) turns out to be much harder. This talk will describe a new approach to matrix factorization that maps it to neural network models of associative memory: each pattern found in the associative memory corresponds to one factor of the matrix decomposition. A detailed theoretical analysis of this new approach shows that matrix factorization in the extensive rank regime is possible when the rank is below a certain threshold. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2024-01-03. Recording duration: 00:44:31.
Computational NeuroscienceMatrix Algebra+1 moreSeries: van Vreeswijk Theoretical Neuroscience SeminarVideo
November 2023
Mean Field Approaches to Learning Dynamics in Deep Networks
Blake Bordelon· Harvard University
Wed, Nov 29 · 16:00 UTC
Deep neural network learning dynamics are very complex with large numbers of learnable weights and many sources of disorder. In this talk, I will discuss mean field approaches to analyze the learning dynamics of neural networks in large system size limits when starting from random initial conditions. The result of this analysis is a dynamical mean field theory (DMFT) where all neurons obey independent stochastic single site dynamics. Correlation functions (kernels) and response functions for the features and gradients at each layer can be computed self-consistently from these stochastic processes. Depending on the choice of scaling of the network output, the network can operate in a kernel regime or a feature learning regime in the infinite width limit. I will discuss how this theory can be used to analyze various learning rules for deep architectures (backpropagation, feedback alignment based rules, Hebbian learning etc), where the weight updates do not necessarily correspond to gradient descent on an energy function. I will then present recent extensions of this theory to residual networks at infinite depth and discuss the utility of deriving scaling limits to obtain consistent optimal hyperparameters (such as learning rate) across widths and depths. Feature learning in other types of architectures will be discussed if time permits. Lastly, I will discuss open problems and challenges associated with this theoretical approach to neural network learning dynamics. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2023-11-29. Recording duration: 00:41:18.
Computational NeuroscienceDynamical Systems+1 moreSeries: van Vreeswijk Theoretical Neuroscience SeminarVideo
Mathematical and computational modelling of ocular hemodynamics: from theory to applications
Giovanna Guidoboni· University of Maine
Tue, Nov 14 · 13:00 UTC
Changes in ocular hemodynamics may be indicative of pathological conditions in the eye (e.g. glaucoma, age-related macular degeneration), but also elsewhere in the body (e.g. systemic hypertension, diabetes, neurodegenerative disorders). Thanks to its transparent fluids and structures that allow the light to go through, the eye offers a unique window on the circulation from large to small vessels, and from arteries to veins. Deciphering the causes that lead to changes in ocular hemodynamics in a specific individual could help prevent vision loss as well as aid in the diagnosis and management of diseases beyond the eye. In this talk, we will discuss how mathematical and computational modelling can help in this regard. We will focus on two main factors, namely blood pressure (BP), which drives the blood flow through the vessels, and intraocular pressure (IOP), which compresses the vessels and may impede the flow. Mechanism-driven models translates fundamental principles of physics and physiology into computable equations that allow for identification of cause-to-effect relationships among interplaying factors (e.g. BP, IOP, blood flow). While invaluable for causality, mechanism-driven models are often based on simplifying assumptions to make them tractable for analysis and simulation; however, this often brings into question their relevance beyond theoretical explorations. Data-driven models offer a natural remedy to address these short-comings. Data-driven methods may be supervised (based on labelled training data) or unsupervised (clustering and other data analytics) and they include models based on statistics, machine learning, deep learning and neural networks. Data-driven models naturally thrive on large datasets, making them scalable to a plethora of applications. While invaluable for scalability, data-driven models are often perceived as black- boxes, as their outcomes are difficult to explain in terms of fundamental principles of physics and physiology and this limits the delivery of actionable insights. The combination of mechanism-driven and data-driven models allows us to harness the advantages of both, as mechanism-driven models excel at interpretability but suffer from a lack of scalability, while data-driven models are excellent at scale but suffer in terms of generalizability and insights for hypothesis generation. This combined, integrative approach represents the pillar of the interdisciplinary approach to data science that will be discussed in this talk, with application to ocular hemodynamics and specific examples in glaucoma research.
Mathematical ModelingApplied Mathematics+3 moreSeries: Mathematical and Computational OphthalmologyVideo
Prediction Models for Brains and Machines
Kimberly Stachenfeld· Google Deep Mind
Wed, Nov 1 · 15:00 UTC
Humans and animals learn and plan with flexibility and efficiency well beyond that of modern Machine Learning methods. This is hypothesized to owe in part to the ability of animals to build structured representations of their environments, and modulate these representations to rapidly adapt to new settings. In the first part of this talk, I will discuss theoretical work describing how learned representations in hippocampus enable rapid adaptation to new goals by learning predictive representations. I will also cover work extending this account, in which we show how the predictive model can be adapted to the probabilistic setting to describe a broader array of generalization results in humans and animals, and how entorhinal representations can be modulated to support sample generation optimized for different behavioral states. I will also talk about work applying this perspective to the deep RL setting, where we can study the effect of predictive learning on representations that form in a deep neural network and how these results compare to neural data. In the second part of the talk, I will overview some of the ways in which we have combined many of the same mathematical concepts with state-of-the-art deep learning methods to improve efficiency and performance in machine learning applications like physical simulation, relational reasoning, and design. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2023-11-01. Recording duration: 00:50:44.
Computational NeuroscienceCognition+1 moreSeries: van Vreeswijk Theoretical Neuroscience SeminarVideo
October 2023
Multimodal units fuse-then-accumulate evidence across channels
Dan Goodman· Imperial college
Wed, Oct 25 · 15:00 UTC
We continuously detect sensory data, like sights and sounds, and use this information to guide our behaviour. However, rather than relying on single sensory channels, which are noisy and can be ambiguous alone, we merge information across our senses and leverage this combined signal. In biological networks, this process (multisensory integration) is implemented by multimodal neurons which are often thought to receive the information accumulated by unimodal areas, and to fuse this across channels; an algorithm we term accumulate-then-fuse. However, it remains an open question how well this theory generalises beyond the classical tasks used to test multimodal integration. Here, we explore this by developing novel multimodal tasks and deploying probabilistic, artificial and spiking neural network models. Using these models we demonstrate that multimodal units are not necessary for accuracy or balancing speed/accuracy in classical multimodal tasks, but are critical in a novel set of tasks in which we comodulate signals across channels. We show that these comodulation tasks require multimodal units to implement an alternative fuse-then-accumulate algorithm, which excels in naturalistic settings and is optimal for a wide class of multimodal problems. Finally, we link our findings to experimental results at multiple levels; from single neurons to behaviour. Ultimately, our work suggests that multimodal neurons may fuse-then-accumulate evidence across channels, and provides novel tasks and models for exploring this in biological systems. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2023-10-25. Recording duration: 00:30:13.
Computational NeuroscienceNeuroscience+1 moreSeries: van Vreeswijk Theoretical Neuroscience SeminarVideo
Learning with multimodal enrichment
Katharina von Kriegstein· Technical University Dresden
Thu, Oct 5 · 16:00 UTC
September 2023
Self as Processes (BACN Mid-career Prize Lecture 2023)
Jie Sui· University of Aberdeen, UK
Wed, Sep 13 · 15:15 UTC
An understanding of the self helps explain not only human thoughts, feelings, attitudes but also many aspects of everyday behaviour. This talk focuses on a viewpoint - self as processes. This viewpoint emphasizes the dynamics of the self that best connects with the development of the self over time and its realist orientation. We are combining psychological experiments and data mining to comprehend the stability and adaptability of the self across various populations. In this talk, I draw on evidence from experimental psychology, cognitive neuroscience, and machine learning approaches to demonstrate why and how self-association affects cognition and how it is modulated by various social experiences and situational factors
Foundation models in ophthalmology
Pearse Keane· University College London and Moorfields Eye Hospital NHS Foundation Trust
Wed, Sep 6 · 12:00 UTC
Abstract to follow.
June 2023
Reduced label complexity for tight linear regression
Alex Gittens· Rensselaer Polytechnic Institute
Thu, Jun 29 · 18:30 UTC · Providence, USA · In person
Alex Gittens studies how many data points must be labelled to fit a linear regression model with nearly the predictive power of a fully labelled dataset. Existing coreset and iterative approaches handle constant-factor approximations, but tighter approximations that improve with dataset size need different methods. The talk presents a polynomial-time algorithm that reduces label complexity by an additive O(sqrt(n)), using a sharp analysis of regression error for a coreset formed by backward selection.
Linear AlgebraComputational Mathematics+2 moreSeries: Institute for Computational and Experimental Research in Mathematics (ICERM), Brown UniversityVideo
Diverse applications of artificial intelligence and mathematical approaches in ophthalmology
Tiarnán Keenan· National Eye Institute (NEI)
Tue, Jun 6 · 14:00 UTC
Ophthalmology is ideally placed to benefit from recent advances in artificial intelligence. It is a highly image-based specialty and provides unique access to the microvascular circulation and the central nervous system. This talk will demonstrate diverse applications of machine learning and deep learning techniques in ophthalmology, including in age-related macular degeneration (AMD), the leading cause of blindness in industrialized countries, and cataract, the leading cause of blindness worldwide. This will include deep learning approaches to automated diagnosis, quantitative severity classification, and prognostic prediction of disease progression, both from images alone and accompanied by demographic and genetic information. The approaches discussed will include deep feature extraction, label transfer, and multi-modal, multi-task training. Cluster analysis, an unsupervised machine learning approach to data classification, will be demonstrated by its application to geographic atrophy in AMD, including exploration of genotype-phenotype relationships. Finally, mediation analysis will be discussed, with the aim of dissecting complex relationships between AMD disease features, genotype, and progression.
March 2023
Epilepsy surgery is a safe but underutilised treatment for drug-resistant focal epilepsy. One challenge in the presurgical evaluation of patients with drug-resistant epilepsy are patients considered “MRI negative”, i.e. where a structural brain abnormality has not been identified on MRI. A major pathology in “MRI negative” patients is focal cortical dysplasia (FCD), where lesions are often small or subtle and easily missed by visual inspection. In recent years, there has been an explosion in artificial intelligence (AI) research in the field of healthcare. Automated FCD detection is an area where the application of AI may translate into significant improvements in the presurgical evaluation of patients with focal epilepsy. I will provide an overview of our automated FCD detection work, the Multicentre Epilepsy Lesion Detection (MELD) project and how AI algorithms are beginning to be integrated into epilepsy presurgical planning at Great Ormond Street Hospital and elsewhere around the world. Finally, I will discuss the challenges and future work required to bring AI to the forefront of care for patients with epilepsy.
February 2023
Understanding Machine Learning via Exactly Solvable Statistical Physics Models
Lenka Zdeborová· EPFL
Wed, Feb 8 · 05:00 UTC
The affinity between statistical physics and machine learning has a long history. I will describe the main lines of this long-lasting friendship in the context of current theoretical challenges and open questions about deep learning. Theoretical physics often proceeds in terms of solvable synthetic models, I will describe the related line of work on solvable models of simple feed-forward neural networks. I will highlight a path forward to capture the subtle interplay between the structure of the data, the architecture of the network, and the optimization algorithms commonly used for learning.
January 2023
Geometry of concept learning
Haim Sompolinsky· The Hebrew University of Jerusalem and Harvard University
Wed, Jan 4 · 05:00 UTC
Understanding Human ability to learn novel concepts from just a few sensory experiences is a fundamental problem in cognitive neuroscience. I will describe a recent work with Ben Sorcher and Surya Ganguli (PNAS, October 2022) in which we propose a simple, biologically plausible, and mathematically tractable neural mechanism for few-shot learning of naturalistic concepts. We posit that the concepts that can be learned from few examples are defined by tightly circumscribed manifolds in the neural firing-rate space of higher-order sensory areas. Discrimination between novel concepts is performed by downstream neurons implementing ‘prototype’ decision rule, in which a test example is classified according to the nearest prototype constructed from the few training examples. We show that prototype few-shot learning achieves high few-shot learning accuracy on natural visual concepts using both macaque inferotemporal cortex representations and deep neural network (DNN) models of these representations. We develop a mathematical theory that links few-shot learning to the geometric properties of the neural concept manifolds and demonstrate its agreement with our numerical simulations across different DNNs as well as different layers. Intriguingly, we observe striking mismatches between the geometry of manifolds in intermediate stages of the primate visual pathway and in trained DNNs. Finally, we show that linguistic descriptors of visual concepts can be used to discriminate images belonging to novel concepts, without any prior visual experience of these concepts (a task known as ‘zero-shot’ learning), indicated a remarkable alignment of manifold representations of concepts in visual and language modalities. I will discuss ongoing effort to extend this work to other high level cognitive tasks.
December 2022
Can a single neuron solve MNIST? Neural computation of machine learning tasks emerges from the interaction of dendritic properties
Ilenna Jones· University of Pennsylvania
Wed, Dec 7 · 15:00 UTC
Physiological experiments have highlighted how the dendrites of biological neurons can nonlinearly process distributed synaptic inputs. However, it is unclear how qualitative aspects of a dendritic tree, such as its branched morphology, its repetition of presynaptic inputs, voltage-gated ion channels, electrical properties and complex synapses, determine neural computation beyond this apparent nonlinearity. While it has been speculated that the dendritic tree of a neuron can be seen as a multi-layer neural network and it has been shown that such an architecture could be computationally strong, we do not know if that computational strength is preserved under these qualitative biological constraints. Here we simulate multi-layer neural network models of dendritic computation with and without these constraints. We find that dendritic model performance on interesting machine learning tasks is not hurt by most of these constraints and may synergistically benefit from all of them combined. Our results suggest that single real dendritic trees may be able to learn a surprisingly broad range of tasks through the emergent capabilities afforded by their properties.
Connecting performance benefits on visual tasks to neural mechanisms using convolutional neural networks
Grace Lindsay· New York University (NYU)
Wed, Dec 7 · 05:00 UTC
Behavioral studies have demonstrated that certain task features reliably enhance classification performance for challenging visual stimuli. These include extended image presentation time and the valid cueing of attention. Here, I will show how convolutional neural networks can be used as a model of the visual system that connects neural activity changes with such performance changes. Specifically, I will discuss how different anatomical forms of recurrence can account for better classification of noisy and degraded images with extended processing time. I will then show how experimentally-observed neural activity changes associated with feature attention lead to observed performance changes on detection tasks. I will also discuss the implications these results have for how we identify the neural mechanisms and architectures important for behavior.