Machine Learning seminars
September 2026
AI agents for therapeutic reasoning across biological contexts
Michelle M. Li· Carnegie Mellon University
Tue, Sep 1 · 14:30 UTC · Massachusetts, online recording
Michelle M. Li examines how computational analyses can preserve the biological context of a proposed treatment, including cell type, disease state, genetic background and patient characteristics. She introduces Medea, an AI system that combines biological software, predictive models and literature retrieval while checking intermediate steps and reconciling evidence. The seminar presents evaluations involving cell-specific target selection, cancer-cell synthetic lethality and immunotherapy response. A separate yeast experiment tests predictions against previously unpublished measurements of gene-pair interactions under DNA-damaging treatments. The research addresses whether an agent can transfer useful evidence between contexts while recognizing when that transfer is unsupported. Reported comparisons cover predictive performance, computational failures and the ability to abstain. This recording retains the original seminar date.
Computational BiologyArtificial Intelligence+1 moreSeries: Microsoft Research New England Generative Modeling & Sampling SeminarVideo
July 2026
Learning Genetic Perturbation Effects at Single-Cell Resolution for Virtual Cells
Jiaqi Zhang· MIT at the seminar; incoming Assistant Professor, Columbia University
Tue, Jul 14 · 14:30 UTC · Massachusetts, online recording
Jiaqi Zhang examines how computational models can learn the effects of genetic interventions from single-cell experiments. Such experiments reveal causal relationships, but their high-dimensional measurements are costly to collect and difficult to interpret. The seminar connects identifiable causal representations with a predictive method for previously unseen perturbations. The approach incorporates prior biological knowledge and changes in data distributions to estimate responses at individual-cell resolution. It also uses predictions to guide subsequent experiments. An application identifies and experimentally validates previously unknown T-cell regulators with potential relevance to cancer immunotherapy. The recording follows the original July seminar; the series lists Zhang at MIT, while the recording biography describes her incoming Columbia appointment.
Computational GenomicsGenetics+3 moreSeries: Microsoft Research New England Generative Modeling & Sampling SeminarVideo
June 2026
Geometry and Information in Precision Collider Physics
Benoit Assi· University of Cincinnati
Thu, Jun 25 · 15:00 UTC · Waterloo, Canada
Benoit Assi examines limits on reliable theoretical predictions for present and future particle colliders. Low-order simulations of QCD radiation carry large uncertainties, hadronization is commonly modeled rather than derived, and effective theories contain more operators than measurements can constrain. Information theory and machine learning can incorporate improved calculations into simulations, quantify uncertainty, choose informative observables and advance hadronization theory. Geometry of effective-theory field space combines infinite operator families into finite physical quantities, while information measures identify the combinations experiments can resolve. The talk develops these complementary approaches to extracting the available information from LHC and future-collider data.
March 2026
Unsupervised representation learning by amortised neural message-passing
Lior Fox· Gatsby Computational Neuroscience Unit
Wed, Mar 4 · 16:00 UTC
Useful internal representations should explain the patterns of regularities and dependencies among observations. Probabilistic graphical models promise a principled way to uncover latent factors as such, but they are hard to scale to handle high-dimensional sensory observations and complicated dependencies structures. Neural-networks, on the other hand, excel at approximating complicated high-dimensional functions, but their internal representations do not easily lend themselves to a probabilistic interpretation. Despite some successes, a general unified approach is still missing for integrating the two approaches. I will describe a novel approach towards merging adaptive neural-network components into a probabilistic framework, based on three core ideas. The first is to train a set of networks to collectively perform inference, leveraging the ability of pattern-recognition methods to amortise complicated transformations. The second is to constrain the way in which the outputs of these networks are interpreted, transformed, and combined together. These constraints, together with the learning objective itself, are derived directly from probabilistic considerations encoded in a graphical model. Finally, the third core idea is that of recognition-parametrisation, allowing the inference ("recognition") procedure to directly define the model itself, without requiring an explicit "generative" decoder. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2026-03-04. Recording duration: 00:48:26.
Computational NeuroscienceArtificial Intelligence+1 moreSeries: van Vreeswijk Theoretical Neuroscience SeminarVideo
February 2026
1. Can we reconstruct images that a person saw, directly from their fMRI brain recordings? 2. Can we reconstruct the training data that a deep-network trained on, directly from the parameters of the network? The answer to both of these intriguing questions is “Yes!” In this talk I will present some of our work in both domains. I will then show how combining the power of Brains and Machines can lead to significant breakthroughs in both areas, and potentially bridge the gap between Minds and Machines. Finally, I will show how combining the power of Multiple Brains (with NO shared data) may lead to new breakthrough discoveries in Brain-Science, and allow mapping of information between different brains. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2026-02-18. Recording duration: 00:50:24.
Computational NeuroscienceBrain Imaging+1 moreSeries: van Vreeswijk Theoretical Neuroscience SeminarVideo
The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm
Robert Gower· Flatiron Institute
Fri, Feb 6 · 15:30 UTC · Providence, USA · In person
Robert Gower introduces Polar Express for the polar decomposition and matrix sign function, motivated by Muon neural-network training. Using only matrix-matrix products makes the method suited to high-throughput GPUs. Each iteration adapts its polynomial update through minimax optimization, building on Chen and Chow and Nakatsukasa and Freund. Worst-case error minimization gives rapid initial and asymptotic convergence. The talk addresses finite-precision implementation in bfloat16 and reports improved validation loss when training GPT-2 on one billion FineWeb tokens across several learning rates.
Linear AlgebraDeep Learning+3 moreSeries: Institute for Computational and Experimental Research in Mathematics (ICERM), Brown UniversityVideo
Structured Matrix Learning from Matrix-Vector Products
Chris Musco· New York University
Wed, Feb 4 · 21:30 UTC · Providence, USA · In person
Chris Musco studies how to approximate an unknown matrix by a structured one using a limited, adaptively chosen sequence of matrix-vector products. This models operator learning in scientific machine learning as well as computational algorithms. Randomized SVD provides strong guarantees for low-rank targets; analogous results for sparse and hierarchical structures are less developed. The talk presents progress on efficient algorithms for these classes and a broader complexity theory. Joint work with Noah Amsel, Pratyush Avi, Tyler Chen, Prathamesh Dharangutte, Chinmay Hegde, Feyza Duman Keles, Diana Halikias, Cameron Musco, and David Persson.
Linear AlgebraComputational Mathematics+2 moreSeries: Institute for Computational and Experimental Research in Mathematics (ICERM), Brown UniversityVideo
Matrix-Mimetic Tensor Algebra: Optimal Decompositions and Equivariant Learning
Lior Horesh· IBM Research
Wed, Feb 4 · 16:30 UTC · Providence, USA · In person
Lior Horesh presents tensor-tensor algebra designed to retain key properties of matrix algebra while representing multidimensional correlations. An Eckart–Young-like tensor representation theorem underpins computationally feasible, provably optimal decompositions. Matrix-mimetic operations allow existing computational workflows to be adapted to tensors. Examples include tensorized neural-network structures and tensor graph convolutional networks for time-evolving graphs. The discussion concludes with tensor group symmetries and extensions to equivariant learning.
Linear AlgebraDeep Learning+2 moreSeries: Institute for Computational and Experimental Research in Mathematics (ICERM), Brown UniversityVideo
December 2025
Learning mechanistic models that link cells, circuits, and computations
Jakob Macke· Tubingen University
Wed, Dec 17 · 16:00 UTC
Modern experimental techniques now reveal the structure and function of neural circuits at unprecedented scale and resolution. How can we use this wealth of data to understand how cells and circuits implement computations underlying behaviour? Achieving this goal requires models that are consistent with biophysical mechanisms and circuit dynamics, yet flexible enough to capture behaviourally relevant computations. We develop simulation-based machine learning methods that address this challenge. I will show how these approaches—in combination with connectomic measurements—make it possible to build large-scale mechanistic models of the fruit fly visual system. Our methods generalize across systems and scales, defining a new way to study biological systems by algorithmically learning interpretable models that reveal how structure and dynamics gives rise to behaviour. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2025-12-17. Recording duration: 00:45:20.
November 2025
Latent-aligned generative models uncover shared structure in spontaneous whole-brain dynamics
Georges Debrégeas· CNRS, Paris
Wed, Nov 5 · 16:00 UTC
Assessing how brain activity generalizes across individuals is a central challenge in experimental neuroscience. Traditional task- or stimulus-driven approaches align data through trial averaging and anatomical registration, but these methods fail for spontaneous activity, where no shared temporal reference exists. In this talk, I will introduce a statistical framework, called latent-aligned Restricted Boltzmann Machines, to build a common representational space from whole-brain recordings of spontaneous activity in multiple zebrafish larvae. This shared latent space, composed of spatially localized co-activation motifs or cell assemblies, allows bidirectional mapping of brain states: activity patterns from one fish can be encoded and decoded into another. The translated activity patterns retain their original spatial structure and show high plausibility within the recipient brain. We further use this shared space to segment spontaneous activity into discrete brain states and we quantify their Markovian transition statistics. Remarkably, these state-to-state dynamics are stereotyped across individuals, suggesting that spontaneous activity reflects intrinsic computational priors of neural processing. Together, these results demonstrate how probabilistic generative modeling can bridge individual variability and reveal conserved organizational principles of vertebrate brains. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2025-11-05. Recording duration: 00:36:37.
October 2025
What is So Interesting About Reinforcement Learning?
Andrew Barto· University of Massachusetts Amherst
Wed, Oct 29 · 15:00 UTC
This talk aims to answer these questions along four dimensions. First is history. RL was the basis of AI long before the term AI was introduced in 1956. The first machine learning (ML) systems were based on RL even before digital computers existed. Despite notable early successes of ML based on RL, RL essentially disappeared from ML until relatively recently. A second reason for renewed interest in RL is the clarification of some misunderstandings that have been prevalent in the ML community. A third, and most important, reason for this resurgence is that new, or rediscovered, algorithms and connections to well developed mathematical and engineering methods have been worked out. Finally, a fourth reason for the renewed interest in RL is its strong links to animal reward systems, in particular, to the role that dopamine plays in motivation and learning. VVTNS Sixth Season Opening Lecture. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2025-10-29. Recording duration: 00:55:21.
Computational NeuroscienceNeuroscience+1 moreSeries: van Vreeswijk Theoretical Neuroscience SeminarVideo
May 2025
From neurons to Newtons: Brain evolution as a machine learning problem
Alexei Koulakov· Cold Spring Harbor Laboratory
Wed, May 21 · 15:00 UTC
We have entered a golden age of artificial intelligence research, driven mainly by the advances in the artificial neural networks over the last several decades. Applications of these techniques—to machine vision, speech recognition, autonomous vehicles, natural language, and many other domains—are coming so quickly that many observers predict that the long-elusive goal of “Artificial General Intelligence” (AGI) is within our grasp. However, we still cannot build a machine capable of building a nest, stalking prey, or loading a dishwasher. I will describe how evolution may have shaped the algorithms that the brain is using to solve some of these challenging problems. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2025-05-21. Recording duration: 00:44:20.
Computational NeuroscienceNeuroscience+1 moreSeries: van Vreeswijk Theoretical Neuroscience SeminarVideo
Energy efficient learning in neural networks
Mark van Rossum· University of Nottingham
Wed, May 7 · 15:00 UTC
The brain is one of the most energy intense organs. Some of this energyis used for neural information processing, however, fruitfly experiments have shown that also learning is metabolically costly. We will present estimates of this cost and introduce a general model of this cost, and compare it to costs in computers. Next, we turn to a supervised artificial network setting and explore a number of strategies that cansave energy need for plasticity. Either by modifying the objective function, by restricting plasticity, or by using less costly transient forms of plasticity. Finally, we will discuss adaptive strategies and possible relevance for biological learning. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2025-05-07. Recording duration: 00:34:58.
April 2025
Learning generative dynamical systems models from multi-modal and multi-animal neuro-data
Daniel Durstewitz· Central Institute of Mental Health, Mannheim
Wed, Apr 23 · 15:00 UTC
For decades dynamical systems theory played a pivotal role in theoretical and computational neuroscience, as it links biophysical and biochemical processes to neural computation. In fact, dynamical systems are computationally universal. Rather than hand-crafting computational theories of neural function based on dynamical systems, recent developments in scientific machine learning (ML) and AI suggest that we may be able to infer such dynamical-computational models directly from neurophysiological and behavioral observations. This is called dynamical systems reconstruction (DSR), the learning of generative surrogate models of the underlying dynamics, including its long-term temporal and geometrical properties, from time series data. In my talk I will cover recent ML/AI architectures, training algorithms, and validation procedures for DSR. I will discuss specifically how recent AI architectures for DSR can integrate neuroscience data from multiple modalities (like multiple single-unit recordings and behavioral choices), across diverse time scales, and across many different animals and task designs, into a joint DSR model. This provides first steps toward dynamical systems based AI foundation models for neuroscience. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2025-04-23. Recording duration: 00:53:24.
Computational NeuroscienceNeuroscience+1 moreSeries: van Vreeswijk Theoretical Neuroscience SeminarVideo
February 2025
Active learning of neural population dynamics
Matthew Golub· University of Washington
Wed, Feb 5 · 16:00 UTC
Recent advances in techniques for monitoring and perturbing neural populations have greatly enhanced our ability to study circuits in the brain. In particular, two-photon holographic optogenetics now enables precise photostimulation of experimenter-specified groups of individual neurons, while simultaneous two-photon calcium imaging enables the measurement of ongoing and induced activity across the neural population. Despite the enormous space of potential photostimulation patterns and the time-consuming nature of photostimulation experiments, very little algorithmic work has been done to determine the most effective photostimulation patterns for identifying the neural population dynamics. Here, I will discuss ongoing development of active learning techniques to efficiently select which neurons to stimulate such that the resulting neural responses will best inform a dynamical model of the neural population activity. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2025-02-05. Recording duration: 00:42:29.
Computational NeuroscienceDynamical Systems+1 moreSeries: van Vreeswijk Theoretical Neuroscience SeminarVideo
October 2024
On finding what you’re (not) looking for: prospects and challenges for AI-driven discovery
André Curtis Trudel· University of Cincinnati
Thu, Oct 10 · 14:00 UTC
Recent high-profile scientific achievements by machine learning (ML) and especially deep learning (DL) systems have reinvigorated interest in ML for automated scientific discovery (eg, Wang et al. 2023). Much of this work is motivated by the thought that DL methods might facilitate the efficient discovery of phenomena, hypotheses, or even models or theories more efficiently than traditional, theory-driven approaches to discovery. This talk considers some of the more specific obstacles to automated, DL-driven discovery in frontier science, focusing on gravitational-wave astrophysics (GWA) as a representative case study. In the first part of the talk, we argue that despite these efforts, prospects for DL-driven discovery in GWA remain uncertain. In the second part, we advocate a shift in focus towards the ways DL can be used to augment or enhance existing discovery methods, and the epistemic virtues and vices associated with these uses. We argue that the primary epistemic virtue of many such uses is to decrease opportunity costs associated with investigating puzzling or anomalous signals, and that the right framework for evaluating these uses comes from philosophical work on pursuitworthiness.
July 2024
Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models that natively support multilinguality, coding, reasoning, and tool usage. Our largest model is a dense Transformer with 405B parameters and a context window of up to 128K tokens. This paper presents an extensive empirical evaluation of Llama 3. We find that Llama 3 delivers comparable quality to leading language models such as GPT-4 on a plethora of tasks. We publicly release Llama 3, including pre-trained and post-trained versions of the 405B parameter language model and our Llama Guard 3 model for input and output safety. The paper also presents the results of experiments in which we integrate image, video, and speech capabilities into Llama 3 via a compositional approach. We observe this approach performs competitively with the state-of-the-art on image, video, and speech recognition tasks. The resulting models are not yet being broadly released as they are still under development.
June 2024
A Bi-metric Framework for Fast Similarity Search
Piotr Indyk· Massachusetts Institute of Technology
Fri, Jun 21 · 17:00 UTC · Berkeley, USA
Nearest-neighbor indexes usually rely on a single distance function, but accurate comparisons can be expensive. This talk proposes a bi-metric framework: a cheap proxy metric builds the index, while the query procedure uses a limited number of evaluations of both the proxy and an expensive ground-truth metric. The theory applies to DiskANN and Cover Tree. When the proxy approximates the ground-truth metric within a bounded factor, the resulting structure can achieve arbitrarily good approximation guarantees under the accurate metric. Experiments on text retrieval using models with very different computational costs show improved accuracy-efficiency tradeoffs on almost all MTEB datasets compared with alternatives such as reranking. Joint work with Haike Xu and Sandeep Silwal.
Artificial IntelligenceApplied Mathematics+2 moreSeries: Simons Institute for the Theory of ComputingVideo
Learning and prediction in artificial deep neural networks: scaling, data manifolds, and universality
Yasaman Bahri· Google DeepMind
Wed, Jun 19 · 15:00 UTC
Developing scientifically-grounded theories for representation learning and generalization in artificial deep neural networks remains a grand challenge of fundamental interest to theoretical neuroscience and machine learning. I will discuss our work on one facet of this challenge — namely understanding generalization or “scaling laws” in learned neural networks as a function of basic control variables. I’ll discuss a taxonomy we develop that classifies different regimes of scaling behavior. We identify regimes where generalization exhibits universal scaling behavior and others where it can be traced back to properties of the data and neural architecture. The theoretical analysis is enabled by leveraging exactly solvable models of deep neural networks that arise naturally in the limit of large hidden layers. Along the way, I’ll also discuss our work on these theoretical models, which have been a useful starting point for theoretical descriptions of neural network dynamics. Finally, I’ll discuss our findings connecting generalization in neural networks to properties of the learned data manifold. I’ll close by discussing future directions and new hypotheses that emerge from our findings Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2024-06-19. Recording duration: 00:46:53.
Computational NeuroscienceMathematical Modeling+1 moreSeries: van Vreeswijk Theoretical Neuroscience SeminarVideo
Using ML tools in neuroscience to define optimality in complex natural behavior
Stephanie Palmer· University of Chicago
Wed, Jun 5 · 15:00 UTC
Biological systems must selectively encode partial information about the environment, as dictated by the capacity constraints at work in all living organisms. For example, we cannot see every feature of the light field that reaches our eyes; temporal resolution is limited by transmission noise and delays, and spatial resolution is limited by the finite number of photoreceptors and output cells in the retina. Classical efficient coding theory describes how sensory systems can maximize information transmission given such capacity constraints, but it treats all input features equally. Not all inputs are, however, of equal value to the organism. Our work quantifies whether and how the brain selectively encodes stimulus features, specifically predictive features, that are most useful for fast and effective movements. We have shown that efficient predictive computation starts at the earliest stages of the visual system in the retina. We borrow techniques from machine learning, statistical physics, and information theory to assess how we get terrific, predictive vision from these imperfect (lagged and noisy) component parts. In broader terms, we aim to build a more complete theory of efficient encoding in the brain, and along the way have found some intriguing connections between approaches to coarse graining in biology, machine learning, and physics. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2024-06-05. Recording duration: 00:41:40.
Computational NeuroscienceNeuroscience+1 moreSeries: van Vreeswijk Theoretical Neuroscience SeminarVideo