Learning and prediction in artificial deep neural networks: scaling, data manifolds, and universality
Google DeepMind
Hosted by van Vreeswijk Theoretical Neuroscience Seminar
Recording
Abstract
Developing scientifically-grounded theories for representation learning and generalization in artificial deep neural networks remains a grand challenge of fundamental interest to theoretical neuroscience and machine learning. I will discuss our work on one facet of this challenge — namely understanding generalization or “scaling laws” in learned neural networks as a function of basic control variables. I’ll discuss a taxonomy we develop that classifies different regimes of scaling behavior. We identify regimes where generalization exhibits universal scaling behavior and others where it can be traced back to properties of the data and neural architecture. The theoretical analysis is enabled by leveraging exactly solvable models of deep neural networks that arise naturally in the limit of large hidden layers. Along the way, I’ll also discuss our work on these theoretical models, which have been a useful starting point for theoretical descriptions of neural network dynamics. Finally, I’ll discuss our findings connecting generalization in neural networks to properties of the learned data manifold. I’ll close by discussing future directions and new hypotheses that emerge from our findings