Topic: Large hidden layers

Seminar
1 seminar

Domains featuring this topic

Explore the domains where this topic appears.

SeminarComputational NeuroscienceRecording

Learning and prediction in artificial deep neural networks: scaling, data manifolds, and universality

Yasaman Bahri
Google DeepMind
Jun 19, 2024

Developing scientifically-grounded theories for representation learning and generalization in artificial deep neural networks remains a grand challenge of fundamental interest to theoretical neuroscience and machine learning. I will discuss our work on one facet of this challenge — namely understanding generalization or “scaling laws” in learned neural networks as a function of basic control variables. I’ll discuss a taxonomy we develop that classifies different regimes of scaling behavior. We identify regimes where generalization exhibits universal scaling behavior and others where it can be traced back to properties of the data and neural architecture. The theoretical analysis is enabled by leveraging exactly solvable models of deep neural networks that arise naturally in the limit of large hidden layers. Along the way, I’ll also discuss our work on these theoretical models, which have been a useful starting point for theoretical descriptions of neural network dynamics. Finally, I’ll discuss our findings connecting generalization in neural networks to properties of the learned data manifold. I’ll close by discussing future directions and new hypotheses that emerge from our findings Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2024-06-19. Recording duration: 00:46:53.

We use essential cookies to run the site. Analytics cookies are optional and help us improve World Wide. Learn more.