On the implicit bias of SGD in deep learning
Tel Aviv University
Recording
Event Information
Recording
Available
Host
van Vreeswijk TNS
Duration
70 minutes
Abstract
Tali's work emphasized the tradeoff between compression and information preservation. In this talk I will explore this theme in the context of deep learning. Artificial neural networks have recently revolutionized the field of machine learning. However, we still do not have sufficient theoretical understanding of how such models can be successfully learned. Two specific questions in this context are: how can neural nets be learned despite the non-convexity of the learning problem, and how can they generalize well despite often having more parameters than training data. I will describe our recent work showing that gradient-descent optimization indeed leads to 'simpler' models, where simplicity is captured by lower weight norm and in some cases clustering of weight vectors. We demonstrate this for several teacher and student architectures, including learning linear teachers with ReLU networks, learning boolean functions and learning convolutional pattern detection architectures.
Topics
Related Job Opportunities
PhD Studentship: Mitochondrial Metabolism and Novel Therapeutic Strategies for Metabolic Dysfunction-Associated Steatotic Liver Disease (MASLD) (Fixed Term)
Supervisors: Professor Andrew Murray, Department of Physiology, Development and Neuroscience, University of Cambridge Dr Ross Lindsay, Novo Nordisk Funding: Fully funded PhD studentship (Home/UK…
Research Associate (Fixed Term)
We seek a highly motivated Postdoctoral Research Associate to join the laboratory of Professor Kathy Niakan. We are based in the Loke Centre for Trophoblast Research (LCTR), in the Department of…
Research Assistant/Associate (Fixed Term)
Applications are invited for a postdoctoral research associate position to study the neural mechanisms of visual learning in mice, in the laboratories of Professor Ole Paulsen and Dr Jasper Poort…