On the implicit bias of SGD in deep learning
Tali's work emphasized the tradeoff between compression and information preservation. In this talk I will explore this theme in the context of deep learning. Artificial neural networks have recently revolutionized the field of machine learning. However, we still do not have sufficient theoretical understanding of how such models can be successfully learned. Two specific questions in this context are: how can neural nets be learned despite the non-convexity of the learning problem, and how can they generalize well despite often having more parameters than training data. I will describe our recent work showing that gradient-descent optimization indeed leads to 'simpler' models, where simplicity is captured by lower weight norm and in some cases clustering of weight vectors. We demonstrate this for several teacher and student architectures, including learning linear teachers with ReLU networks, learning boolean functions and learning convolutional pattern detection architectures.
Appearance-based impression formation
Despite the common advice “not to judge a book by its cover”, we form impressions of character within a second of seeing a stranger’s face. These impressions have widespread consequences for society and for the economy, making it vital that we have a clear theoretical understanding of which impressions are important and how they are formed. In my talk, I outline a data-driven approach to answering these questions, starting by building models of the key dimensions underlying impressions of naturalistic face images. Overall, my findings suggest deeper links between the fields of face perception and social stereotyping than have previously been recognised.