Big Boxes, Not Black Boxes: What we can compute about LLMs, and what it may say about AGI
Computer Science seminar by Zohar Ringel, Hebrew University of Jerusalem
Hosted by Perimeter Institute for Theoretical Physics
Wednesday 14:00–15:30 Toronto (GMT-4)
Starts in 3 days
Abstract
Deep networks are often thought of as black boxes. Their ability to encompass vast swathes of knowledge indeed makes them hard to explain. Yet many of their behaviours — generalization under overparametrization, grokking, OOD failures, neural scaling laws — recur across architectures and scales, and each, however surprising, can be reproduced and explained in controlled settings. I will review these efforts to identify and explain the universal phenomena of deep learning, and suggest that an LLM may amount to a sum of such tractable sub-phenomena, interpolative in nature. Finally, leaving scientific rigor aside, I'll argue that what separates this prosaic picture from the apparent magic of LLMs may well be the industrial
scale of compute and human labour behind it, and that AGI in its deeper extrapolative sense may be much further away than claimed.
Topics
Related seminars
Learning and prediction in artificial deep neural networks: scaling, data manifolds, and universality
More on generalization
Tuts, a Talk and AGI !!
Related research
Can a single neuron solve MNIST? Neural computation of machine learning tasks emerges from the interaction of dendritic properties
Related research