The Limits of Causal Reasoning in Human and Machine Learning
A key purpose of causal reasoning by individuals and by collectives is to enhance action, to give humans yet more control over their environment. As a result, causal reasoning serves as the infrastructure of both thought and discourse. Humans represent causal systems accurately in some ways, but also show some systematic biases (we tend to neglect causal pathways other than the one we are thinking about). Even when accurate, people’s understanding of causal systems tends to be superficial; we depend on our communities for most of our causal knowledge and reasoning. Nevertheless, we are better causal reasoners than machines. Modern machine learners do not come close to matching human abilities.
Exploration beyond bandits
Machine learning researchers frequently focus on human-level performance, in particular in games. However, in these applications human (or human-level) behavior is commonly reduced to a simple dot on a performance graph. Cognitive science, in particular theories of learning and decision making, could hold the key to unlock what is behind this dot, thereby gaining further insights into human cognition and the design principles of intelligent algorithms. However, cognitive experiments commonly focus on relatively simple paradigms such as restricted multi-armed bandit tasks. In this talk, I will argue that cognitive science can turn its lens to more complex scenarios to study exploration in real-world domains and online games. I will show in one large data set of online food delivery orders and across many online games how current cognitive theories of learning and exploration can describe human behavior in the wild, but also how these tasks demand us to expand our theoretical toolkit to describe a rich repertoire of real-world behaviors such as empowerment and fun.
Multitask performance humans and deep neural networks
Humans and other primates exhibit rich and versatile behaviour, switching nimbly between tasks as the environmental context requires. I will discuss the neural coding patterns that make this possible in humans and deep networks. First, using deep network simulations, I will characterise two distinct solutions to task acquisition (“lazy” and “rich” learning) which trade off learning speed for robustness, and depend on the initial weights scale and network sparsity. I will chart the predictions of these two schemes for a context-dependent decision-making task, showing that the rich solution is to project task representations onto orthogonal planes on a low-dimensional embedding space. Using behavioural testing and functional neuroimaging in humans, we observe BOLD signals in human prefrontal cortex whose dimensionality and neural geometry are consistent with the rich learning regime. Next, I will discuss the problem of continual learning, showing that behaviourally, humans (unlike vanilla neural networks) learn more effectively when conditions are blocked than interleaved. I will show how this counterintuitive pattern of behaviour can be recreated in neural networks by assuming that information is normalised and temporally clustered (via Hebbian learning) alongside supervised training. Together, this work offers a picture of how humans learn to partition knowledge in the service of structured behaviour, and offers a roadmap for building neural networks that adopt similar principles in the service of multitask learning. This is work with Andrew Saxe, Timo Flesch, David Nagy, and others.
The Structural Anchoring of Spontaneous Analogies
It is generally acknowledged that analogy is a core mechanism of human cognition, but paradoxically, analogies based on structural similarities would rarely be implemented spontaneously (e.g. without an explicit invitation to compare two representations). The scarcity of deep spontaneous analogies is at odds with the demonstration that familiar concepts from our daily-life are spontaneously used to encode the structure of our experiences. Based on this idea, we will present experimental works highlighting the predominant role of structural similarities in analogical retrieval. The educational stakes lurking behind the tendency to encode the problem’s structures through familiar concepts will also be addressed.