Topic: Speech Recognition

Seminar
2 seminars
SeminarComputational NeuroscienceRecording

From neurons to Newtons: Brain evolution as a machine learning problem

Alexei Koulakov
Cold Spring Harbor Laboratory
May 21, 2025

We have entered a golden age of artificial intelligence research, driven mainly by the advances in the artificial neural networks over the last several decades. Applications of these techniques—to machine vision, speech recognition, autonomous vehicles, natural language, and many other domains—are coming so quickly that many observers predict that the long-elusive goal of “Artificial General Intelligence” (AGI) is within our grasp. However, we still cannot build a machine capable of building a nest, stalking prey, or loading a dishwasher. I will describe how evolution may have shaped the algorithms that the brain is using to solve some of these challenging problems. Presented in the van Vreeswijk Theoretical Neuroscience Seminar series (formerly WWTNS) on 2025-05-21. Recording duration: 00:44:20.

SeminarNeuroscienceRecording

Encoding and perceiving the texture of sounds: auditory midbrain codes for recognizing and categorizing auditory texture and for listening in noise

Monty Escabi
University of Connecticut
Oct 1, 2021

Natural soundscapes such as from a forest, a busy restaurant, or a busy intersection are generally composed of a cacophony of sounds that the brain needs to interpret either independently or collectively. In certain instances sounds - such as from moving cars, sirens, and people talking - are perceived in unison and are recognized collectively as single sound (e.g., city noise). In other instances, such as for the cocktail party problem, multiple sounds compete for attention so that the surrounding background noise (e.g., speech babble) interferes with the perception of a single sound source (e.g., a single talker). I will describe results from my lab on the perception and neural representation of auditory textures. Textures, such as a from a babbling brook, restaurant noise, or speech babble are stationary sounds consisting of multiple independent sound sources that can be quantitatively defined by summary statistics of an auditory model (McDermott & Simoncelli 2011). How and where in the auditory system are summary statistics represented and the neural codes that potentially contribute towards their perception, however, are largely unknown. Using high-density multi-channel recordings from the auditory midbrain of unanesthetized rabbits and complementary perceptual studies on human listeners, I will first describe neural and perceptual strategies for encoding and perceiving auditory textures. I will demonstrate how distinct statistics of sounds, including the sound spectrum and high-order statistics related to the temporal and spectral correlation structure of sounds, contribute to texture perception and are reflected in neural activity. Using decoding methods I will then demonstrate how various low and high-order neural response statistics can differentially contribute towards a variety of auditory tasks including texture recognition, discrimination, and categorization. Finally, I will show examples from our recent studies on how high-order sound statistics and accompanying neural activity underlie difficulties for recognizing speech in background noise.

We use essential cookies to run the site. Optional analytics and public-page session replay help us improve World Wide. Learn more.