Natural Language Processing seminars
March 2025
Using Machine Learning and Digital Technology to Identify Challenges and Improve Outcomes for Labor Market Transitions
Susan Athey, Bentley MacLeod, Suresh Naidu, Joseph Stiglitz· Stanford University
Mon, Mar 10 · 22:00 UTC · New York, United States
Susan Athey examines how machine learning and digital interventions can help explain and improve workers’ transitions between jobs. One project uses Swedish administrative data to identify groups whose earnings and employment are less resilient after layoffs, revealing substantial differences among workers within the same firms and labour markets. A second line of work uses transformer models and large language models to represent careers and analyse gender wage gaps, identifying settings in which large unexplained differences persist. Two further projects develop and evaluate digital interventions intended to help disadvantaged workers enter expanding occupations in information technology and data science. The lecture connects new methods for measuring labour-market disadvantage with evidence about practical interventions.
EconomicsArtificial IntelligenceSeries: Columbia University — Program for Economic Research and Center for Political EconomyVideo+1 more
December 2024
Towards open meta-research in neuroimaging
Kendra Oudyk· ORIGAMI - Neural data science - https://neurodatascience.github.io/
Mon, Dec 9 · 05:00 UTC · Online
When meta-research (research on research) makes an observation or points out a problem (such as a flaw in methodology), the project should be repeated later to determine whether the problem remains. For this we need meta-research that is reproducible and updatable, or living meta-research. In this talk, we introduce the concept of living meta-research, examine prequels to this idea, and point towards standards and technologies that could assist researchers in doing living meta-research. We introduce technologies like natural language processing, which can help with automation of meta-research, which in turn will make the research easier to reproduce/update. Further, we showcase our open-source litmining ecosystem, which includes pubget (for downloading full-text journal articles), labelbuddy (for manually extracting information), and pubextract (for automatically extracting information). With these tools, you can simplify the tedious data collection and information extraction steps in meta-research, and then focus on analyzing the text. We will then describe some living meta-research projects to illustrate the use of these tools. For example, we’ll show how we used GPT along with our tools to extract information about study participants. Essentially, this talk will introduce you to the concept of meta-research, some tools for doing meta-research, and some examples. Particularly, we want you to take away the fact that there are many interesting open questions in meta-research, and you can easily learn the tools to answer them. Check out our tools at https://litmining.github.io/
July 2024
Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models that natively support multilinguality, coding, reasoning, and tool usage. Our largest model is a dense Transformer with 405B parameters and a context window of up to 128K tokens. This paper presents an extensive empirical evaluation of Llama 3. We find that Llama 3 delivers comparable quality to leading language models such as GPT-4 on a plethora of tasks. We publicly release Llama 3, including pre-trained and post-trained versions of the 405B parameter language model and our Llama Guard 3 model for input and output safety. The paper also presents the results of experiments in which we integrate image, video, and speech capabilities into Llama 3 via a compositional approach. We observe this approach performs competitively with the state-of-the-art on image, video, and speech recognition tasks. The resulting models are not yet being broadly released as they are still under development.
April 2024
Improving Language Understanding by Generative Pre Training
Tue, Apr 23 · 11:00 UTC · Online
Natural language understanding comprises a wide range of diverse tasks such as textual entailment, question answering, semantic similarity assessment, and document classification. Although large unlabeled text corpora are abundant, labeled data for learning these specific tasks is scarce, making it challenging for discriminatively trained models to perform adequately. We demonstrate that large gains on these tasks can be realized by generative pre-training of a language model on a diverse corpus of unlabeled text, followed by discriminative fine-tuning on each specific task. In contrast to previous approaches, we make use of task-aware input transformations during fine-tuning to achieve effective transfer while requiring minimal changes to the model architecture. We demonstrate the effectiveness of our approach on a wide range of benchmarks for natural language understanding. Our general task-agnostic model outperforms discriminatively trained models that use architectures specifically crafted for each task, significantly improving upon the state of the art in 9 out of the 12 tasks studied. For instance, we achieve absolute improvements of 8.9% on commonsense reasoning (Stories Cloze Test), 5.7% on question answering (RACE), and 1.5% on textual entailment (MultiNLI).
March 2024
Large Language Models (LLMs) have recently demonstrated remarkable capabilities in natural language processing tasks and beyond. This success of LLMs has led to a large influx of research contributions in this direction. These works encompass diverse topics such as architectural innovations, better training strategies, context length improvements, fine-tuning, multi-modal LLMs, robotics, datasets, benchmarking, efficiency, and more. With the rapid development of techniques and regular breakthroughs in LLM research, it has become considerably challenging to perceive the bigger picture of the advances in this direction. Considering the rapidly emerging plethora of literature on LLMs, it is imperative that the research community is able to benefit from a concise yet comprehensive overview of the recent developments in this field. This article provides an overview of the existing literature on a broad range of LLM-related concepts. Our self-contained comprehensive overview of LLMs discusses relevant background concepts along with covering the advanced topics at the frontier of research in LLMs. This review article is intended to not only provide a systematic survey but also a quick comprehensive reference for the researchers and practitioners to draw insights from extensive informative summaries of the existing works to advance the LLM research.
November 2022
Do large language models solve verbal analogies like children do?
Claire Stevenson· University of Amsterdam
Thu, Nov 17 · 04:00 UTC
Analogical reasoning –learning about new things by relating it to previous knowledge– lies at the heart of human intelligence and creativity and forms the core of educational practice. Children start creating and using analogies early on, making incredible progress moving from associative processes to successful analogical reasoning. For example, if we ask a four-year-old “Horse belongs to stable like chicken belongs to …?” they may use association and reply “egg”, whereas older children will likely give the intended relational response “chicken coop” (or other term to refer to a chicken’s home). Interestingly, despite state-of-the-art AI-language models having superhuman encyclopedic knowledge and superior memory and computational power, our pilot studies show that these large language models often make mistakes providing associative rather than relational responses to verbal analogies. For example, when we asked four- to eight-year-olds to solve the analogy “body is to feet as tree is to …?” they responded “roots” without hesitation, but large language models tend to provide more associative responses such as “leaves”. In this study we examine the similarities and differences between children's and six large language models' (Dutch/multilingual models: RobBERT, BERT-je, M-BERT, GPT-2, M-GPT, Word2Vec and Fasttext) responses to verbal analogies extracted from an online adaptive learning environment, where >14,000 7-12 year-olds from the Netherlands solved 20 or more items from a database of 900 Dutch language verbal analogies.
March 2022
Understanding Natural Language: Insights From Cognitive Science, Cognitive Neuroscience, and Artificial Intelligence
James McClelland· Stanford University
Thu, Mar 17 · 21:00 UTC
November 2021
Language, Cognition, Biology
Cedric Boeckx· Catalan Institute for Advanced Studies (ICREA)
Tue, Nov 16 · 23:00 UTC
July 2021
Probabilistic Analogical Mapping with Semantic Relation Networks
Hongjing Lu· UCLA
Thu, Jul 1 · 16:00 UTC
Hongjing Lu will present a new computational model of Probabilistic Analogical Mapping (PAM, in collaboration with Nick Ichien and Keith Holyoak) that finds systematic correspondences between inputs generated by machine learning. The model adopts a Bayesian framework for probabilistic graph matching, operating on semantic relation networks constructed from distributed representations of individual concepts (word embeddings created by Word2vec) and of relations between concepts (created by our BART model). We have used PAM to simulate a broad range of phenomena involving analogical mapping by both adults and children. Our approach demonstrates that human-like analogical mapping can emerge from comparison mechanisms applied to rich semantic representations of individual concepts and relations. More details can be found https://arxiv.org/ftp/arxiv/papers/2103/2103.16704.pdf
February 2021
Kamala Harris and the Construction of Complex Ethnolinguistic Political Identity
Nicole Holliday· University of Pennsylvania
Fri, Feb 26 · 06:00 UTC
Over the past 50 years, sociolinguistic studies on black Americans have expanded in both theoretical and technical scope, and newer research has moved beyond seeing speakers, especially black speakers, as a monolithic sociolinguistic community (Wolfram 2007, Blake 2014). Yet there remains a dearth of critical work on complex identities existing within black American communities as well as how these identities are reflected and perceived in linguistic practice. At the same time, linguists have begun to take greater interest in the ways in which public figures, such as politicians, may illuminate the wider social meaning of specific linguistic variables. In this talk, I will present results from analyses of multiple aspects of ethnolinguistic variation in the speech of Vice President Kamala Harris during the 2019-2020 Democratic Party Primary debates. Together, these results show how VP Harris expertly employs both enregistered and subtle linguistic variables, including aspects of African American Language morphosyntax, vowels, and intonational phonology in the construction and performance of a highly specific sociolinguistic identity that reflects her unique positions politically, socially, and racially. The results of this study expand our knowledge about how the complexities of speaker identity are reflected in sociolinguistic variation, as well as press on the boundaries of what we know about how speakers in the public sphere use variation to reflect both who they are and who we want them to be.
December 2020
Machine Learning as a tool for positive impact : case studies from climate change
Alexandra (Sasha) Luccioni· University of Montreal and Mila (Quebec Institute for Learning Algorithms)
Thu, Dec 10 · 15:00 UTC
Climate change is one of our generation's greatest challenges, with increasingly severe consequences on global ecosystems and populations. Machine Learning has the potential to address many important challenges in climate change, from both mitigation (reducing its extent) and adaptation (preparing for unavoidable consequences) aspects. To present the extent of these opportunities, I will describe some of the projects that I am involved in, spanning from generative model to computer vision and natural language processing. There are many opportunities for fundamental innovation in this field, advancing the state-of-the-art in Machine Learning while ensuring that this fundamental progress translates into positive real-world impact.
October 2020
Abstract semantic relations (e.g., category membership, part-whole, antonymy, cause-effect) are central to human intelligence, underlying the distinctively human ability to reason by analogy. I will describe a computational project (Bayesian Analogy with Relational Transformations) that aims to extract explicit representations of abstract semantic relations from non-relational inputs automatically generated by machine learning. BART’s representations predict patterns of typicality and similarity for semantic relations, as well as similarity of neural signals triggered by semantic relations during analogical reasoning. In this approach, analogy emerges from the ability to learn and compare relations; mapping emerges later from the ability to compare patterns of relations.
July 2020
Predicting Patterns of Similarity Among Abstract Semantic Relations
Nick Ichien· UCLA
Thu, Jul 9 · 16:00 UTC
In this talk, I will present some data showing that people’s similarity judgments among word pairs reflect distinctions between abstract semantic relations like contrast, cause-effect, or part-whole. Further, the extent that individual participants’ similarity judgments discriminate between abstract semantic relations was linearly associated with both fluid and crystallized verbal intelligence, albeit more strongly with fluid intelligence. Finally, I will compare three models according to their ability to predict these similarity judgments. All models take as input vector representations of individual word meanings, but they differ in their representation of relations: one model does not represent relations at all, a second model represents relations implicitly, and a third model represents relations explicitly. Across the three models, the third model served as the best predictor of human similarity judgments suggesting the importance of explicit relation representation to fully account for human semantic cognition.
End of results.