SQI Seminar Series: Tal Linzen, NYU
Tal Linzen, an associate professor of linguistics and data science at New York University and a research scientist at Google, will speak in MIT's Siegel Family Quest for Intelligence seminar series. His work combines behavioral experiments and computational methods to study language learning and comprehension, alongside large-language-model post-training, evaluation, and interpretability.
Young Investigator, Open Language Models for Biology
Ai2 is recruiting a postdoctoral Young Investigator for CellOLMo, a collaboration using open language models with single-cell brain data. The researcher will develop multimodal models that combine gene-expression data and text, evaluate biological language grounding on the SEA-AD Alzheimer's disease atlas, build donor- and brain-region-level representations, and release resulting models, code, data, and research artifacts openly.
Research Scientist, Asta
Ai2's Asta team is recruiting a research scientist to advance open agentic systems for scientific reasoning. The role spans language-model post-training, data generation and curation, control strategies, evaluation, literature synthesis, hypothesis formation, scientific coding, computational experiments, and data analysis in collaboration with human scientists.
Llama 3.1 Paper: The Llama Family of Models
Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models that natively support multilinguality, coding, reasoning, and tool usage. Our largest model is a dense Transformer with 405B parameters and a context window of up to 128K tokens. This paper presents an extensive empirical evaluation of Llama 3. We find that Llama 3 delivers comparable quality to leading language models such as GPT-4 on a plethora of tasks. We publicly release Llama 3, including pre-trained and post-trained versions of the 405B parameter language model and our Llama Guard 3 model for input and output safety. The paper also presents the results of experiments in which we integrate image, video, and speech capabilities into Llama 3 via a compositional approach. We observe this approach performs competitively with the state-of-the-art on image, video, and speech recognition tasks. The resulting models are not yet being broadly released as they are still under development.