Skip to content
SeminarStarts in 9 daysMachine Learning

Large-Scale Multilingual Evaluation of Large Language Models on Real-World Clinical Data

Jie Yang

Harvard Medical School

Hosted by Broad Institute — Machine Learning for Health (ML4H) Seminar Series

· 60 minutes
Online

Abstract

Jie Yang presents BRIDGE, a multilingual evaluation benchmark built from real clinical tasks and more than one million samples derived from electronic health records. The benchmark addresses the gap between simplified examination-style tests and the complexity of clinical data, while enabling comparison of rapidly changing language models. Evaluation of 95 models involved over 24,000 experiments and 39 million predictions. Results vary substantially by model family, task and language, with some open-source systems matching proprietary models. The study also finds that chain-of-thought prompting frequently reduces accuracy on these tasks and examines stigmatizing language produced during model reasoning at large scale.

Topics

We use essential cookies to run the site. Optional analytics and public-page session replay help us improve World Wide. Learn more.