Skip to content

Broad Institute — Machine Learning for Health (ML4H) Seminar Series

Seminars and recordings

October 2026

Large-Scale Multilingual Evaluation of Large Language Models on Real-World Clinical Data

Jie Yang· Harvard Medical School

Starts today

Wed, Oct 7 · 18:00 UTC · Online

Jie Yang presents BRIDGE, a multilingual evaluation benchmark built from real clinical tasks and more than one million samples derived from electronic health records. The benchmark addresses the gap between simplified examination-style tests and the complexity of clinical data, while enabling comparison of rapidly changing language models. Evaluation of 95 models involved over 24,000 experiments and 39 million predictions. Results vary substantially by model family, task and language, with some open-source systems matching proprietary models. The study also finds that chain-of-thought prompting frequently reduces accuracy on these tasks and examines stigmatizing language produced during model reasoning at large scale.

Machine LearningInformatics+3 more
End of results.

We use cookies for analytics.