Large-Scale Multilingual Evaluation of Large Language Models on Real-World Clinical Data
Jie Yang · Harvard Medical School
Wed, Oct 7, 2026 · 18:00 UTC
Jie Yang presents BRIDGE, a multilingual evaluation benchmark built from real clinical tasks and more than one million samples derived from electronic health records. The benchmark addresses the gap between simplified examination-style tests and the complexity of clinical data, while enabling comparison of rapidly changing language models. Evaluation of 95 models involved over 24,000 experiments and 39 million predictions. Results vary substantially by model family, task and language, with some open-source systems matching proprietary models. The study also finds that chain-of-thought promptin