Why AI Doctors Struggle Across Languages: A Polish Medical Study Reveals a Critical Gap
Large language models (LLMs) are increasingly used to help with medical problem-solving, but most research focuses on English-language contexts, leaving a blind spot for how these AI systems perform in other languages and healthcare systems. A new study from researchers at DeepSense introduces the first benchmark dataset based on Polish medical licensing exams, revealing significant performance gaps when AI models encounter non-English medical knowledge and cross-lingual translation challenges.
What Does This Polish Medical Benchmark Actually Test?
The research team created a novel dataset using real Polish medical licensing and specialization exams, including questions from the LEK (general medical licensing exam), LDEK (specialist exam), and PES (postgraduate specialization exam). The dataset also includes a subset of professionally translated Polish-English medical questions, allowing researchers to evaluate how well AI models handle translation between languages while maintaining medical accuracy.
The benchmark includes state-of-the-art LLMs spanning three categories: general-purpose models like GPT-4o and Gemini, domain-specific medical models, and Polish-specific language models. Researchers then compared AI performance against human medical students and practicing doctors to establish a realistic performance baseline.
How Do Current AI Models Perform on Medical Knowledge Across Languages?
The findings paint a mixed picture. Models like GPT-4o achieve near-human performance on medical exams, suggesting that cutting-edge AI systems have absorbed substantial medical knowledge during training. However, this strong performance largely holds true for English-language medical content. When the same models encounter Polish medical questions or must translate medical concepts between languages, their accuracy drops noticeably.
The study highlights two persistent challenges that undermine AI reliability in global healthcare contexts:
- Cross-Lingual Translation Gaps: Even when AI models understand medical concepts in English, translating those concepts accurately into Polish while preserving clinical precision proves difficult, potentially leading to misdiagnosis or treatment errors.
- Domain-Specific Understanding Disparities: Different medical specialties and healthcare systems use terminology and diagnostic approaches that vary by region; AI models trained primarily on English-language medical data struggle to adapt to these local variations.
- Language-Specific Model Limitations: Polish-specific language models, while designed to handle the language better, often lack the depth of medical training data available to larger English-trained models, creating a trade-off between linguistic accuracy and medical knowledge.
Why Should Healthcare Systems Care About This Research?
The implications are significant for any healthcare organization considering deploying AI systems globally. As hospitals and clinics in non-English-speaking countries explore AI tools for diagnostic support, clinical decision-making, and patient education, this research reveals that current models may not perform reliably outside their primary training language. The study emphasizes critical ethical considerations: deploying an AI system that performs well in English but struggles in Polish could create a false sense of security, potentially putting patients at risk.
The researchers note that these findings highlight broader disparities in model performance across languages and medical specialties. This disparity reflects a fundamental imbalance in AI training data; most large language models are trained predominantly on English-language text, including medical literature, clinical notes, and exam materials. Non-English medical knowledge remains underrepresented, creating a performance cliff when models encounter it.
Steps to Evaluate AI Medical Models for Your Healthcare System
- Test in Your Primary Language: Do not assume that an AI model performing well in English will translate directly to your language; conduct local validation using medical questions, case studies, and clinical scenarios relevant to your healthcare system.
- Benchmark Against Local Experts: Compare AI model performance not just against general accuracy metrics, but against the actual performance of medical professionals in your region, accounting for local diagnostic practices and terminology.
- Evaluate Specialty-Specific Performance: Test AI models across the medical specialties your organization uses most frequently; performance may vary significantly between general medicine, radiology, cardiology, and other fields.
- Assess Translation Accuracy for Multilingual Contexts: If your healthcare system serves multilingual patient populations, verify that AI systems can accurately translate medical information without losing clinical precision or introducing errors.
The Polish medical benchmark study adds to a growing body of research examining how AI models perform in specialized domains and non-English contexts. Previous research has shown that large language models can struggle with medical problem-solving in general, but this work specifically isolates the language and localization problem as a distinct challenge.
As AI adoption accelerates in healthcare globally, this research serves as a cautionary tale: impressive performance on English-language benchmarks does not guarantee reliable performance in other languages or healthcare systems. Organizations deploying AI medical tools must conduct rigorous local validation before relying on these systems for clinical decision-making. The stakes are too high for assumptions.