Logo
FrontierNews.ai

Europe's New AI Model Tackles a Problem Big Tech Ignores: Rare Languages

The European Commission is developing a specialized artificial intelligence model designed to support all 24 official EU languages, addressing a critical gap that commercial AI systems have largely ignored. Most large language models (LLMs), the AI systems that power chatbots and translation tools, are trained primarily on English and a handful of major languages, leaving minority languages like Irish, Latvian, and Maltese severely underrepresented in the data used to teach these systems how to understand and generate text.

Why Are Minority Languages Disappearing From AI?

The problem is straightforward but consequential. When AI models are trained on massive datasets scraped from the internet, they reflect what's actually out there. Common Crawl, one of the most widely used training datasets in the industry, contains only 0.09% Latvian content, 0.07% Irish, and 0.03% Maltese. The least-represented half of all EU official languages combined account for merely 2.4% of the dataset. This means AI systems perform poorly when asked to work with these languages, creating a vicious cycle where digital tools become less useful, discouraging people from using them online, which further reduces the amount of training data available.

For a multilingual continent like Europe, this is more than an academic concern. Public administrations, small businesses, and civil society organizations across the EU need AI tools that can handle the full linguistic diversity of their populations. When commercial models fail to support these languages adequately, European institutions lose the ability to deploy AI solutions that work for all their citizens.

How Is the EU Building a Better Multilingual AI Model?

The European Commission's Directorate-General for Translation is constructing what it calls the EU Institutional LLM, a specialized AI model built from the ground up to handle Europe's linguistic complexity. Rather than starting from scratch, the team is enhancing an existing open-source European model created by Mistral AI, using two critical assets unique to EU institutions:

  • Supercomputing Infrastructure: Access to Europe's most powerful computers through the European High Performance Computing Joint Undertaking, including systems like MeluXina in Luxembourg, Leonardo in Bologna, and MareNostrum 5 in Barcelona, which provide the computational muscle needed to train large models efficiently.
  • Institutional Data: The European Advanced Multilingual Information System (Euramis), a vast collection of high-quality multilingual texts from EU institutions, carefully curated to meet strict quality standards and free from copyright issues.
  • Language Expertise: Direct access to professional translators and language specialists who can evaluate the model's performance and provide feedback during development, ensuring the final product actually meets real-world needs.

The results so far are striking. In early testing on EU institutional texts, the new model dramatically outperformed the original across all EU languages tested. Irish nearly quadrupled its performance score, Estonian improved by almost 80%, Greek nearly doubled its result, and both Latvian and Lithuanian recorded gains of around 70 to 75%. These aren't marginal improvements; they represent the difference between a tool that barely works and one that's genuinely useful.

What Makes This Model Different From Commercial AI?

The EU Institutional LLM takes a fundamentally different approach than the proprietary models offered by major tech companies. Each of the 24 EU official languages is represented with at least 1 billion tokens, or units of text, ensuring no language is left behind. The model is built on European technology, trained on European infrastructure, and aligned with European values and legal standards. It's already powering eSummary, one of the EU's AI-based translation tools, and future versions will support additional services.

Critically, the model is freely available to any EU-based legal entity, including public administrations, small businesses, academia, and nonprofit organizations. Two versions have been released: a base model for continued training and an instruction-tuned version optimized for specific tasks. This open approach stands in sharp contrast to the closed, proprietary models that dominate the global AI market.

How Are Researchers Measuring Success Across Languages?

Building a model that covers all EU languages is only half the challenge. The other half is having reliable tools to measure whether the model actually performs well across those languages. Most AI benchmarks used today were developed in English and don't reflect European educational, cultural, or societal contexts, making it difficult to assess how a model truly performs in other languages.

To address this gap, the Commission has released the EU MMLU, a new multilingual benchmarking dataset designed specifically for evaluating AI models across Europe's linguistic diversity. Unlike most existing multilingual benchmarks that rely on machine translation, the EU MMLU uses a human-centered approach. The team partnered with student translators and project managers from the European Master's in Translation network to manually translate and revise over 1,000 benchmark questions across seven subject areas relevant to the EU, including science, technology, mathematics, law, ethics, economics, and public affairs. This ensures the questions retain the same meaning, difficulty, and testing value across all languages, not just a rough machine-translated approximation.

Steps to Access and Use the EU Institutional LLM

  • Check Eligibility: Confirm your organization qualifies as an EU-based legal entity, which includes public administrations, small businesses, academic institutions, and nonprofit organizations across the 27 EU member states.
  • Download From the European Language Data Space: Access either the base model for continued pretraining or the instruction-tuned version optimized for specific use cases, both available through the official EU repository.
  • Evaluate Performance: Use the EU MMLU benchmarking dataset to test how well the model performs in your specific language and domain, ensuring it meets your organization's accuracy and reliability requirements before deployment.
  • Deploy for Multilingual Services: Integrate the model into your organization's AI-based tools and services, whether for translation, document summarization, or other language-dependent applications that serve EU citizens and stakeholders.

This initiative represents a strategic shift in how Europe approaches AI development. Rather than relying entirely on models built by American and Chinese tech companies, the EU is investing in its own technological sovereignty while solving a problem that commercial incentives have largely ignored. The focus on minority languages isn't just about fairness; it's about ensuring that digital tools remain accessible across Europe's diverse populations and that the continent retains the ability to build AI systems tailored to its own needs and values.

The EU Institutional LLM is positioned as the first step in a longer-term program, with future iterations planned to expand capabilities and support additional use cases. As Europe continues to implement the EU AI Act and other regulatory frameworks, having access to AI models that are transparent, controllable, and aligned with European standards becomes increasingly important for both compliance and innovation.