Open and European language models: building digital sovereignty in government

Published on 1 September 2026

Generative AI offers new opportunities for the Dutch government. It can help improve public services for citizens and businesses, while also enabling government organisations to work more efficiently. At the same time, recent geopolitical developments have highlighted the importance of critically assessing digital dependencies. To strengthen the Dutch government's digital sovereignty, encouraging the adoption of open-source models and European alternatives is essential.

However, terms such as open source, open AI and European are often interpreted in various ways in relation to AI, and lack universally accepted definitions. Providers largely determine for themselves how they classify their language models, creating uncertainty for users.

Report: open and European language models

However, terms such as open source, open AI and European are often interpreted in various ways in relation to AI, and lack universally accepted definitions. Providers largely determine for themselves how they classify their language models, creating uncertainty for users.

What is an open language model?

The meaning of the term open has become the subject of growing debate among AI developers, researchers and policymakers. The terms open and open source are frequently used interchangeably, yet they do not necessarily mean the same thing. While open source has a relatively well-established definition in the software domain, within AI it also carries legal implications under the EU AI Act.

The Act recognises the importance of open source as a driver of innovation and grants certain exemptions to free and open-source (FOS) general-purpose AI models and AI systems under specific conditions.

Increasingly, open is described as a moving target: a concept that continues to evolve and is difficult to define precisely. Some large language model developers engage in what has become known as open washing by highlighting selected aspects of openness while withholding information that is essential for scientific scrutiny or legal assessment.

To address this ambiguity, several frameworks have been developed to evaluate openness in AI systems. Drawing on the Model Openness Framework (MOF) and the European Open Source AI Index (EOSAI), language models can be positioned along five levels of openness, with closed source, open-weight and fully open source representing the most common categories.

When is a language model European?

Determining whether a language model can be considered European is less straightforward than is often assumed. The country in which the developer is based provides an important first indication, but it does not tell the whole story.

To incorporate factors such as technological sovereignty and public values, a broader definition is required. According to the report, a model can be regarded as European based on three interconnected dimensions.

1. Origin and governance of the developer

The model is developed, governed, or substantively directed by an organisation established within Europe and subject to European law.

2. Data sovereignty and infrastructure

The origin, storage and processing of training and pre-training data should comply with all applicable EU legislation and regulations. From a data sovereignty perspective, datasets should ideally be collected or managed primarily within Europe, and the model should run on European infrastructure or through certified providers that comply with EU requirements.

3. Training on European languages, values and development choices

Training data in European languages not only shapes a model's linguistic capabilities but may also reflect the cultural values embedded within those languages.

Key considerations when selecting open and European language models

Deploying language models requires informed and responsible choices. The report presents an overview of the key dimensions that organisations should consider when selecting and deploying language models. These dimensions are grouped into four categories: openness, European characteristics, technical specifications and intended use.

The framework is designed as a practical tool to help decision-makers assess which factors matter most when choosing a language model, taking into account the extent to which the model is open and/or European.

Meet the authors

Recent articles