> Bron: https://neuralex.nl/en/blog/embeddings-nederlandse-tekst
> There is no embedding model you can call objectively "best for Dutch" without testing it yourself. E5-NL, Qwen3 Embedding, Jina Embeddings v5, BGE-M3 and multilingual-e5 are candidates — MTEB-NL and your own retrieval test should decide the winner.

[Back to Insights](/en/blog)

RAG ·August 29, 2026 ·7 min read

# Choosing embeddings: which model understands Dutch best?

There is no embedding model you can call objectively "best for Dutch" without testing it yourself. E5-NL, Qwen3 Embedding, Jina Embeddings v5, BGE-M3 and multilingual-e5 are candidates — MTEB-NL and your own retrieval test should decide the winner.

There is currently no embedding model that can be declared universally "best for Dutch" without testing it on your own data. For a Dutch-language RAG knowledge base, models specifically adapted to Dutch, such as the recent E5-NL family, are obvious candidates, but strong multilingual models including Qwen3 Embedding, Jina Embeddings v5, BGE-M3 and multilingual-e5 should also be evaluated. For managed APIs, OpenAI text-embedding-3-large, Cohere Embed v4 and Voyage 4 are relevant options. The winner should be determined by retrieval quality on your documents rather than by a single leaderboard position.

This distinction has become easier to measure because Dutch now has a dedicated MTEB-NL benchmark. It evaluates Dutch embeddings across retrieval, clustering, classification, reranking, semantic similarity and other tasks. The accompanying ACL 2026 research also introduced the E5-NL models and specifically identified the relative underrepresentation of Dutch in broader multilingual embedding resources. General MTEB rankings remain useful, but MTEB-NL provides a much more relevant signal for Dutch applications. Rankings in this field change quickly, so any leaderboard should be checked again at the time of deployment.

## Why embeddings determine RAG quality

An embedding model converts text into a numerical vector intended to represent its meaning. In a [RAG system](/en/blog/rag-uitgelegd), documents or chunks are embedded and stored in a vector index. When a user asks a question, that query is embedded as well, allowing the system to retrieve passages whose vectors are semantically close to the query.

This makes the embedding model an upstream component of the generative model. If retrieval selects the wrong passages, even an excellent language model may fail because the information required for the answer never reaches its context window. Chunking, metadata filtering, hybrid search and reranking all matter, but they cannot fully compensate for consistently poor semantic retrieval.

## Dutch is not simply translated English

A model that performs extremely well on English retrieval does not automatically perform equally well on Dutch. English usually represents a much larger share of training material, while Dutch introduces its own syntax, word order, terminology and long compound words. A tokenizer may successfully split a word such as arbeidsongeschiktheidsverzekering into useful subword units, but that alone does not guarantee that the resulting embedding captures its semantic relationship to concepts such as income protection or long-term disability coverage.

Domain language adds another complication. Legal, medical, financial and technical knowledge bases contain terminology that may be rare even in large multilingual corpora. Dutch and Flemish phrasing can differ as well. "Multilingual support" therefore means that a model can process a language; it does not imply equal retrieval performance across every supported language.

## English-only, multilingual and Dutch-specific models

English-only embedding models still exist. Cohere, for example, distinguishes explicitly between English-only and multilingual embedding families. Such models may still produce usable vectors for Dutch because modern tokenizers and pretraining often provide some cross-lingual capability, but they should not be the default for a primarily Dutch corpus unless your own evaluation shows an advantage.

The next category is multilingual embeddings. multilingual-e5-large was trained from a multilingual XLM-RoBERTa foundation and supports roughly one hundred languages. BGE-M3 also supports more than one hundred languages and is particularly interesting for RAG because it can perform dense, sparse and multi-vector retrieval.

More recent families broaden the choice. Qwen3 Embedding is available in several model sizes, supports more than one hundred languages and offers configurable output dimensions. Jina Embeddings v5-text is another current multilingual family; its small model supports more than one hundred languages and allows embeddings to be shortened to several predefined dimensions.

For Dutch specifically, E5-NL deserves attention. The CLiPS collection includes models such as e5-small-trm-nl, e5-base-trm-nl and e5-large-trm-nl, developed and evaluated alongside MTEB-NL. The published collection reports strong results within their respective model-size categories. They are therefore compelling open-weight candidates for Dutch retrieval and local deployment, but that still does not make them automatically superior on every company knowledge base.

## Which hosted models are relevant?

For teams that prefer an embedding API, several mature options exist. OpenAI describes text-embedding-3-large as its most capable embedding model for both English and non-English tasks. It can produce vectors up to 3072 dimensions and supports reducing their size through the dimensions parameter.

Cohere Embed v4 supports more than one hundred languages and allows several output dimensions. Voyage currently positions voyage-4 and voyage-4-large as general-purpose multilingual retrieval models, again with multiple embedding dimensions available.

These are sensible candidates for a benchmark, not automatic winners. A large proprietary model may provide stronger general retrieval, while a smaller Dutch-adapted model may perform better on a narrowly Dutch corpus.

## Build your own Dutch retrieval evaluation

The most reliable selection method is to create a small, fixed evaluation set. Collect representative user questions and define which document or passage should be retrieved for each one. Include realistic language: synonyms, abbreviations, spelling variation and questions where the wording differs substantially from the target passage.

Embed exactly the same document chunks with each candidate model and compare their results. Metrics such as Recall@k, precision@k, MRR and nDCG can quantify whether relevant passages appear near the top. Manual inspection remains useful as well. A model may retrieve passages containing similar vocabulary without actually resolving the intended meaning.

Evaluate the complete retrieval stack too. A slightly weaker dense embedding model combined with BM25, good metadata filters and a reranker can outperform a nominally stronger model used in isolation.

## API versus on-prem

API-hosted embeddings are operationally convenient. There is no model serving infrastructure to maintain, and usage is normally charged according to the amount of text processed, commonly expressed as a price per million tokens. The trade-offs are external data processing, network dependency, possible latency and continuing usage charges.

Open models such as E5-NL, multilingual-e5, BGE-M3, Qwen3 Embedding and some Jina models can instead be deployed locally or in a controlled private environment — comparable to the trade-offs behind [air-gapped AI](/en/blog/air-gapped-ai). That can be important when source documents should not leave the organisation. Local inference removes the external per-request token bill, but compute, memory, deployment, monitoring and engineering effort become your responsibility.

For privacy-sensitive workloads, deployment requirements may therefore eliminate several candidates before benchmark quality is even considered.

## Dimensions and cost matter too

Embedding dimensions have direct infrastructure consequences. Larger vectors consume more storage and RAM in a vector index and can increase the computational work required for nearest-neighbour search. Modern embedding models increasingly support Matryoshka-style or otherwise configurable vector sizes, allowing teams to trade some representational capacity for smaller indexes and potentially faster retrieval.

For hosted services, compare retrieval quality alongside the current price per million tokens. For an on-prem model, translate that calculation into hardware requirements, throughput and operational cost. A small gain in retrieval quality may not justify substantially larger vectors or significantly more expensive inference hardware.

Changing models also has a migration cost. In most cases, switching embedding models means regenerating the embeddings for the entire corpus because vectors produced by different models do not share the same semantic space. Choosing an embedding model is therefore not merely a model-selection decision; it becomes part of the system architecture.

## Conclusion

There is no permanent number-one embedding model for Dutch RAG. E5-NL deserves a prominent place on any local shortlist because it is explicitly developed and benchmarked for Dutch, while strong multilingual families such as Qwen3 Embedding, Jina Embeddings v5, BGE-M3 and multilingual-e5 provide credible alternatives. OpenAI, Cohere and Voyage are relevant managed API options. Use MTEB and especially MTEB-NL to narrow the field, then let a representative Dutch retrieval evaluation determine the final choice. And check the leaderboards again when making that decision: in the embedding market, a ranking that is only a few months old may already describe yesterday's field.

RAG that understands what's being asked

## Not sure which embedding model fits your Dutch knowledge base?

We build and evaluate RAG architectures on your own documents and questions — including a real retrieval evaluation instead of a generic leaderboard claim.

[Get in touch](/contact) [Read how many documents you need](/en/blog/hoeveel-documenten-voor-rag)

---
Volledige (opgemaakte) versie: https://neuralex.nl/en/blog/embeddings-nederlandse-tekst
