> Bron: https://neuralex.nl/en/blog/rag-kwaliteit-meten
> Without a fixed evaluation set, improving a RAG system is largely guesswork. How to measure retrieval precision and recall, answer quality, latency and cost — and how to prevent regressions when you change chunking, embeddings or prompts.

[Back to Insights](/en/blog)

RAG ·24 September 2026 ·7 min read

# How do you measure whether your RAG system is improving?

Without a fixed evaluation set, improving a RAG system is largely guesswork. How to measure retrieval precision and recall, answer quality, latency and cost — and how to prevent regressions when you change chunking, embeddings or prompts.

You measure whether a [RAG system](/en/blog/rag-uitgelegd) is improving by evaluating more than its final answers. The entire pipeline matters: does it retrieve the right information, use that information correctly, and produce the answer quickly enough and at an acceptable cost? A useful evaluation framework therefore combines retrieval metrics such as precision and recall with answer quality, latency, and cost.

The foundation is a stable evaluation set, often called a golden set: a collection of representative questions for which you already know which sources or passages are relevant and what a satisfactory answer should contain. Running every version of the RAG system against the same set makes it possible to determine whether changes to chunking, embeddings, prompts, or the language model actually improve the system — without introducing regressions elsewhere.

## Why judging a few good answers is not enough

With conventional software, correctness can often be tested deterministically: a given input should produce a specific output. RAG systems are less straightforward. An answer may sound convincing even when the retrieval stage selected the wrong documents. Conversely, retrieval may return exactly the right passages while the language model produces an incomplete or inaccurate answer from them.

It therefore helps to evaluate at least two stages separately:

-   **Retrieval** — did the system find the right information?
-   **Generation** — did the model produce a good, well-supported answer from that information?

Operational characteristics such as speed, reliability, and cost should be evaluated alongside them.

## How do you measure retrieval quality?

Every RAG system depends on retrieval. If the required information never reaches the model's context, even a capable language model cannot reliably produce the correct answer. Two concepts remain central even after [reranking](/en/blog/wat-is-reranking): precision and recall.

### Retrieval precision

Precision measures how much of the retrieved material is actually relevant to the question. If a retriever returns five passages but only two are useful, the model receives a considerable amount of noise. Low precision can increase the chance that the model focuses on irrelevant passages or combines information that happens to look semantically similar but does not actually answer the question.

### Retrieval recall

Recall focuses on the other side of the problem: how much of the relevant information was successfully retrieved. If three passages are required to answer a question properly but the retriever finds only one, the final response may be incomplete even if that one passage is highly relevant.

Precision and recall may pull in different directions: retrieving more chunks increases the likelihood of finding all relevant information, but also adds irrelevant material. The goal is therefore not to maximise one metric in isolation, but to find an appropriate balance for the application. It also helps to measure whether relevant passages appear near the top of the ranking, particularly when only a limited number of chunks are sent to the language model.

## How do you measure answer quality?

Correct retrieval does not guarantee a correct final response. The generation stage needs its own evaluation. One important property is *faithfulness*, sometimes called grounding: are the claims in the answer actually supported by the retrieved source material? A RAG system should not quietly supplement its answer with details from the model's training data when those details cannot be supported by the supplied context.

Other useful dimensions include:

-   **Relevance** — does the response actually answer the user's question?
-   **Completeness** — does it include the information required for a satisfactory answer?
-   **Citation correctness** — do citations support the claims they are attached to?
-   **Abstention behaviour** — does the system acknowledge when the knowledge base does not contain enough evidence?

Some of these properties can be evaluated automatically or with another language model acting as an evaluator. Human review remains useful for important applications, particularly where correctness depends on nuance or specialist interpretation.

## Build a golden set that resembles real usage

Without a fixed set of test questions, improving a RAG system becomes largely subjective. A golden set should contain questions representative of how the system will actually be used. For every question, you can record:

-   which documents or passages are relevant;
-   which facts a satisfactory answer must contain;
-   which citations should support the answer;
-   whether the correct behaviour is to decline to answer.

Do not limit the set to easy questions where the exact wording appears in a single document. Include paraphrases, synonyms, ambiguous wording, questions requiring information from multiple sources, and questions for which the knowledge base contains no answer. Production failures are particularly valuable — when a meaningful failure occurs, add it to the evaluation set so future versions are tested against it.

## Use offline evaluation before deployment

Many changes to a RAG pipeline can be evaluated before they reach users. Suppose you change the chunk size, switch embedding models, introduce hybrid search, or add a reranker. Run the old and new configurations against exactly the same golden set, then compare retrieval quality, answer quality, latency, and cost. This makes it much harder to mistake a handful of impressive examples for a genuine improvement, when the change may be degrading performance across many other queries.

## When should you use A/B testing?

Offline evaluation cannot capture every real-world behaviour — users may ask questions your test set does not contain. With an A/B test, some traffic is routed to version A and some to version B. You can then compare practical signals such as:

-   how often users need follow-up questions;
-   how often they rephrase their query;
-   explicit positive or negative feedback;
-   how often a human needs to intervene;
-   latency and cost per completed request.

A/B testing is particularly helpful when two configurations perform similarly in offline evaluations but may behave differently in actual use.

## Do not ignore latency and cost

A system that produces slightly better answers is not necessarily a better product. Retrieving more chunks, using a larger reranker, or calling a more expensive language model may improve answer quality while substantially increasing response time and operating costs. Track at least:

-   retrieval latency;
-   total response latency;
-   input and output token usage;
-   model and embedding costs;
-   number of retrieval and model calls per query.

The right configuration is usually a trade-off between quality, speed, and cost based on the use case. For an internal knowledge assistant, a somewhat slower answer may be acceptable if source grounding improves significantly. For a customer-facing assistant, latency may carry more weight.

## How do you prevent regressions?

RAG pipelines contain multiple interacting components — a seemingly small change can have unexpected side effects. A different chunking strategy might improve retrieval from long policy documents while performing worse on short support articles. Every meaningful change should therefore be evaluated against the same baseline set. For each version, record at least:

-   chunking configuration;
-   embedding model;
-   retrieval method and reranking;
-   number of retrieved chunks;
-   system prompt;
-   generation model and relevant settings;
-   evaluation results.

Where possible, turn the most important evaluations into automated regression tests: a new configuration should only be promoted when it meets predefined quality requirements and does not introduce unacceptable degradation elsewhere.

## Practical checklist

A useful RAG evaluation process does not need to be complicated at the start. At minimum:

-   maintain a fixed golden set of representative questions;
-   define which sources are relevant for each question;
-   evaluate retrieval precision and recall;
-   check answer relevance and grounding;
-   test whether the system abstains when evidence is missing;
-   measure latency and cost per query;
-   compare every significant change with the previous configuration;
-   add production failures to the evaluation set.

The key shift is to stop judging a RAG system by a few answers that happen to look good, and instead make quality measurable and repeatable. Only then can you demonstrate that a new embedding model, chunking strategy, prompt, or language model has genuinely improved the system.

## Frequently asked questions

**How many questions should a RAG evaluation set contain?** There is no universally correct number. Coverage matters more than an arbitrary target. Start with a manageable set representing different question types, document types, and failure modes, then expand it with meaningful cases from production.

**Can RAG evaluation be fully automated?** A substantial part can be automated, including retrieval tests, latency, cost measurements, and some forms of answer evaluation. Human review remains valuable for nuance, factual interpretation, and cases where several different answers may all be acceptable.

**Should you rerun evaluations after every prompt change?** Yes. Even a small prompt change can affect completeness, use of sources, answer style, and abstention behaviour. Running the same evaluation set again shows whether the change caused an improvement or a regression.

Measurable quality

## Do you know whether your RAG system is actually getting better?

Neuralex helps set up a golden set and evaluation framework for RAG — from retrieval metrics to regression tests before every release.

[From demo to production](/en/blog/rag-van-demo-naar-productie) [Ask your question](/en/contact)

---
Volledige (opgemaakte) versie: https://neuralex.nl/en/blog/rag-kwaliteit-meten
