> Bron: https://neuralex.nl/en/blog/chunking-uitgelegd
> Good chunking determines what your RAG system can retrieve in the first place. Why chunking goes wrong, how to choose chunk size and overlap, and how to chunk PDFs, tables and code differently.

[Back to Insights](/en/blog)

RAG ·21 August 2026 ·6 min read

# Chunking explained: how do you split a document without breaking its meaning?

Good chunking determines what your RAG system can retrieve in the first place. Why chunking goes wrong, how to choose chunk size and overlap, and how to chunk PDFs, tables and code differently.

Good chunking means dividing a document into pieces that are small enough to retrieve precisely but large enough to retain meaningful context on their own. That means you generally do not want to insert a split mechanically every few hundred words. Wherever possible, chunk boundaries should follow the structure of the content: a paragraph, section, table, function or coherent topic.

In a [RAG](/en/blog/rag-uitgelegd) system, chunking matters because retrieval usually searches the individual chunks created from a document rather than the original document as a whole. If an answer is split across poorly chosen chunks, even a strong embedding model may retrieve the wrong fragment. Chunks that are too large create the opposite problem: they bring back too much irrelevant context. The right strategy therefore depends on document structure, user queries and the type of information you need to retrieve.

## What is chunking?

Chunking is the process of dividing source material into smaller units before indexing it in a RAG system.

Suppose you have a sixty-page technical manual. You could store the entire manual as a single embedding, but that one vector would then represent dozens of unrelated topics. A query about resetting a password would have to compete with information about installation, billing, permissions and error handling.

Instead, you create smaller units such as:

-   one paragraph per chunk;
-   several related paragraphs;
-   a section below a single heading;
-   a table together with its explanation;
-   one function or class from a source file.

Each chunk will usually receive its own embedding and metadata such as the document name, page number, chapter or URL. The goal is not to make chunks as small as possible. The goal is to create useful retrieval units.

## Why chunking goes wrong

The most common mistake is splitting text without considering meaning. Consider this text:

*"Accounts using two-factor authentication follow a different reset procedure. If the authentication device is lost, contact the administrator."*

If a chunk boundary falls exactly between those sentences, a search for "lost phone with authenticator" may retrieve only the second sentence. The missing first sentence contains the context explaining that the instruction applies specifically to two-factor authentication.

The opposite problem occurs when chunks are too large. Imagine a single chunk containing an entire chapter on installation, access control, logging and backups. The relevant answer may be present, but the embedding represents several subjects at once. Retrieval also sends a large amount of unrelated text into the language model.

Poor chunking therefore creates two distinct problems:

-   related information is separated;
-   relevant information is buried inside too much irrelevant text.

Both reduce retrieval quality.

## Chunk size and overlap

There is no universal ideal chunk size. The correct value depends on the documents and the questions users ask. Smaller chunks are useful when users search for highly specific facts: higher precision, but more likely to lack necessary context. Larger chunks preserve more relationships, but are more likely to contain multiple topics.

This is why overlap is often introduced: part of the end of chunk A is repeated at the beginning of chunk B. A chunk ending with "The system retains audit logs for the configured retention period. Administrators can change this period for each environment." can repeat that last sentence as the opening of the next chunk, before new information about, say, compliance requirements. The overlap reduces the chance that an important transition is lost exactly at a boundary.

Overlap is not, however, a substitute for good segmentation. If you split arbitrarily through headings, tables and paragraphs, adding more overlap mainly creates duplicate content in the vector store. A better order is: use natural document boundaries first, then add overlap only where it is useful.

## Fixed-size versus semantic chunking

With **fixed-size chunking**, text is divided according to a fixed maximum length, typically measured in tokens or characters. It is simple and predictable — for relatively uniform prose such as long reports, it can provide a useful baseline. Its weakness is that content has no influence on where the boundary falls: a chunk can end halfway through an argument.

With **semantic or structural chunking**, boundaries are chosen based on the content itself: chapters, headings, paragraphs, lists, topic changes, or structural information from HTML or Markdown. A section such as `## Resetting a password` followed by the explanation below it should normally remain together rather than separating the heading from the text that belongs with it.

Semantic chunking is not automatically superior. If a section is extremely long, it still needs to be divided further. In practice, a hybrid approach often works well: first split using natural structure, then subdivide blocks that exceed a chosen maximum size.

## Chunking by document type

**PDFs.** With PDFs, chunking effectively starts before the split itself: during extraction. PDF is primarily a visual format — text may be extracted in the wrong order, headers and footers may appear repeatedly and columns can become interleaved. A better pipeline is therefore:

`pdf → reconstruct text and structure → identify headings, paragraphs and tables → create chunks → generate embeddings`

If extraction is wrong, no chunking strategy can fully repair the damage later. It is also useful to preserve metadata such as document name, page number and section title — this makes source citation and debugging significantly easier.

**Tables.** Tables should not be treated as ordinary prose. A table loses its meaning if the column headings end up in a different chunk from the values:

Product

Retention period

Region

A

30 days

EU

B

90 days

EU

A table chunk should therefore usually contain both the column names and the relevant rows. Large tables can be divided by row, but the headers should then be repeated or added structurally. Context around the table also matters: a table titled "Retention periods for closed accounts" means something different from the same values without that title.

**Code.** Code requires another type of boundary. Cutting a Python function in half simply because it reached a token limit often creates a chunk with little standalone value. Natural chunking units for code include a function, method, class, module or configuration block — where possible, retain surrounding context such as the filename, class name, imports or function signature. For source-code repositories, language-aware parsing can therefore produce much better chunks than a generic text splitter.

## How do you test whether your chunking works?

Do not evaluate chunking only by looking at average chunk length. Test whether real user questions retrieve the right information. Create a small set of representative queries for which you already know where the correct answer appears — for example "how can a user reset their account if they no longer have access to their authenticator?" with an expected source of Manual > Account management > Two-factor authentication. Then run retrieval and inspect the retrieved chunks before looking at the language model's final answer.

Check:

-   does the correct chunk appear among the top results?
-   does it contain enough context to answer the question?
-   is there a large amount of irrelevant text around it?
-   is essential information split across several chunks?
-   does overlap cause nearly identical chunks to appear repeatedly?

You can then compare alternative strategies: smaller chunks, larger chunks, less overlap, structural splitting or different metadata. This matters because chunking and retrieval form one system — a theoretically elegant chunker that performs poorly on your real questions is not a good chunker.

## Chunking is part of retrieval, not just preprocessing

Chunking may look like a simple preprocessing step, but it determines what your retriever can find in the first place. An embedding model cannot reliably reconstruct a relationship that was completely broken during chunking. Start with the natural structure of the document, keep related information together and use fixed limits primarily as upper bounds. Then test the result using real search questions instead of treating a single chunk size as a universal standard.

For the broader process of retrieving sources and generating grounded answers, continue with [RAG explained](/en/blog/rag-uitgelegd). If your chunks are being retrieved but the best passages still do not rank highly enough, [What is reranking and when do you need it?](/en/blog/wat-is-reranking) is the logical next topic.

Retrieval that holds up

## Not sure if your documents are split correctly?

Neuralex takes a no-obligation look at your chunking and retrieval strategy — from document extraction to the final answer quality.

[View the showcase](/en/showcase/rechtspraak) [Ask your question](/en/contact)

---
Volledige (opgemaakte) versie: https://neuralex.nl/en/blog/chunking-uitgelegd
