> Bron: https://neuralex.nl/en/blog
> Clear pieces on RAG, on-prem AI, GEO and automation — written to inform and to be cited by AI assistants.

Insights

# Knowledge that's correct — and gets cited

No content mill, just pieces that explain how reliable AI really works. Written for people who want to understand it properly — and structured so AI assistants pick it up correctly (GEO).

[

RAG ·10 June 2026 ·6 min read

## Why RAG without citations is worthless

Everyone says "RAG". But without verifying the source, you still get a model that lies convincingly. This is the difference between a demo and a system you can build on.

Read the article

](/en/blog/rag-zonder-hallucinatie)

## More from Insights

[

Agentsread

### When does a human need to approve? A decision tree for AI agent autonomy

Not every action an AI agent can take deserves the same amount of freedom. A practical framework based on reversibility, impact, scale and input trustworthiness — with a concrete decision tree for when a human should approve first.

Read](/en/blog/wanneer-moet-een-mens-goedkeuren)[

GDPRread

### How do you anonymise documents before they enter a model?

Documents should be anonymised before AI processing by detecting personal information and removing, masking or replacing it — ideally locally, before text reaches an external API. Pseudonymisation is often more practical, but the resulting data remains personal data under the GDPR.

Read](/en/blog/documenten-anonimiseren-voor-ai)[

RAGread

### How do you measure whether your RAG system is improving?

Without a fixed evaluation set, improving a RAG system is largely guesswork. How to measure retrieval precision and recall, answer quality, latency and cost — and how to prevent regressions.

Read](/en/blog/rag-kwaliteit-meten)[

Infrastructureread

### Monitoring AI services: which signals actually matter?

An AI service can be technically "up" while users get slow, expensive or low-quality answers. Which signals beyond latency, cost, drift and hallucinations your monitoring actually needs to catch.

Read](/en/blog/ai-monitoring)[

AI Actread

### Chatbot disclosure: how do you say it without scaring visitors off?

A chatbot disclosure doesn't need to sound like a legal warning. Where the notice belongs, five example sentences from subtle to explicit, and when to be more explicit in healthcare, finance or with minors.

Read](/en/blog/chatbot-disclosure-formuleren)[

GEOread

### Writing for Humans and Models: A Structure That Serves Both

Content that works well for people often works well for language models too — if you write with the same discipline. Why the answer should come first, when a list or table helps, and how to write content that stays easy to retrieve without sounding robotic.

Read](/en/blog/schrijven-voor-mens-en-model)[

Agentsread

### How do you see what your agent did last night?

An agent that runs autonomously overnight should be traceable the next morning. What to log, how tracing with something like OpenTelemetry helps, and how to build a simple audit trail and dashboard.

Read](/en/blog/agent-observability)[

GDPRread

### Data residency: what does an 'EU region' really mean with a US cloud provider?

An EU cloud region does not automatically mean European control. What data storage location, access, the US CLOUD Act and cloud sovereignty really mean for your AI project.

Read](/en/blog/data-residency-eu-regio)[

Infrastructureread

### Rate limits and quotas: how do you stop your service from grinding to a halt?

Rate limits and quotas can slow down or interrupt AI services without warning. Retries, queues, model routing, budget guards and monitoring together turn temporary scarcity into controlled degradation instead of an outage.

Read](/en/blog/rate-limits-en-quota)[

Municipalitiesread

### AI for Municipalities: Making Policy Documents Searchable

Policy papers, council records, permit files and local regulations become searchable by meaning instead of just keywords with RAG, with citations back to the original document. Why the Woo and data sovereignty matter, and when an application can qualify as high-risk under the AI Act.

Read](/en/blog/ai-voor-gemeenten)[

RAGread

### When Should an AI Assistant Say 'I Don't Know'?

An AI assistant that never refuses is predictably unreliable. How relevance thresholds, missing-context detection and citation verification let a RAG system refuse exactly when it should.

Read](/en/blog/wanneer-moet-ai-weigeren)[

GEOread

### Why Does Your FAQ Page Rank Worse Than One Good Question Per Page?

A long FAQ page makes several topics compete for attention, while search and retrieval systems often process content in smaller passages. When does a question deserve its own knowledge-base page, and when does it belong in the FAQ?

Read](/en/blog/faq-pagina-of-losse-paginas)[

Agentsread

### Prompt injection: how a single email can hijack your agent

An AI agent that reads email and takes action can be manipulated by text inside that email — no password, no hack required. What indirect prompt injection is, and how least privilege and human-in-the-loop limit the damage.

Read](/en/blog/prompt-injectie-uitgelegd)[

GDPRread

### What is a DPIA and when do you need one for AI?

You need a DPIA for an AI system when its processing of personal data is likely to result in a high risk to individuals. AI itself is not the legal trigger — the combination of data, purpose, scale and consequences is.

Read](/en/blog/dpia-voor-ai)[

RAGread

### Choosing embeddings: which model understands Dutch best?

There is no embedding model that is objectively "best for Dutch". E5-NL, Qwen3 Embedding, Jina Embeddings v5, BGE-M3 and multilingual-e5 are candidates — your own retrieval test should decide the winner.

Read](/en/blog/embeddings-nederlandse-tekst)[

Infrastructureread

### Caching for AI: which answers are safe to reuse?

AI responses are safest to reuse when the same request, under the same conditions, should produce the same valid result. Time-sensitive, personalised or creative answers usually aren't — and semantic caching has its own risks.

Read](/en/blog/caching-voor-ai)[

Healthcareread

### AI in healthcare: why on-prem is almost always the answer

Health data is a special category of personal data under GDPR. That's why on-prem or a tightly controlled private environment is usually the most defensible architecture for AI working with patient records — not an absolute ban on cloud AI, but a much higher bar.

Read](/en/blog/ai-in-de-zorg)[

AI Actread

### AI literacy (Article 4): what does the law expect from your staff?

Article 4 of the AI Act requires providers and deployers to take measures supporting AI literacy among staff — proportionate to role and risk, with no mandatory certificate. In force since 2 February 2025, amended in 2026.

Read](/en/blog/ai-geletterdheid-artikel-4)[

GEOread

### Which AI crawlers visit your site — and how do you spot them in your logs?

GPTBot, ClaudeBot, PerplexityBot and more do not all do the same thing: some gather training data, some build a search index, some fetch a page live at a user's request. How to filter and verify them in your logs.

Read](/en/blog/ai-crawlers-herkennen)[

Agentsread

### How do you give an AI agent safe access to your files?

Keep the scope as small as possible: an explicit allowlist instead of a denylist, separate read and write permissions, backups before every mutation, and a tool layer that enforces the boundary outside the model.

Read](/en/blog/agent-toegang-tot-bestanden)[

GDPRread

### Are you allowed to run customer data through ChatGPT?

Under GDPR it's allowed under conditions: a lawful basis, data minimisation, and — when OpenAI acts as a processor — a data processing agreement. But a personal ChatGPT account is not the same as ChatGPT Business, Enterprise or the API.

Read](/en/blog/klantgegevens-door-chatgpt)[

RAGread

### How many documents do you need before RAG is worth it?

There is no minimum. RAG earns its complexity once documents no longer fit a single context, content changes often, many users need to search, or answers need citations — not at a fixed file count.

Read](/en/blog/hoeveel-documenten-voor-rag)[

Infrastructureread

### Fallback chains: what does your system do when the provider goes down?

A fallback chain combines alternative AI providers or models with health checks, timeouts, retries and circuit breakers. Downtime is not the only failure mode — measurable quality degradation counts too.

Read](/en/blog/fallback-ketens)[

Accountancyread

### AI for accountants: where's the upside, and where's the risk?

The upside lies in faster document processing, journal-entry suggestions and periodic close support. The risk lies in hallucinated figures, tax interpretation, GDPR and liability.

Read](/en/blog/ai-voor-accountants)[

AI Actread

### What is a high-risk AI system — and are you running one?

Not every advanced AI system is high-risk under the EU AI Act, and not every simple one is exempt. Article 6 and Annex III decide it — along with the deadlines the Digital Omnibus pushed back in mid-2026.

Read](/en/blog/hoog-risico-ai-systeem)[

GEOread

### Should you block AI crawlers or let them in?

Blocking every AI crawler is usually too blunt. GPTBot, ClaudeBot, PerplexityBot, Google-Extended and CCBot each serve a different purpose — how to configure robots.txt for training versus visibility.

Read](/en/blog/block-ai-crawlers-or-let-them-in)[

Agentsread

### n8n or code: when is a workflow tool enough?

Where workflow tools are strong, when custom code fits better, and why combining n8n with code is often the strongest architecture for AI automation.

Read](/en/blog/n8n-or-code)[

AI Actread

### Article 50 explained: when do you have to disclose that it's AI?

Article 50 does not require every use of AI to be disclosed. The transparency duty applies specifically to chatbots, AI-generated content, deepfakes and emotion recognition — in force since 2 August 2026.

Read](/en/blog/artikel-50-uitgelegd)[

AI Actread

### Do You Need to Label AI-Generated Content?

The AI Act does not require a blanket visible label on all AI content. Article 50 distinguishes between technical marking, deepfakes and AI text on matters of public interest.

Read](/en/blog/ai-content-labelen)[

Infrastructureread

### Model routing: why not every request should go to the most expensive model

Not every AI task needs the biggest model. How to route on task complexity, latency, cost and reliability — with classifiers, rules, fallbacks and cost ceilings.

Read](/en/blog/model-routing-uitgelegd)[

Legalread

### Can AI search case law without making things up?

A language model without source grounding can produce convincing but non-existent judgments and ECLI numbers. How RAG, mandatory grounding and a link to the actual judgment reduce that risk.

Read](/en/blog/ai-en-jurisprudentie)[

On-premread

### Ollama, vLLM or llama.cpp: which runtime fits your situation?

llama.cpp is the C/C++ inference engine, Ollama the user-friendly layer around it for local experimentation, and vLLM the production-serving engine for GPUs with many concurrent users. Which one fits your workload.

Read](/en/blog/ollama-vllm-of-llamacpp)[

GEOread

### How do ChatGPT, Perplexity and Gemini decide which source to cite?

Ranking well in Google is not the same as being cited in an AI answer. How ChatGPT Search, Perplexity and Gemini retrieve, assess and link sources to specific claims — and what that means for GEO.

Read](/en/blog/hoe-kiezen-ai-modellen-bronnen)[

Agentsread

### What is MCP (Model Context Protocol) and why is it changing agents?

The open protocol that lets AI agents connect to tools, data and systems in a standardized way. How it relates to APIs, function calling and RAG, and what it means for security and on-premises AI.

Read](/en/blog/wat-is-mcp)[

On-premread

### What hardware do you need to run a language model locally?

RAM, VRAM, quantization and storage determine which LLMs you can run locally. Practical memory rules for 7B to 70B+ models on CPUs, GPUs and Apple Silicon.

Read](/en/blog/hardware-voor-lokale-ai)[

RAGread

### What is reranking and when do you need it?

Reranking is the second stage in a RAG pipeline that reorders candidate documents by relevance using a cross-encoder. When does it actually add value, and when is it unnecessary complexity?

Read](/en/blog/wat-is-reranking)[

RAGread

### Chunking explained: how do you split a document without breaking its meaning?

Good chunking determines what your RAG system can retrieve in the first place. Why chunking goes wrong, how to choose chunk size and overlap, and how to chunk PDFs, tables and code differently.

Read](/en/blog/chunking-uitgelegd)[

Infrastructureread

### What does AI really cost? From API token to your own server

APIs bill by the token, a server you own bills in purchase price, electricity and upkeep. What a million tokens actually costs at Anthropic and OpenAI, when owned hardware pays for itself, and the hidden line items that inflate the bill.

Read](/en/blog/wat-kost-ai-echt)[

Legalread

### AI in legal practice: what can a law firm actually use it for today?

Case law and legislative research, contract analysis, case-file search and draft correspondence already work — provided the output stays verifiable. Why hallucination is the core risk, and what RAG with source-grounded citation fixes.

Read](/en/blog/ai-in-de-juridische-praktijk)[

AI Actread

### The EU AI Act for SMEs: what do you need to have in place right now?

Since 2 August 2026 the transparency rules apply and enforcement has begun. The high-risk system rules have been postponed to 2027/2028 — what that concretely means for SMEs.

Read](/en/blog/ai-act-voor-het-mkb)[

GEOread

### From a score of 36 to 85+: what actually raises an agent-readability score

An online store scored 36 on AI readability and climbed to 85+ without changing a single line in the CMS. Which fixes made the difference, and why they shipped through an edge worker.

Read](/en/blog/score-36-naar-85-agent-leesbaarheid)[

GEOread

### Measure your AI visibility: is it you in the answer, or your competitor?

A ranking in Google no longer says anything about whether ChatGPT, Perplexity or Gemini mention your business. How to measure, yourself, whether you show up in AI answers — and what to do when the answer is no.

Read](/en/blog/meet-je-ai-zichtbaarheid)[

RAGread

### RAG: from demo to production

A demo answers ten questions well. Production has to answer ten thousand, including the ones nobody anticipated. What breaks along the way, and what it takes to hold up.

Read](/en/blog/rag-van-demo-naar-productie)[

Agentsread

### AI agents with guardrails: letting a system act without losing control

The brain may only propose, the hands may only act within limits set in advance. On whitelists, deny-lists, cooldowns and the approval card.

Read](/en/blog/ai-agents-met-vangrails)[

Costsread

### Keeping AI costs under control without crippling your system

Most AI bills grow because nobody measures per call. On routing to the cheapest model that fits, caching, and knowing the cost of every request.

Read](/en/blog/ai-kosten-in-de-hand-houden)[

GEOread

### Getting found by AI assistants: what actually works

Which crawlers matter, what robots.txt and llms.txt do for them, and the checks that took one site from 36 to 85+.

Read](/en/blog/gevonden-worden-door-ai-assistenten)[

GEOread

### Schema.org for AI: which structured data do assistants actually understand?

The five types that actually pay off, how to tie them into a single entity with @id, and the four mistakes that turn markup against you.

Read](/en/blog/schema-org-voor-ai)[

GEOread

### llms.txt: the new robots.txt — what it is and whether you need one

A proposed file in your site root that tells AI systems in plain text what your site is and where things live. Origin, structure, and whether your site needs one.

Read](/en/blog/llms-txt-uitgelegd)[

Agentsread

### Multi-agent orchestration: how do you get several AI agents to work together without chaos?

Adding more agents rarely fixes the problem — it moves it to the coordination between them. On one shared playbook, fixed roles and one log as the single source of truth.

Read](/en/blog/multi-agent-orkestratie)[

RAGread

### RAG explained: how to build an AI that doesn't lie

Retrieval, context, generation and citation verification — the four ingredients that separate a RAG demo from a system you can actually trust.

Read](/en/blog/rag-uitgelegd)[

On-premread

### On-prem AI: when is it worth the investment?

Cloud is easy, until your bill and your DPO start paying attention. An honest trade-off.

Read](/en/blog/on-prem-ai-wanneer-de-moeite)[

Agentsread

### Why your AI agents start over every session — and how to fix it

The most expensive problem in agent work is memory loss. On the memory layer that turns separate sessions into one continuous brain.

Read](/en/blog/continu-geheugen-agents)[

GEOread

### GEO: how to get cited by ChatGPT and Gemini

SEO got you into Google. GEO gets you into the answer the user reads next.

Read](/en/blog/geo-geciteerd-door-ai)[

Modelsread

### Training your own model? Start with RAG.

Why "our own LLM" is in practice usually RAG plus light fine-tuning — and that’s good news.

Read](/en/blog/eigen-model-trainen-begin-met-rag)[

GDPRread

### AI and GDPR: what's allowed, what isn't, and how to get it right

Using AI with personal data is allowed under the GDPR — provided you have a legal basis, apply data minimisation, and know where the data goes. From legal basis to DPIA, and where the EU AI Act adds requirements.

Read](/en/blog/ai-en-de-avg)[

Agentsread

### AI agents for SMEs: what actually works today (and what doesn't yet)

Searching documents and fixed workflows already work reliably. Fully autonomous decision-making doesn't yet. What works, what doesn't, and how an SME gets started.

Read](/en/blog/ai-agents-voor-het-mkb)[

On-premread

### Air-gapped AI: can a language model run without the internet?

Yes, it can — but the question is whether your threat model justifies that level of isolation. What it takes technically, and why most businesses only need on-prem.

Read](/en/blog/air-gapped-ai)[

Modelsread

### Local models in 2026: Llama, Qwen and Gemma compared fairly

Licensing, sizes and context of Llama 4, Qwen 3.5/3.6 and Gemma 4, laid out plainly — and why "which one is best" is the wrong question in a field that shifts monthly.

Read](/en/blog/lokale-modellen-2026-vergeleken)

Don't miss out

## A question that should be answered here?

We'd rather write about what our clients actually run into. Ask your question — it might become the next article, or a conversation right away.

[Ask your question](/en/contact) [Read our approach](/en/aanpak)

---
Volledige (opgemaakte) versie: https://neuralex.nl/en/blog
