Getting found by AI assistants: what GEO actually is
Classic SEO optimizes for a results list. GEO optimizes for being cited as a source in a generated answer. What that means technically and editorially — including checklists.
More and more people no longer put their question to a search engine, but to an assistant. They get back no list of ten links, but one formulated answer — with a handful of sources underneath it. That shifts the question of what your site should be optimized for. No longer: am I in that list? But: am I used in that answer?
What changes when the answer replaces the results list?
A results list is a menu: ten options, the user clicks. A generated answer is a summary: the model has already chosen, read and rephrased. At most, the user clicks through to check something. Two things follow from that.
First, the unit of visibility is no longer the page, but the passage. A model does not take over a whole page; it takes the few sentences that carry its answer. Second, the selection is stricter. Out of ten links a user can pick the messy but usable one; a model picks the passage it can repeat without risk. Ambiguity is not forgiven, it is skipped.
What exactly is GEO — and how does it differ from SEO?
Classic SEO optimizes for a results list: the goal is a position. GEO (Generative Engine Optimization) optimizes for being cited as a source in a generated answer: the goal is a passage a model is willing to take over.
That difference is not a nuance — it determines how you write. SEO rewards coverage: more pages, more terms, more internal links. GEO rewards precision: a claim that is correct, can be read on its own and is verifiable. The two do not exclude each other. A page that scores well in classic search is usually easy for a model to find as well. But the last step — actually being cited — has its own requirements, and those are editorial at least as much as they are technical.
Put practically: SEO is about being found, GEO is about being used.
What has to be right technically?
Before a model can cite anything, it has to be allowed to fetch your page and be able to understand it. That is the technical layer. These are the four means we use ourselves on neuralex.nl.
robots.txt — do you grant AI crawlers access?
AI crawlers are different user agents than the classic search bots. A robots.txt that only covers Googlebot says nothing about them — and many default configurations and firewall rules even block them implicitly. We allow them explicitly: GPTBot (OpenAI), PerplexityBot, ClaudeBot (Anthropic) and Google-Extended. This is the first check, because if it fails, the rest of your work is invisible.
llms.txt — what is the short version of your site?
An llms.txt in the root of your domain is a readable file that describes in
structured text what your site is, which pages matter and where which information sits.
Where robots.txt states what a bot is allowed to do, llms.txt states
where a bot should look. It saves a model the work of combing through an entire
site to find out what you do.
schema.org — what kind of page is this, in machine language?
Schema.org markup (as JSON-LD in the page) states explicitly what a model would
otherwise have to infer: this is an article, this is the author, this is the publication
date, this is a service, this is an organization. It is not a ranking trick but a
translation: you put the meaning a human derives from the layout into a form a machine
does not have to guess at. This page itself contains a BlogPosting block —
you can view it in the source.
Agent cards via DNS — how does an agent find you without reading your site?
The newest layer is the least known: machine-readable cards published via DNS, so that an agent can discover what a domain offers and where it can go — without parsing HTML first. We use this on neuralex.nl. If you want to see what such a setup looks like in practice, it is worked out in the agent showcase.
How do you write in a way that can be cited on its own?
This is the part most sites skip, and where the biggest gain sits. The structural lesson from our own work:
An AI assistant cites a passage that holds up on its own. A paragraph that only has meaning in the context of the previous three paragraphs does not get cited.
Concretely that means: no paragraphs that open with "this means that…" or "as described above". Put the subject in the sentence itself. Repeat a term rather than using a reference. A block of three to five sentences that you can cut out of the page and that still holds up — that is the building block.
The second lesson is about the content of those blocks. Vague claims do not get cited, verifiable ones do. "We are the market leader" is useless to a model: it cannot be checked and cannot be repeated without risk. "X costs Y and works like this" is citable, because it is a verifiable assertion. Numbers, methods, definitions and explicit conditions survive the jump into a generated answer; superlatives do not.
A practical editorial checklist per page:
- H2s are phrased as the question a user actually asks.
- Under every H2, the answer is in the first two sentences, not only in the conclusion.
- Every definition appears once, explicitly, in one block — not scattered across the text.
- No paragraph starts with a reference to something that came before it.
- Claims contain a number, a method or a condition — or they go.
- What you cannot substantiate, you do not write down.
How do you measure whether it works?
GEO is largely checkable, and that is the pleasant part. You do not have to wait for a ranking report: you can run most of the checks on your own site today.
- Fetch
/robots.txtand check whether GPTBot, PerplexityBot, ClaudeBot and Google-Extended have access. - Check whether
/llms.txtexists and whether its contents still match the site. - Validate your JSON-LD: is it valid schema.org, and does it describe what the page really is?
- Read a random paragraph separately from the rest — does it still hold up without the context before it?
- Search your pages for vague claims without a number, method or condition.
We have bundled this into our own measurement tool that runs a site through 15+ checks and attaches a score to it. In one measured case, a site went from score 36 to 85+. The interesting part is not the score itself but what explains it: the largest share of the gain sat in things that were simply missing — no llms.txt, blocked crawlers, missing or incorrect markup. What you measure usually turns out to be straightforward to repair. How we approach a project like that is set out under our services.
What GEO is not
GEO is not a trick and not a new form of keyword stuffing. Cramming a block of text full of terms you hope a model will latch onto is counterproductive: it makes the passage less suitable to take over, because it no longer reads as an assertion that can hold up. Nor is there a hidden instruction you can put in your HTML to convince a model — and if there were, it would break with the next model version.
What is left is less exciting and more durable: make sure crawlers are allowed in, that your structure is explicit, and that your assertions still hold up separately from their context. That is largely just good writing, with a technical layer underneath it. A model that cites you cites you because your passage covers the answer — not because you talked it into it.
In summary
GEO is optimizing your site to be cited as a source in a generated answer, instead of appearing as a position in a results list. Technically that comes down to accessibility and explicit structure: a robots.txt that allows AI crawlers, an llms.txt, schema.org markup and agent cards via DNS. Editorially it comes down to two rules: write blocks that hold up on their own, and make claims verifiable. Both parts are measurable — and that is exactly why this is work and not a gamble.