> Bron: https://neuralex.nl/en/blog/hoe-kiezen-ai-modellen-bronnen
> Ranking well in Google is not the same as being cited in an AI answer. How ChatGPT Search, Perplexity and Gemini retrieve, assess and link sources to specific claims — and what that means for GEO.

[Back to Insights](/en/blog)

GEO ·20 August 2026 ·7 min read

# How do ChatGPT, Perplexity and Gemini decide which source to cite?

Ranking well in Google is not the same as being cited in an AI answer. How ChatGPT Search, Perplexity and Gemini retrieve, assess and link sources to specific claims — and what that means for GEO.

ChatGPT, Perplexity and Gemini do not simply cite a page because it ranks number one in Google. For search-based questions, these systems typically retrieve a set of potential sources first, then determine which passages are useful for supporting the answer. Relevance to the specific query, accessibility, freshness, reliability and how well a passage supports a concrete claim all play a role.

Their exact ranking formulas are not public. However, OpenAI, Perplexity and Google publish enough information to reconstruct the general process. For companies trying to improve their visibility in AI-generated answers, one distinction matters in particular: **AI visibility is not only about being found, but also about being selected as evidence for a specific part of an answer.**

## An AI answer usually starts with retrieval

A language model can answer many questions from knowledge acquired during training. For current, specific or verifiable information, systems such as ChatGPT, Perplexity and Gemini can also retrieve information from the web — a process that conceptually resembles [retrieval-augmented generation (RAG)](/en/blog/rag-uitgelegd). Relevant information is retrieved first, the language model receives it as context, and the final answer is generated from it.

The user normally sees only the final response. Behind the scenes, several steps may take place: interpreting the question, generating one or more search queries, retrieving candidate web pages, selecting relevant passages, combining information from different sources, generating the answer, and linking sources to the claims they support. The stages between retrieval and generation explain why a page may be found during search but still not appear as a cited source in the final answer.

## How ChatGPT selects sources

OpenAI explains that ChatGPT Search can rewrite a user's question into one or more targeted search queries. After the initial results, the system may run additional searches to fill gaps in the available information, after which the answer is generated from the retrieved material and may include inline citations.

That means a question such as "what Dutch rules apply when processing personal data locally with an AI model?" does not necessarily become one literal search query. ChatGPT may break the information need into searches around GDPR, local AI inference, personal data, processors and applicable regulations. A page is therefore not competing only for the exact wording of the original question — it can also become a candidate source for one of the subqueries ChatGPT generates while researching the answer.

OpenAI also states that ranking in ChatGPT Search depends on multiple factors, and that no method allows a publisher to guarantee a top placement. One technical prerequisite, however, is that the site is accessible to OAI-SearchBot and that its infrastructure does not block OpenAI's crawler. For GEO, this leads to an important conclusion: a strong page does not necessarily have to be the best broad article on a topic — a clear passage that answers one specific question precisely can be more valuable.

## How Perplexity selects sources

Perplexity was designed from the outset as an answer engine. It searches the web in real time, retrieves information from multiple sources and uses that information to generate a direct answer with citations. With Pro Search, Perplexity can run multiple search operations and collect information from sources including websites, academic publications, forums and videos.

Perplexity also uses explicit source labels: domains can, for example, be marked as Government, Academic or Trusted, based on factors such as author attribution, correction policies and the separation between editorial content, advertising and opinion. That evaluation happens at domain level — it does not mean every individual page on a domain is automatically correct, nor is it proof that a label automatically produces a higher ranking. It does show that source quality is explicitly considered. For companies, the implication is that publishing a large volume of content is not enough: a clearly identifiable organisation, transparent authorship and solid sourcing make a site easier to evaluate as a credible source.

## How Gemini selects sources

Google describes the process particularly clearly in its documentation for Gemini with Google Search grounding. When Google Search is available as a tool, Gemini first analyses the prompt and determines whether a search would improve the answer. If so, it generates one or more search queries, sends them through Google Search, processes the results and produces a grounded response — and can record exactly which source supports which segment of the generated text. Citation therefore does not operate only at document level; it can be linked down to passage level.

For Google's broader AI search experiences, such as AI Overviews and AI Mode, Google also describes a technique called query fan-out: multiple searches across different subtopics at once, with additional supporting pages discovered while the answer is being generated. The underlying principle is the same across products: one user question can result in several retrieval queries. Google stresses that no special "AI SEO" markup is required — a page mainly needs to be indexable and eligible to appear in Google Search with a snippet.

## A citation is not the same as a traditional ranking

This is probably the most important difference between SEO and GEO. Traditional SEO ranks a list of documents; a generative answer requires a system to also decide which information inside those documents is useful for constructing the answer.

Imagine ten websites explain what RAG is. One has the strongest domain and ranks highly in Google. A smaller specialist site, however, contains an unusually clear paragraph explaining the difference between vector retrieval and [reranking](/en/blog/wat-is-reranking). If someone specifically asks why reranking is necessary, that smaller page may be the more attractive source — not because the article as a whole ranks better, but because one passage matches the information need particularly well.

## Why the same question can produce different citations

A citation should not be treated as a permanent position comparable to a traditional number-one ranking. The sources used can change because the query is phrased differently, the system generates different subqueries, newer information becomes available, or language and location affect retrieval. That is why testing a single prompt once and concluding that a company is "visible" or "not visible" in AI is of limited value — [GEO](/en/blog/geo-geciteerd-door-ai) should be measured across a set of realistic questions.

## What makes a page easier to cite?

Public documentation from the three platforms reveals no secret formula for forcing citations, but it does reveal a fairly consistent pattern: a page first needs to be technically accessible to the retrieval system, its content must be sufficiently relevant to the query, and the system must be able to extract usable support for the answer. For topics where reliability matters, the credibility of the source becomes more important too.

A practical review therefore includes questions such as:

-   Can the relevant crawler access, render and index the page?
-   Does the text answer the query clearly, without requiring the system to infer it from five paragraphs?
-   Does the page cover specific subquestions an AI system might search for independently?
-   Are important claims supported by primary or otherwise reliable sources?
-   Is it clear who publishes the information and what their expertise is?
-   Does the page contain original information or analysis not found in identical form across dozens of other sites?
-   Has outdated information been updated where freshness matters for the query?

None of these factors guarantee citations — they mainly increase the likelihood that a page is considered both retrievable and usable during retrieval.

## GEO is more than SEO for ChatGPT

It is tempting to treat GEO as traditional SEO with a few extra optimisations, but that does not fully reflect how generative search systems work. SEO remains part of the foundation: a crawler must be able to reach the site, and the content must be strong enough to enter the candidate set in the first place. On top of that sits a second problem: can the generative system actually use the page to construct an answer? Content that answers specific questions unambiguously, contains self-contained passages and supports claims with evidence helps here. A company does not need a separate article for every imaginable prompt — a well-designed topical cluster can cover many retrieval queries when its individual sections each address a clear information need.

## Conclusion

ChatGPT, Perplexity and Gemini do not publish the full formulas behind their citation choices. What they do disclose points to a broadly similar process: the question is interpreted, potentially split into several search queries, relevant information is retrieved, and sources are selected to support specific parts of the generated answer. For AI visibility, this means that "ranking highly" alone is too narrow a goal. A page must be discoverable, but also relevant, clear, credible and citable enough to actually appear as a source in an AI-generated answer.

Want to know where you stand?

## Does your company show up in ChatGPT, Perplexity and Gemini answers?

The Agentic Search Optimizer measures it objectively — 15+ checks, a score and a roadmap. Or ask your question directly through the contact form.

[View the Optimizer](/en/showcase/agent) [Ask your question](/en/contact)

---
Volledige (opgemaakte) versie: https://neuralex.nl/en/blog/hoe-kiezen-ai-modellen-bronnen
