BlogAI search

How ChatGPT decides which pages to cite

Benjamin Libor10 min read

When ChatGPT answers with search, it decides whether to look things up, turns your one question into several searches, fetches a handful of pages, picks the passages that answer best and writes a reply with citations attached to the sentences those passages support. OpenAI has not published the ranking behind each step, but the shape of the process is visible in the product and in OpenAI's own documentation, and it points to three things a brand can control: being findable for the sub-questions, having pages that answer one thing clearly, and being present on the third-party pages that get fetched.

In short

  • ChatGPT answers from two sources: what the model learned in training, and what it fetches from the web when it decides to search.
  • A searched answer is built in steps: decide to search, split the question into several searches, fetch pages, choose passages, write, then cite.
  • Training knowledge still shapes which brands get named, even when ChatGPT searches, because the model forms its shortlist before and while it reads.
  • Citations change from day to day because the searches, the search results, the sampling and the model itself all change.
  • For a brand that means: rank for the sub-questions, build pages that answer one thing plainly, and be present on the comparison and review pages that get fetched.

ChatGPT can answer a question without touching the web. The model has read a very large amount of text during training and holds a compressed memory of it, including which companies exist, what they do and how they are described. Ask "what is a good scheduling tool for agencies" with search off and you get an answer written entirely from that memory, with no citations, and with the brands the model happens to associate with the question.

With search on, ChatGPT can run web searches and read pages before it writes. The answer then carries citations, and the brands it names come partly from what it read just now. OpenAI's documentation describes a dedicated crawler for this, OAI-SearchBot, separate from GPTBot, which gathers training data, and from ChatGPT-User, which fetches a page when a user asks for it directly. That separation is the clearest public statement of how the two sources are kept apart: blocking the training crawler does not stop your site appearing in search answers, and allowing the search crawler is what makes it possible.

ChatGPT decides on its own whether a question needs search, and what we can observe is that it searches for things that look time-sensitive, specific or comparative, which covers most buying questions. A user can also force it.

Step by step: how a searched answer is built

What follows is the process as it can be observed from the outside, cross-checked against what OpenAI has said. The exact weights are not public.

The model reads the question and judges whether its own knowledge is enough. "Who founded Fernwick" is probably a search; "explain what a webhook is" probably is not. Questions about prices, versions, recent events and "best" or "vs" comparisons are the ones that most reliably trigger a search.

2. Turn one question into several searches

A buying question is really a bundle of smaller ones, and ChatGPT runs several searches to cover them. This is often called query fan-out. ChatGPT sometimes shows the searches it ran in its activity panel, and when it does, they look like rewrites of the question from different angles: a general version, a version with a year, a version with a platform name, a version aimed at reviews.

This step is the one most marketers miss. You do not need to rank for the question as typed. You need to be in the results for the searches it fans out into.

3. Fetch the pages

The search results come from a web index. OpenAI's announcement of ChatGPT search said it draws on third-party search providers as well as content from publishing partners, and Microsoft Bing was named at launch. ChatGPT then fetches a selection of the results. How many is not published; the sources panel on an answer typically shows a handful to a few dozen pages, not all of which end up cited.

4. Choose passages

The model does not use whole pages. It looks inside each fetched page for the parts that bear on the question, and keeps those. A page that answers the question in its first paragraph is easy to use; a page that gets there after six hundred words of preamble may be fetched and still contribute nothing.

5. Write, then cite

The answer is written as prose, and citations are attached to the claims that came from a fetched passage. Claims that came from the model's own knowledge usually carry no citation at all, which is one way to tell the two sources apart when you read an answer: the named brand with no link next to it came from memory.

A worked example: one question, several searches

Suppose an agency owner asks ChatGPT: "What's the best scheduling tool for a 30-person agency that uses HubSpot?"

We cannot see the exact queries ChatGPT runs for that, but when the activity panel shows them, they tend to look like this set:

  • best scheduling software for agencies
  • scheduling tools with HubSpot integration
  • Slotwise vs Fernwick for teams
  • scheduling app reviews agencies
  • team scheduling tool pricing 2026

Each search returns its own results, and the pages fetched will be a mix: a "best scheduling tools for agencies" list from a marketing publication, HubSpot's own integrations directory, a Reddit thread where agency owners compare tools, a G2 category page, and one or two vendor pages that happen to rank for the HubSpot query. The answer will then name the brands that appear across those pages, with the HubSpot integration page and the comparison list cited next to the relevant sentences.

Fernwick, a made-up scheduling tool, is in this answer or not depending on three things: whether it ranks for any of those five searches, whether the pages that do rank mention it, and whether the model already knew it. If Fernwick has a HubSpot integration but no page that says so in its first line, it fails the second search. If the "best tools" list left it out, it fails the first. The question as typed never mattered.

Why training knowledge still shapes the answer

Searching does not make the model forget. When ChatGPT decides what to search for, it uses what it already believes about the category, so the fan-out queries can include competitor names it knows and leave out ones it does not. When it reads the fetched pages, it weighs them against what it already believes. And when the fetched pages disagree or are thin, it falls back on memory to fill the gaps.

The practical effect is a baseline. A brand that was well documented across the web before the model's training cut-off starts every answer with a seat at the table; a newer or quieter brand has to earn it through search every time. This is also why a brand can be absent from ChatGPT's answer and present in Perplexity's for the same question: different models, different memories, different fetched pages. How to get cited by Perplexity covers the differences.

You cannot edit the training data. You can make sure the text that goes into the next model describes you accurately and often, which is the same work as being present on the pages that get fetched.

Why the same question gives different citations on different days

Monday's answer cites a comparison list, Thursday's cites a Reddit thread, and a competitor absent on Monday is recommended on Thursday. Four things are moving.

  1. The fan-out differs. The searches ChatGPT runs are themselves generated by the model, so they vary from run to run. A different set of searches fetches a different set of pages.
  2. The search results differ. The underlying index updates constantly, and a page that was third last week may be eighth now.
  3. The writing differs. The model samples its words with some randomness. Given the same fetched pages, it may still name three brands one time and four the next.
  4. The product and the model change. OpenAI updates models and the search product without a changelog you can act on. A shift across all your prompts at once is usually this, not something you did.

Personalisation adds a fifth: a signed-in user with memory and chat history gets answers shaped by that history, so checking from your own account tells you little about what a stranger sees. This is why measuring AI visibility means a fixed set of prompts, in clean sessions, repeated on a schedule: one answer is one sample, and only the share across many samples is a measurement.

What this means for a brand

The mechanics point to a short list of work, roughly in order of leverage.

Be findable for the sub-questions

List the questions buyers ask, then list the searches each one would fan out into: the "best" version, the "with [platform]" version, the "vs" version, the reviews version, the pricing version. Those are keywords in the ordinary sense, and ranking for them in Bing and Google is how you get fetched. Many are long-tail and cheap to win because nobody else is treating them as a target.

Have pages that answer one thing clearly

For each sub-question, one page, with the answer in the first lines, a heading structure that mirrors the question, and facts stated plainly with dates. Integration pages, pricing pages and comparison pages are the ones most often fetched for buying questions and most often missing or vague. The page-level checklist for getting cited by ChatGPT goes through what a page needs.

Be present on the third-party pages that get fetched

The comparison lists, review platforms and forum threads that come back for the fan-out searches are the referees. Read the citations across your prompts, make a list of the sites that keep appearing, and work through them: complete the profiles, correct the entries, pitch the lists, and answer honestly in the threads. Why AI answers lean on Reddit, reviews and forums explains why these sites carry so much weight.

Let the crawler in

Check robots.txt for OAI-SearchBot and make sure nothing upstream, such as a bot-protection rule, is blocking it. OpenAI's documentation is explicit that this is the crawler for search answers. A page it cannot fetch cannot be cited, no matter how well it ranks.

Questions people ask

Does ChatGPT use Google or Bing for its searches?

OpenAI has said ChatGPT search uses third-party search providers alongside its own index and partner content, and Microsoft Bing was named when the feature launched in 2024. The exact mix at any given time is not published. In practice, ranking well in Bing and Google both help, and Bing's index is the one more often overlooked by small teams.

Can I pay to be cited by ChatGPT?

Not at the time of writing: there is no advertising product that places a brand inside ChatGPT's answers or citations. What you can buy is presence on the pages ChatGPT reads, such as a review-platform listing or a sponsored comparison entry, which works only if the page is one that gets fetched for the relevant searches.

How do I know if ChatGPT fetched my page?

Look in your server or CDN logs for the user agents OpenAI documents: OAI-SearchBot for search answers, ChatGPT-User for pages a user asked it to open, and GPTBot for training. A fetch is not a citation, but no fetch means no citation is possible, and the pages that are never fetched tell you where the crawler is stuck.

Why does ChatGPT name my competitor but not link to them?

A brand named without a citation usually came from the model's training knowledge rather than from a fetched page. It means the competitor is well established in the text the model learned from, and the fetched pages did not contradict that. You counter it on the fetched pages, where a clear, current entry can earn both a mention and a link.

Written by Benjamin Libor, founder, echo.

Something wrong or out of date? Write to hello@echo-aeo.com and we'll fix it.

See what AI says about your brand. Free, in about a minute.

The free AI score checks how ready your website is for AI assistants. Echo itself tracks your customers' questions across the assistants every day and turns the answers into actions.