BlogAI search

What is AI visibility? A plain definition, and how to measure it

Benjamin Libor10 min read

AI visibility is the share of AI answers to your customers' questions that name your brand. If a buyer asks ChatGPT, Perplexity, Gemini, Claude or Google's AI Overviews which tool to use and your company is in the answer, you are visible; if it is not, a competitor is. It is measured by running a fixed set of questions on a schedule, across engines, and counting how often you appear. It is not a ranking and it is not traffic.

In short

  • AI visibility is the percentage of answers, to a fixed set of customer questions, that mention the brand by name.
  • It is not a ranking: there is no position 1, only named or not named, and several brands can be named in one answer.
  • Answers change from run to run, so one question asked once tells you little; a measurement is a set of prompts, repeated on a schedule, across engines.
  • The first measurement is most useful for what it shows you are missing: the questions where you are absent, the competitors named instead, and the sites the answers cite.

What AI visibility means

AI visibility is a share. The numerator is the number of answers that name your brand. The denominator is the number of answers you looked at: every tracked question, on every engine, on every run. Ask twenty questions on four engines once a week and the denominator for that week is eighty.

A brand is named when the answer mentions it by name, whether or not it links to your website. A link without the name is a citation; a name without a link still counts, because the reader now knows you exist.

AI visibility = answers that name the brand ÷ all answers to the tracked questions, over a period, usually shown as a percentage.

The questions matter more than the maths. "What is Fernwick?" will name Fernwick every time and tells you nothing. "What's the best scheduling tool for a 30-person agency?" is the question a buyer asks before they know you exist, and that is where visibility is won or lost. The short version of how to choose the prompts to track: mostly unbranded questions, in the customers' phrasing, grouped by topic.

It is not a ranking

Search results are a list, and a page is first, fifth or nowhere. An AI answer is prose: it might name three tools in one sentence, recommend one in the next and never mention a fourth. There is no first place to win; there is being in the answer or not. The nearest thing to a rank is position: whether you are the first brand named, the second, or an afterthought.

It is not traffic

AI referral traffic, the visits GA4 attributes to chatgpt.com or perplexity.ai, is real and worth tracking, but it measures something else. Most AI answers are read and never clicked, and a buyer who sees your brand recommended three times and then types your name into Google arrives as branded search, not as an AI referral. Visibility measures whether you were in the conversation; traffic measures whether someone clicked out of it (how to see AI traffic in GA4).

The metrics that sit next to it

Visibility answers "how often are we in the answer?". Four more metrics answer the questions that follow.

MetricThe question it answersHow it is worked out
AI visibilityHow often are we in the answer?Answers naming the brand ÷ all answers
Share of voiceWhen brands are named, how much of that is us?Mentions of the brand ÷ mentions of all tracked brands, yours and the competitors'
PositionWhen we are named, how prominently?Where the first mention falls: first brand named, second, later
SentimentWhen we are named, how?Whether the answer recommends, describes neutrally or warns
CitationsWhich pages does the answer lean on?The links attached to the answer, grouped by site

Share of voice is the competitive view: two brands can both have 40% visibility, but if one is always named alongside five others and the other usually alone, their shares differ a lot (share of voice, explained). Position matters because an answer that opens with Slotwise and ends with "Fernwick is another choice" has named both, but not equally. Sentiment is the difference between being recommended and being named as the thing people complain about. Citations are the raw material for action: they tell you which pages the engine read before it wrote, and therefore where you need to be.

Why the same question gives different answers

If you have asked ChatGPT the same thing twice and got two different lists, you have seen the central measurement problem. There are four causes, and none of them is going away.

  1. Sampling. The models generate text with a degree of randomness. The same prompt on the same day produces slightly different wording, and sometimes a different set of brands.
  2. The engines' own searches. ChatGPT with search, Perplexity, Gemini and AI Overviews run web searches and read pages before they answer. Which pages they fetch depends on the search results of the moment, and when the sources shift, the named brands shift with them (how ChatGPT decides which pages to cite).
  3. Personalisation and context. A signed-in user with chat history, a saved location and memory gets a different answer from a clean session. Your own checks in your own account are the least representative sample there is.
  4. Model and product updates. The engines change models and retrieval logic without notice. A drop one week may be the engine, not you.

What that means for measuring

A single answer is one draw from a distribution, so measuring means sampling often enough that the share you report is stable. Four rules follow.

Track a fixed set of prompts. Decide the questions once and keep the wording identical. Changing a prompt resets its history; add a new one instead. Twenty to fifty prompts is enough for a first programme.

Run them on a schedule. Weekly is the usual cadence for a small team; daily if you are changing things fast. The point is comparable runs, not maximum runs.

Run them across engines. The engines disagree more than most people expect, because they read different sources and start from different models. A brand can be well known to Perplexity and invisible to Gemini. Until you know where your buyers ask, track all of them.

Repeat. Each run of each prompt is one sample. A prompt that names you in two of four runs this month and three of four next month has not clearly improved; one going from one in twelve to nine in twelve has. Report over periods rather than run against run, and use clean sessions: no history, no account, a fixed country.

A worked example

Fernwick is a made-up scheduling tool for agencies. Its two-person marketing team tracks 20 prompts across four engines, weekly. After a month, five of the prompts look like this. Each cell is the number of runs, out of four, in which Fernwick was named; the figures are invented for the example.

PromptChatGPTPerplexityGeminiAI Overviews
best scheduling tool for agencies2/44/41/40/4
scheduling software that integrates with HubSpot0/41/40/40/4
Slotwise alternatives for teams3/43/42/41/4
how to stop no-shows for client calls0/40/40/40/4
Fernwick vs Slotwise4/44/44/42/4

Across these five prompts and 80 answers, Fernwick is named in 36, so its visibility is 45%. That one number is the least interesting thing in the table. What the team can actually read:

  • Perplexity knows them; Gemini and AI Overviews barely do. Perplexity reads the web heavily and Fernwick's pages and reviews are being fetched. Gemini and Google's answers lean on different sources, probably comparison lists where Fernwick is missing.
  • The HubSpot prompt is a gap with an obvious fix. Buyers ask it, competitors are named, and Fernwick has a HubSpot integration but no page that says so plainly. That is one page to build this week.
  • The no-shows prompt names nobody. The answers list tactics, not tools. Being absent there is a content opportunity: a guide that AI answers can quote.
  • The branded prompt is inflating the average. "Fernwick vs Slotwise" will almost always name Fernwick. Report branded and unbranded prompts separately, or the headline number hides the gaps.

What a good first measurement looks like

A first measurement is a baseline, so its job is to be repeatable rather than impressive.

  • 20 to 50 prompts, grouped by topic, mostly unbranded, written in the customers' own words.
  • Your brand plus three to five competitors, so that every answer is also scored for who else is named.
  • All the engines your buyers might use. Drop one later if it never matters.
  • Four weeks of weekly runs before you draw conclusions, so sampling noise settles.
  • A changelog of what changed on the website and in your marketing during those weeks, so you can later connect a movement to a cause.

What you do with it

The baseline is a list of problems ranked by how many buyers meet them. Work through three lists.

The questions where you are absent. For each, ask whether you have a page that answers it directly, whether the AI crawlers can reach it, and whether anyone else on the web says you are an answer to it. Usually at least one of the three is missing.

Who is named instead. The competitors in your answers are the real competitive set, which often differs from the one in your sales deck. Note which ones appear on which engines and whether they are recommended or merely listed. This is the starting point for source gap analysis: the pages that cite them and not you.

Which sites get cited. The citations across all your prompts usually cluster on a short list: a review platform, two or three comparison lists, a Reddit thread, a few publications. Those sites are where the engines form their opinion of the category, and being accurately present there moves visibility faster than most changes to your own site.

Each item becomes an action with an owner and a before-and-after check: a page to build, a listing to claim, a fact to correct, a thread to join. Re-run the prompts, and read the change over weeks, not days.

Questions people ask

Is AI visibility the same as share of voice?

No. Visibility is the share of answers that name you at all. Share of voice is your share of all brand mentions within those answers, so it depends on how many competitors are named alongside you. Both are useful; visibility tells you whether you are in the conversation, share of voice tells you how much of it you own.

How many prompts do I need to track?

Enough to cover the questions buyers ask at each stage, usually 20 to 50 for a product with one or two use cases. More prompts add coverage; more runs per prompt add reliability. For a small team, fewer prompts run every week beats hundreds run once.

Can I measure AI visibility by hand?

Yes, for a few weeks. Use a clean browser session with no account, ask each prompt on each engine, and log which brands and which sites appear. The limits are time and consistency: a check you skip one week leaves a hole in the series.

Does a citation count as visibility?

Only if the answer also names the brand. A link to your page under an answer that never says your name is a citation without a mention: good for traffic, a sign the engine trusts the page, but the reader has not learned who you are. Treat the gap between the two as a copy problem on the cited page.

How often should I run the prompts?

Weekly is enough to see trends and cheap enough to sustain. Run daily for a few weeks around a launch or a big content push, when you want to see whether a change registered. Whatever you choose, keep it constant, because a change in cadence changes the sample.

Written by Benjamin Libor, founder, echo.

Something wrong or out of date? Write to hello@echo-aeo.com and we'll fix it.

See what AI says about your brand. Free, in about a minute.

The free AI score checks how ready your website is for AI assistants. Echo itself tracks your customers' questions across the assistants every day and turns the answers into actions.