How to choose an AI visibility tool: 12 questions to ask
Benjamin Libor10 min read
Choose an AI visibility tool by asking twelve questions before you look at a demo: which engines it covers and whether Google's AI Overviews is included, how often prompts run, whether you choose the prompts and markets, whether you see full answers and citations, whether competitors run on the same prompts, whether it shows sentiment and position, finds wrong facts, turns findings into actions, connects to GA4 and Search Console, fits a budget you understand, fits a small team, and leaves your data yours. Then run a two-week trial with your own prompt set before deciding.
In short
- Engine coverage, including Google's AI Overviews, and how often prompts run are the two questions that decide whether the numbers mean anything.
- You should see the full answers and the citations, not a score; a score you cannot trace to a sentence cannot be acted on.
- Competitors must be measured on the same prompts, with sentiment and position, or share of voice is not comparable.
- A tool earns its keep by turning findings into owned actions and connecting visibility to visits in GA4 and Search Console.
- Decide on a two-week trial with your own prompts, your own competitors and your own team doing the work.
Before the questions: what the tool is for
An AI visibility tool runs a set of questions (prompts) against AI assistants on a schedule, records whether your brand and your competitors are named, where, how, and which pages the answers cite, and shows the change over time. What is AI visibility defines the measures these tools report.
The questions below are in rough order of how much they change the answer. Ask the first six of every vendor; the rest depend on your team.
The 12 questions
1. Which engines, and is Google's AI Overviews included?
Why it matters. Each engine reads different sources and gives different answers, and Google's AI Overviews appear on a large share of informational searches. A tool that covers ChatGPT and nothing else measures one engine, not AI search.
A good answer. ChatGPT, Perplexity, Gemini and Claude at least, with Google's AI Overviews (and AI Mode where offered) as a named part of the product, not a roadmap item.
2. How often do prompts run, and are answers repeated?
Why it matters. AI answers vary from run to run. One answer per prompt per week is an anecdote. A tool that runs each prompt several times, or daily, and reports the share of runs in which you appear is measuring something; a tool that shows one answer is showing a screenshot.
A good answer. A stated schedule (daily is common), a stated number of repetitions or a visibility figure defined as a share of runs, and the dates visible on every number. Ask whether you can trigger a run after you ship a fix.
3. Can you choose your own prompts, markets and languages?
Why it matters. The prompts are the product. Suggested prompts from a vendor's generator are a start, but your buyers ask specific questions in their own words and often in their own language. If you cannot edit the set, add a market or run in German, the tool is measuring someone else's funnel.
A good answer. Unlimited editing of prompts within your plan, prompts grouped by topic or stage, a location and language per prompt or per group, and suggestions you can accept or reject. How to choose the prompts to track is worth reading before a demo so you can test with real ones.
4. Do you see the full answers and the citations, not just a score?
Why it matters. A visibility percentage tells you that something changed. The answer text tells you what the engine said, and the citations tell you which page said it first. Without both, you cannot find the source of a caveat or a wrong fact, and you cannot brief anyone to fix it.
A good answer. Every answer stored in full with its date and engine, every brand mention highlighted, every cited URL listed and linked, and a view of citations across prompts: which domains and pages are cited most, for you and for competitors.
5. Are competitors tracked on the same prompts?
Why it matters. Share of voice only means something when every brand is measured on the same questions, at the same time, on the same engines. Some tools track competitors on a separate, smaller set, or count any mention anywhere, which distorts the comparison.
A good answer. You name the competitors, the tool measures them on your full prompt set within a stated limit, and the share-of-voice table shows the same columns for every brand.
6. Does it show sentiment and position?
Why it matters. Being named is not the same as being recommended. "Fernwick is powerful but expensive" is a mention with a caveat; being listed fifth of six is a mention with little weight. Brand sentiment in AI answers explains what to look for.
A good answer. A sentiment label on each mention with the quoted sentence as evidence, position within the answer, and both aggregated per prompt, per engine and per competitor. Be wary of a single sentiment score with no quotes behind it.
7. Does it find wrong facts?
Why it matters. Engines get prices, features, locations and company facts wrong, and present the error with confidence. These are the cheapest problems to fix and the most damaging to leave, and finding them by reading hundreds of answers by hand is slow.
A good answer. A way to enter the facts an engine should know (pricing, integrations, who it is for) and a report of the answers that contradict them, with the source page where the wrong fact appears.
8. Does it turn findings into actions with owners?
Why it matters. The failure mode of visibility tools is a beautiful dashboard that nobody acts on. A small team needs a short list: fix this page, write that page, get onto this third-party site, each with proof, an owner and a date.
A good answer. Actions generated from the data with the evidence attached (the answer, the citation, the gap), a field for the owner and due date, and a before-and-after once the action ships.
9. Does it connect to GA4 and Search Console?
Why it matters. Visibility is an input. The outcome is visits, sign-ups and deals. Tying the two together, which pages the answers cite and which pages the AI referrals land on, is how you prove the work mattered and decide what to do next.
A good answer. A read-only connection to GA4 and Search Console, AI referral sessions split by assistant and shown next to the cited pages, and a monthly report that puts visibility, citations and visits on one page.
10. Does it stay inside a budget you understand?
Why it matters. The cost scales with prompts, engines, repetitions, markets and competitors, and some pricing pages hide the multiplication. A plan that looks affordable at 20 prompts on one engine can be a different number at 60 prompts on five engines in three markets.
A good answer. A price you can compute yourself from the number of prompts, engines and runs, a stated limit on each, no per-seat charge for a small team, and a clear statement of what happens when you reach a limit.
11. Does it fit a small team, or does it need an agency?
Why it matters. Enterprise platforms assume someone will maintain the prompt set, configure the dashboards and interpret the output. If your marketing team is two people, that someone is one of them.
A good answer. Set-up in an afternoon: the tool suggests prompts and competitors from your website, you edit them, and the first results arrive the next day. A named person on your side can run it in an hour or two a week. The routine for a two-person team is a good test: if the tool makes that routine easier, it fits.
12. Is your data yours?
Why it matters. Your prompt set, the answers, the citations and the actions are a record of your market over time.
A good answer. Export of prompts, answers and results as CSV at any time, an API for the same data, a plain statement that your prompts and answers are not used to train anyone's models or shared with other customers, and a clear answer on what happens to it when you cancel.
The kinds of tools, and who each suits
The market sorts roughly into three kinds. The descriptions are about the category, not about any vendor's current features, and none of them is right for every team.
Enterprise platforms, such as Profound, are built for large brands and agencies: many markets, many brands under one roof, deep reporting, enterprise integrations and a sales process to match. They suit a company with a dedicated SEO or content team that will maintain a large prompt set, and a budget that treats this as a line item rather than a decision.
Mid-market tools, such as Peec AI, focus on the visibility and citation data itself: prompts, engines, competitors, share of voice and sources, presented clearly for a team that wants a precise instrument for this one job. They suit a team with a marketer who owns AI search as a channel and acts on the data elsewhere.
Lighter tools that bundle visibility with the rest of a small team's marketing work, such as Echo, treat AI visibility as one data source among several (SEO, mentions, competitor changes, performance from GA4) and put the emphasis on the to-do list that comes out of it: actions, briefs, pages and reports in one place. They suit a 10 to 200 person company where two or three people run all of marketing and would rather have one list than four dashboards.
A reasonable rule: if you have a person whose job is AI search, look at the first two kinds; if AI search is a slice of one person's week, look at the third. Whichever you pick, the twelve questions apply equally.
Run a two-week trial with your own prompts
Do not decide on a demo. A demo shows the vendor's best brand on the vendor's best prompts. Instead:
- Bring your own prompt set. Twenty to forty real questions from sales calls, support tickets and Search Console, in the language and market your buyers use. Load the same set into every tool you trial.
- Name the same competitors in each. Five to eight who appear in the answers, not the ones on your battlecard.
- Let it run for two weeks. Long enough to see the day-to-day variation and to tell whether the share-of-voice number is stable enough to report.
- Do one real action in each tool. Find a wrong fact or a missing page, fix it, and see whether the tool helps you notice the change.
- Have the person who will own it do the trial. If they cannot run the weekly review in the tool in 90 minutes by week two, the tool is wrong for the team regardless of its data.
- Export at the end. Question 12, tested.
Then compare. The numbers will not match across tools, because engines vary and tools sample differently. What matters is whether each tool's citations lead to real pages and whether your team actually used it.
Questions people ask
What should an AI visibility tool track?
Whether your brand and your competitors are named in answers to a set of prompts across ChatGPT, Perplexity, Gemini, Claude and Google's AI Overviews, with position, sentiment and the pages cited, on a schedule, so you can see change over time and trace each finding to a source.
Is one AI visibility score enough?
No. A score tells you that something moved, not what or why. Ask for the full answers, the highlighted mentions and the cited pages behind every number, so that each change can be traced to a page and fixed.
How long should I trial an AI visibility tool?
Two weeks with your own prompts, your own competitors and the person who will run it doing the work. That is long enough to see how much answers vary, whether the numbers are stable enough to report, and whether the tool helps you find and act on one real fix.
Written by Benjamin Libor, founder, echo.
Something wrong or out of date? Write to hello@echo-aeo.com and we'll fix it.