Brand sentiment in AI answers: what it is and how to track it
Benjamin Libor9 min read
Brand sentiment in AI answers is how an answer characterises your company when it names it: recommended, neutral, with reservations, or negative, together with the adjectives and caveats attached to the name. It matters more than social sentiment because the answer is the recommendation: a caveat such as "expensive" or "steep learning curve" sits between the reader and your sign-up page. You track it by classifying each mention in each answer on a simple scale, keeping the quotes, and reporting the recurring phrases with counts and a trend rather than a single score.
In short
- Sentiment in an AI answer is the framing attached to your name, from "the best choice for small teams" to "powerful but expensive", not whether the answer is positive about the topic.
- It matters more than sentiment in social listening because the reader is asking for a recommendation and the answer is the recommendation.
- The caveats come from review sites, forum threads, comparison articles, your own pricing page and old news, and each source needs a different fix.
- Measure it by classifying each mention on a four-point scale, by hand at first, and aggregating per prompt, per engine and per competitor, with the quotes kept.
- Report the recurring phrases with counts and a trend, never one number.
What sentiment means in an AI answer
Brand sentiment in an AI answer is the way the answer characterises your brand at the point where it names it: the verdict (recommended, neutral, with reservations, negative) and the words the verdict is built from.
Take a prompt like "best scheduling tool for a small agency". An answer might say:
- "Fernwick is the simplest option here and the one most small agencies end up with." That is recommended.
- "Fernwick also supports round-robin booking." That is neutral: a factual mention with no verdict.
- "Fernwick is powerful, though several users mention a steep learning curve and pricing that climbs quickly with seats." That is with reservations: named, half-endorsed, with two caveats a buyer will remember.
- "Fernwick has had reliability complaints and is probably not the right pick for a growing team." That is negative.
Sentiment is about your name, not the topic. And notice that the adjectives do the work. "Powerful", "simple", "affordable", "enterprise-grade", "dated", "expensive" and "steep learning curve" are the units of sentiment, and they are what you will end up counting.
Why it matters more here than in social listening
In social listening, a negative post is one voice among many, and the reader can scroll past it. An AI answer is different in three ways.
The answer is the recommendation. The person asked "which should I pick" and the assistant replied. There is no feed, no second opinion and often no click. If the reply says "expensive", the reader has been told your price is high before they have seen it.
Caveats are presented as consensus. "Users report a steep learning curve" reads as a summary of everyone's experience, even when the model found it in one review and one forum comment.
It repeats. The same caveat can appear in the answer to the same prompt every day, for every person who asks, on every engine that read the same source.
So the right frame is conversion, not reputation. "Expensive" attached to your name in a buying question is a pricing-page problem and a sales-objection problem at the same time.
Where the framing comes from
AI answers do not invent sentiment. They compress what they read, and the sources are predictable. Knowing which one produced a phrase tells you where to fix it.
| Source | What it tends to produce | Example phrase |
|---|---|---|
| Review sites (G2, Capterra, Trustpilot, app stores) | Pros-and-cons language, star-rating summaries | "highly rated for ease of use, but reviewers mention slow support" |
| Forum threads (Reddit, Hacker News, Stack Overflow, niche communities) | Blunt verdicts from a few vocal users | "most people in the thread recommend switching to…" |
| Comparison articles and "best of" lists | Category labels and one-line positioning | "a lightweight option for small teams" |
| Your own pricing and product pages | Price-tier language, missing features inferred from silence | "starts at a higher price point than…" |
| Old news, funding announcements, changelogs | Outdated facts dressed as current | "recently raised funding and is expanding into…" |
Two deserve extra attention. Forums are over-weighted by most engines because they look like candid, first-hand experience; the article on why AI answers lean on Reddit and reviews explains the mechanics. And your own pages are a bigger source of caveats than most teams expect: if your pricing page lists the enterprise tier first, the answer may describe you as enterprise-priced.
How to measure it
You do not need a model to start. You need a prompt set, the answers, and a spreadsheet.
1. Collect the answers
Run your tracked prompts against each engine you care about (ChatGPT, Perplexity, Gemini, Claude, and Google's AI Overviews where they appear) and save the full text of each answer. If you have not chosen prompts yet, how to choose the prompts to track covers that first.
2. Classify every mention
For each answer, find every place your brand and each tracked competitor is named, and give each mention one label on a four-point scale:
- Recommended: the answer endorses the brand for the question asked.
- Neutral: the brand is named as an option or for a fact, with no verdict.
- With reservations: the brand is named with a caveat ("but", "though", "however", "some users report").
- Negative: the answer steers the reader away.
Next to the label, paste the sentence that contains the mention and underline the adjectives and caveats. The quote is the valuable part.
Doing this by hand for 30 prompts across four engines takes an afternoon the first time and an hour a week after that. Once you trust your own labels, a model can apply the same scale; check a sample of its labels each month so the scale does not drift. Tools built for this do the classification for you; in Echo it sits on the Prompts page next to position and citations.
3. Aggregate three ways
- Per prompt: the share of answers to that prompt where you are recommended, neutral, with reservations or negative. This finds the questions where your framing is weakest.
- Per engine: the same split by ChatGPT, Perplexity, Gemini and Claude. Engines read different sources, so a caveat that only appears in Perplexity usually traces to a page Perplexity cites and the others do not.
- Per competitor: the same scale applied to each rival on the same prompts. This is what turns sentiment from a mood into a comparison: if a competitor is "easy to use" in four out of five answers and you are "powerful but complex", that is the positioning gap you are fighting.
4. Count the phrases
Group the underlined adjectives and caveats into a short list: "expensive", "steep learning curve", "limited integrations", "great support", "good for small teams". Count the answers each appears in this month. That list is the heart of the report.
What to look for
Once the sheet exists, patterns appear fast.
The recurring caveat. A phrase that shows up in a quarter of your answers on a quarter of your prompts is not noise, it is your reputation in AI search. Usually it traces to two or three sources that every engine reads.
The competitor who is always "easy". When one rival gets the same flattering adjective across prompts and engines, they did not get lucky. Somewhere there are review titles, a comparison page or a forum thread that repeat that word, and the engines learned it. You can learn where from a source gap analysis.
A wrong fact dressed as an opinion. "Fernwick lacks a mobile app" is a fact, and a checkable one. "Fernwick is less suited to mobile-first teams" is the same error wearing an opinion. Both come from the same missing page. Treat them as a fact to fix, not a sentiment to accept.
What to do about each
Each kind of finding has a fix, and the fixes belong in an actions list with an owner and a date, not in a slide.
- A wrong fact: correct it on your own site first, in plain text, on the page the engines already cite. Then correct the third-party page where the error lives, by asking the author, updating your review-site profile or replying in the thread. The full process is in finding and fixing the facts ChatGPT gets wrong.
- A true but unaddressed objection ("expensive", "complex"): answer it on your own pages with proof. A pricing page that states what a typical small team pays, a "set up in an afternoon" page with a real timeline, and a comparison page that puts the caveat in context give the engines something to quote instead of the complaint.
- A competitor's flattering adjective: get customer stories and reviews that say the opposite about you, using the word you want. If you want "easy", ask customers who found it easy to say so on G2 and in a story on your site.
- A forum thread that keeps getting cited: be present in it. A plain, non-promotional reply from a named person at the company, correcting a fact or adding context, becomes part of what the engines read.
Allsite, a made-up website builder, finds "limited e-commerce" attached to its name in a third of answers about online shops, traced to a two-year-old review and its own features page, which never mentioned payments. It updates the features page, adds a short page on selling online with two customer examples, and asks three e-commerce customers for reviews. The phrase fades from the answers as the engines re-read the pages; there is no guarantee on timing, but the sequence is the right one.
How to report it
Resist the single score. "Sentiment: 72" tells nobody what to do and moves for reasons nobody can explain. Report this instead, monthly, on one page:
- The four-point split for your brand and each competitor across all prompts, with last month's split beside it.
- The recurring phrases table: each phrase, the number of answers it appeared in this month and last month, the prompts it appears on, and the source it traces to.
- Three quotes: the worst caveat, the best endorsement and the one that changed most since last month.
- The actions taken and planned, each tied to a phrase.
A table like this is enough:
| Phrase | Answers this month | Last month | Prompts | Likely source | Action |
|---|---|---|---|---|---|
| "steep learning curve" | 14 | 19 | 6 | G2 reviews, one Reddit thread | onboarding page, 4 new reviews |
| "good for small teams" | 22 | 17 | 9 | own site, comparison article | keep |
| "expensive" | 9 | 9 | 4 | pricing page, Capterra | pricing page rewrite |
The numbers are illustrative; the shape is what a leadership team can read in two minutes and a marketer can act on in a week.
Questions people ask
What is brand sentiment in AI search?
It is how an AI answer characterises your brand when it names it: recommended, neutral, with reservations or negative, plus the adjectives and caveats attached. It measures the framing of your name, not whether the answer is positive about the topic.
How is AI answer sentiment different from social media sentiment?
A social post is one voice in a feed. An AI answer is the recommendation itself, presented as consensus, repeated for everyone who asks the same question. A caveat in an answer therefore acts like a sales objection the reader hears before talking to you.
Can I track sentiment in ChatGPT by hand?
Yes. Run your prompts, save the answers, label every mention of your brand and your competitors on a four-point scale, and keep the sentence as evidence.
How do I fix negative sentiment in AI answers?
Find the source of each phrase from the answer's citations, then fix the fact on your own pages, answer true objections with proof, collect reviews and customer stories that use the words you want, and reply in the threads that keep being cited. Track the phrase count monthly.
Written by Benjamin Libor, founder, echo.
Something wrong or out of date? Write to hello@echo-aeo.com and we'll fix it.