TL;DR Most brands have no idea what AI chatbots say about them, and a single check in ChatGPT tells you almost nothing. AI is wildly inconsistent, with less than a 1% chance of giving you the same answer twice. This guide gives you a repeatable workflow: set up monitoring, measure sentiment the right way, interpret the results, and act on them. Along the way you'll learn why sentiment in AI is more than just "positive or negative," how to separate mentions from citations, and why most AI sentiment is actually neutral. By the end, you'll be able to build your own tracking system and measure sentiment with useful accuracy.
Why Tracking What AI Chatbots Say About Your Brand Matters Right Now
Here's a number that should stop you in your tracks: 37% of consumers now start their searches with AI instead of Google. And by 2028, an estimated $750 billion in consumer spending will flow through AI-powered search. This isn't a future problem. It's happening right now.
The catch? AI is describing your brand whether you know it or not. Roughly 88% of businesses are invisible in ChatGPT, but the ones that do show up are being described based on sources they don't control. Around 85% of brand mentions in AI come from third-party pages, not your own website.
Here's where it gets tricky. There's a big difference between a mention and a citation. AI can use your content as a source without ever naming your brand. In fact, 62% of AI citations are "ghost citations," where the source is linked but your brand is never mentioned. If you only track mentions, you'll miss a huge part of the picture.
Despite all this, only 22% of marketers actively track their AI visibility. That's the opportunity. The people who start now get a head start that's hard to catch up on.
What You Need Before You Start Tracking AI Brand Mentions
Before you dive in, get your basics ready. Here's your checklist:
- A list of brands and products to track. Start with your own, then add competitors later.
- A set of prompts that real customers actually ask. Keep them short and conversational. Short queries generate 30 to 50 times more brand mentions than long, structured ones.
- A choice of AI platforms to monitor. Start with ChatGPT and Gemini, which together cover roughly 88% of the market.
- A monitoring tool. This can be anything from a manual spreadsheet to a dedicated platform like Semly.
Now, a warning. Manual checking is unreliable. AI is non-deterministic, meaning there's less than a 1% chance you'll get the same answer twice from the same prompt. A single measurement is basically worthless. You need repetition and consistency.
Also, know when this approach won't work. If you have fewer than about 50 prompts in your panel, your results will be too random to draw any real conclusions. More on that in a moment.
Step 1 – Set Up Your AI Brand Monitoring System
Start by choosing which platforms to monitor. ChatGPT holds about 78% of the market, with Gemini at roughly 10%. Those two are your foundation. Add Perplexity, Claude, and Copilot as you grow.
Next, design your prompt panel. Use short, conversational questions that mimic how people actually talk to AI. Think "best [category] for [need]," "what's a good [product type]," or "who makes the best [product]." Aim for a minimum of 50 prompts, but push toward 100 or more for better results.
Frequency matters too. Measure daily, on a 24-hour cycle, always in fresh sessions. AI changes its answers even with the same prompt, so a single check is a waste of time. You need repeated measurements over time.
As for tools, you can start with a manual spreadsheet, but it's labor-intensive and error-prone. Platforms like Semly automate the whole cycle. Semly runs a 24-hour measurement cycle with repeated, standardized queries across fresh sessions, tracking results per prompt and per model. That kind of consistency is exactly what you need for reliable data.
Step 2 – Measure Sentiment Accurately (Not Just Once)
Sentiment in AI isn't just "positive or negative." Real-world data shows the distribution looks like this: 28% endorsement, 41% neutral, 19% cautious or hedged, and 12% hallucinated attributes. Most of what AI says about brands is neutral, and that's normal.
For measurement, keep it simple. Use a three-point scale: +1 for a positive recommendation, 0 for a neutral mention, and -1 for a negative or warning tone. That's enough to get started.
Here's the critical part: never rely on a single measurement. With less than a 1% chance of getting the same answer twice, one result means nothing. For a statistically meaningful picture, you need around 384 prompts for a ±5% margin of error at 95% confidence. Practically, that means running each prompt 3 to 5 times per cycle and averaging the sentiment across those runs.
Citation quality adds another layer. Think of it on a 1 to 5 scale, from a simple name drop at level 1 to being the exclusive source at level 5. Five mentions at level 1 can be worth less than a single mention at level 4. Semly uses exactly this approach: a +1/0/-1 sentiment scale, a Citation Quality score from 1 to 5, and a Response Stability metric that shows how consistent your results are across runs.
Step 3 – Interpret Your Results and Spot the Real Problems
First, know what normal looks like. Most mentions will be neutral, around 41%. Some will be endorsements, about 28%. A chunk, around 19%, will be cautious or hedged. Don't panic when you see neutral mentions. That's the baseline.
Now the red flags. The 12% hallucination rate is the big one. AI can attribute wrong features, prices, or names to your brand. That's a real reputational risk. Also watch for a high share of negative sentiment in specific prompts, and for big swings in sentiment between platforms. Each platform is its own ecosystem, with only about 11% overlap between ChatGPT and Perplexity.
Source analysis is where the real insight lives. Check which domains AI cites when it describes your brand. If your sentiment is bad, the sources are probably the problem, especially since 85% of mentions come from third-party pages.
Run through these checkpoints: Is your sentiment stable between cycles? Does it differ drastically between platforms? Are the hallucinations repeating? If you can answer those three, you'll know whether your measurement is trustworthy.
Step 4 – Take Action Based on What You Find
If your sentiment is neutral with accurate attributes, that's a realistic goal. Don't chase enthusiastic recommendations. AI rarely gives them.
If you're seeing hallucinations, dig into the sources. Which pages is AI citing when it generates wrong information? Fix your own content, build a knowledge base with correct data, and report bad information where you can.
If competitors are winning, look at which sources support them. Study their content strategy, then build better, more citable material of your own.
Monitoring is just the beginning. Platforms like Semly, with its Leon AI Agent, go further by generating GEO-optimized content and publishing it to a knowledge base automatically. That closes the loop from insight to improvement.
And remember, this is continuous work. Around 40 to 60% of cited domains change every month. Monitoring has to be ongoing, not a one-time audit.
Common Mistakes When Tracking AI Brand Sentiment (and How to Fix Them)
1. Checking once in ChatGPT and treating it as "the truth." AI is wildly inconsistent. Fix: run each prompt 3 to 5 times, on a daily cycle, in fresh sessions.
2. Using only long, detailed prompts. Short, conversational queries generate 30 to 50 times more brand mentions. Fix: design your panel around natural customer questions.
3. Ignoring the mention vs. citation distinction. AI can cite your page without naming your brand. Fix: track both metrics separately.
4. Treating "AI visibility" as one number. Each platform is its own ecosystem, with only 11% overlap between ChatGPT and Perplexity. Fix: measure per platform.
5. Skipping source analysis. You won't know where AI gets its information. Fix: always check sources with every mention.
6. Using too small a prompt sample. Fifty prompts isn't enough for meaningful conclusions. Fix: aim for at least 100, and push toward 400+ for a ±5% margin of error.
7. Doing a one-time audit instead of continuous monitoring. Around 40 to 60% of citations rotate monthly. Fix: make monitoring an ongoing process.
What to Do Next – Building a Long-Term AI Visibility Strategy
Once you've mastered the basics, scale up. Expand your prompt panel to 400+ for statistical significance. Add more platforms like Claude, Perplexity, and Copilot. Start tracking competitors and measuring your share of voice.
Then move from monitoring to action. Build a knowledge base, optimize content for AI citation, and use tools that automate the work, like Semly's Leon AI Agent. Shift your focus from visibility metrics to business metrics: AI-referred traffic, conversions from AI, and AI's share of registrations or sales.
One honest note: if your category has low AI adoption, focus on traditional SEO and social listening first, and treat AI monitoring as an early warning signal.
If you want a zero-friction start, grab a free AI visibility report at report.semly.ai. It gives you a multi-model analysis across ChatGPT, Gemini, and Perplexity in about two minutes, with no payment required. It's the easiest way to see where you stand before you build your full system.