AI visibility tracking measures whether AI assistants – ChatGPT, Perplexity, Gemini, Google’s AI Overviews mention or cite your brand when someone asks a question you’d want to win. It’s the AI search equivalent of rank tracking: a fixed set of prompts, run across the AI platforms on a schedule, with brand mentions, citations and share of voice logged over time. You’ll also hear it called LLM visibility or AI brand monitoring, same thing.
Before you read on: I run an AI automation agency and I built RankX AI, an AI visibility platform. I have an obvious interest in you deciding this category matters. So this guide does the one thing a vendor guide usually won’t it teaches you to run the whole thing yourself, free, in enough detail that you may never need a tool. Mine included.
This post sits under our guide to generative engine optimisation that covers how to earn AI citations; this one covers how to measure them.
What does AI visibility tracking actually measure?
One number sits at the centre: the share of your tracked prompts where your brand appears in the AI’s answer. Everything else hangs off it.
| Layer | What it measures | The question it answers |
|---|---|---|
| Traditional SEO tracking | Rankings, clicks, impressions | Are we winning search traffic? |
| AI visibility tracking | Mentions, citations, share of voice in AI answers | Do the AI engines know we exist? |
| AI performance tracking | AI-referred sessions, leads, revenue | Is any of it worth money? |
Five terms, brand mentions, citations, share of voice, sentiment, positioning get muddled constantly, and the muddle costs buyers money:
- A mention is your brand named in the answer text.
- A citation is your page linked as a source.
- Share of voice is how often you appear against a named competitor set “AI share of voice” on most dashboards.
- Sentiment is whether the answer describes you positively, neutrally or with caveats.
- Positioning is how it frames you market leader, budget option, niche specialist which shapes buyer expectations before they ever reach your site.
Most platforms roll the first three into a single AI visibility score. That’s fine as a headline as long as you can see the per-engine numbers underneath it.
These numbers move independently, and the gap between them is where the insight is. Semrush’s 2026 AI Visibility Index 126 million US prompts analysed between January and April 2026 found that 62% of all AI citations are “ghost citations”: the engine uses a brand’s content as a source without naming the brand at all. You can be doing the work and getting none of the credit, and only citation-level tracking shows you.
The same index found the AI models barely agree with each other. ChatGPT cites around 15 sources per response; Gemini cites three. Only 36 brands held top-100 visibility across all four tracked engines every month. A tool that reports one blended score across engines is averaging away the only interesting information.
How do you track AI visibility yourself, step by step?
You can run a full AI visibility audit yourself, free, before spending anything. We call it the 20-Prompt Test, and it’s the first thing we run for any client who asks about AI visibility. Here is the whole method, not the summary.
Step 1: Build the prompt set
Write 20 questions a real buyer would ask before choosing you. Not your keywords questions, phrased the way people actually talk to an assistant. Split them across four groups:
| Group | How many | Example (for an answering service) |
|---|---|---|
| Category questions | 6 | “What’s the best out-of-hours answering service for a small business?” |
| Comparison questions | 5 | “AI answering service vs a human call centre which is better for a letting agent?” |
| Problem questions | 6 | “We keep missing tenant calls at night what are my options?” |
| Local or industry questions | 3 | “Answering service for UK care homes that can handle safeguarding calls” |
The problem questions matter most and get skipped most. Buyers early in research don’t know the category name yet they describe the pain. If you only test category prompts, you’re measuring the market’s endgame and missing where buyers actually start.
Keep this set fixed. Changing prompts between runs destroys the trend, and the trend is the product.
Step 2: Run it clean
Ask each prompt in ChatGPT, Perplexity and Google (for AI Overviews and AI Mode). Free tiers are fine. Two hygiene rules, because both can quietly rig your results:
- Strip personalisation. Use a logged-out or incognito session, and in ChatGPT turn memory off. If you’ve been researching your own brand all month, a logged-in assistant will flatter you.
- Note your location. AI Overviews vary by country. If you sell in the UK and the US, run the Google prompts from both (a VPN does it) or at least record which side you measured.
Step 3: Log it
A spreadsheet is enough. One row per prompt per platform, with these columns:
| Column | What goes in it |
|---|---|
| Date | Run date you’ll compare month to month |
| Prompt | The exact wording, never paraphrased |
| Platform | ChatGPT / Perplexity / Google AIO |
| Mentioned? | Y/N – your brand named in the answer |
| Cited? | Y/N – your page linked as a source |
| Who appeared | Every brand the answer named, in order |
| Framing | One line: how it described you (or the winner) |
Keep the AI responses themselves, paste them into a second sheet or export the chat. When a number moves next month, you’ll want to read what changed in the answer, not just that it changed.
Step 4: Score it
Three numbers fall straight out of the sheet:
- Visibility rate = prompts where you’re mentioned ÷ total prompts. This is your AI visibility score, minus the vendor branding.
- Citation rate = prompts where your page is linked ÷ total prompts.
- Share of voice = your mentions ÷ all mentions of you-plus-named-competitors.
Expect low numbers first time. Most SMEs score 0–10% visibility on their first run that’s normal, and it’s the baseline that makes month two meaningful.
Step 5: Repeat monthly, and read trends, not runs
One caveat that vendors under-explain: AI answers are not deterministic. Ask the same prompt twice and you can get different brands back. A single run is weather; the pattern across 20 prompts, monthly, is climate. That’s also why a fixed 20-prompt set run monthly beats five prompts run daily, you want breadth across buyer questions, not noise-sampling on a few.
When we ran this on our own market, the answer was uncomfortable. AI Overviews appear on queries representing 73% of US search volume for after-hours answering services, and 41.8% of the UK equivalent (our own SEMrush market analysis, 5 August 2026) and at the time we were mentioned in 1 of the 143 AI Overviews we tracked. That number, not a vendor’s pitch deck, is why we take this category seriously.
If the test shows AI answers barely appear for your market’s questions, stop there. You’ve spent an afternoon and saved a subscription.
Which AI platforms should you track?
Not all of them. Track where your buyers actually ask, and use each platform’s citation behaviour to decide effort:
| Platform | Why it matters | Citation behaviour (Semrush 2026 index) |
|---|---|---|
| Google AI Overviews / AI Mode | Sits inside the search results your buyers already use the largest reach by default | Appears on a large share of commercial queries; sources skew to established sites |
| ChatGPT | The biggest standalone assistant audience | ~15 sources per answer; leans on community and reference sites, Reddit and Wikipedia heavily |
| Perplexity | Smaller but research-heavy users strong for B2B shortlists | Cites by design on every answer, so citation tracking is cleanest here |
| Gemini | Android and Workspace reach | ~3 sources per answer from a narrow pool , Wikipedia, Reddit, YouTube so single citations carry more weight |
| Copilot / Claude | Track only if your buyers live in Microsoft 365 or developer tooling | Optional for most SMEs |
Start with Google AI Overviews plus ChatGPT. Add Perplexity if your sales cycle involves someone building a shortlist at a desk. The 615x variation in citation rates between platforms that Semrush measured means a platform you ignore isn’t “roughly the same” as the ones you track, it’s a different country.
How do AI visibility tools actually work?
Worth understanding before you buy, because it explains both the pricing and the disagreements.
Every tool in this category does a robot version of the 20-Prompt Test: it runs your prompt set against each AI platform on a schedule through official APIs, browser automation, or both from clean, un-personalised sessions, usually several times per prompt to smooth out the non-determinism. It stores every answer, then extracts mentions, citations, sentiment and competitor names from the text. The dashboard is a spreadsheet with better graphs and no manual typing.
Three practical consequences:
- Two tools will give you two different scores for the same brand. Different prompt phrasings, different sampling schedules, different locations. Neither is lying; they’re measuring different samples. Pick one tool and judge the trend, never compare absolute scores across tools.
- Ask vendors whether they query the API or the consumer product. API responses can differ from what a logged-in human sees, model versions and retrieval settings aren’t always identical. A vendor who can’t answer that question plainly hasn’t earned the subscription.
- Pricing scales with prompts × platforms × frequency, because every check is a paid model call on their side. That’s why entry tiers cap prompts hard, and why daily tracking of 500 prompts costs enterprise money. Most SMEs need neither.
Which AI visibility tools are worth looking at in 2026?
Declared interest again: RankX is mine. The table reflects how the market actually tiers by price and buyer, from published pricing as of August 2026, prices in this category move quickly, so check before you rely on them. Most offer a free trial, and several run a free AI visibility checker as a lead magnet, fine for a one-off look, too thin for tracking.
| Tool | Rough price | Built for | Honest note |
|---|---|---|---|
| Otterly.AI | From ~$29/mo | Small teams starting out | Cheapest credible entry; prompt limits arrive fast |
| Peec AI | From €70/mo | Mid-market and agencies | Strong tracking depth for the price |
| Semrush AI Toolkit | ~$99/mo add-on | Teams already on Semrush | Convenient; shallower than the specialists |
| Ahrefs Brand Radar | ~$199/platform | One-off research | Better for a research exercise than daily tracking |
| Profound | From $99/month | Large brands, reputation-led | The category leader; priced like it |
| Scrunch AI | $250/month | Large brands | Enterprise workflows, enterprise sales process |
| RankX AI | $49/month | SMEs and agencies | Mine. Judge it against the checklist below like the rest |
Whatever you shortlist, apply the same four checks:
- Per-engine reporting, not a blended score. The engines disagree too much to average.
- Prompt-level history. Prompt tracking is the product; if you can’t see how one prompt’s answer changed over time, you can’t trust the trend line.
- Citation tracking separated from mention counting. Because of the ghost-citation problem above.
- An export path. Scoring methods in this category are young and mostly opaque; if you can’t pull your data out including plain visibility reports you can hand to a client or a board, you can’t check the maths or leave.
If a vendor can’t explain their scoring in plain English in a demo, keep looking. I say that as someone who has to pass that test too.
Is AI visibility tracking worth it for your business?
The demand side is real. Adobe’s data reported by Digital Commerce 360 on 17 June 2026 shows AI-referred traffic to US retail sites up 138% year on year in May 2026, converting 54% better than non-AI traffic, and up 1,324% since Adobe began tracking it in October 2024. Meanwhile Semrush’s index found 45% of marketing leaders can’t measure their brand’s AI visibility at all, and only 9% have tooling that covers all the metrics they care about. The gap between “this channel is growing” and “we can see our brand visibility in it” is the whole case for AI search visibility tracking.
A caution on the famous forecast, though: Gartner predicted in February 2024 that traditional search volume would fall 25% by 2026. It’s 2026. It hasn’t. Search is shifting shape rather than shrinking which argues for measuring AI answers alongside rankings, not for abandoning SEO tracking.
When it’s worth paying for:
- Buyers research you before contact high-consideration services, software, anything with a shortlist.
- The 20-Prompt Test showed AI answers on a meaningful share of your buyer questions.
- A competitor keeps appearing where you don’t.
When it isn’t and this is the part that costs me money to write:
- Your leads come from referrals, repeat business or outbound. An AI engine’s opinion of you is a curiosity, not a channel.
- The 20-Prompt Test came back quiet. Track manually each quarter and spend the budget on content instead.
- You haven’t fixed basic measurement yet. If you can’t see which enquiries come from where today, an AI visibility dashboard adds a second thing you can’t act on.
One more honest caveat: this category is roughly two years old. Scoring methods aren’t standardised, tools appear and vanish, and any of the prices above may be wrong within a quarter. Start small, keep your prompt set stable, and judge tools on whether the data changed a decision not on the dashboard.
How do you improve AI visibility once you’ve measured it?
Measurement without a fix is a subscription to bad news. Four levers move these numbers, in rough order of effort-to-impact for an SME:
1. Make your pages citable. The Muck Rack study covered by Nieman Lab in July 2025 over a million AI citations analysed, found journalistic content took 27% of all citations, rising to 49% on queries needing recent information. The pattern generalises: engines cite specific, dated, attributable material. A page that says “answering services cost £0.90–£2 per minute (provider rate cards, August 2026)” gets lifted; a page that says “costs vary” doesn’t exist as far as an AI answer is concerned.
2. Make them extractable. Answer the question in the first 60 words. Use real tables, not images of tables. Add an FAQ with self-contained 40–60-word answers and FAQ schema. Every one of those is a primitive an engine can lift without needing the surrounding page, full method in our guide to writing content AI engines actually cite.
3. Build the third-party footprint. Remember where ChatGPT and Gemini actually look: Reddit, Wikipedia, YouTube, review sites, trade press. The 62% ghost-citation figure cuts both ways, engines lean heavily on pages about you that you don’t own. Genuine answers in the subreddits your buyers read, a presence on the comparison sites the engines already cite, one well-placed trade article these often move AI mentions faster than anything on your own domain.
4. Keep your entity consistent. Same brand name, same one-line description of what you do, everywhere it appears site, directories, LinkedIn, press. Models build their idea of “who is this company” from the aggregate; if half the web calls you an agency and half a software product, the answer engines hedge, and hedged brands don’t get recommended. Our GEO optimization strategies guide covers this and the schema work that supports it.
Can you see the AI traffic you’re already getting?
Partly and it’s worth wiring up before you pay for anything, because it’s the closest thing to revenue attribution this channel has.
In GA4: AI assistants that link out show up as ordinary referrers. Build a segment (or custom channel group) matching referrers containing chatgpt.com, perplexity.ai, gemini.google.com and copilot.microsoft.com, and you have an AI-referral report from data you already own. Watch conversion rate on that segment, not just sessions Adobe’s finding that AI-referred visitors convert 54% better held in their sample; check whether it holds in yours.
In Search Console: harder, and there’s a trap. Google doesn’t label AI Overview impressions separately, and AI Mode inflates your impression counts with synthetic “fan-out” sub-queries that can never produce a click. When we dug into our own collapsed click-through rate, roughly 96% of our reported impressions turned out to be fan-out queries our real human CTR was 1.3–1.9%, not the 0.07% the dashboard implied. If you judge content on raw Search Console CTR in 2026, you’ll make bad calls.
Frequently asked questions
What is AI visibility tracking?
AI visibility tracking is the practice of running a fixed set of buyer questions through AI assistants ChatGPT, Perplexity, Gemini, Google AI Overviews on a schedule, and logging whether your brand is mentioned, whether your pages are cited as sources, and how you compare with competitors over time.
How is it different from SEO rank tracking?
Rank tracking measures your position in a list of links. AI visibility tracking measures presence in a synthesised answer, where there are no positions you’re either in the answer, cited beneath it, or absent. The two disagree often enough that you need both views.
Which AI platforms should I track first?
Google AI Overviews and ChatGPT first between them they cover search reach and the largest assistant audience. Add Perplexity if your buyers build research shortlists. Gemini, Copilot and Claude are worth tracking only when your audience demonstrably uses them.
How much do AI visibility tools cost?
As of August 2026: entry tools start around $29/month (Otterly.AI), mid-market platforms run roughly €70/mo by prompt volume (Peec AI), SEO-suite add-ons cost $99–199/month (Semrush, Ahrefs), and enterprise platforms like Profound reach five figures annually. Prices in this category move quickly.
Can I track AI visibility for free?
Yes. Run the 20-Prompt Test: 20 real buyer questions through ChatGPT, Perplexity and Google monthly, logging brand mentions and citations in a spreadsheet. Any free AI visibility tracker will be prompt-capped anyway; the manual audit costs nothing and tells you whether paid tracking is justified.
How often should I check AI visibility?
Monthly is enough for most SMEs AI answers move, but not daily, and a stable monthly prompt set gives you a trend you can trust. Move to weekly only if AI answers are already a measurable source of enquiries or you’re in a reputation-sensitive market.
If you’d rather someone ran the 20-Prompt Test for you and told you honestly whether tracking is worth paying for including “no” book a discovery call. It’s the same test either way; the only difference is who does the typing.
