Skip to content
RankX AI

AI Visibility Measurement

How to Track Brand Mentions in AI Search

To track brand mentions in AI search, run a fixed set of buyer questions against each AI assistant on a schedule and record, for every answer, whether your brand was named, where it ranked, which sources were cited and how it was described. One answer is a single draw, so only the aggregate is worth reading.

By Asif Syed, Founder & CEOPublished 14 minute read

On this page
  1. How do you track brand mentions in AI search?
  2. What’s the difference between an AI mention and an AI citation?
  3. Why can’t you set an alert for AI answers?
  4. How many checks do you need before the number means anything?
  5. How do you tell “not mentioned” from “we couldn’t check”?
  6. How do you stop a tracker counting mentions that aren’t mentions?
  7. Can you track how an assistant talks about your brand?
  8. Which AI assistants can you track, and what does each expose?
  9. Five questions to ask any AI brand-monitoring tool
  10. What do you do when the mention is wrong?
  11. How RankX AI records a brand mention

Tracking brand mentions in AI search means running a fixed set of prompts against each AI assistant on a schedule and reading every answer that comes back. There is no index to query and no alert to subscribe to. The only way to know what an assistant says about your brand is to ask it and record what it said.

A working setup has four parts, and the order matters:

  1. A prompt set. The questions your buyers actually ask, not your keywords. A keyword is what someone types into Google; a prompt is a whole question, and prompt research is a different exercise from keyword research.
  2. A schedule. Every prompt asked repeatedly, so a result can be read as a trend rather than an anecdote.
  3. A record per answer. Whether the brand was named, where it sat among the brands listed, which sources the answer cited, and which competitors appeared beside it.
  4. An honest denominator. How many checks ran, over what period, and how many produced no readable answer at all.

Most guides stop after the third part. The fourth is the one that decides whether the first three produce a number worth acting on, and it is where this article spends its time. Getting started is easy. Knowing whether your tracker is telling you the truth is not.

What’s the difference between an AI mention and an AI citation?

An AI mention is an answer that names your brand in its text. An AI citation is an answer that links one of your pages as a source. They are different events, they are recorded separately, and a brand-monitoring setup that conflates them will misread its own data.

They come apart constantly, in both directions. An answer can cite your page while recommending a competitor, which reads as a win in a citation report and is a loss in the market. An answer can name your brand from the model’s training with no link at all, which a link-based tracker never sees.

For brand monitoring, the mention is the event that matters. Being named in the answer is what a reader takes away; a citation is a link in a list most readers never open. Pick the AI mention as your primary metric and treat citations as evidence of how the mention was earned.

Being cited still pays, and the size of it is measurable. Seer Interactive found organic click-through rate fell 61 per cent on queries carrying an AI Overview. But brands cited inside the Overview earned 35 per cent more organic clicks and 91 per cent more paid clicks than uncited brands on those same queries. That is the argument for tracking both, and for never merging them into one figure.

It is also why an unlinked brand mention is worth more in AI search than it was in classic SEO. The assistant doesn’t need a link to name you.

Why can’t you set an alert for AI answers?

You can’t set an alert on an AI answer because the answer isn’t a document. Google Alerts, Mention and Brandwatch all watch a corpus of published pages: a page exists at a URL, a crawler finds it, a match fires. An assistant’s answer has no URL, is generated at the moment it is asked, and is never published anywhere a crawler can reach.

Three consequences follow, and each breaks a habit carried over from classic brand monitoring:

  • There is no backfill. A page published last March can be found today. An answer given last March was never stored by anyone and cannot be recovered. Monitoring starts the day you start asking, which is the argument for starting before you need the data.
  • Coverage is a choice, not a given. A web monitor sees whatever it crawls. An AI monitor sees only the questions it thought to ask, so your prompt set *is* your coverage, and a question nobody put to the assistant has no answer at all.
  • The same question doesn’t give the same answer. Asking twice is two samples, not one confirmed reading.

That third point is not a rounding error. Ahrefs tracked more than 43,000 keywords carrying at least sixteen recorded AI Overviews each. It found an AI Overview has a 70 per cent chance of changing between observations, that 45.5 per cent of its cited sources are entirely new when it does, and that the average one is rewritten every 2.15 days.

Assistants behave the same way. Vismore’s 750-response audit across five engines in March 2026 found the same prompt asked three times returned a meaningfully different brand list in 38 per cent of runs. SparkToro and Gumshoe put the stricter version at under a 1 in 100 chance that two runs of one prompt return the identical brand list.

So the question isn’t whether to sample repeatedly. It’s how many samples you need, and that has an arithmetic answer rather than an opinion.

How many checks do you need before the number means anything?

More than almost anyone recommends. Published advice on this ranges from quarterly to daily and is, as far as we can find, asserted every single time rather than derived. It is derivable, and the answer is uncomfortable.

Start with the base rate, because it is the input everything else needs and almost nobody publishes it. Measured across 394 prompt and run units in RankX AI production on 27 August 2026, 346 of them, or 87.8 per cent, contained no answer mentioning the brand at all. Not a failed check. A real answer, read, with the brand absent. A typical mention rate in that data sits near 12 per cent.

A mention rate is a proportion, so its margin of error is the standard one: plus or minus 1.96 times the square root of p(1-p)/n, where p is the mention rate and n is the number of checks. At a 12.2 per cent mention rate that gives:

Checks in the window

95% confidence interval

Smallest change you can actually detect

25

plus or minus 12.8 points

18.1 points

50

plus or minus 9.1 points

12.8 points

100

plus or minus 6.4 points

9.1 points

200

plus or minus 4.5 points

6.4 points

400

plus or minus 3.2 points

4.5 points

800

plus or minus 2.3 points

3.2 points

1,600

plus or minus 1.6 points

2.3 points

One check is one prompt put to one assistant once, so the arithmetic is prompts multiplied by assistants multiplied by runs. The widely repeated advice of thirty prompts, checked monthly, on a handful of assistants lands somewhere near 100 to 150 checks a month. Read the table: that carries roughly plus or minus 6 to 9 points of noise, and needs a swing near 9 points before a month-on-month move is distinguishable from nothing happening.

Here is why that matters more than it sounds. In our own data, a rival’s share of voice moved 8.1 points with nothing changing but the check schedule. At the sample size the standard advice produces, a move that large is inside the noise band. You would not be able to tell that artefact from a real competitive shift, and most reporting does not try.

Two honest caveats, because a table like this gets quoted. The maths assumes checks are independent, and they are not quite: the same prompt repeated correlates with itself, so real intervals are a little wider than shown. And a mention rate far from 12 per cent shifts the numbers, most at 50 per cent and least near the extremes. Treat the table as a floor on how much data you need, not a ceiling.

The practical read: 40 prompts across 5 assistants, run weekly, is about 200 checks a week and is roughly where week-on-week reporting starts being defensible. Pool four weeks and you are at 800 checks, where a 3-point move is real. Below about 100 checks in a window, report the direction and the count, and don’t quote a percentage to one decimal place.

How do you tell “not mentioned” from “we couldn’t check”?

You store them as different values. This is the single most important design decision in any brand-mention tracker, and collapsing the two is how a provider outage renders as your brand disappearing from AI search.

Four states are needed, not two:

State

What it means

Mentioned

The answer named the brand, and the record can point at where.

Not mentioned

An answer came back, it was read, and the brand is not in it. This is a real finding.

Not assessed

The check did not run, or ran and produced no readable verdict. This is not a zero.

Pending

The check is in flight and the answer has not returned.

The third state is not a rare edge case. In RankX AI production data about 12 per cent of scheduled checks fail, and the worst single prompt fails 54 per cent of the time. Assistants rate-limit, time out, refuse, and occasionally return something no parser can read.

The fix lives in the arithmetic rather than the chart: a mention rate must divide by the checks that produced a verdict, not by the checks that were attempted. One measurement found 96 checks on a single assistant of which only 31 carried a verdict. Dividing by 96 rather than 31 would have reported a mention rate around a third of its true value and presented a guess as a measurement. Where nothing carries a verdict, the honest output is no number at all, not a zero.

So when you evaluate a tool, ask this first: show me a day when a provider was down, and show me what the dashboard said that day. If the chart dipped, the tool is measuring its own reliability and labelling it your visibility.

From RankX AIAI VisibilitySee how often AI assistants name your brand.See your mention rate

How do you stop a tracker counting mentions that aren’t mentions?

You bound what the tracker is allowed to record, in code, rather than trusting a language model’s judgement about what it just read. The failure mode isn’t the model missing a mention. It’s a model, or a naive string match, confidently recording mentions that never happened.

Two live examples from RankX AI data show the shape of it. Both are false positives generated by the question itself:

  • A tracked competitor named Leeds, on a Leeds web-design project whose prompts read “best web design agencies in Leeds”. It matched on 26 of 26 runs, every check ever run, while genuine competitors matched on one to three. The answers contained the word because the question did.
  • A tracked competitor named Automation Agency, on a project whose prompts read “which AI automation agency should I hire”. It matched on 288 of 575 runs, catching ordinary prose like “if you’re hiring an AI automation agency in the UK”, which names no company at all.

Neither is fixable with a stoplist, because the offending words are perfectly good brand names on a different project. “Leeds” is distinctive for a brand that isn’t in Leeds. The rule that does work is contextual: a term that is ambient to the check, meaning it appears in the prompt itself or is the brand’s own city or market, carries no signal when it turns up in the answer.

There is a deliberate trade-off in that rule and it is worth stating. On a comparison prompt such as “how does X compare to Acme?”, Acme is ambient, so a mention of Acme in the answer does not count. That loses a real mention. It is still right: a model naming a brand it was explicitly told to discuss is not evidence of organic visibility, and counting it inflates exactly the prompts a vendor would most like to look good on.

Can you track how an assistant talks about your brand?

Yes, and the honest version is narrower than most tools imply. Sentiment in AI answers is worth tracking, but only when it is anchored to text that is genuinely about your brand, and only when “we can’t tell” is an available answer.

The naive approach sends the whole answer to a model and asks for a label. It produces a number for every check and it is close to meaningless, because an answer that praises three competitors and lists your brand in a table is not a positive mention of your brand.

RankX AI resolves brand sentiment from verbatim spans, and a proposed span clears six deterministic gates before it counts. The model proposes a label; the gates decide whether it stands:

  • Verbatim. The span appears word for word in the stored answer.
  • Anchored. The span is tied to the brand rather than floating in the answer.
  • Subject. The brand is what the span is about, so a sentence praising a rival is discarded.
  • Not a listing. A brand name in a table cell with punctuation round it is not an opinion.
  • Length. The span carries enough content to be informative.
  • Confidence. The proposal clears a floor, below which the model is guessing.

Two design choices follow, and both are worth copying whatever tool you use. There is no neutral bucket, because a large confident “62 per cent neutral” slice teaches nobody anything; the fourth outcome is undetermined, and it is a real stored answer rather than a failure. And there is no coverage floor, because a floor is an instruction to guess. AI brand sentiment covers what the measurement can and can’t support.

One further test, and it comes from our own mistake rather than someone else’s. RankX AI once carried a sentiment field on the answer row that nothing ever wrote, and a screen rendered it as a confident “0% positive”. A field with no writer looks exactly like a working one until somebody checks. That field is now deliberately dead, a test keeps it unwritten, and the real sentiment lives elsewhere. Ask any tool what its sentiment figure is computed from, and what it shows when it can’t tell.

Which AI assistants can you track, and what does each expose?

Track the assistants your buyers actually use, and expect them to disagree. The important caveat first: the assistants are not reached the same way, and citation coverage differs with the method, so no single figure describes all of them.

RankX AI supports six AI assistants, and five are on by default for a new website: ChatGPT, Gemini, Claude, Perplexity and Grok, with Microsoft Copilot available and switched on per website rather than by default.

What differs between them, and what to plan around:

  • Overlap is low, and it is the reason to pick rather than to cover everything. Across 100,000 identical prompts put to both, only 11 per cent of the domains cited by ChatGPT were also cited by Perplexity. Being named by one assistant is weak evidence about the others.
  • Citation coverage varies widely. Some assistants return structured sources on almost every answer, some rarely, some never. A citation-based metric will look strong on one and empty on another for reasons that have nothing to do with your visibility.
  • Market coverage is not universal. Not every assistant serves every market, and an unserved market should be recorded as unsupported rather than quietly answered from somewhere else.

Google AI Overviews is deliberately not in that list. An AI Overview is triggered by a search rather than prompted like an assistant, so RankX AI tracks it through the rank-tracking pipeline against your tracked keywords, and the free AI Overview Checker shows whether a keyword triggers one and who is cited inside it. Treating a search feature and an assistant as one surface produces a number nobody can interpret.

Five questions to ask any AI brand-monitoring tool

Every vendor in this category will show you a share-of-voice chart. These five questions separate the ones measuring your brand from the ones measuring their own pipeline. We would expect a good answer to all five, including from us.

  1. What does the chart do on a day a provider is down? If it dips, failed checks are being counted as absences. Ask to see a real outage day.
  2. What is the denominator? A mention rate must divide by checks that produced a verdict. Ask what happens when none did: the honest answer is no number, not zero.
  3. Can a brand be recorded as mentioned when its name is not in the answer text? If yes, a model’s guess is being stored as a fact. Ask how fabrications are bounded.
  4. Is a competitor credited when the prompt named that competitor? If yes, comparison prompts are inflating everyone, including you.
  5. How many checks sit behind this number? Cross-check it against the table above. Under about 100 in the window, a percentage to one decimal place is decoration.

A sixth, if the tool reports sentiment: ask what it is computed from and what it displays when the evidence does not support a verdict.

What do you do when the mention is wrong?

Treat a wrong mention as a source problem rather than an AI problem. An assistant saying something false about your brand almost always read it somewhere, and the record of which sources an answer cited is what turns a complaint into a task.

The sequence that works: find the answers carrying the error, read the sources those answers cited, correct the claim at source where you can, and publish a clearer statement of the fact on a page of your own that is easy to extract. Then keep checking, because a correction only surfaces once the assistant next retrieves that source. Given the volatility figures above, that is days rather than minutes, and it is not instant even then.

This is reputation work rather than measurement work, and online reputation management covers the full sequence, including what to do when the source is one you don’t control.

How RankX AI records a brand mention

RankX AI asks each assistant the question and records the answer it gives. One check is one prompt, put to one assistant, once. Scheduled checks run at 1, 3, 7, 14 or 30 day intervals, set per prompt rather than per website, so a handful of high-value questions can run daily while a long tail runs monthly. Daily is available on every paid plan.

From each stored answer, RankX AI records whether the brand was mentioned, where it sat among the brands named, which sources the answer cited, which competitors appeared beside it, and how the answer talked about the brand. The stored answer sits behind every one of those numbers, so any figure can be checked against the text it came from.

Three properties decide whether a number is trustworthy, rather than whether a dashboard looks full:

  • A brand whose name does not appear in the answer text cannot be recorded as mentioned. That is structural in the code, not a filter someone remembers to switch on.
  • “We did not look” and “we looked and you were not there” are different values everywhere in the product, and they are drawn differently.
  • Cadence and assistant targeting are set per prompt, so the set behind any two periods can differ, and a comparison states the set it is actually comparable across.

AI Visibility carries the assistant side of this, and AI share of voice covers how individual mentions aggregate into a number you can report. If you are starting from nothing, start with the prompts rather than the tooling. A small set of questions your buyers genuinely ask, checked consistently, beats a large set checked once.

From RankX AITwo free checksThe AI Readiness Score grades a single page against the extraction rules. The AI Crawler Access Checker reads the robots.txt half.Run both, no account needed

Questions about AI Visibility Measurement

How often should you check brand mentions in AI search?

Often enough that the sample supports the claim. At a typical 12 per cent mention rate, 200 checks in a window gives roughly plus or minus 4.5 points, and you need about a 6-point swing to call a change real. Forty prompts across five assistants, run weekly, is about where week-on-week reporting becomes defensible.

Can you track brand mentions in AI search for free?

Partly. A free check tells you what one assistant said once, which is a snapshot rather than tracking, and the same question asked again can give a different answer. Free tools are useful for a first look and for spot-checking a claim. Tracking means the same prompts asked repeatedly over time, and that part needs automation.

How many prompts should you track?

Enough to cover the questions that decide a purchase, which is usually dozens rather than hundreds. Coverage matters more than volume: a prompt set that misses the comparison question your buyers actually ask will report a healthy number while you lose the deal. Start with what your sales team hears most often.

Does RankX AI track Google AI Overviews as an AI assistant?

No. An AI Overview is triggered by a search rather than prompted like an assistant, so RankX AI tracks it through the rank-tracking pipeline against your tracked keywords instead. The two surfaces answer different questions, and combining them into one visibility figure produces a number nobody can interpret.

How many AI assistants does RankX AI track?

Six are supported and five are on by default for a new website: ChatGPT, Gemini, Claude, Perplexity and Grok. Microsoft Copilot is the sixth and is switched on per website rather than by default. The assistants are not all reached the same way, and what each exposes, citations especially, differs with the method.

Is a brand mention in an AI answer worth more than a backlink?

They are different things on different surfaces, so the comparison is imperfect. What is measurable is that being named in the answer is what a reader takes away, while a citation is a link most readers never open. Track the mention first and use citations as evidence of how the mention was earned.

Related reading

Written by

Asif Syed · Founder & CEO

Asif Syed is the founder and CEO of RankX AI, the AI search visibility platform. He builds the product and writes here about GEO, AI search measurement and WordPress.

Sources

  1. RankX AI production measurement, 27 August 2026: 346 of 394 prompt and run units contained no mentioning answer (87.8%) Checked 2026-09-14.
  2. RankX AI production measurement: approximately 12% of scheduled checks fail, worst single prompt 54% Checked 2026-09-14.
  3. RankX AI production measurement: 96 checks on one assistant carrying only 31 verdicts, the reason a mention rate divides by known checks Checked 2026-09-14.
  4. RankX AI production measurement: a tracked competitor matching on 26 of 26 runs, and another on 288 of 575, both from terms present in the prompt itself Checked 2026-09-14.
  5. RankX AI production measurement: a rival's share of voice moved 8.1 points with no change other than the check schedule Checked 2026-09-14.
  6. Ahrefs, AI Overviews change every 2 days: 70% change rate between observations, 45.5% citation turnover, 43,000+ keywords (opens in a new tab) Checked 2026-09-14.
  7. Seer Interactive, AI Overview impact on Google CTR: cited brands earn 35% more organic and 91% more paid clicks (opens in a new tab) Checked 2026-09-14.
  8. SparkToro and Gumshoe, AI brand-recommendation consistency study, 2,961 runs (opens in a new tab) Checked 2026-09-14.
  9. Profound, citation overlap between ChatGPT and Perplexity across 100,000 identical prompts (vendor study) (opens in a new tab) Checked 2026-09-14.
  10. Vismore, 50x5x3 AI mention audit, 750 responses, March 2026: 38% of repeated runs returned a meaningfully different brand list (vendor study, competitor) (opens in a new tab) Checked 2026-09-14.

CoversThis article covers the AI Visibility Measurement topic, the AI Visibility feature and the AI Overview Checker tool.

Terms usedAI Mention, AI Citation, AI Brand Sentiment, Unlinked Brand Mention, AI Share of Voice, Citation Rate, Prompt Volume and Synthetic Query.

Read this page asMarkdown: /blog/track-brand-mentions-ai-search.md.

All articlesEverything RankX AI publishes is listed on the blog index.

Ask an assistantAsk ChatGPT (opens in a new tab), Ask Claude (opens in a new tab) or Ask Perplexity (opens in a new tab).

Preferred sourceIf Google is your front door, you can add RankX AI as a preferred source (opens in a new tab), which asks your own results to surface more of what we publish.

Back to the top

The 30-day plan

Fix your AI visibility in thirty measured days.

RankX AI publishes its full 30-Day AI Visibility Plan free: four evidence-graded weeks from baseline to first citations, with a printable workbook, a 48-prompt starter pack and eight paced emails for whoever wants them.

Read the 30-Day Plan

The whole plan is on the page. No email needed to read it.