AI visibility tracking means running a defined set of questions against generative engines on a schedule and recording your brand's presence in the answers. The difference from rank tracking is precise: what is measured is not a position in a list, but whether you made it into the answer at all.
That distinction has a practical consequence. An Ahrefs analysis from October 2025 found that 28.3% of the pages ChatGPT cites most often have zero organic visibility on Google. Your search console can look healthy while you are absent from AI answers entirely — and you will not notice until you measure it.
What exactly does it measure?
Tracking records three separate signals rather than collapsing them. A mention is your brand name appearing in the answer; it reflects awareness. A citation is a specific page of yours named as the source; it carries a clickable link and measurable traffic. Share of voice is your portion of all brand mentions across the tracked question set.
Reported together, they give no direction. High mentions with low citations means the problem is not your content but its citability. The reverse means one page is doing all the work, which is fragile.
Every scan also records whether the answer came from live web search or from the model's training memory. Without that split, you cannot measure the effect of a content change you made today.
Which engines are covered?
Eight engines are tracked: ChatGPT, Claude, Gemini, Perplexity, Google AI Overview, Bing Copilot, Grok and DeepSeek. Each selects sources differently, so a result on one engine does not generalise to the rest.
ChatGPT search leans on the Bing index — a page Bing has not indexed never enters the candidate pool. Perplexity uses its own index and is the most freshness-sensitive of the group. Google AI Overview splits a question into sub-questions and looks for a source for each. DeepSeek does not perform live search at all; mentions there come from training data and cannot be changed in the short term.
The report does not flatten those differences. Where you lose ground on one engine, you read it against that engine's own logic.
Which questions should I track?
A good prompt set has three groups. Category questions ("best X tool", "which software for X") sit closest to purchase intent and are where you are directly compared with competitors. Problem questions ("how do I solve X") are top of funnel and shape your content plan. Brand questions ("is X reliable", "what is X") measure reputation.
Tracking only brand questions is the most common mistake: of course you appear when your own name is the query, while your real standing inside the category stays invisible.
Keeping the set fixed and dating any changes matters too. When the set changes, measurements stop being comparable with history.
Can I trust the numbers?
Generative models are probabilistic: the same question can return two different answers on the same day. That is the design, not a defect. A single scan is therefore not a trend.
Reliability rests on three variables: how many questions are tracked, how many engines are swept, and how many times each question is repeated per month. Running the same set on different dates is the only way to separate real movement from noise.
For that reason the report never shows a bare percentage; every number arrives with the sample it came from and the dates it was measured on.
What do I do after the measurement?
Every scan produces an action list: which of your pages could have been cited for a given question but was not, and why. The reasons are a short list — the page is not in Bing's index, crawler access is blocked, the content is not written answer-first, or no page on the topic exists at all.
Recommendations are page-level. Instead of "improve content quality", you get which heading needs which kind of answer block.
Lost questions also feed the content plan. A question where competitors appear and you do not is, directly, the subject of the next article.
What must be in place before tracking starts?
Tracking measures visibility; it does not create it. If three technical blockers are not cleared, your first report will mostly document your own infrastructure problems.
Crawler access: OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot and Google-Extended must be explicitly allowed in robots.txt, and your CDN or firewall must not return 403 to them. Server-rendered content: these agents do not run a browser, so JavaScript-assembled content looks empty to them. The Bing index: because ChatGPT search depends on it, a site not verified in Bing Webmaster Tools starts from zero on the ChatGPT side by construction.
The free analyzer checks all three without a signup, which makes it the fastest first step before tracking begins.
Frequently Asked Questions
How is AI visibility tracking different from rank tracking?
Rank tracking measures your position in a list; AI visibility tracking measures whether you made it into the answer. The two are largely independent: 28.3% of the pages ChatGPT cites most often have zero organic visibility on Google (Ahrefs, October 2025).
How many questions should I track?
For a meaningful trend, at least 20–30 questions with three repeats per month. Measuring a small set frequently beats measuring a large set once, because generative models answer the same question differently and a single measurement is noise.
How long before I see results?
Once technical blockers (crawler access, Bing index) are cleared, first changes usually appear within a few weeks. Content-driven gains are slower, and brand mentions can take months because they depend on third-party sources.
Can I try it for free?
Yes. The free analyzer needs no signup and scores a single page's AI readiness, surfacing crawler-access and content-structure blockers. Continuous multi-engine tracking carries real per-query API cost, so it is part of the paid plans.
Can I track competitors too?
Yes, and you define the comparison set. Share of voice is only read comparatively; a percentage on its own means nothing without knowing which brands it was measured against.