An audit is a
photograph. This is the film.
An AI visibility
audit tells you where you stand today: cited or not, against which
competitors, for a fixed set of buying questions. It is a one-off
diagnosis. AI visibility tracking is the ongoing measurement that
follows it, the same prompt set, run on a schedule, so you can see
whether your changes actually moved the number. If you have not run a
baseline yet, start with the audit. This page is for after that.
The five metrics we track
Mention rate. How often your brand is named anywhere
in the answer, whether or not it is linked. This is the widest measure
and the first one to move when entity clarity improves.
Citation rate. How often you are named with a source
link attached, rather than mentioned in passing. A citation is worth
more than a mention because it sends traffic and signals the engine
trusts a specific page of yours enough to point at it.
Position in answer. Where you land in the response: named first, buried in a list of nine, or the sole recommendation. Being
cited third out of three reads differently to a buyer than being cited
once among nine alternatives.
Sentiment and accuracy. What the model says about
you when it does answer, and whether the facts are right: correct
pricing, correct service scope, correct sector. A wrong fact repeated
confidently across engines is worse than silence, and this is the metric
most likely to surface it.
Share of voice vs competitors. The same prompt set
run against the two or three brands you compete with, so your trend has
a comparison point. A rising mention rate means nothing if your
competitor’s is rising faster.
What’s included
- A fixed, re-runnable prompt set, built from your
actual buying questions and locked so every run is measuring the same
thing. We do not change the questions mid-tracking. That is how you
lose the ability to compare month to month. - Monthly runs across the major engines (ChatGPT,
Google AI Overviews and AI Mode, Gemini, Perplexity, Claude and
Microsoft Copilot) logged per engine, because they behave differently
and a single blended score hides which one has the problem. - Competitor tracking on the same prompt set, so
movement is read against the market, not in isolation. - A dashboard view of the trend over time, not a
fresh PDF you have to compare by hand against last month’s. - A ranked action list each period, flagging what to
fix next based on where the citation gap actually sits, on-site,
off-site or entity confusion.
The dashboard
Our GEO dashboard is where the
tracking output lives: multi-platform monitoring across Google AI
Overviews, ChatGPT, Perplexity and Gemini in one view, with trend
analysis so you can see which change moved a given metric, query-level
detail on what triggers a mention, and competitor visibility alongside
your own. It runs automated checks on a schedule and lets you trigger a
manual one after a specific change, so you are not waiting a month to
see whether something worked.
How we set the baseline
Before we track anything, we establish where you start. That baseline
uses the same method as our AI visibility audit: a
structured prompt set covering category questions (the buyer knows the
problem, not the brands), comparison questions (they are choosing
between two or three named options) and branded questions (what the
models say about you directly, including whether it’s accurate). The set
is sized to be repeatable, enough prompts to distinguish a real pattern
from noise, few enough to run identically across seven engines every
period. Once that baseline is set, tracking is simply re-running it and
reporting the delta.
Proof
Official London
Theatre is our clearest tracked result: 85% AI search visibility,
the outcome of a five-year search and influencer partnership. We did not
arrive at that figure from a single run. It is the product of a
baseline set early in the relationship and tracked consistently against
it since, which is the same discipline this page describes, applied over
years rather than months.
We also do GEO tracking work for Sons. We are not publishing
prompt-grid or before/after numbers for that account until they are
sourced and cleared for release. This page names the relationship
without the figures, in line with house policy on unverified client
numbers.
Reporting cadence
Tracking clients get a monthly report against the fixed prompt set (mention rate, citation rate, position, sentiment and share of voice, per
engine, with the delta from the previous period) plus a quarterly
review where we step back from the monthly noise and look at the trend:
what has actually moved over three months, and what the ranked action
list should focus on next. This sits inside our normal terms: a 24-hour
response on any query you raise between reports, and 60 days’ notice on
either side if you want to stop, at any point.
What it costs
AI visibility tracking runs on a GEO retainer, from £1,250 a month ex VAT. There is no minimum term and either side can give 60 days’ notice at any point. Scope is set by how many engines, markets and competitors you want covered, and whether tracking runs alongside a full SEO and GEO programme or as a standalone measurement engagement. The GEO pricing page sets out what each tier includes. Book a call and we will price it against your brief.
Related services
- GEO: the full search-and-citation
workstream this tracking sits inside. - AI Visibility Audit: the one-off baseline you run before ongoing tracking starts.
- GEO Visibility Score: a
free, instant first read, no sign-up.
Latest GEO Insights
Stay ahead with our expert content
AI Visibility: Does ChatGPT Recommend Your Brand?
To find out whether ChatGPT recommends your brand, ask it the questions your buyers ask, run each one several times in a fresh session,…
GEO Audit Guide: How to Check Your Brand’s AI Visibility
What Is a GEO Audit and Why Does Your Brand Need One? A GEO audit is a structured process for measuring and improving how…
Schema Markup and AI Search: Why Structured Data Matters More Than Ever
What Is Schema Markup and Why Does It Matter for AI Search? Schema markup is a standardised vocabulary of structured data that tells AI…
ChatGPT vs Google: What It Means for Your Brand’s Visibility
What Is the Difference Between ChatGPT and Google for Brand Visibility? ChatGPT and Google represent two distinct search ecosystems, each requiring a different strategy…
Frequently asked questions
Partly, and it is worth being precise about which part. Tracking proves whether your visibility in AI answers moved, against a fixed prompt set, with the raw answers attached so anyone can check. It does not by itself prove revenue. What we do is report the citation trend alongside organic sessions and enquiries for the same query set, so the board can see whether the questions you are winning are the ones that produce business. Where the link cannot be evidenced we say so rather than imply it.
Yes. The same prompt set is scored for the competitors you name, so every report shows citation share across the set rather than your own count in isolation. That matters, because your score can fall while your position relative to the market improves, or rise in a category where everyone rose. We usually track three to five named competitors. More than that and the report becomes a spreadsheet nobody reads. You can change the list at any point and we will rebase the comparison.
Usually not. These systems are non deterministic, so a single month moving a few points is normal variance rather than a signal. We look at the three month trend and at whether the drop is concentrated in one engine or spread across all of them. A fall in one engine alone usually means that engine changed something. A fall everywhere at once, right after a release on your side, is worth investigating properly. The report says which of those we think it is, and why.
Tell us and we will check it that week. If our data is wrong we correct the report and say what happened. If the data is right and the model is wrong about you, that is a finding rather than an error, and it goes on the fix list. Incorrect model output about a brand is nearly always traceable to something ambiguous or outdated on your own site or on a third party profile. Those are fixable, and they are usually the highest value thing in the report.
The report is enough for most clients. The dashboard earns its place when several people need to look at the numbers between reports, or when you are shipping changes often and want to check the effect without waiting for month end. It also lets you trigger a manual run after a specific release. If only one person reads the numbers once a month, take the report and skip the dashboard. It is included in the retainer either way, so this is about attention rather than cost.
The audit is a one-off diagnosis: where you stand
today, against whom, and why. Tracking is the same method run on a
schedule, so you can see whether a fix actually worked and catch it when
an engine’s behaviour changes. Most clients run the audit first and move
to tracking once they have a baseline worth defending.
Monthly,
against a fixed prompt set, with a quarterly review that looks at the
three-month trend rather than a single month’s noise. You can also
trigger a manual check through the dashboard after a specific
change.
ChatGPT, Google AI
Overviews and AI Mode, Gemini, Perplexity, Claude and Microsoft Copilot.
We report each separately, because a brand can be strong on one and
invisible on another, and a blended average hides that.
Because a changing question set breaks the comparison. If
we swap prompts mid-tracking, you cannot tell whether a metric moved
because of your work or because we asked a different question. We only
change the set if your buying questions themselves genuinely change.
Yes. We
run the same prompt set against the two or three competitors you name,
so your trend has a comparison point rather than sitting in
isolation.
Not
on its own. These systems are non-deterministic. The same prompt can
return a different answer to different users on different days. That is
why we track a share across repeated runs and a trend over months, not a
single snapshot.
That is exactly what the sentiment and accuracy
metric is for. If a model states an inaccurate fact (wrong pricing,
wrong scope, wrong sector) we flag it in the report and prioritise the
entity fix that is most likely to correct it.
The monthly report is the record; the dashboard is where you can look
between reports, trigger a manual check after a change, or check a
specific query without waiting for the next reporting date.
GET IN TOUCH
Ready to grow your brand? Get in touch with our team to discuss how we can help you achieve your goals.