0422 428 584
AI SEO

How Accurate Are AI Visibility Tools? Why Most Cannot Show You the Full Picture

Most AI visibility tools measure one narrow surface and report it as your whole visibility. Here is what they miss, why, and how we test AI search visibility differently at Search Scope.

How accurate are AI visibility tools: the same question returning three different result sets across three surfaces

TL;DR: how accurate are AI visibility tools?

  • Most AI visibility tools run their checks through provider APIs, which is not the product your customers use. The API is stateless, has no memory, no custom instructions and a different system prompt to the consumer app.
  • Petra Labs ran 900 trials on the same prompt across three OpenAI access surfaces in one day and found a 32 percentage point visibility swing for a single brand, with only 25% overlap between the top five brands in paid ChatGPT versus the API.
  • Even holding the surface constant, the same model does not repeat itself. Thinking Machines Lab sent 1,000 identical prompts at temperature 0 and got 80 unique completions.
  • On the Google side, most tools read an SEO SERP feed. Google says AI Overviews and AI Mode may fan one query out into multiple searches, so the keyword you tracked is often not the query that sourced the answer.
  • The honest position is not that tools lie. It is that they measure one narrow surface accurately and sell it as the whole picture. Use them for direction, not for a number you can put in a board pack.

Someone shows me an AI visibility dashboard most weeks now. It reports a clean figure. Thirty-four percent visibility in ChatGPT, up four points this month.

The first question I ask is: thirty-four percent of what, measured where, on whose account, in which browser, logged in or logged out, free plan or paid, with web browsing on or off?

Nobody can answer it, because the tool does not surface it. And that gap is not a footnote. It is usually larger than the movement the dashboard is reporting.

What AI visibility tools are actually measuring

Almost every AI visibility tool on the market works the same way underneath. It sends your prompts to a provider API, parses the response for your brand name and any cited URLs, stores the result, and charts it.

On the Google side it is the same shape with different plumbing. The tool calls a SERP data provider, the kind of feed built for rank tracking, pulls whatever AI Overview block comes back, and counts citations.

Both approaches are built on infrastructure designed for classic SEO measurement. That infrastructure carries four assumptions baked in from rank tracking:

  1. There is one correct result set for a query.
  2. That result set is stable enough to sample once and trust.
  3. The query you track is the query the system answers.
  4. Two people in the same city see roughly the same thing.

AI answers break all four. Not occasionally. Structurally.

One probe checking one surface returns one tidy result set, while probes across four surfaces return four different result sets

Four reasons the number on your dashboard is not what your buyers see

1. The API is a different product to the app

This is the big one, and it is the least understood.

OpenAI’s own API documentation states plainly that “each text generation request is independent and stateless”. No memory between calls. The ChatGPT app your customers use is the opposite: it carries saved memories, chat history, custom instructions, a plan tier, and OpenAI’s own system prompt that you as a developer never see and cannot replicate.

Same underlying model. Different product. Different answer.

Petra Labs put a number on this. On 26 February 2026 they ran 900 trials of a single product research prompt with forced web search across three OpenAI access surfaces: logged-in paid ChatGPT, logged-out free ChatGPT, and the Responses API. One brand swung 32 percentage points across those three surfaces on the same day with the same prompt. The overlap between the top five brands surfaced in paid ChatGPT and in the API was 25%.

Read that last figure again. Three out of four brands in the API’s recommendation set were not in the paid app’s recommendation set. If your tool checks the API, it is reporting on a shortlist your buyers are largely not seeing.

Their surface-level variance measured 10.8% to 16.1% RMSE, against 4.5% to 7.1% within a single surface. The differences between surfaces were substantially larger than the noise inside any one of them, which is the part that matters. It is not randomness. It is a real, structural gap.

2. The same model does not give the same answer twice

Set the surface question aside and there is still a floor of instability underneath.

Thinking Machines Lab published a study on nondeterminism in LLM inference that should be required reading for anyone selling AI visibility data. They sent 1,000 identical prompts to Qwen3-235B at temperature 0, the setting that is supposed to make output deterministic, and got 80 unique completions. The responses started diverging at token 103.

The cause is not mysticism about model creativity. It is that inference servers batch requests together, and the maths changes slightly depending on how many other people’s requests got batched with yours. Their words: “the primary reason nearly all LLM inference endpoints are nondeterministic is that the load (and thus batch-size) nondeterministically varies.”

So the answer you get partly depends on how busy the server was at that moment. A tool that samples once per day per prompt is reading noise and charting it as a trend.

3. Google’s AI answers are not built from the keyword you tracked

On the Google side, the mismatch is structural in a different way.

Google’s own documentation on AI features confirms that “Both AI Overviews and AI Mode may use a ‘query fan-out’ technique, issuing multiple related searches across subtopics and data sources, to develop a response.”

Query fan-out: one typed question splitting into several hidden sub-queries that are then assembled into a single cited answer

Your tool tracked “emergency plumber Perth”. Google may have fanned that out into a dozen internal queries about after-hours callout fees, burst pipe response times, licensing and suburb coverage, then assembled an answer from those results. The citations in that answer came from queries nobody tracked.

This hits local businesses hardest, because the fan-out queries are the ones about your suburb, your callout window and your licensing, none of which appear in the head term you are paying a tool to watch. It is the single biggest reason AI SEO in Perth is not just national AI SEO with a city name bolted on.

Google also says directly that “AI Mode and AI Overviews may use different models and techniques, so the set of responses and links they show will vary.” Two surfaces, same search box, different answers, confirmed by the vendor. Most tools report one number for “Google AI visibility”.

Google did add generative AI performance reports to Search Console in June 2026, which is a genuine improvement on guessing from the outside. It is also first-party data about your own site, so it tells you nothing about which competitors are being recommended instead of you.

4. Personalisation is the variable nobody controls

Location changes AI Overviews. Account history changes ChatGPT. Plan tier changes which model actually serves the request. Whether browsing fired changes whether the answer came from the live web or from training data.

A tool running from a data centre IP on a clean API key has none of that. It is measuring a hypothetical user who does not exist: no history, no location signal, no preferences, no plan.

Your actual buyer is a 52 year old business owner in Subiaco who has been chatting to ChatGPT about their business for eight months and has a paid plan. The tool cannot see that person’s results, and neither can we, fully. But one of us is pretending otherwise.

This is a measurement problem, not a fraud problem

I am not going to name tools and call them scams. I have paid for several and tested more, and they are all built by people solving a genuinely hard problem with the tools available.

The criticism is narrower and, I think, fairer: they measure one surface accurately and present it as total visibility. The methodology is usually not disclosed on the dashboard, so a business owner reads “34% visibility in ChatGPT” as a fact about their market, when it is a fact about one API endpoint on one account with browsing configured one way.

There is a commercial reason it happens. APIs are cheap, fast, stable and legal to automate at scale. Driving hundreds of real browser sessions across real accounts is slow, expensive and breaks constantly. Vendors optimise for what scales. That is rational. It just produces a number that does not mean what the buyer thinks it means.

The reasonable ask is disclosure. Tell me which surface you sampled, how many times, from where, on what account type, and what the variance was. Nobody puts that on the dashboard, because a confidence interval is harder to sell than a percentage.

How we test AI visibility at Search Scope

Our answer to this is not clever technology. It is mostly refusing to take the cheap path, and it costs us more.

Many varied test conditions sampled separately, then averaged into a single trend line rather than reported as one spot number

We crowdsource the human side. We post micro jobs on crowdsourcing platforms and pay real people to run our prompts and send back what they saw. Different models, different browsers, different countries, different devices, some on free accounts and some on paid. That gives us answers from real consumer surfaces on real accounts with real usage history behind them, which is the exact thing an API cannot produce.

We use the API layer too, but not as the source of truth. We run API checks through multiple keys tied to different accounts rather than hammering one key, and we treat that output as one surface among several, not as the answer.

We run a Google account pool for AI Overviews. We maintain around 50 Gmail accounts, each seasoned with its own browsing and search history inside an antidetect browser so the profiles stay separated at the fingerprint level rather than just the cookie level. We build them out to match the buyer profile of a specific campaign, so the account querying an AI Overview about roof repairs has actually behaved like someone who needs a roof repaired. This is the newest and least mature part of what we do. We are still working out what genuinely changes the result and what is just noise.

We run deep AI brand visibility audits. Across multiple accounts and environments, we prompt the models about a client’s brand in two modes: with browsing on, and from memory only. The first tells you what the live web is feeding the answer. The second tells you what the model has absorbed about the brand, which is much slower to change and much more valuable when it is working for you.

Both layers run a few times a week. Not once a month, and not once. AI answers are dynamic enough that a single reading is worthless.

We report averages and direction, not a spot number. The output we care about is which domains get cited most often across many conditions, and whether the client is in that set or not. That is a durable finding. “You were at 34% on Tuesday” is not.

Where our approach falls short

It would be cheap to spend 1,500 words on everyone else’s methodology and skip our own, so here is the honest ledger against what we do.

Crowdsourced testers are not instruments. People misread instructions, screenshot the wrong pane, run the wrong model, or quietly reuse an account they used yesterday. We can require screenshots and check for internal consistency, but we cannot fully verify what happened on someone else’s machine. Every crowdsourced result carries a trust discount that an API response does not.

Our sample sizes are small. Petra Labs ran 900 trials to make one claim about three surfaces. We are running dozens of observations across many more variables. That is broader coverage with thinner evidence in each cell, which means we can see a pattern and still be wrong about the size of it. Wider net, weaker per-cell confidence. That trade is real and we should say so.

Seasoned accounts are proxies, not buyers. An account we built to look like a Perth homeowner is not a Perth homeowner. It has no eight month conversation history about their renovation. The personalisation layer that most affects your real customer is the one layer nobody outside that customer can reproduce.

More variables, less attribution. When our numbers move, isolating why is harder than for a single-surface tool. A tool measuring one endpoint has a clean signal about one endpoint. We have a messier signal about a wider reality, and messy is harder to act on.

It does not scale like a dashboard. Paying humans and maintaining accounts costs more per data point than an API call by a wide margin. We do this for clients where AI visibility genuinely affects revenue. For everyone else it would be an expensive hobby.

The Google account work is early. I would not present it as a finished method. It is an experiment we are running in public with client work, and we say so to clients.

None of that makes an API-only number better. It makes ours less wrong in a way we can describe, which is a lower bar than most of this industry currently clears.

Measurement is only half the job. The other half is the work the measurement is meant to inform, which is what our AI SEO and GEO service actually does: making a site citable enough that it shows up regardless of which surface the buyer happens to be on.

What to ask before you pay for an AI visibility tool

Whether you are an agency choosing a stack or a business owner being sold a subscription, these questions sort the serious vendors from the rest quickly.

  • Which surfaces do you sample: API, logged-out web, logged-in free, logged-in paid?
  • How many times per prompt per day, and do you report the variance or only the mean?
  • Are the accounts you sample from real accounts with usage history, or clean keys?
  • What location and account signals are set, and can I change them?
  • For Google, are you reading AI Overviews from a SERP feed or from a rendered browser session?
  • Do you capture AI Mode separately from AI Overviews?
  • What happens to the number if I ask you to re-run the same prompt an hour later?

A vendor with a defensible methodology will enjoy those questions. A vendor selling a percentage will change the subject to their feature list.

What to actually do with AI visibility data

Use it for direction, not precision.

Track whether your brand appears at all in answers to the prompts your buyers actually type, and whether that is trending up over months rather than days. Pay more attention to which domains keep getting cited across every condition you test, because that list is your real competitive set in AI search, and it is often not the list from your organic rank tracker.

Then fix the things that hold across every surface: being genuinely citable, having clear factual pages a model can lift an answer from, and having your entity data consistent everywhere. That work pays off regardless of which surface a given customer is on, which is exactly why it is the work worth doing while the measurement layer is still this immature.

If you want the practical version of that, our LLM visibility checklist covers the on-page and entity work, and we go deeper on the retrieval side in our guides to optimising for Brave Search and optimising for Bing and Copilot visibility, both of which feed AI answers more than most people realise. For the bigger commercial picture, AI is coming for local search covers what this shift does to lead flow.

FAQ

Are AI visibility tools accurate?

They are accurate about the narrow surface they measure, usually a provider API, and inaccurate as a description of what real buyers see. Petra Labs ran 900 trials across three OpenAI access surfaces on the same day with the same prompt and recorded a 32 percentage point swing in one brand’s visibility, with only 25% overlap between the top five brands shown in paid ChatGPT and in the API. A single-surface number cannot represent that spread.

Why do AI visibility tools use APIs instead of the real chat interfaces?

APIs are cheap, fast, stable and legal to automate at scale. Logging into hundreds of real accounts across browsers, plans and locations is slow, expensive and fragile. Vendors optimise for the thing that scales, which is a reasonable business decision but produces a measurement that does not match the product your customers use.

Does ChatGPT give different answers in the app than through the API?

Yes. OpenAI’s own documentation states that each API text generation request is independent and stateless, so there is no memory between calls. The ChatGPT app layers memory, chat history, custom instructions and its own system prompt on top of the model. Same model, different product, different answer.

Why do AI Overviews change every time I check them?

Partly because Google builds them differently to a classic result. Google’s documentation confirms AI Overviews and AI Mode may use a query fan-out technique, issuing multiple related searches behind your one query, and that AI Mode and AI Overviews may use different models and techniques so the responses and links they show will vary. Location, account state and ongoing testing add more variance on top.

What should I actually measure for AI search visibility?

Direction and pattern, not a precise percentage. Useful questions: across many conditions, which domains get cited most for the prompts our buyers actually type, is our brand present in the answer at all, and is that share going up or down over months. Setting a KPI on a single visibility percentage from a single tool is how you end up optimising for a measurement artefact.

Where this leaves you

AI search visibility is worth measuring. It is not yet worth measuring to one decimal place, and any tool presenting it that way is selling confidence it does not have.

If you want to know how your business actually shows up across AI search, and get a straight answer about what the data can and cannot tell you, . We will show you the method before we show you a number, including the parts of it we are still working out.

GBP Insider Newsletter

Get ahead of Google instead of reacting to it.

Frontline updates from the Google Business Profile and AI search era: what changed this week, what to action, and what to ignore. Written by Dorian.

  • New GBP suspension patterns and how to dodge them
  • AI Overview / map-pack ranking shifts as they happen
  • Tactical playbooks before they leak into the SEO mainstream

1–2 emails max per quarter. Value-packed emails. No spam, unsubscribe in one click.

Your subscription could not be saved. Please try again.
You're in. Watch your inbox for the next GBP Insider.