Home/Insights/How to measure AI visibility when the answers keep changing

How to measure AI visibility when the answers keep changing

AI answers change between runs. Here is how many questions, engines and repeat runs you need before an AI visibility number means anything.

BY Esh, FOUNDER, ETHOS ENGINEPUBLISHED OCTOBER 3, 20268 MIN READ
KEY TAKEAWAYS
  1. One answer is an observation. A visibility figure needs a fixed question set, several engines and repeat runs.
  2. The studies that recommend 3, 7 or 400 runs are answering different questions. Pick the sample that matches your decision.
  3. A panel of 180 answers cannot confirm a 10 point change. Treat small monthly moves as noise until they repeat.
  4. Cited sources turn over much faster than brand mentions, so they need weekly snapshots and a separate count.

A screenshot of one ChatGPT answer tells you almost nothing about your AI visibility. Ask the same question again and the list of brands changes. The figure worth reporting is a rate: across a fixed set of questions, engines and repeat runs, what share of answers name you. The open question is how large that sample has to be. The published research gives answers ranging from 3 runs to 400, and each of them is right for a different job.

A single answer is an anecdote

SparkToro and Gumshoe had 600 volunteers run 12 prompts through ChatGPT, Claude and Google's AI results 2,961 times in November and December 2025. The study, published in January 2026, found less than a 1 in 100 chance of getting the same list of brands twice, and roughly 1 in 1,000 of getting the same list in the same order. Rand Fishkin's verdict on position tracking was blunt: "any tool that gives a 'ranking position in AI' is full of baloney."

The same study also showed why measurement is still possible. The lists varied, but the rates did not vary much. One hospital appeared in 69 of 71 answers to a regional cancer care question. A brand's share of answers is a property you can estimate. Its position in any one answer is not.

This matters for digital authority, because AI visibility is one of its five pillars alongside press coverage, search presence, brand entity and organic authority. A pillar scored from one answer per question is scored from noise.

Why the studies disagree on how many runs you need

Four sources, four numbers.

Pratyush Kumar's preprint, covering 102 brands and 102,025 responses across five engines between March and May 2026, says three runs. Researchers at the University of St. Gallen, in a preprint submitted in April 2026, say at least seven runs per prompt per day for brands and eight for sources. Gumshoe's February 2026 guidance says about 100 answers for a margin of plus or minus 10 points and about 400 for plus or minus 5. cloro's August 2026 analysis lands close to that: about 300 answers to reach plus or minus 5 points on a rate near 25%.

They are not in conflict. They are measuring different things.

Kumar looked at each combination of brand, prompt and engine and asked whether the result flips. Mostly it does not. In his data 77.5% of combinations were strictly always or never mentioned, and only 6.8% sat in the unstable band between 30% and 70%. Three runs are enough to sort a combination into "named", "not named" or "unstable". They are not enough to put a number on the unstable ones. Note that Kumar co-founded the tracking platform the data came from, and his brand sample leans towards software and fintech.

The St. Gallen team asked a harder question: what is one brand's mention rate for one prompt? With a single run, the 95% interval on that rate was plus or minus 72 points, which is no information at all. At three runs it was plus or minus 37. At seven it was plus or minus 16. So even their recommended minimum gives a loose estimate. Their data came from four engines, four industries and servers in Switzerland, so treat the exact figures as indicative.

Gumshoe and cloro are doing arithmetic on a pooled rate across all prompts. That arithmetic is standard, and it is where the 100 and 400 come from.

Repeats are worth less than new questions

There is a catch in the pooled arithmetic. It assumes every answer is independent. Ten runs of one prompt are not ten independent readings. They are ten samples of one question.

Answer Engine Land's July 2026 guide calls treating every run as independent the most common statistical mistake in this field, and gives the standard correction: the more the runs of one prompt agree with each other, the fewer effective observations you have. Dmitrij Zatuchin's July 2026 preprint points the same way from a different angle. Across 12,933 responses he found that a repeat beyond the fifth added almost nothing, while adding models or languages reduced error several times more. His outcome was sentiment, not mentions, and he is affiliated with a tracking vendor, so the precise numbers do not transfer. The direction does.

The practical reading: once you have three to five runs per question per engine, spend the next unit of budget on another distinct buyer question, not another repeat.

How many answers each decision needs

Match the sample to the decision.

What you want to knowMinimum sampleBasis
Is the brand named at all for this question on this engine3 runsKumar, 2026
One brand's mention rate for one question7 or more runs, spread over 2 to 4 weeksUniversity of St. Gallen, 2026
Overall visibility rate within 10 pointsAbout 100 answersGumshoe, 2026
Overall visibility rate within 5 pointsAbout 300 to 400 answerscloro and Gumshoe, 2026
Whether a 20% rate has really moved to 30%About 294 answers in each periodcloro, 2026
Which publications are the stable cited sourcesWeekly snapshots for at least a monthSISTRIX, 2026 (the one month floor is our reading)
Position within the answerDo not track itSparkToro, 2026

A sampling method you can run every month

  1. Fix the question set. Twenty to fifty distinct buyer questions, written the way buyers ask them. Variations of one question count as one question.
  2. Fix the engines and keep them separate. Zen Media reported in September 2026 that one campaign's coverage ranged from 15.5% to 59.5% depending on the engine, across a set of 900 prompts and five engines. An average of those two figures describes neither.
  3. Run each question at least three times per engine, on different days. The St. Gallen data shows a 7 day window still leaves a wide interval and a 28 day window a narrow one. Spreading runs across the month costs nothing.
  4. Record three things per answer. Was the brand named. Was the brand's own site cited. Which other domains were cited. Semrush's June 2026 study of 115 prompts found that 62% of citations came without a mention of the cited brand, so mentions and citations need separate counts.
  5. Report the rate with its denominator. "Named in 54 of 180 answers" is a measurement. "30% visibility" on its own is not.
  6. Change nothing mid-series. An analysis summarised by Elmo in September 2026 showed that reweighting the same three groups of prompts moved an aggregate score from 39.7% to 70.5%. If you change the questions or the weights, start a new baseline.

A worked example, with hypothetical numbers

This is an illustration of the arithmetic, not a result from any real brand.

Take a panel of 20 questions, three engines and three runs each. That is 180 answers a month. Suppose the brand is named in 54 of them, a 30% rate. If all 180 answers were independent, the 95% margin would be about plus or minus 7 points. They are not independent, because each question was asked three times. If runs of the same question agree with each other moderately, the panel behaves more like 90 answers, and the margin is closer to plus or minus 9.5 points.

Next month the brand is named in 72 of 180, a 40% rate. It looks like a 10 point gain. The margin on the difference between two such months is about 10 points in the best case and nearer 14 once the repeats are accounted for. The gain is inside the noise.

That does not make a 180 answer panel useless. It reliably separates a brand at 10% from one at 40%. It shows which questions never name you, which is the list to act on. And a rise that holds for three months running is far more convincing than one jump. What it cannot do is confirm a single month's 10 point move. For that you need roughly 300 answers per period, or patience.

Cited sources drift faster than brand names

Brand mentions are the steadier signal. In the St. Gallen data, the overlap between one day's brand set and the next was 0.45 to 0.59. For cited sources it was 0.34 to 0.42.

SISTRIX's study of 82,619 prompts over 17 weeks to April 2026 puts numbers on the turnover. Weekly, 74% of cited domains in ChatGPT search were new, and 56% in Google AI Mode. Yet in AI Mode, 86.5% of prompts kept a stable core of one to five domains throughout. Alex Birkett's August 2026 snapshot of roughly 100 prompts found the same shape a month apart: 49% of the top URLs survived, but 65% of the domains did.

Here the sources disagree, and it is worth saying so. SISTRIX found Google AI Overviews almost static, with 5% weekly churn. Profound's 2025 comparison of about 80,000 prompts per platform found 59.3% of AI Overviews domains changed in a month. Ahrefs found in late 2025 that 45.5% of cited URLs were new from one observation to the next. The studies use different units, dates and definitions, and they do not reconcile neatly. SISTRIX offers the most useful clue: in its data, 53% of AI Overviews prompts never changed a single source in 17 weeks, while a minority rotated heavily. An average hides two populations. You have to look per question.

The consequence for measurement is direct. The publications an engine keeps retrieving for a question are the ones that appear week after week, not the ones in today's answer. Take weekly snapshots, count how many weeks each domain appears, and treat the domains present in most weeks as your cited sources. Those are the targets for press coverage. The rotating remainder is not.

Check your sample against the platforms' own counts

Two first-party sources now exist that do not depend on your sampling. Microsoft added an AI Performance report to Bing Webmaster Tools in February 2026, showing how often your pages are cited across Copilot and Bing's AI summaries, and added a citation share figure per grounding query in June 2026. Google's generative AI reports in Search Console reached all sites on 31 August 2026, according to Search Engine Journal.

Both have limits. Bing aggregates across its surfaces and samples the grounding queries. Google shows impressions only, without clicks or queries. Neither tells you whether your brand was named. Use them as a cross-check on direction: if your sampled citation rate rises and the platform's count falls, trust neither until you know why.

Measurement is the last step of the five-pillar checklist, and it is the step that tells you whether the others worked. If you would like a baseline across the five pillars, the Digital Authority report is free with an account.

Questions

How many times should I run each prompt?

Three runs are enough to see whether a brand is always, never or sometimes named. For a mention rate on one question, a University of St. Gallen study from April 2026 suggests at least seven runs. For an overall rate, total answers matter more than repeats.

Why does my AI visibility score differ between tools?

Tools ask different questions, weight them differently, sample different engines and run each prompt a different number of times. Engines also disagree with each other. Compare a score only with its own earlier readings, taken with the same questions and rules.

Can I track my ranking position inside an AI answer?

Not reliably. SparkToro's study, published in January 2026, found the same brands in the same order about once in 1,000 runs. Whether you are named, and how often, is measurable. The order is mostly noise.

How often should I measure?

Monthly is a reasonable reporting rhythm if the sampling is spread across the month. Answers and cited sources shift from day to day, so a reading taken on one afternoon describes that afternoon. Two to four weeks of spread sampling gives a steadier figure.

Sources

  1. SparkToro, NEW Research, AIs are highly inconsistent when recommending brands or products sparktoro.com
  2. arXiv (Schulte, Bleeker, Kaufmann, University of St. Gallen), Don't Measure Once, Measuring Visibility in AI Search arxiv.org
  3. arXiv (Kumar, Ranqo), Generative Engine Optimization at Scale, Measuring Brand Visibility Across AI Search Engines arxiv.org
  4. Gumshoe, How Much Data Do You Need to Measure AI Visibility with Confidence gumshoe.ai
  5. cloro, AI Visibility Sample Size, How Many Prompts and Runs cloro.dev
  6. Answer Engine Land, AI Search Visibility Confidence Intervals answerengineland.com
  7. arXiv (Zatuchin), Where Does the Noise Come From, A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers arxiv.org
  8. SISTRIX, AI Citation drift, How stable are sources in AI search results sistrix.com
  9. Profound, AI Search Volatility, Why AI search results keep changing tryprofound.com
  10. Ahrefs, AI Overview change study ahrefs.com
  11. Alex Birkett, Citation Drift, How Stable Are AEO Citation Sources alexbirkett.com
  12. Zen Media via GlobeNewswire, One Brand's AI Visibility Tracking Ranged From 15.5% to 59.5% Depending on the AI Engine globenewswire.com
  13. Semrush, Why 62% of AI citations don't lead to brand mentions semrush.com
  14. Elmo, Why AI Visibility Tools Disagree on Your Score elmohq.com
  15. Microsoft Bing Webmaster Blog, Introducing AI Performance in Bing Webmaster Tools Public Preview blogs.bing.com
  16. Microsoft Bing Blog, New AI Visibility Insights in Bing Webmaster Tools blogs.bing.com
  17. Search Engine Journal, Google Search Console AI Reports Rolled Out Worldwide searchenginejournal.com
E
Esh

Founder, Ethos Engine. Builds the Digital Authority Score and runs the placement marketplace. esh@ethosengine.io

NEXT STEP

See where your brand stands.

Your Digital Authority Score, who AI recommends for your target questions, and the placements that close each gap. Your first report is free.

Get your free report →
MOREINSIGHTS