Obility Editorial · · 4 min read
The same question rarely returns the same list twice
In January 2026 SparkToro published a study it ran with Gumshoe, a company that sells AI visibility tracking. Six hundred volunteers ran the same twelve recommendation prompts through ChatGPT, Claude and Google’s AI Overviews or AI Mode, producing 2,961 answers over November and December 2025. The prompts covered chef’s knives, headphones, cancer hospitals, cloud providers for startups, marketing consultants and more.
SparkToro reports less than a one in 100 chance that ChatGPT or Google’s AI returned the same list of brands in any two answers, and roughly one in 1,000 for the same list in the same order. The number of items moved too, from two or three in some answers to more than ten in others. The volunteers were told to use their normal settings, history and devices, on purpose, so the spread reflects what real people see.
The practical consequence is simple. A screenshot of one answer, whether it shows you winning or missing, is a single draw from a distribution. It is evidence that an outcome is possible, not a measurement of how often it happens.
Your position inside an answer is close to random
The clearest example in the study is City of Hope, which appeared in 69 of 71 ChatGPT answers about cancer care hospitals on the US West Coast. It was named first in only 25 of them. A report that tracked its position would have shown it bouncing between first place and lower slots while nothing about the hospital or the web had changed.
Ordering has a second problem. If one answer lists three brands and the next lists ten, third place means something different in each. SparkToro’s conclusion was blunt: a tool that reports a ranking position in AI answers is not measuring anything stable. An average position computed from such lists inherits the same noise and hides it behind a decimal point.
How often you appear behaves like a real measurement
Presence held up much better than order. Google’s AI named one e-commerce marketing agency in 85 of 95 answers. In a separate test, volunteers wrote 142 prompts of their own about buying headphones for a family member who travels, and the wording barely overlapped, with an average semantic similarity of 0.081 between pairs. Yet across the 994 answers those prompts produced, Bose, Sony, Sennheiser and Apple each appeared between 55 and 77 percent of the time.
So the useful number is a rate: out of many answers to a group of related buying questions, how many named you. SparkToro moved from skeptic to calling that rate a reasonable metric, while listing what the study did not settle, including how many runs are enough and whether answers collected through an API match what people see in the apps. Read the result with its conflict in mind too: Gumshoe sells this kind of tracking, and SparkToro says so in the write-up.
A handful of runs cannot tell you much
The arithmetic of a rate is unforgiving at small sample sizes. Take an illustrative brand that appears in three of five answers. That is 60 percent, but a standard 95 percent confidence interval runs from about 23 to 88 percent. At 30 of 50 answers the interval is roughly 46 to 72 percent. At 60 of 100 it is still about 50 to 69. To tell a 40 percent rate from a 50 percent rate at conventional confidence and statistical power, you need close to 390 answers in each period you compare.
You reach those numbers by pooling, not by asking one prompt 400 times. Group prompts that share an intent, such as discovery questions for one product line in one market, and count answers across the group and across days. Keep engines apart, because ChatGPT, Claude and Google’s AI behave differently and a blended rate describes none of them. Keep the wording fixed once a series starts, or the change you see may be your own edit.
Report a rate with its range and keep the conditions attached
A visibility report a team can act on states the rate, the number of answers behind it and the range around it, per engine and per group of questions. Each answer stays attached to its prompt wording, engine, date, language and location, which is the approach behind Obility’s Answer Engine Insights: prompts are grouped by audience, buying stage and market so that like is compared with like, and the surrounding answer, the competitors named and the sources cited are kept with each observation.
There are limits to what even a good rate tells you. The engines change their models and retrieval without notice, so a rate can move while you did nothing; treat a change as real only when the ranges stop overlapping, and look for a cause before claiming credit. A high rate does not mean an engine judged you the best option, and a mention is not a recommendation or a sale. The study itself is from late 2025, and newer models may vary more or less. What it does settle is the question of method: measure how often you appear across many answers, and stop reporting where you appeared in one.