AI viewability tracking data is not entirely reliable. Because generative models often produce different answers, quote shares and rankings on your dashboard are just snapshots of an ever-changing target, not fixed facts.
A difference between you and a competitor may be real or simply a fluctuation between measurements. A new paper from IQRush, scheduled for publication next week (we had early access) provides a method for distinguishing them, showing that no fixed amount of data can definitively settle the question.
The article is written by Ron Sielinski, co-founder of IQRush, which sells software that measures AI visibility as the article claims. The reason it’s worth doing is that a separate team published a similar repeated measures result in April, so IQRush isn’t the only one making this point.
How far these numbers move
Repeatedly querying SearchGPT, Gemini, or Perplexity with the same question may produce different sources each time. They’re designed to add a bit of randomness to each answer, so that each quote is just one of many possible URLs it could have pulled. A previous article by the same author explored this variability, showing that, for example, when testing SearchGPT on running trains, Tom’s Guide accounted for approximately 9.5% of citations, while Runner’s World accounted for approximately 6.0%. On the dashboard, Tom’s Guide appeared more often, but the large margin of error caused the numbers to overlap. With just one sample, it was not accurate to say that Tom’s Guide outperformed Runner’s World, as the 3.5 point difference was within the margin of error. The new paper aims to avoid this mistake by addressing a simple but often overlooked question: How much data is needed for rankings to be truly meaningful?
When a ranking deserves confidence
The answer has two parts, and both must be true for a ranking to be reliable. First, the order must stop changing.
At first, rankings may change frequently as new answers are added, because no site has a clear advantage yet. Only after collecting enough responses do the top sites begin to clearly stand out, allowing the order to stabilize. It is also important that the best sites are well spaced; if they are very close, the rankings might not make sense, because close competition doesn’t really show who is really ahead. The article examines whether the difference between the top sites is greater than each person’s margin of error. When it is, the ranking reflects a real difference. If not, it’s probably just statistical noise. Both conditions must be met at the same time, but neither is sufficient. In 30 platform topic tests, the number of answers needed for both conditions to be met ranged from 33 to 94, counting only answers with quotes.
Three out of 30 people didn’t reach this point even after 125 questions, all on SearchGPT, where the top sites were too similar to tell apart. There is no single threshold applicable everywhere; what works for one platform and topic may not work for another.
We circled this
In January I discussed The discovery of SparkToro that AI tools give a different list of recommended brands more than 99% of the time you ask the same question. This article left one question unanswered: how many times should the question be asked before the results stabilize? This article offers the clearest answer I have encountered.
Rand Fishkin, who led this study, shares some helpful tips. Before spending money to track AI visibility, he suggests making sure your vendor “shows the math.” The IQRush paper is a great way to do this because it provides a simple stopping rule, so you don’t have to rely solely on your intuition to know how many runs are enough.
This also matches a series of studies covered by SEJ over the past year, each reporting AI citation numbers as if corrected. He turns around, looks at the measurement itself, and wonders if these numbers are stable enough to be able to compare in the first place.
What this changes for your reporting
The number on your dashboard is just a sample. Before trusting it, check if your tracker performs the same check repeatedly and shows a range, or if it pulls the data just once and shows a clear number. A clean number may actually be a warning sign, not insurance.
A gain after a content change is easy to misinterpret. For example, a three-point increase in your SearchGPT citation share may seem like proof that your efforts have paid off, but such a change may fall within the natural variability of successive runs, depending on the data in the original article.
To win, measure before and after more than once each. A single before and after reading cannot separate your currency from ordinary noise.
The platform you measure changes the amount of data you need, and not in the ways you might guess. It depends on how much independent information each answer contains, not how many citations it gives you. Gemini stacks quotes on the same handful of sites into a single answer, so a lot of those quotes tell you the same thing. SearchGPT gives fewer citations per answer but spreads them out, so each answer contains more independent information than the raw count suggests. The same number of answers on two engines does not buy the same confidence, and a budget that rules Gemini can leave you guessing on SearchGPT.
Sometimes the honest answer is that you can’t say it yet. Three of the 30 tests never clearly separated their top sites within the budget. For those, the right decision is to maintain the ranking, not to publish a ranking that the data cannot support. A tracker that can tell you “not enough data” is worth more than a tracker that prints a confident command every time you ask.
The top of the leaderboard is the part you can defend the most. With enough answers, leaders move away from the middle and tail, even if they are not accurate. The margins of error widen quickly below the front, until neighboring positions are a toss-up, and even the top 10 were not flawless, with the typical margin of error at a top 10 site being around five positions and one in five wider than 10. Trust the leaders, treat the middle and bottom as approximate, and don’t report exact positions beyond the start of the list.
What the document does not prove
None of this comes from a completed, peer-reviewed study. This is a pre-publication built on 30 platform topic tests across three engines, using ChatGPT-generated questions rather than real user searches, on just one part of the collection. The exact numbers will not transfer clearly to your subjects, so treat them as the shape of the problem, not a lookup table.
These counts only include answers with citations, which matters most on SearchGPT because a portion of its questions return no citations. In one topic, 125 questions produced 104 usable answers, a failure rate of 17%, so you will need to submit more questions than these totals suggest.
Control of the method is also internal. The journal compares the classification it established at the beginning with the final classification of this same collection, and not with an external ground truth. This tests whether the stopping rule is consistent with itself, which is why the unaffiliated team’s match result does a real job here. The authors of this April article, Julius Schulte, Malte Bleeker and Philipp Kaufmann, are researchers at the University of St. Gallen. They analyzed a separate data set and came to the same verdict, which is that a single reading is unreliable and that you have to sample an engine repeatedly to trust what it’s telling you.
Where is this going
The newspaper falls short of what most people want, which is a way to know your operating budget before you start collecting. Sielinski leaves this for future work and notes that the number depends on the shape of each platform’s citation pattern, so a single universal budget is probably not forthcoming.
The biggest change is that AI viewability reporting is moving in the same direction as advertising and analytics reporting, toward numbers with a margin of error instead of a false decimal point. This happens while basic plumbing is still missing, from Search Console I still won’t tell you which clicks come from AI. In the meantime, it’s your responsibility to check more than once and report the range, not the unique number your dashboard gives you.
More resources
Featured Image: Sticky/Shutterstock





