Zahir Hasan He didn’t need to tell me his company’s numbers were wrong.
I had sent Hasan, COO of the Oslo-based search company Clovion AIa list of methodological questions about “Surviving the AI Funnel,” Clovion’s new study on how Claude, ChatGPT, and Gemini recommend brands during a conversation. The tenth question was a routine question, the kind of thing every research team is asked. The report states that the three AI assistants flatly contradict each other on brand facts 15% of the time, based on 33 verified contradictions. Was 33 really enough to support a claim that which model tends to undersell a brand’s features and which tends to oversell them?
Hasan’s response was not a defense of the number. It was a correction. “The real number is 330,” he replied. “A designer left a zero in the layout.” The same slipped decimal, he said, had also turned 2,040 marks into “204” on page seven of the PDF sent to me before it was published. A revised version is coming out this week. So I got the corrected numbers first.
It’s a strange way to begin a column on an AI research report, first of all admitting that the draft report contained an error. But it’s the most honest way to come in, because the correction says something the study’s main statistics could never say. Properly reading AI responses, whether you’re a marketer trying to determine whether ChatGPT recommends your product or a researcher building a study about it, is like typing the decimal point before strategizing on it.
The funnel, summarized
Put aside the typo for a moment and the underlying research holds up. Clovion conducted 69,120 multi-round conversations with the three assistants across 36 B2B and fintech software categories, asking an opening question like “best CRM tools?” then a single realistic follow-up. By asking the same question again, 90% of the recommended list was preserved. Adding a regular buyer detail, something as simple as “for a small team,” only left 28%. Sixty-two percent of brands that responded to the first response disappeared following the second.
I asked Hasan if the “small team” had been selected to produce this drop. This was not the case. His team also tested “for a large company” and got an almost identical churn rate, about 72% in both cases, compared to about 10% when the question was simply repeated. The list is not unstable. It’s responsive, and especially if the model has decided who a brand is really for.
This is the part worth sitting down with if you do SEO or branding to make a living. Being named in an AI response is not the same as being trusted by it. A model that puts you in its first CRM list may still remove you as soon as a buyer gets specific, and Clovion’s data indicates that this happens most of the time, not some of the time.
The correction changes the form of the smallest and most cited number
This is where the fixed decimal is really important for how you should read this study. The old figure, 33 verified contradictions, was small enough that any model assertions based on that figure were on thin ice. Corrected, it is 330, and the distribution by model shared by Hasan is much more revealing than the overall figure of 15% reached in the draft report: Claude understates the specific characteristics of a brand 160 times compared to 10 overstatements. ChatGPT underwrites 70 times and never overclaims. Gemini goes the other way, claiming 80 times too much versus claiming 30 times too little.
Hasan’s working theory, drawn from a separate, as-yet-unpublished Clovion study on where each model comes from for its responses, is that Gemini relies more on marketing materials and video, so it tends to attribute whatever it claims to a brand. Claude and ChatGPT rely more on documentation and product pages, describe the core product accurately, and guard against “don’t have it” when a newer feature is not well documented. If this holds up in the study that Clovion has yet to release, it means that the direction of an AI assistant’s error regarding your product depends on what type of content you have placed in front of it and where that content is located.
I’ve spent over 20 years explaining to my clients that good grading and accurate description are two different issues. This is the clearest evidence I’ve seen that this is now the same problem, happening within a single conversation, and that the solution depends on which assistant is doing the misdescription.
Why no one catches the missing zero
Frédéric Vallaeys has a story in his book”The marketer amplified by AI“This explains exactly why a dropped decimal survives until publication. An automated report once reported “excellent performance” on a keyword because its cost per acquisition was much higher than the target. Somewhere in the system, high had been replaced for good, while a high CPA is bad news, not good news. Anyone skimming the summary would have nodded, because the phrase read smoothly even if its meaning had been reversed.
Vallaeys connects this to research on predictive processing, the idea that fluent readers don’t decode every word, they predict what’s next based on context and move on. That’s how “te” reads when “the” and a missing “no” slide right in front of you. As Vallaeys says, our mental model of the sentence overrides the text in front of us. A reliable, well-formatted PDF is the easiest place in the world for this to happen, and a deleted zero in a layout file is a much smaller, much more forgivable version of the same failure.
This is also why the fix is not to “trust the report less.” It’s about “keeping a human pilot in the know who checks the number instead of the mood of the paragraph around it.” Thirty-three contradictions and 330 contradictions differ not only by a factor of ten. They support completely different levels of confidence in the reality on a model-by-model basis. Two hundred and four brands and 2,040 brands do not constitute the same study. If Clovion hadn’t figured it out, and if I hadn’t asked, the smallest, flimsiest numbers would have continued to circulate as fact, cited by exactly the kind of trade press that’s supposed to detect this.
What Clovion Doesn’t Claim and Why It’s the Most Honest Part
The report is careful to clarify that the link between how a model perceives your suitability and whether it recommends you is “a strong, consistent coupling, not a proven causal law.” I pushed Hasan on what a true causal test would look like. His answer: change one thing, the content of a brand’s public positioning, leave everything else alone and see if the behavior of models evolves compared to brands that no one has touched. Clovion has not yet carried out this test. He also directly conceded the more uncomfortable possibility that a brand’s actual positioning in the real world likely determines both how the model describes it and whether it is recommended, which would make the actual lever positioning and “perception” of the model merely a symptom, not a cause.
That’s an unusually candid response from a company selling AI viewability monitoring, and it’s exactly why I trust the rest of what Hasan told me. It also didn’t have data on how quickly the AI’s perception of a brand changes after that brand changes its own content. “We didn’t do a before and after test,” he said. “Consider it worth testing, not guaranteed in X weeks.” Anyone who tells you they can promise a specific timetable for changing Claude’s or Gemini’s opinion of your brand is guessing, by Clovion’s own admission.
What to actually do about it
There are three things you should do, based on what Hasan told me and what the corrected data supports.
First, follow the entire conversation, not the first response. If you monitor AI visibility With a single-prompt check, you’re measuring the top of a funnel that loses 62% of its content one sentence later. Build your monitoring around the follow-up questions your real buyers actually ask.
Second, attach the helpers one by one, in order. Hasan made it clear that a single content change would not move all three models at once, as they come from different sources. His suggested order: Fix flat factual errors first, because those are cheap wins, then find the segment adjustment combinations that matter most to your pipeline, checking each helper over multiple runs rather than trusting a single answer.
Third, don’t cite a statistic whose source you haven’t traced, including this one. Clovion’s own report needed a correction on its most technical and cited figure. Before creating a column, client presentation, or content summary around an AI research percentage, ask where the underlying count came from and if anyone has checked the calculations since they left the design software.
I’ve seen SEO go through a few of these moments, from Panda to mobile-first indexing to the slow purge of zero-click search. Everyone rewarded practitioners who checked the primary source instead of repeating the title number. AI visibility is evolving in the same way. The brands that win the disappearing act documented by Clovion won’t be the ones with the best press release on their AI Overviews strategy. They will be the ones who read the report carefully enough to wonder what a “33” really means, and who continue to ask that question after this one.
Zahir Hasan is COO of Clovion AI, based in Oslo, Norway. The corrected version of Clovion’s “Surviving the AI Funnel,” reflecting the numbers in this column, is expected this week.





