AI Visibility Metrics Need to Move Beyond Mentions, GEO Experiments Show
New GEO experiments indicate that third-party citations play a central role in AI visibility. The more useful question is which buyer queries and source surfaces create measurable business value.
Brand mentions in AI answers are becoming a useful visibility signal, but they are not a business outcome on their own. Two Generative Engine Optimization, or GEO experiments reported by Search Engine Land show why measurement needs to go further: brands must understand which buyer questions generate visibility, which external sources support it, and whether that visibility connects to commercial results.
The experiments, conducted by Zeeshan Yaseen, challenge the assumption that publishing more content on a company website is the main route to appearing in AI-generated answers. Across the tests, third-party articles, reviews, editorial roundups and partner pages had a much larger role in citations than branded listicles. That makes a simple report such as “AI mentions us” incomplete. It says little about where a brand is visible, why an AI engine selected it, or whether the appearance helped attract qualified prospects.
What the GEO experiments found
The Search Engine Land report on the GEO experiments describes two tests of how brands appear across AI engines. The first ran from May 30 to June 28, 2026 and tracked 15 keywords across four engines before expanding. It found that owned content could contribute to visibility, but it was not the dominant citation factor. External sources remained influential.
The second experiment was a 30-day cold-start test covering 15 commercial-intent SaaS buyer queries across ChatGPT, Claude, Gemini, Perplexity, Google AI Mode and Grok. Search Engine Land's reporting summarized the outcome as roughly 85 to 86% of AI citations coming from earned third-party placements, compared with roughly 14% from branded, owned listicles.
The results do not mean companies should abandon their websites. Owned pages remain a foundation for explaining products, services and proof points. The more important conclusion is that a GEO plan built mainly around self-published content may miss the source ecosystem that AI systems actually retrieve and cite.
| Experiment | Scope | What it indicated |
|---|---|---|
| Experiment 1 | May 30 to June 28, 2026; 15 keywords; four engines initially | Owned content could help, but external sources remained highly influential. |
| Experiment 2 | 30-day cold start; 15 commercial SaaS buyer queries; six AI engines | Earned third-party placements produced the substantial majority of citations, while owned listicles produced a minority share. |
Retrieval is necessary, but it is not the finish line
A brand first needs to be retrieved before it can be cited in an AI response. But retrieval does not guarantee a mention. The experiments point to a second step: AI engines draw from a wider set of sources, and the surrounding evidence can determine whether a retrieved brand is ultimately presented to the user.
That distinction matters for reporting. A team that only measures whether its own pages are indexed or surfaced may conclude it is well positioned, even when independent sources are absent from the answers buyers see. It also explains why a single placement cannot be treated as a universal win. The tests found platform variation, so a mention that improves visibility in one engine may not create the same result in another.
A more useful framework for AI visibility measurement
The practical unit of measurement is not a generic mention count. It is a specific buyer question, on a specific AI engine, supported by identifiable source surfaces. For example, a company should distinguish between appearing for a broad educational question and appearing when a prospect asks for software recommendations, alternatives, reviews or implementation options.
A practical reporting framework can include:
- Question coverage: the commercial and informational questions where the brand appears or is absent.
- Mention rate: how often the brand is named in responses for a defined set of queries.
- Citation rate: how often the brand is supported by citations, where the engine provides them.
- Source surfaces: the third-party pages, reviews, roundups and partner pages associated with visibility.
- Engine-level patterns: differences across ChatGPT, Claude, Gemini, Perplexity, Google AI Mode and Grok.
- Business attribution: the downstream enquiries, visits, trials or other outcomes a company can connect to the relevant visibility work through its own analytics.
The first four measures help explain how AI visibility is being created. The final measure prevents teams from treating visibility as the objective. The supplied experiments test retrieval and citations, not revenue attribution, so no universal conversion effect can be inferred from them. Businesses need their own measurement setup to determine whether a stronger presence on high-intent questions produces meaningful commercial results.
This also changes how teams prioritize work. If a question has commercial value but a brand is absent, the response may not be another self-published article. It may be improving product information on owned pages while also pursuing credible third-party coverage that accurately explains the company, its category and its use cases. If the brand is already cited but outcomes are weak, the issue may lie in the query set, message, offer or conversion path rather than citation volume.
For businesses investing in content, PR, partnerships or review presence, this is a more disciplined way to allocate effort. It identifies where evidence is missing, reveals which associations are actually visible in AI answers, and avoids treating every mention as equally valuable.
AI search visibility is difficult to manage when reporting stops at a total mention count. Scalevise can help map the questions potential buyers ask, the engines where your brand appears and the source surfaces associated with those answers, so marketing effort can focus on gaps with practical commercial relevance. Use the Scalevise AI Visibility and GEO Checker to turn scattered AI mentions into a clearer visibility baseline and start an AI visibility scan.
Frequently Asked Questions
What did the GEO experiments show about owned content?
Owned content could help brands appear in AI answers, but the experiments found that it was not the dominant citation driver. Third-party sources remained highly influential.
Why are third-party citations important for AI visibility?
The experiments indicate that editorial articles, reviews, roundups and partner pages can be important source surfaces for AI citations. A brand may need this broader ecosystem of evidence in addition to its own website.
Does retrieval guarantee that an AI engine will cite a brand?
No. Retrieval is required for a brand to appear, but the experiments found that retrieval alone does not guarantee a citation or mention in the final AI answer.
Should businesses measure AI visibility separately for each AI engine?
Yes. The experiments found platform-to-platform variation, meaning a placement that improves visibility in one AI engine may not have the same effect in another.
Can AI mentions be treated as proof of business results?
No. A mention is a visibility signal, not proof of commercial impact. Businesses need to connect query-level visibility work to their own analytics and downstream outcomes.
Conclusion
The GEO experiments support a more rigorous view of AI visibility: it is shaped by retrieval, citations and a brand's wider network of third-party evidence, not owned content alone. The most useful reports therefore connect specific buyer questions, AI engines and source surfaces to measurable business outcomes, rather than celebrating an undifferentiated count of AI mentions.