Gemini’s Local Search Volatility Shows Why AI Visibility Needs Multi-Engine Measurement
Steady Demand’s local-search study finds that generative AI recommendations can shift sharply between identical prompts. Businesses need repeated testing across engines, not one-off AI visibility reports.
Generative AI is becoming a new route into local business discovery, but its answers may be far less repeatable than established local search results. A Steady Demand field study of Gemini and ChatGPT found that Gemini returned the same top local-business recommendation only about 7% of the time when an identical prompt was repeated. By comparison, Google’s traditional Local Pack returned the same top listing roughly 90% of the time.
The result changes what an AI visibility metric can credibly represent. A single prompt response is not a durable ranking position. Instead, it is one observation from a variable system whose cited sources, leading recommendation, and source ecosystem can change across runs and differ substantially between AI engines.
What the local AI citation study found
Steady Demand’s August 6, 2026 study, What Gemini and ChatGPT Cite in Local Search, analyzed 1,487 identical local-intent prompts across 50 U.S. metropolitan areas and 10 service categories. The research produced 14,472 citations from Google Gemini and OpenAI ChatGPT.
The study describes this instability as Grounding Drift. Repeating the same query word for word produced overlapping cited sources only about 40% of the time. Gemini’s top-brand result was especially variable, even when the query itself did not change. Steady Demand’s methodology and interactive dashboards attribute the observed pattern to the AI generation process rather than changes in the underlying Google Business Profile or local data.
| Measurement | Google Gemini | Google Local Pack |
|---|---|---|
| Consistency of the top local result on repeated identical queries | About 7% | Roughly 90% |
| Nature of the result | Generated local-business recommendation | Traditional local listing result |
Grounding Drift makes one-off reporting unreliable
Traditional local-rank tracking is built around repeated observation of a comparatively stable result set. That approach does not translate directly to an AI answer surface where the leading business can change from one identical request to the next. Reporting that a brand ranked first in one Gemini response may be accurate for that response, yet misleading if presented as a stable position.
The practical unit of analysis should therefore be a distribution of outcomes, not an isolated answer. Teams need to know how often a business is cited, how often it leads recommendations, which domains appear alongside it, and how those measures vary by engine, market, prompt, and service category.
AI engines draw from different source ecosystems
The volatility is only part of the measurement challenge. The study found limited overlap between Gemini and ChatGPT: about 8% of cited domains overlapped, while only about 4.2% of top-1 brands overlapped. A business that appears prominently in one assistant should not assume it will receive comparable exposure in another.
Source patterns also varied by vertical. Gemini cited a business’s own website in 60% of citations, while Reddit and directories played outsized roles in different categories. Those findings argue against a universal optimization checklist. Local AI discovery depends on the sources each engine selects and how those source ecosystems behave for a particular service category.
For enterprise teams, the study supports several operational conclusions:
- Test repeatedly because one response cannot establish a reliable visibility position.
- Monitor multiple engines because Gemini and ChatGPT can surface different businesses and domains.
- Segment reporting by metro, service category, prompt type, and engine rather than using a single blended score.
- Separate observation from action by documenting sample sizes, prompt design, timing, and decision rules before changing content or local-search strategy.
Local AI discovery is becoming a measurement and governance problem, not simply an SEO reporting exercise. Scalevise helps teams establish repeatable, multi-engine visibility baselines, interpret volatile recommendation patterns, and connect findings to responsible content and operating decisions. Its AI Visibility / GEO Checker gives stakeholders a practical starting point for examining how brands appear across AI answers rather than treating one response as definitive. Start an AI Visibility scan.
How enterprises can govern AI visibility measurement
A useful governance model begins by defining what the organization is measuring. Citation presence, top recommendation frequency, source-domain representation, and consistency across repeat tests are distinct metrics. Combining them without clear definitions can obscure whether an apparent change reflects an actual pattern, a change in prompt construction, or normal generative variation.
Build a repeatable sampling framework
A repeatable framework should use standardized local-intent prompts and run them multiple times across the engines relevant to the business. Results should retain the prompt wording, engine, location, service category, run timing, cited domains, and recommendation position. This creates an auditable record and allows teams to compare distributions rather than anecdotes.
The aim is not to assume that AI answers will reproduce traditional search rankings. It is to understand the conditions under which a brand is represented, the sources associated with that representation, and the degree of uncertainty decision-makers should attach to the result.
Tie actions to evidence, not isolated outputs
Governance also requires thresholds for action. If a brand is absent from one AI response, that alone is weak evidence for a major content, directory, or local-listing intervention. Repeated patterns across relevant queries and engines provide a more defensible basis for prioritization.
This distinction matters for larger organizations managing multiple markets and categories. Without a common testing protocol, separate teams may draw conflicting conclusions from different AI responses. A documented measurement approach helps ensure that AI visibility reporting informs decisions without overstating what any individual generated answer means.
Frequently Asked Questions
What is Grounding Drift in local AI search?
Grounding Drift is Steady Demand’s term for the variation in AI citations and recommendations when the same local-intent query is repeated. In the study, identical prompts produced overlapping cited sources only about 40% of the time.
How consistent were Gemini’s top local recommendations?
Gemini returned the same top local-business recommendation only about 7% of the time across repeated identical queries in Steady Demand’s study.
How did Gemini compare with Google’s Local Pack?
Google’s traditional Local Pack returned the same top listing roughly 90% of the time on identical repeated queries, substantially more consistently than Gemini’s generated top recommendation.
Why should businesses measure more than one AI engine?
Gemini and ChatGPT showed limited citation overlap in the study, with about 8% of domains and about 4.2% of top-1 brands overlapping. Multi-engine measurement helps reveal engine-specific visibility patterns.
Conclusion
Steady Demand’s findings show that AI-driven local discovery cannot be assessed with the assumptions used for conventional rank tracking. Gemini’s volatile top recommendations, limited overlap with ChatGPT, and vertical-specific source patterns make repeated, multi-engine measurement essential. Businesses that treat AI visibility as a governed distribution of outcomes will have a firmer basis for interpreting results and deciding where to focus.