ChatGPT and Gemini Rarely Agree on Top Local Businesses, Study Finds
Cross-engine research into local-service searches finds that AI visibility varies sharply between ChatGPT and Gemini, making single-platform measurement unreliable for SMBs.
AI visibility is not a single score that a business can measure once and treat as settled. A cross-engine study of local-service searches found that ChatGPT and Gemini named the same top business in only 4.2% of identical queries. For small businesses trying to be discovered through AI assistants, that gap means a strong result in one engine may say very little about how another assistant presents the market.
The research, published by Steady Demand in its AI Citation Ledger, examined 1,487 queries across 50 U.S. metropolitan areas and 10 service verticals. It focused on prompts such as “best plumber near me,” tracking the businesses named and the sources used to ground responses. Its central finding is practical: AI-driven discovery is fragmented by engine, source mix, location, and category.
That does not prove that AI responses drive more leads than conventional local search. The study measures citations and top-name outcomes, not conversions or overall ranking quality. But it provides a useful baseline for marketers because it shows why checking a brand in one AI assistant is not enough to understand its broader AI visibility.
What the cross-engine data shows
The study compared how Gemini and ChatGPT answered the same local-business prompts. Their differences extended beyond the final recommendation. The systems often drew on different source ecosystems, which helps explain why they surface different businesses.
| Measure | Gemini | ChatGPT |
|---|---|---|
| Exact top-business match between engines | 4.2% of identical queries produced the same top business | |
| Typical citation mix | About 60% of citations were business websites | More reliance on Reddit and traditional directories |
| Overlap in cited domains | About 8% overlap | |
| Repeated-query source alignment | About 40% alignment, described as grounding drift | |
| Top-result repeatability benchmark | About 7% top-match stability in AI-generated results | Not specified separately in the supplied research |
The contrast with Google’s local pack is notable. In the study’s benchmark, Google local-pack top results reappeared about 90% of the time, compared with roughly 7% top-match stability for Gemini’s AI-generated results. That does not make one channel inherently better. It does show that an AI answer can be more variable than the familiar local-search result set businesses have historically monitored.
The source patterns also matter. Gemini’s heavier use of business websites suggests that clear, accessible on-site information can be especially important in its answers. ChatGPT’s greater use of Reddit and directories means third-party discussion and listing accuracy may have more influence there. Neither pattern supports a universal checklist, because the source mix changes by service category.
For example, the research found different citation behavior across verticals including personal injury law, dentistry, and auto repair. National platforms such as Angi, the Better Business Bureau, and Reddit can serve as common reference points across metros, while regional differences still affect which businesses and sources appear.
Why repeated prompts can produce different evidence
A second challenge is repeatability. When the researchers ran queries word for word again, cited sources aligned only about 40% of the time. The report calls this grounding drift: variation in the sources an AI system uses to support a response over repeated runs.
For a business, this means a one-off screenshot of a favorable recommendation is weak evidence of durable visibility. The same applies to a negative or missing mention. A useful benchmark needs repeated prompts, a defined location, and separate reporting for each engine.
A practical measurement approach for SMBs
The findings point toward a more disciplined way to evaluate AI discovery. Rather than asking whether a company “ranks in AI,” businesses can track whether they are consistently named and grounded across the assistants their customers may use.
A practical monitoring process should include:
- Test multiple engines separately, at minimum ChatGPT and Gemini when they are relevant to the audience.
- Use representative prompts that reflect real customer language, services, locations, and decision stages.
- Repeat the same prompts over time to distinguish a one-off answer from a recurring pattern.
- Record both the business named and the cited sources, because sources reveal where each engine is finding evidence.
- Segment results by service category and market, since a tactic that helps one vertical or metro may not transfer to another.
This is not a call to chase every mention. It is a way to identify gaps that can be addressed with evidence-based work. If Gemini frequently cites business websites but a company has thin service pages or inconsistent location details, improving those pages may make its information easier to ground. If ChatGPT repeatedly draws on directories or community discussions, the business can review whether core listings are accurate and whether public information about its services is clear.
Structured data can be part of that work when it accurately represents the business and its services. However, the study does not establish that a particular markup implementation, directory, or content tactic guarantees inclusion in either assistant. Its evidence supports monitoring the source environment, not promising a universal optimization formula.
For marketing and customer-support teams, the same principle applies to research workflows. AI assistants can be useful for discovering how consumers may encounter a brand, but their responses should not be treated as a stable market ranking or as a substitute for verified local-search data. Teams should document the engine, prompt, location, date, cited sources, and repeat runs when using outputs in decisions.
AI discovery can create visibility opportunities, but unmanaged variation can also hide missed demand. Scalevise helps SMBs assess how their brands appear across AI answer engines, identify source and content gaps, and turn findings into a focused visibility plan rather than disconnected experiments. Use the AI Visibility and GEO Checker to establish a cross-engine baseline for the queries that matter to your customers, then start an AI Visibility scan.
Frequently Asked Questions
Why did ChatGPT and Gemini name the same top business only 4.2% of the time?
The study found that the two systems often relied on different cited sources. Gemini skewed toward business websites, while ChatGPT relied more on Reddit and traditional directories, producing different top-name outcomes.
What is grounding drift in AI search?
Grounding drift is variation in the sources cited when an AI assistant receives the same prompt again. In this study, repeated word-for-word queries had about 40% source alignment.
Should local businesses optimize only for Gemini or ChatGPT?
No. The research indicates that visibility differs substantially between the two engines. Businesses should monitor relevant AI assistants separately rather than treating one engine as representative of all AI discovery.
Does this study show that AI visibility increases conversions?
No. The research measures citations and businesses named in answers. It does not measure downstream conversions or establish the overall quality of rankings.
Do the findings apply outside U.S. local-service businesses?
The study covers 10 local-service categories in 50 U.S. metros, so its findings may not generalize to other regions or industries. It is best used as evidence that engine-specific monitoring is necessary, not as a universal forecast.
Conclusion
Steady Demand’s research shows that AI visibility is an engine-specific and variable form of consumer discovery. ChatGPT and Gemini can cite different evidence, name different businesses, and change their grounding across repeated prompts. For SMBs, the practical response is not a single optimization tactic. It is consistent, cross-engine measurement that reveals where brand information is being found, where it is absent, and which gaps are worth addressing.