• Skip to primary navigation
  • Skip to main content
Steady Demand Google Business Profile and Local Services Ads management and Local SEO agency.

Steady Demand Google Business Profile and Local Services Ads management and Local SEO agency.

Google Business Profile and Local Services Ads management and Local SEO agency.

  • Services
    • AI Optimization Services – AEO – GEO – AIO. Prove to Help – Backed by Data
    • GBP Management
    • GBP Listing Reinstatement
    • GBP Listing Assurance
    • Local SEO Services
    • Google Business Profile (GBP) Fake Listing Spam Reporting
    • Social Media Services
    • Local Services Ads (LSA) Management
    • LocalPics
  • Whitelabel
  • Free AI and GBP Tools
    • Steady Demand Research Index
    • Free AI Optimization Audit Tool
    • Free GBP Audit Tool
    • Free AI Text Detection Tool – Improve Human Score
  • About
    • Where Ben and Steady Demand speak and write
    • Testimonials
    • Case Studies
  • Blog
  • Contact
  • 888-778-0401
  • Clients
  • Show Search
Hide Search

Grounding Drift: Why AI Search Citations Change Every Time in ChatGPT and Gemini

Ben Fisher · August 6, 2026 ·

SteadyDemand
Full Report Dashboard Grounding Drift Raw Data Gemini vs. ChatGPT Methodology What Changed
The Citation Ledger · Deep Dive

Grounding Drift

Ask Gemini the identical local-business question twice and only about 43% of the cited domains overlap. The September retest confirms the instability, but it does not confirm our earlier claim that overlap gets worse as more time passes. This page shows the original experiment, the larger correction, the local-pack control, and what the result does—and does not—mean for measurement.

13 min read 1,500 queries, 50 metros analyzed
Updated Sept 9, 2026 Reddit has now fallen out of both major AI engines — 13.7% to 1.7% of Gemini’s citations, five weeks after the same thing happened in ChatGPT. Directories absorbed the share in both, and a re-test retired the Grounding Drift page’s next-day decay finding. Read what changed →

Major findings

Repeat instability 43.0%

Average cited-domain overlap for identical questions repeated the same day across 497 queries.

Jump to the experiment →
Correction: no time decay 43.0% vs. 42.7%

Same-day and next-day overlap differed by 0.2 percentage points; 95% CI −0.9 to +1.2. The July decay claim did not survive the larger test.

Jump to the correction →
Gemini versus local pack 5.1% vs. 83.3%

Top-business repeat rate across five rounds. Loose list overlap was 38.3% for Gemini and 83.9% for Google’s local pack.

Jump to the control →
Gemini versus ChatGPT 8.3%

Average cited-domain overlap on 1,500 matched queries. The same top business appeared on 4.9% of the 783 queries where both engines named one.

Jump to the comparison →
Wording does not explain it 41.2%

Cross-phrasing citation overlap was close to the 43.0% identical-repeat result. Phrasing and repeat-run instability are separate measurements.

Jump to the experiment →
Search strings varied more 1.1%–5.6%

Overlap among Gemini’s exposed grounding search strings in the original four rounds, versus 26.5%–46.3% for cited domains.

Jump to the evidence →
Supporting: local pack moved too 90.2% → 83.3%

Google local-pack top-business repeatability declined between July and September. It remained far more stable than Gemini.

Jump to the control →
Supporting: Gemini list overlap rose 17.7% → 38.3%

Gemini’s loose list overlap increased between waves, while strict top-business repeatability moved from 7.9% to 5.1%.

Jump to the wave table →
Supporting: one check is not a verdict 57%

With 43% citation overlap, a repeated identical query changes most of the combined cited-domain set. Presence or absence once is weak evidence of persistent visibility.

Jump to the implications →
Supporting: broader research agrees on instability 1.55M snapshots

SISTRIX found substantial week-to-week citation churn across AI Overviews, AI Mode, and ChatGPT Search. Its metric and time scale differ from this experiment.

Jump to prior research →
43.0%
How much two byte-identical repeat queries agree on cited sources — same text, same day. The September 2026 test repeated 497 queries across ten same-day round pairs. Cross-phrasing agreement was similar at 41.2%, but it is a separate measure.
On This Page
  1. The Experiment
  2. Why It Happens
  3. Is This Just Gemini, or All of Local Search?
  4. Is This Just Gemini, or Every AI Engine?
  5. We’re Not the Only Ones Seeing This
  6. What This Actually Costs a Business
  7. How to Measure Through the Noise
  8. Sources
01

The Experiment

The main dataset already turned up something odd: the same underlying question, worded three different ways (“best X in Y,” “top rated X near Y,” “who is a good X in Y”), only agreed on cited sources 41% of the time on average. The obvious explanation is that different phrasing pulls up different sources — that’s exactly what “phrasing sensitivity” should look like. So we built a follow-up to isolate it: six fixed queries (dentist, auto repair, house cleaning, HVAC, plumber, and personal injury lawyer, all “best X in New York, NY”), called with zero wording change at all, multiple times.

  • 6 fixed queries, repeated
  • 4 separate rounds, same wording every time
  • 72 total repeat-test calls logged

We measured two things on every call: which domains got cited (Jaccard similarity between the cited-domain sets of each pair of repeats), and which search strings the model’s own grounding tool generated internally to go find those domains (web_search_queries, exposed directly in Gemini’s grounding metadata).

Cited domains, back-to-back repeats
0.463
Jaccard similarity, same-day immediate repeats
Model’s own search strings, same repeats
0.056
Jaccard similarity — barely any overlap

Then, to rule out “maybe the live web just changed in the meantime,” we ran the same six queries a third time, roughly 3.5 hours later the same day:

Cited domains, ~3.5 hours later
0.410
Search strings, ~3.5 hours later
0.020
Cross-round comparison
0.401
Domains, immediate round vs. 3.5-hours-later round, directly compared

The elapsed-time number (0.410 / 0.401) doesn’t sit meaningfully below the back-to-back number (0.463) — if anything, it’s in the same band. Whatever’s driving the disagreement, it isn’t the live web changing between calls. It just isn’t.

That test only covered a gap of hours, still inside the same calendar day. This is meant to be a live, checked-not-just-trusted study, so we ran the same six queries again the following day — roughly 19–20 hours after the original rounds — to see whether a real day boundary (a fresh crawl cycle, a different day of the week, more time for the live web to actually move) changes the picture.

Cited domains, ~19–20 hours later
0.380
Cross-round comparison vs. the original immediate round
Repeat-consistency, next-day round itself
0.265
Three next-day repeats compared against each other

The cross-day figure (0.380, or 0.363 against the delayed round specifically) lands close to the same-day cross-round number (0.401) — not a sharp drop-off. What looked more telling at the time was the second number: the next day’s own three repeats agreed with each other only 0.265 of the time, against 0.332, 0.410 and 0.463 the day before. I read that as a day boundary making things worse.

Corrected — September 2026: that reading didn’t survive a bigger sample Two things about the paragraph above. First, I originally quoted only the two highest of the previous day’s three rounds (0.463 and 0.410), which made 0.265 look like more of an outlier than it was — the third round was 0.332. Second, and more important: the whole thing rested on six queries in one city. The September 2026 wave re-ran it properly, and the effect is not there.

Instead of six queries three times, the re-test compares full rounds of 497 queries against each other — ten pairs of rounds separated by two to ten hours, and five pairs separated by a day and a bit. If elapsed time degrades agreement, those two groups must differ.

Rounds 2–10 hours apart 43.0%
Rounds 30–40 hours apart 42.7%

Citation-set agreement by how far apart the two rounds ran. 497 queries per comparison, September 2026 wave. Difference +0.002, 95% CI -0.009 to +0.012.

0.429 against 0.427. The gap is +0.002 with a confidence interval of -0.009 to +0.012 — centred on nothing, at roughly eighty times July’s sample size. Ask Gemini the same question two hours apart or a day and a half apart and you get the same amount of disagreement either way.

So the July next-day number was real in its own sample — all six queries moved together, and it holds up statistically within that sample — but it doesn’t generalise. Six queries in one city across a single night and the following day isn’t enough to establish a decay curve, and I presented it with more confidence than it had earned. Time is not the variable. The disagreement is large and it is constant, which if anything is the more useful finding: there’s no quiet period to wait for and no freshness window to exploit. It’s like this all the time.

02

Why It Happens

The search-string numbers are the tell. If two identical prompts disagreed on cited domains but agreed on the underlying search strings the model used to find them, that would point to something downstream — ranking, or the live index shifting under a fixed query. That’s not what happened. The search strings themselves barely overlapped at all (0.011–0.056 across all four rounds), consistently far less than the domains that came out the other end (0.265–0.463). The observed variation appears upstream: Gemini exposed different grounding search strings for identical prompts before returning different citation sets.

Put plainly: in this test, asking the same question twice did not produce the same grounding searches. The input was byte-identical, but the exposed search strings differed. The data is consistent with a stochastic query-generation step; it does not reveal the complete internal mechanism.

What this rules out
  • Not phrasing sensitivity. These were the same words, not different ones.
  • Not the live web changing. The 3.5-hour-later round didn’t disagree any more than the back-to-back one did.
  • The evidence points to query generation. Search strings varied more than the resulting citations, placing at least part of the observed variation before retrieval.
03

Is This Just Gemini, or All of Local Search?

  • 500 combinations, identical to the Gemini side
  • 4,940 pairwise round comparisons, Google local pack (complete, 5 rounds)
  • 4,996 pairwise round comparisons, Gemini (complete, 5 rounds)

This is the noise floor — every wave-over-wave comparison in this series gets measured against it.

Wave Gemini repeat
same query, 5 rounds
Google local pack repeat Cross-phrasing
July 20267.9% strict / 17.7% loose90.2% strict / 88.4% loose40.3%
August 2026pendingpendingpending
September 20265.1% strict / 38.3% loose83.3% strict / 83.9% loose41.2%

Maybe local search is just noisy and Gemini is reflecting ordinary churn? Fair question, so we ran the same 500 metro/vertical combinations on the same repeated cadence against Google’s own local pack. Google returned the same top business 83.3% of the time (0.839 average loose overlap). Gemini managed 5.1% (0.383) — the instability belongs to the AI answer, not to local search.

Worth saying plainly, because it cuts against my own framing: the control moved too. Google’s local pack was at 90.2% in July and is at 83.3% now. The contrast with Gemini is still about 16x, so the finding holds — but “one surface holds still while the other doesn’t” is an overstatement I’m not going to keep making. Neither surface is fixed. One is just far less unfixed than the other.

→ Read the full control test in the Report (Finding Eight)

04

Is This Just Gemini, or Every AI Engine?

  • 1,500 identical queries answered by both engines
  • 3,996 ChatGPT citations classified into the same domain taxonomy used throughout this study
  • 783 queries where both engines named at least one recognizable business, for a direct recommendation comparison

The tested Gemini surface is much less repeatable than the local pack. To see whether the broader disagreement extended beyond Gemini, we ran the identical 1,500-query design through ChatGPT’s search mode and compared head to head. The two engines read different parts of the web: Gemini sends 49.6% of its citations to the business’s own site, ChatGPT just 10.5%, while ChatGPT leans on general business directories instead (46.0% vs. 17.7% for Gemini). Community forums used to be the sharpest split of all — 41.7% of ChatGPT’s July citations against 14.4% of Gemini’s — but both engines have since dropped them, to 0.1% and 3.9%. On the same question, only 8.3% of cited domains match, the same top business gets named 4.9% of the time, and the full recommended lists overlap at 0.052 (loose Jaccard). Their returned source and recommendation sets differ substantially.

→ Read the full comparison in the Report (Finding Nine) · Browse every matched query side by side · Full methodology for this comparison

05

We’re Not the Only Ones Seeing This

Our repeat-test was small and deliberately narrow — 72 calls, one AI engine, one city, six verticals. Before I treat it as a real phenomenon rather than a quirk of our own pipeline, it’s worth checking whether anyone else, using a completely different method, found the same underlying instability. As of mid-2026, several independent groups have — including one that landed on a strikingly similar name for a related pattern, which is worth addressing head-on.

SISTRIX — Johannes BeusPublished May 2026

AI Citation Drift: How Stable Are Sources in AI Search Results?

A far bigger study: 82,619 prompts and 1,548,213 snapshots, across six countries, tracked weekly for 17 weeks (December 2025–April 2026), across three platforms — Google AI Overviews, Google AI Mode, and ChatGPT Search. SISTRIX measures something different from what we did (week-over-week domain churn, not same-day identical-repeat overlap), but the headline finding rhymes with ours: instability that doesn’t level off over time. Weekly domain churn-in ran 5% for Google AI Overviews (the most stable of the three), 56% for Google AI Mode, and 74% for ChatGPT Search — and the study found no sign of these rates settling down across all 17 weeks measured. They call it a structural feature of the platforms, not a temporary launch effect.

Naming note: this study predates ours and uses the term “Citation Drift” for its week-over-week churn metric. We deliberately call our own finding Grounding Drift instead — not to distance ourselves from SISTRIX’s work (we’re citing it here precisely because it’s relevant), but because the two measure genuinely different things at different time scales, and reusing an existing published term for a different metric would misrepresent both. Our name also points at what our data actually isolates: the instability traces back to Gemini’s grounding step specifically, not citation churn in general.

Grossman, Liu, Chen, Smith, Borcea & ChenSIGIR 2026 (peer-reviewed)

How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews

Accepted to the 49th ACM SIGIR Conference (Melbourne, July 2026), this is peer-reviewed academic research comparing Google Search, Gemini, and AI Overviews directly — the same three systems this project’s methodology touches most. Its existence matters as much as any single number in it: this isn’t a fringe question. It’s an active, credentialed research area, and academic work in the space has separately concluded that generative search systems show materially lower response consistency than traditional search — the same direction as every other study cited here.

GEO research field2023–2026 critical survey, arXiv 2607.14035

Optimizing Visibility in Generative Engines: A Critical Survey of GEO

A 2026 survey reviewing 45 Generative Engine Optimization studies since the field’s 2023 founding paper (Princeton’s original “GEO: Generative Engine Optimization”) reaches a practical conclusion that lines up with ours from a completely different direction: optimizing for a single query is an unstable target, because it can’t capture the topic-level citation preferences that actually persist across semantic variations. Their advice to practitioners — define a query set, run multiple repeats, log citations, and only then change anything — is basically the same experimental discipline this repeat-test was built around.

06

What This Actually Costs a Business

The question most local business owners actually lose sleep over isn’t “is Grounding Drift real?” It’s some version of “does ChatGPT even know I exist?” Drift makes that question harder to answer honestly, in both directions.

The visibility gap is already real, before drift enters the picture

SOCi’s 2026 Local Visibility Index analyzed over 350,000 business locations across 2,751 brands and found that ChatGPT currently recommends just 1.2% of local business locations outright. The same research found only about 45% overlap between businesses that rank well in traditional Google results and those that also show up in AI recommendations — a strong Google ranking, on its own, is no longer a reliable proxy for AI visibility.

Grounding Drift makes that gap worse, not better. A business that checks once, gets no citation, and concludes “I’m invisible to AI” might be wrong — our own data shows a second identical check has better than even odds of coming back different. And a business that checks once, sees itself cited, and relaxes might be just as wrong — that result wasn’t necessarily earned by anything stable, and a second check isn’t guaranteed to repeat it. Neither a good result nor a bad one, taken from a single check, tells you as much as it feels like it does.

What this means
  • A single “am I cited” test is closer to a coin flip with better-than-even odds than a verdict. Treat it as one data point, not the answer.
  • The businesses actually at risk aren’t the ones who got a bad result once. They’re the ones who never check twice, in either direction.
  • A real Google ranking no longer means AI visibility follows automatically. They’re two separate things to manage now.
07

How to Measure Through the Noise

This study tested repeatability, not an optimization intervention. It does not show that adding pages, listings, or mentions causes more citations. It does show that a single check is too unstable to support a confident visibility claim.

  • Repeat identical questions. That separates run-to-run instability from wording changes.
  • Test a set of phrasings. One prompt cannot represent a topic or the ways customers ask about it.
  • Keep engines separate. Gemini and ChatGPT cite different source pools and should not be folded into one score.
  • Compare dated waves. Use the same query set and method each time, then interpret movement against observed repeat-run variability.
Scope note: The original repeat-test behind most of this page is a small, targeted experiment — 72 calls across four rounds spanning two days, one metro, six verticals, one AI engine (Gemini’s grounding API) — built to test repeatability under fixed prompts, not to re-estimate the drift rate at scale. The control test above it is a separate, larger follow-up covering all 500 (metro, vertical) combinations from the main study, run across five timed rounds on both Gemini and, via a third-party SERP data provider, Google’s own local pack. The Gemini-vs-ChatGPT comparison is a third, separate follow-up covering the full 1,500-query main-study design against a second AI engine (ChatGPT’s search mode, via the same third-party provider), rather than repeated rounds of the same engine. The 41% cross-phrasing agreement figure comes from the full 1,500-query, 50-metro dataset described in the main report and its methodology page.
Sources

Further Reading

  • Beus, J. — “AI Citation Drift: How Stable Are Sources in AI Search Results?” SISTRIX, published May 1, 2026. 82,619 prompts, 1,548,213 snapshots, 6 countries, 3 platforms, 17-week tracking window.
  • Grossman, R., Liu, S., Chen, M.K., Smith, M., Borcea, C. & Chen, Y. — “How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews” Proceedings of the 49th International ACM SIGIR Conference, Melbourne, July 2026.
  • “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)” arXiv:2607.14035, reviewing 45 GEO studies since the field’s 2023 founding paper.
  • SOCi — 2026 Local Visibility Index 350,000+ business locations analyzed across 2,751 brands. 1.2% ChatGPT recommendation rate and ~45% Google/AI ranking overlap figures as reported via MarketingCode, 2026.
SteadyDemand

Need AI Optimization help?

Contact Us →
The Citation Ledger — a live study of AI local-search citations. Full Report · Dashboard · Raw Data · Gemini vs. ChatGPT · Methodology

See the stats/study in context: Steady Demand Research Index

See the stats/study in context: Steady Demand Research Index

AI

About Ben Fisher

As a specialist in local SEO, Ben has been helping businesses grow their online presence since 1994. Thanks to his contributions to the Google Business Profile Forum, Ben has been hand-picked by Google as a Google Business Profile Diamond Product Expert. Ben is also a contributor to the annual Moz Local Search Ranking Factors Study, and a regular contributor to BrightLocal.

Ben is the co-founder of Steady Demand, a local SEO company. The team at Steady Demand specializes in helping clients fight map spam, navigate the most complex Google My Business issues, and troubleshoot ranking issues on Google.

Contact Steady Demand

If you'd rather give us a call, or schedule an appointment we can be reached at (888) 778-0401.

Get the Latest Marketing Tips and Resources!

Get our weekly social media and content marketing advice, tips and resources sent straight to your inbox! No Spam, guaranteed!

On Social Media

  • Facebook
  • LinkedIn
  • Pinterest
  • Twitter
  • YouTube

© 2026 · Steady Demand, LLC

  • Home
  • Privacy
  • Terms
  • Contact