Grounding Drift
Ask Gemini the identical local-business question twice and only about 43% of the cited domains overlap. The September retest confirms the instability, but it does not confirm our earlier claim that overlap gets worse as more time passes. This page shows the original experiment, the larger correction, the local-pack control, and what the result does—and does not—mean for measurement.
Major findings
Average cited-domain overlap for identical questions repeated the same day across 497 queries.
Jump to the experiment →Same-day and next-day overlap differed by 0.2 percentage points; 95% CI −0.9 to +1.2. The July decay claim did not survive the larger test.
Jump to the correction →Top-business repeat rate across five rounds. Loose list overlap was 38.3% for Gemini and 83.9% for Google’s local pack.
Jump to the control →Average cited-domain overlap on 1,500 matched queries. The same top business appeared on 4.9% of the 783 queries where both engines named one.
Jump to the comparison →Cross-phrasing citation overlap was close to the 43.0% identical-repeat result. Phrasing and repeat-run instability are separate measurements.
Jump to the experiment →Overlap among Gemini’s exposed grounding search strings in the original four rounds, versus 26.5%–46.3% for cited domains.
Jump to the evidence →Google local-pack top-business repeatability declined between July and September. It remained far more stable than Gemini.
Jump to the control →Gemini’s loose list overlap increased between waves, while strict top-business repeatability moved from 7.9% to 5.1%.
Jump to the wave table →With 43% citation overlap, a repeated identical query changes most of the combined cited-domain set. Presence or absence once is weak evidence of persistent visibility.
Jump to the implications →SISTRIX found substantial week-to-week citation churn across AI Overviews, AI Mode, and ChatGPT Search. Its metric and time scale differ from this experiment.
Jump to prior research →The Experiment
The main dataset already turned up something odd: the same underlying question, worded three different ways (“best X in Y,” “top rated X near Y,” “who is a good X in Y”), only agreed on cited sources 41% of the time on average. The obvious explanation is that different phrasing pulls up different sources — that’s exactly what “phrasing sensitivity” should look like. So we built a follow-up to isolate it: six fixed queries (dentist, auto repair, house cleaning, HVAC, plumber, and personal injury lawyer, all “best X in New York, NY”), called with zero wording change at all, multiple times.
- 6 fixed queries, repeated
- 4 separate rounds, same wording every time
- 72 total repeat-test calls logged
We measured two things on every call: which domains got cited (Jaccard similarity between the
cited-domain sets of each pair of repeats), and which search strings the model’s own
grounding tool generated internally to go find those domains (web_search_queries,
exposed directly in Gemini’s grounding metadata).
Then, to rule out “maybe the live web just changed in the meantime,” we ran the same six queries a third time, roughly 3.5 hours later the same day:
The elapsed-time number (0.410 / 0.401) doesn’t sit meaningfully below the back-to-back number (0.463) — if anything, it’s in the same band. Whatever’s driving the disagreement, it isn’t the live web changing between calls. It just isn’t.
That test only covered a gap of hours, still inside the same calendar day. This is meant to be a live, checked-not-just-trusted study, so we ran the same six queries again the following day — roughly 19–20 hours after the original rounds — to see whether a real day boundary (a fresh crawl cycle, a different day of the week, more time for the live web to actually move) changes the picture.
The cross-day figure (0.380, or 0.363 against the delayed round specifically) lands close to the same-day cross-round number (0.401) — not a sharp drop-off. What looked more telling at the time was the second number: the next day’s own three repeats agreed with each other only 0.265 of the time, against 0.332, 0.410 and 0.463 the day before. I read that as a day boundary making things worse.
Instead of six queries three times, the re-test compares full rounds of 497 queries against each other — ten pairs of rounds separated by two to ten hours, and five pairs separated by a day and a bit. If elapsed time degrades agreement, those two groups must differ.
Citation-set agreement by how far apart the two rounds ran. 497 queries per comparison, September 2026 wave. Difference +0.002, 95% CI -0.009 to +0.012.
0.429 against 0.427. The gap is +0.002 with a confidence interval of -0.009 to +0.012 — centred on nothing, at roughly eighty times July’s sample size. Ask Gemini the same question two hours apart or a day and a half apart and you get the same amount of disagreement either way.
So the July next-day number was real in its own sample — all six queries moved together, and it holds up statistically within that sample — but it doesn’t generalise. Six queries in one city across a single night and the following day isn’t enough to establish a decay curve, and I presented it with more confidence than it had earned. Time is not the variable. The disagreement is large and it is constant, which if anything is the more useful finding: there’s no quiet period to wait for and no freshness window to exploit. It’s like this all the time.
02Why It Happens
The search-string numbers are the tell. If two identical prompts disagreed on cited domains but agreed on the underlying search strings the model used to find them, that would point to something downstream — ranking, or the live index shifting under a fixed query. That’s not what happened. The search strings themselves barely overlapped at all (0.011–0.056 across all four rounds), consistently far less than the domains that came out the other end (0.265–0.463). The observed variation appears upstream: Gemini exposed different grounding search strings for identical prompts before returning different citation sets.
Put plainly: in this test, asking the same question twice did not produce the same grounding searches. The input was byte-identical, but the exposed search strings differed. The data is consistent with a stochastic query-generation step; it does not reveal the complete internal mechanism.
- Not phrasing sensitivity. These were the same words, not different ones.
- Not the live web changing. The 3.5-hour-later round didn’t disagree any more than the back-to-back one did.
- The evidence points to query generation. Search strings varied more than the resulting citations, placing at least part of the observed variation before retrieval.
Is This Just Gemini, or All of Local Search?
- 500 combinations, identical to the Gemini side
- 4,940 pairwise round comparisons, Google local pack (complete, 5 rounds)
- 4,996 pairwise round comparisons, Gemini (complete, 5 rounds)
This is the noise floor — every wave-over-wave comparison in this series gets measured against it.
| Wave | Gemini repeat same query, 5 rounds |
Google local pack repeat | Cross-phrasing |
|---|---|---|---|
| July 2026 | 7.9% strict / 17.7% loose | 90.2% strict / 88.4% loose | 40.3% |
| August 2026 | pending | pending | pending |
| September 2026 | 5.1% strict / 38.3% loose | 83.3% strict / 83.9% loose | 41.2% |
Maybe local search is just noisy and Gemini is reflecting ordinary churn? Fair question, so we ran the same 500 metro/vertical combinations on the same repeated cadence against Google’s own local pack. Google returned the same top business 83.3% of the time (0.839 average loose overlap). Gemini managed 5.1% (0.383) — the instability belongs to the AI answer, not to local search.
Worth saying plainly, because it cuts against my own framing: the control moved too. Google’s local pack was at 90.2% in July and is at 83.3% now. The contrast with Gemini is still about 16x, so the finding holds — but “one surface holds still while the other doesn’t” is an overstatement I’m not going to keep making. Neither surface is fixed. One is just far less unfixed than the other.
→ Read the full control test in the Report (Finding Eight)
04Is This Just Gemini, or Every AI Engine?
- 1,500 identical queries answered by both engines
- 3,996 ChatGPT citations classified into the same domain taxonomy used throughout this study
- 783 queries where both engines named at least one recognizable business, for a direct recommendation comparison
The tested Gemini surface is much less repeatable than the local pack. To see whether the broader disagreement extended beyond Gemini, we ran the identical 1,500-query design through ChatGPT’s search mode and compared head to head. The two engines read different parts of the web: Gemini sends 49.6% of its citations to the business’s own site, ChatGPT just 10.5%, while ChatGPT leans on general business directories instead (46.0% vs. 17.7% for Gemini). Community forums used to be the sharpest split of all — 41.7% of ChatGPT’s July citations against 14.4% of Gemini’s — but both engines have since dropped them, to 0.1% and 3.9%. On the same question, only 8.3% of cited domains match, the same top business gets named 4.9% of the time, and the full recommended lists overlap at 0.052 (loose Jaccard). Their returned source and recommendation sets differ substantially.
→ Read the full comparison in the Report (Finding Nine) · Browse every matched query side by side · Full methodology for this comparison
05We’re Not the Only Ones Seeing This
Our repeat-test was small and deliberately narrow — 72 calls, one AI engine, one city, six verticals. Before I treat it as a real phenomenon rather than a quirk of our own pipeline, it’s worth checking whether anyone else, using a completely different method, found the same underlying instability. As of mid-2026, several independent groups have — including one that landed on a strikingly similar name for a related pattern, which is worth addressing head-on.
AI Citation Drift: How Stable Are Sources in AI Search Results?
A far bigger study: 82,619 prompts and 1,548,213 snapshots, across six countries, tracked weekly for 17 weeks (December 2025–April 2026), across three platforms — Google AI Overviews, Google AI Mode, and ChatGPT Search. SISTRIX measures something different from what we did (week-over-week domain churn, not same-day identical-repeat overlap), but the headline finding rhymes with ours: instability that doesn’t level off over time. Weekly domain churn-in ran 5% for Google AI Overviews (the most stable of the three), 56% for Google AI Mode, and 74% for ChatGPT Search — and the study found no sign of these rates settling down across all 17 weeks measured. They call it a structural feature of the platforms, not a temporary launch effect.
Naming note: this study predates ours and uses the term “Citation Drift” for its week-over-week churn metric. We deliberately call our own finding Grounding Drift instead — not to distance ourselves from SISTRIX’s work (we’re citing it here precisely because it’s relevant), but because the two measure genuinely different things at different time scales, and reusing an existing published term for a different metric would misrepresent both. Our name also points at what our data actually isolates: the instability traces back to Gemini’s grounding step specifically, not citation churn in general.
How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews
Accepted to the 49th ACM SIGIR Conference (Melbourne, July 2026), this is peer-reviewed academic research comparing Google Search, Gemini, and AI Overviews directly — the same three systems this project’s methodology touches most. Its existence matters as much as any single number in it: this isn’t a fringe question. It’s an active, credentialed research area, and academic work in the space has separately concluded that generative search systems show materially lower response consistency than traditional search — the same direction as every other study cited here.
Optimizing Visibility in Generative Engines: A Critical Survey of GEO
A 2026 survey reviewing 45 Generative Engine Optimization studies since the field’s 2023 founding paper (Princeton’s original “GEO: Generative Engine Optimization”) reaches a practical conclusion that lines up with ours from a completely different direction: optimizing for a single query is an unstable target, because it can’t capture the topic-level citation preferences that actually persist across semantic variations. Their advice to practitioners — define a query set, run multiple repeats, log citations, and only then change anything — is basically the same experimental discipline this repeat-test was built around.
What This Actually Costs a Business
The question most local business owners actually lose sleep over isn’t “is Grounding Drift real?” It’s some version of “does ChatGPT even know I exist?” Drift makes that question harder to answer honestly, in both directions.
SOCi’s 2026 Local Visibility Index analyzed over 350,000 business locations across 2,751 brands and found that ChatGPT currently recommends just 1.2% of local business locations outright. The same research found only about 45% overlap between businesses that rank well in traditional Google results and those that also show up in AI recommendations — a strong Google ranking, on its own, is no longer a reliable proxy for AI visibility.
Grounding Drift makes that gap worse, not better. A business that checks once, gets no citation, and concludes “I’m invisible to AI” might be wrong — our own data shows a second identical check has better than even odds of coming back different. And a business that checks once, sees itself cited, and relaxes might be just as wrong — that result wasn’t necessarily earned by anything stable, and a second check isn’t guaranteed to repeat it. Neither a good result nor a bad one, taken from a single check, tells you as much as it feels like it does.
- A single “am I cited” test is closer to a coin flip with better-than-even odds than a verdict. Treat it as one data point, not the answer.
- The businesses actually at risk aren’t the ones who got a bad result once. They’re the ones who never check twice, in either direction.
- A real Google ranking no longer means AI visibility follows automatically. They’re two separate things to manage now.
How to Measure Through the Noise
This study tested repeatability, not an optimization intervention. It does not show that adding pages, listings, or mentions causes more citations. It does show that a single check is too unstable to support a confident visibility claim.
- Repeat identical questions. That separates run-to-run instability from wording changes.
- Test a set of phrasings. One prompt cannot represent a topic or the ways customers ask about it.
- Keep engines separate. Gemini and ChatGPT cite different source pools and should not be folded into one score.
- Compare dated waves. Use the same query set and method each time, then interpret movement against observed repeat-run variability.
Further Reading
- Beus, J. — “AI Citation Drift: How Stable Are Sources in AI Search Results?” SISTRIX, published May 1, 2026. 82,619 prompts, 1,548,213 snapshots, 6 countries, 3 platforms, 17-week tracking window.
- Grossman, R., Liu, S., Chen, M.K., Smith, M., Borcea, C. & Chen, Y. — “How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews” Proceedings of the 49th International ACM SIGIR Conference, Melbourne, July 2026.
- “Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)” arXiv:2607.14035, reviewing 45 GEO studies since the field’s 2023 founding paper.
- SOCi — 2026 Local Visibility Index 350,000+ business locations analyzed across 2,751 brands. 1.2% ChatGPT recommendation rate and ~45% Google/AI ranking overlap figures as reported via MarketingCode, 2026.
Need AI Optimization help?
See the stats/study in context: Steady Demand Research Index
See the stats/study in context: Steady Demand Research Index
