Ahrefs published a finding in December 2025 that has since become one of the most repeated statistics in AI search discussions. Analyzing September 2025 United States data from its Brand Radar tool, Ahrefs reported that AI Overviews and AI Mode cited the same URL only 13.7% of the time across 540,000 query pairs.
The number usually arrives as shorthand for two separate search engines wearing the same logo. That reading is directionally right and it gets flattened almost immediately, because a second number from the same study travels alongside it: 86% semantic similarity. Put together carelessly, the two produce a sentence like “the surfaces only overlap 86% of the time,” which misstates both figures.
The two numbers measure different things. AI Overviews and AI Mode tend to agree on what to say. They rarely agree on which pages deserve credit for saying it. Once you accept that, the useful question stops being “how different are these surfaces” and becomes operational: can your team sample both surfaces reliably enough to know when something actually changed, and what part of that job does a prompt library really do?
Quick answer
AI Overviews and AI Mode cite the same URL only about 13.7% of the time, but reach similarly worded conclusions about 86% of the time. A prompt library standardizes and version-controls the questions you ask both surfaces — it does not query them on a schedule or log citations for you. Tracking dual-surface citations requires a governed prompt set plus a separate execution and logging process, run at least three times per prompt per surface before you trust a change.
Two numbers, two different measurements
Keep these separate in every internal conversation, because they support different decisions.
13.7% is about sources. Ahrefs measured how often the identical URL appeared in both an AI Overview response and an AI Mode response for the same query, across 540,000 query pairs from September 2025 United States data. Narrowing to the top three citations in each response raised the overlap only to 16.3%. Read plainly: most of the time, the two surfaces pulled from different pages to answer the same question.
86% is about conclusions. Using cosine similarity across a separate sample of 730,000 response pairs, Ahrefs reported that roughly 89.7% of pairs scored above 0.8 on a scale where 1.0 means identical meaning, averaging out to 86% semantic similarity. Word-level overlap, a blunter measure of shared phrasing, sat at about 16%, and the two surfaces opened with the same first sentence only 2.51% of the time.
So the surfaces reach similar conclusions using different phrasing and different sources. If your team tracks only AI Overviews and treats that as AI search visibility, the citation half of your picture is missing for the other surface, and the conclusion half tells you very little about which of your pages is doing the work.
The caveat that should change how you sample
Ahrefs is explicit that the comparison covers single generations, meaning one AI Overview and one AI Mode answer captured at one point in time per query. That matters more than it first appears, because separate Ahrefs research found that 45.5% of AI Overview citations change between generations. The same query can return a different citation set on a later run with no change to any page involved.
Two consequences follow, and they point in opposite directions:
- The citation pools behind each surface may overlap more than a single snapshot suggests, since repeated runs can surface pages that one generation happened to skip.
- Any single run of your own monitoring is equally unreliable. A citation that disappears between Tuesday and Thursday may be variance rather than a loss.
Independent analyses point the same way on overlap. SE Ranking measured roughly 10.7% URL overlap on its own large-scale crawl. Victorious found that about 77% of unique cited domains across a 1,540-query sample appeared on only one surface, with no query in that sample returning identical citations on both. Three methodologies landing in the same range is a stronger signal than any one of them alone.
Treat all of these as dated vendor studies with stated samples rather than standing facts, and confirm current figures before you quote them in a board deck.
Victorious measured a related but different metric — the share of unique cited domains appearing on only one surface (about 77%) — so it isn’t plotted on this URL-overlap axis. It points the same direction: most citations do not repeat across surfaces.
What a prompt library actually does, and what it does not
This is where most advice on the topic gets loose, so here is the boundary in plain terms.
A prompt library stores and standardizes prompts. It gives a defined question set one home, keeps wording consistent between people, versions that wording when it changes, and controls who may edit it. In TeamAI, that means folders, search, role-based edit permissions, a chosen model per prompt, and variables written with double curly brackets so one prompt template covers many brands, competitors, or categories. The live walkthrough for that build sits in how to build and refine a prompt library.
A prompt library does not query Google AI Overviews on a schedule. It does not query AI Mode. It does not capture which URLs those surfaces cited, and it cannot correct for generation variance, because storing a question has nothing to do with running it repeatedly and recording answers.
That distinction is not pedantic. Teams buy a library, organize 40 excellent questions into tidy folders, and then discover six weeks later that nobody has actually run them against either surface and nothing has been recorded. The library was never the missing piece. The execution and logging process was.
The two-part system that does work
Split the job explicitly and staff both halves.
The definition layer is what makes the second half trustworthy. If two people phrase the same question slightly differently across weeks, you cannot tell whether a citation change came from the surface or from the prompt. Governance on wording is the control variable in your own experiment.
The execution layer is where the actual measurement happens, and it is the half teams underestimate.
Log at URL level, not domain level
Most monitoring spreadsheets record whether the brand appeared. That is too coarse to act on, because it cannot tell you which page earned the citation or whether the surface is pointing at something outdated.
Use one row per prompt, per surface, per run, with these fields. Click a field to see why it matters.
Two fields carry most of the weight. our_cited_url turns “we are visible” into “this specific page is doing the work.” all_cited_urls turns a brand check into a competitive record you can query later.
The repeated-run stability protocol
Given that citations shift between generations, a single run per prompt is a presence check, not a measurement. Use repetition as the unit of confidence.
- Run every prompt three times per surface per week, on three different days rather than three times in one sitting.
- Compute a stability rate for each prompt and surface: runs where your URL was cited, divided by total runs.
- Classify into bands. Stable is three of three. Intermittent is one or two of three. Absent is zero of three.
- Act on band changes across two consecutive weeks, not on single-run flips. One missing citation in one run is inside the expected noise the 45.5% figure describes.
- Record the band alongside the raw rows so the history stays auditable when someone challenges a decision months later.
The decision rule in step four is the part that saves time. Without it, teams chase noise weekly and treat normal variance as a crisis or a win.
Try it: check your stability band
Ran your three weekly checks? Enter how many of the three runs cited your URL, for each surface.
Cut prompt count before you cut run count
Monitoring budgets are finite, and the tradeoff is usually framed wrongly. Two plans with identical effort answer different questions:
- 25 prompts, 2 surfaces, 3 runs each: 150 observations per week. This answers whether your citations are stable.
- 75 prompts, 2 surfaces, 1 run each: 150 observations per week. This answers whether you appeared once, on one day, for a wider list.
Take the first. With a fixed budget, cut the prompt count before you cut the run count, because a single generation cannot distinguish a lost citation from routine variance, and the variance on these surfaces is common rather than exceptional. A narrower set you can actually trust beats a broad set that produces confident, unreliable readings.
The exception is a genuine discovery pass. If you have never sampled at all, one wide single-run sweep is a reasonable way to find which 25 prompts deserve the three-run treatment. Do that once, then narrow.
Weekly monitoring budget planner
Enter your own weekly observation budget (prompts × surfaces × runs) and see how many stability-grade prompts it actually buys.
A worked example
The table below is an illustrative example with invented numbers, included to show how the bands read in practice. It is not client data.
| Prompt | AI Overviews cited | AI Mode cited | AIO band | AI Mode band | How to read it |
|---|---|---|---|---|---|
| Best tools for category X | 3 of 3 | 0 of 3 | Stable | Absent | Dual-surface gap. The page works on one system and is invisible on the other |
| How does product Y price | 1 of 3 | 2 of 3 | Intermittent | Intermittent | Present on both, stable on neither. Watch two more weeks before acting |
| Category X vs category Z | 0 of 3 | 0 of 3 | Absent | Absent | Neither surface credits you. This is a content gap, not a tracking gap |
Row one is the finding that only dual-surface sampling produces. A team tracking AI Overviews alone would record a clean three of three and conclude everything is fine.
What this system cannot tell you
Be honest about the ceiling, because it affects which decisions the data can support.
It cannot see query fan-out. Google's systems generate their own sub-queries while assembling a response, so your logged prompt and the actual retrieval behind the answer are not the same thing, even when the visible answer looks identical.
It cannot cover every phrasing. A defined prompt set samples a fixed list of wordings while real buyers ask the same question many ways. The gap between your set and reality widens as a topic broadens.
It cannot establish causation. If you change a page and a citation appears three weeks later, you have a correlation inside a system with documented generation-to-generation movement. Repeated runs narrow the uncertainty. They do not remove it.
None of this argues for trying to force citations through manipulation. Keyword-stuffed question blocks, markup that misrepresents a page, and content shaped to game one prompt run afoul of platform spam policies. They also tend to produce citations that do not survive the next generation. The durable lever is content that answers the underlying question well, which is also what the 86% semantic agreement finding implies: both surfaces are converging on substance.
Where TeamAI fits, and where it does not
Here is the product boundary stated directly, because the alternative is a team buying a workspace for a job it does not do.
TeamAI is where the definition layer lives and where the surrounding analysis work happens. The shared prompt library holds the governed question set with folders, search, role-based permissions, and variable syntax, alongside roughly 200 built-in prompts. You can target a specific model per prompt rather than routing everything through one provider. Workflows accept a table of inputs and run them in a single pass, which is how batch work usually gets done. Connections are available through an API, an MCP server, Zapier, Google Workspace, and form collection, so an external process can hand results back into the workspace where your team already works.
What TeamAI does not currently offer is a verified native integration that queries Google AI Overviews or AI Mode on a schedule and stores their citations for you. Do not buy the workspace expecting that. Research Mode returns answers with source URLs, which is useful for research and is not a measurement of what Google's surfaces cite. The surface queries and the log remain yours to run, by hand or through a process you connect.
That boundary is worth stating plainly because the honest version of this workflow is a governed prompt set in a shared workspace plus a logging process your team owns. If someone tells you a workspace subscription alone solves dual-surface citation monitoring, ask them which field in their log stores the cited URL.
Start a TeamAI WorkspaceWhen the diagnosis should not be yours to run
Sometimes the constraint is not method. A team can understand all of the above, agree with the three-run protocol, and still not have anyone who will run 150 observations a week and read them.
That is a staffing question, and it belongs to a different product. The live Growth Loops explainer on AI visibility and GEO audits covers the managed-diagnosis path and who actually needs it.
Two different doors, and they should stay that way. TeamAI is the workspace answer when your team will run the process and needs the question set governed. The linked Growth Loops article carries the separate managed option without turning this platform guide into a second product pitch.
Start with the narrow version
If you take one thing from the two Ahrefs numbers, make it this: the surfaces agree on answers and disagree on sources, so source tracking has to be done per surface or not at all.
Pick 25 prompts that match real buyer questions. Write them once, store them in a governed library so the wording stops drifting. Run them three times a week against both surfaces. Log the cited URL, not just whether your brand appeared. Wait two weeks before you react to anything.
That is a small enough program to sustain and specific enough to act on, which is more than most dashboards deliver.
Frequently Asked Questions
Do AI Overviews and AI Mode cite the same sources?
Rarely. Ahrefs found the two surfaces cited the identical URL only about 13.7% of the time across 540,000 query pairs (September 2025 US data), rising to 16.3% when narrowed to each surface's top three citations. Independent studies from SE Ranking (about 10.7%) and Victorious point the same direction.
What does the 86% semantic similarity number actually mean?
It measures how similar the two surfaces' conclusions are in meaning, not how often they cite the same sources. Ahrefs found about 86% average semantic similarity across 730,000 response pairs, even though word-level overlap was only about 16% and the two surfaces opened with an identical first sentence just 2.51% of the time.
Can a prompt library track AI Overviews and AI Mode citations automatically?
No. A prompt library standardizes and version-controls the questions you ask, but it does not query AI Overviews or AI Mode on a schedule or record which URLs they cite. That requires a separate execution and logging process built alongside the library.
How many times should I run a prompt to trust a citation change?
At least three runs per surface per week, on different days, since Ahrefs found that 45.5% of AI Overview citations change between generations. Act on a band change only after it holds across two consecutive weeks, not after a single run.
Should I monitor more prompts or run each prompt more times?
With a fixed weekly budget, favor more runs per prompt over more prompts. A narrower set tracked three times per surface per week produces a trustworthy stability band; a wider set checked only once per prompt cannot tell a real change from routine variance.