AI Guide

agents
Marketing Agent Workflows on MCP: The Governance Layer Your Approval Log Is Missing AI Agent Development Services Powered by TeamAI How to Build an AI Agent Library: A Powerful Google Agentspace Alternative
AI Automation
Claude vs ChatGPT vs Gemini: 2026 Head-to-Head Comparison Understanding Gemini Models: A Plain-English Guide to Google's AI Family (2026) How to Automate Your Team's Workflows with AI: A Step-by-Step Guide AI Workspace for Teams: Benefits, Features & How to Choose One (2026) Best AI Models for Coding and Agentic Workflows in 2026 Best AI Models for Writing, Business Tasks and General Intelligence (2026) Who's Winning the AI Race in 2026? Claude vs ChatGPT vs Gemini in 2026: Giants, Challengers, and the AI model Showdown AI Model Benchmarks and Provider Comparison for 2026 22 AI Frontier Models Compared for 2026 How to Set Up AI Automated Workflows
AI Collaboration
The AI-Ready Team: How to Drive Adoption Without the Resistance How to Measure the ROI of AI Across Your Team AI Workspace for Teams: Benefits, Features & How to Choose One (2026) Best AI Models for Writing, Business Tasks and General Intelligence (2026) Who's Winning the AI Race in 2026? Claude vs ChatGPT vs Gemini in 2026: Giants, Challengers, and the AI model Showdown AI Model Benchmarks and Provider Comparison for 2026 22 AI Frontier Models Compared for 2026 How to Get My Team to Collaborate with ChatGPT
AI for Sales
Generating Sales Role-Play Scenarios with ChatGPT
AI Integration
Who's Winning the AI Race in 2026? Claude vs ChatGPT vs Gemini in 2026: Giants, Challengers, and the AI model Showdown AI Model Benchmarks and Provider Comparison for 2026 22 AI Frontier Models Compared for 2026 Integrating Generative AI Tools, like ChatGPT, into Your Team's Operations
AI Processes and Strategy
AI Overviews and AI Mode Cite Different Sources Most of the Time. Can a Prompt Library Track Both? AI Brief for AI Max: what each guideline type steers, and who should approve it How to Automate Your Team's Workflows with AI: A Step-by-Step Guide AI Workspace for Teams: Benefits, Features & How to Choose One (2026) Best AI Models for Writing, Business Tasks and General Intelligence (2026) How to Safeguard My Business Against Bad AI Use by Employees Providing Quality Assurance and Oversight of AI Like ChatGPT How to Choose the Right LLM for Your Business in 2026 How to Use ChatGPT & Generative AI to Scale a Team's Impact
Build an AI Agent
Creating a Custom AI Agent for Businesses Creating a Custom AI Marketing Agent Create an AI Agent for Sales Teams
Generative AI and Business
What Is the Cost of GEO in 2026? The 10 Top GEO Agencies for AI Visibility in 2026 Best AI Models for Writing, Business Tasks and General Intelligence (2026) The Benefits of AI for Small Businesses: Leveling the Playing Field Building a Data-Driven Culture With AI: A Practical Guide for Teams AI Terms Everyone Should Know (2026 Edition) Top 13 Alternatives to ChatGPT Teams Top 7 LLMs for Business in 2026: Ranked and Compared Will ChatGPT and LLMs Take My Job? Understanding the Value of ChatGPT and LLMs for Teams and Businesses Why Use ChatGPT & Generative AI for My Business
Large Language Models (LLMs)
Claude vs ChatGPT vs Gemini: 2026 Head-to-Head Comparison Understanding Gemini Models: A Plain-English Guide to Google's AI Family (2026) How to Automate Your Team's Workflows with AI: A Step-by-Step Guide AI Workspace for Teams: Benefits, Features & How to Choose One (2026) AI Model Economics: Choosing by Budget and Scale (2026) Best AI Models for Complex Reasoning Compared: Opus, GPT-5.4, Gemini, and Which to Use in 2026 Best AI Models for Coding and Agentic Workflows in 2026 Best AI Models for Writing, Business Tasks and General Intelligence (2026) Who's Winning the AI Race in 2026? Claude vs ChatGPT vs Gemini in 2026: Giants, Challengers, and the AI model Showdown AI Model Benchmarks and Provider Comparison for 2026 22 AI Frontier Models Compared for 2026 Every Gemini Model, Compared: Pricing, Context Windows & Which to Use DeepSeek R1, V4 Pro, and V4 Flash Compared: Pricing, Use Cases, and the Full 2026 Model Guide Every Claude Model, Compared: Versions, Pricing & Which to Use Every ChatGPT Model, Compared: Versions, Pricing & Which to Use Meet the Riskiest AI Models Ranked by Researchers Why You Should Use Multiple Large Language Models Overview of Large Language Models (LLMs)
LLM Pricing
How to Measure the ROI of AI Across Your Team AI Model Economics: Choosing by Budget and Scale (2026)
Prompt Libraries
How to Measure the ROI of AI Across Your Team How to Automate Your Team's Workflows with AI: A Step-by-Step Guide AI Prompt Templates for HR and Recruiting AI Prompt Templates for Marketers 8-Step Guide to Creating a Prompt for AI  What businesses need to know about prompt engineering How to Build and Refine a Prompt Library

AI Overviews and AI Mode Cite Different Sources Most of the Time. Can a Prompt Library Track Both?

Ahrefs published a finding in December 2025 that has since become one of the most repeated statistics in AI search discussions. Analyzing September 2025 United States data from its Brand Radar tool, Ahrefs reported that AI Overviews and AI Mode cited the same URL only 13.7% of the time across 540,000 query pairs.

The number usually arrives as shorthand for two separate search engines wearing the same logo. That reading is directionally right and it gets flattened almost immediately, because a second number from the same study travels alongside it: 86% semantic similarity. Put together carelessly, the two produce a sentence like “the surfaces only overlap 86% of the time,” which misstates both figures.

The two numbers measure different things. AI Overviews and AI Mode tend to agree on what to say. They rarely agree on which pages deserve credit for saying it. Once you accept that, the useful question stops being “how different are these surfaces” and becomes operational: can your team sample both surfaces reliably enough to know when something actually changed, and what part of that job does a prompt library really do?

Quick answer

AI Overviews and AI Mode cite the same URL only about 13.7% of the time, but reach similarly worded conclusions about 86% of the time. A prompt library standardizes and version-controls the questions you ask both surfaces — it does not query them on a schedule or log citations for you. Tracking dual-surface citations requires a governed prompt set plus a separate execution and logging process, run at least three times per prompt per surface before you trust a change.

Two numbers, two different measurements

Keep these separate in every internal conversation, because they support different decisions.

13.7% is about sources. Ahrefs measured how often the identical URL appeared in both an AI Overview response and an AI Mode response for the same query, across 540,000 query pairs from September 2025 United States data. Narrowing to the top three citations in each response raised the overlap only to 16.3%. Read plainly: most of the time, the two surfaces pulled from different pages to answer the same question.

86% is about conclusions. Using cosine similarity across a separate sample of 730,000 response pairs, Ahrefs reported that roughly 89.7% of pairs scored above 0.8 on a scale where 1.0 means identical meaning, averaging out to 86% semantic similarity. Word-level overlap, a blunter measure of shared phrasing, sat at about 16%, and the two surfaces opened with the same first sentence only 2.51% of the time.

So the surfaces reach similar conclusions using different phrasing and different sources. If your team tracks only AI Overviews and treats that as AI search visibility, the citation half of your picture is missing for the other surface, and the conclusion half tells you very little about which of your pages is doing the work.

Same study, two different questions
Ahrefs, September 2025 US data. Click a tab to switch what’s measured.

The caveat that should change how you sample

Ahrefs is explicit that the comparison covers single generations, meaning one AI Overview and one AI Mode answer captured at one point in time per query. That matters more than it first appears, because separate Ahrefs research found that 45.5% of AI Overview citations change between generations. The same query can return a different citation set on a later run with no change to any page involved.

Two consequences follow, and they point in opposite directions:

  • The citation pools behind each surface may overlap more than a single snapshot suggests, since repeated runs can surface pages that one generation happened to skip.
  • Any single run of your own monitoring is equally unreliable. A citation that disappears between Tuesday and Thursday may be variance rather than a loss.

Independent analyses point the same way on overlap. SE Ranking measured roughly 10.7% URL overlap on its own large-scale crawl. Victorious found that about 77% of unique cited domains across a 1,540-query sample appeared on only one surface, with no query in that sample returning identical citations on both. Three methodologies landing in the same range is a stronger signal than any one of them alone.

Treat all of these as dated vendor studies with stated samples rather than standing facts, and confirm current figures before you quote them in a board deck.

Two studies, one URL-overlap metric
Both measure the share of queries where AI Overviews and AI Mode cited the identical URL

Victorious measured a related but different metric — the share of unique cited domains appearing on only one surface (about 77%) — so it isn’t plotted on this URL-overlap axis. It points the same direction: most citations do not repeat across surfaces.

What a prompt library actually does, and what it does not

This is where most advice on the topic gets loose, so here is the boundary in plain terms.

A prompt library stores and standardizes prompts. It gives a defined question set one home, keeps wording consistent between people, versions that wording when it changes, and controls who may edit it. In TeamAI, that means folders, search, role-based edit permissions, a chosen model per prompt, and variables written with double curly brackets so one prompt template covers many brands, competitors, or categories. The live walkthrough for that build sits in how to build and refine a prompt library.

A prompt library does not query Google AI Overviews on a schedule. It does not query AI Mode. It does not capture which URLs those surfaces cited, and it cannot correct for generation variance, because storing a question has nothing to do with running it repeatedly and recording answers.

That distinction is not pedantic. Teams buy a library, organize 40 excellent questions into tidy folders, and then discover six weeks later that nobody has actually run them against either surface and nothing has been recorded. The library was never the missing piece. The execution and logging process was.

The two-part system that does work

Split the job explicitly and staff both halves.

Definition layer
Owns: The prompt set, exact wording, version history, who may change it, which model a prompt targets, which variables it accepts.
Lives in: A governed prompt library, shared so every run uses identical wording.
Execution and logging layer
Owns: Running each prompt against each surface on a cadence, capturing the cited URLs, writing one row per run, keeping the history.
Lives in: A process you build: manual runs into a sheet, a scripted run, or a tool you connect.

The definition layer is what makes the second half trustworthy. If two people phrase the same question slightly differently across weeks, you cannot tell whether a citation change came from the surface or from the prompt. Governance on wording is the control variable in your own experiment.

The execution layer is where the actual measurement happens, and it is the half teams underestimate.

Log at URL level, not domain level

Most monitoring spreadsheets record whether the brand appeared. That is too coarse to act on, because it cannot tell you which page earned the citation or whether the surface is pointing at something outdated.

Use one row per prompt, per surface, per run, with these fields. Click a field to see why it matters.

run_id
Lets you group a week’s runs and rerun the comparison later.
run_date
Variance analysis needs spacing between runs.
prompt_id
Ties the observation to governed wording.
prompt_version
Prevents attributing a prompt edit to a surface change.
surface
The entire point of dual-surface tracking.
brand_mentioned
Mention without citation is a real and separate state.
our_url_cited
The citation question, kept apart from mention.
our_cited_url
Tells you which page is working, and which stale page is being surfaced.
citation_position
Position movement often precedes disappearance.
all_cited_urls
Competitive set and directory dependence become visible.
answer_conclusion
Lets you separate conclusion agreement from source agreement.
conclusion_matches_other_surface
Your own local version of the 86% finding.

Two fields carry most of the weight. our_cited_url turns “we are visible” into “this specific page is doing the work.” all_cited_urls turns a brand check into a competitive record you can query later.

The repeated-run stability protocol

Given that citations shift between generations, a single run per prompt is a presence check, not a measurement. Use repetition as the unit of confidence.

  1. Run every prompt three times per surface per week, on three different days rather than three times in one sitting.
  2. Compute a stability rate for each prompt and surface: runs where your URL was cited, divided by total runs.
  3. Classify into bands. Stable is three of three. Intermittent is one or two of three. Absent is zero of three.
  4. Act on band changes across two consecutive weeks, not on single-run flips. One missing citation in one run is inside the expected noise the 45.5% figure describes.
  5. Record the band alongside the raw rows so the history stays auditable when someone challenges a decision months later.

The decision rule in step four is the part that saves time. Without it, teams chase noise weekly and treat normal variance as a crisis or a win.

Try it: check your stability band

Ran your three weekly checks? Enter how many of the three runs cited your URL, for each surface.

Out of 3 runs, how many cited your URL?

Cut prompt count before you cut run count

Monitoring budgets are finite, and the tradeoff is usually framed wrongly. Two plans with identical effort answer different questions:

  • 25 prompts, 2 surfaces, 3 runs each: 150 observations per week. This answers whether your citations are stable.
  • 75 prompts, 2 surfaces, 1 run each: 150 observations per week. This answers whether you appeared once, on one day, for a wider list.

Take the first. With a fixed budget, cut the prompt count before you cut the run count, because a single generation cannot distinguish a lost citation from routine variance, and the variance on these surfaces is common rather than exceptional. A narrower set you can actually trust beats a broad set that produces confident, unreliable readings.

The exception is a genuine discovery pass. If you have never sampled at all, one wide single-run sweep is a reasonable way to find which 25 prompts deserve the three-run treatment. Do that once, then narrow.

Weekly monitoring budget planner

Enter your own weekly observation budget (prompts × surfaces × runs) and see how many stability-grade prompts it actually buys.

Your weekly observation budget
Stability-grade plan (recommended)
25 prompts
At 2 surfaces × 3 runs each, this is how many prompts you can track with a real stability band, not a single-run guess.
Wide single-run sweep
75 prompts
Same budget, spread across 2 surfaces × 1 run each. Useful once, to discover which prompts matter — not for tracking change over time.

A worked example

The table below is an illustrative example with invented numbers, included to show how the bands read in practice. It is not client data.

PromptAI Overviews citedAI Mode citedAIO bandAI Mode bandHow to read it
Best tools for category X3 of 30 of 3StableAbsentDual-surface gap. The page works on one system and is invisible on the other
How does product Y price1 of 32 of 3IntermittentIntermittentPresent on both, stable on neither. Watch two more weeks before acting
Category X vs category Z0 of 30 of 3AbsentAbsentNeither surface credits you. This is a content gap, not a tracking gap

Row one is the finding that only dual-surface sampling produces. A team tracking AI Overviews alone would record a clean three of three and conclude everything is fine.

What this system cannot tell you

Be honest about the ceiling, because it affects which decisions the data can support.

It cannot see query fan-out. Google's systems generate their own sub-queries while assembling a response, so your logged prompt and the actual retrieval behind the answer are not the same thing, even when the visible answer looks identical.

It cannot cover every phrasing. A defined prompt set samples a fixed list of wordings while real buyers ask the same question many ways. The gap between your set and reality widens as a topic broadens.

It cannot establish causation. If you change a page and a citation appears three weeks later, you have a correlation inside a system with documented generation-to-generation movement. Repeated runs narrow the uncertainty. They do not remove it.

None of this argues for trying to force citations through manipulation. Keyword-stuffed question blocks, markup that misrepresents a page, and content shaped to game one prompt run afoul of platform spam policies. They also tend to produce citations that do not survive the next generation. The durable lever is content that answers the underlying question well, which is also what the 86% semantic agreement finding implies: both surfaces are converging on substance.

Where TeamAI fits, and where it does not

Here is the product boundary stated directly, because the alternative is a team buying a workspace for a job it does not do.

TeamAI is where the definition layer lives and where the surrounding analysis work happens. The shared prompt library holds the governed question set with folders, search, role-based permissions, and variable syntax, alongside roughly 200 built-in prompts. You can target a specific model per prompt rather than routing everything through one provider. Workflows accept a table of inputs and run them in a single pass, which is how batch work usually gets done. Connections are available through an API, an MCP server, Zapier, Google Workspace, and form collection, so an external process can hand results back into the workspace where your team already works.

What TeamAI does not currently offer is a verified native integration that queries Google AI Overviews or AI Mode on a schedule and stores their citations for you. Do not buy the workspace expecting that. Research Mode returns answers with source URLs, which is useful for research and is not a measurement of what Google's surfaces cite. The surface queries and the log remain yours to run, by hand or through a process you connect.

That boundary is worth stating plainly because the honest version of this workflow is a governed prompt set in a shared workspace plus a logging process your team owns. If someone tells you a workspace subscription alone solves dual-surface citation monitoring, ask them which field in their log stores the cited URL.

Start a TeamAI Workspace

When the diagnosis should not be yours to run

Sometimes the constraint is not method. A team can understand all of the above, agree with the three-run protocol, and still not have anyone who will run 150 observations a week and read them.

That is a staffing question, and it belongs to a different product. The live Growth Loops explainer on AI visibility and GEO audits covers the managed-diagnosis path and who actually needs it.

Two different doors, and they should stay that way. TeamAI is the workspace answer when your team will run the process and needs the question set governed. The linked Growth Loops article carries the separate managed option without turning this platform guide into a second product pitch.

Start with the narrow version

If you take one thing from the two Ahrefs numbers, make it this: the surfaces agree on answers and disagree on sources, so source tracking has to be done per surface or not at all.

Pick 25 prompts that match real buyer questions. Write them once, store them in a governed library so the wording stops drifting. Run them three times a week against both surfaces. Log the cited URL, not just whether your brand appeared. Wait two weeks before you react to anything.

That is a small enough program to sustain and specific enough to act on, which is more than most dashboards deliver.

Frequently Asked Questions

Do AI Overviews and AI Mode cite the same sources?

Rarely. Ahrefs found the two surfaces cited the identical URL only about 13.7% of the time across 540,000 query pairs (September 2025 US data), rising to 16.3% when narrowed to each surface's top three citations. Independent studies from SE Ranking (about 10.7%) and Victorious point the same direction.

What does the 86% semantic similarity number actually mean?

It measures how similar the two surfaces' conclusions are in meaning, not how often they cite the same sources. Ahrefs found about 86% average semantic similarity across 730,000 response pairs, even though word-level overlap was only about 16% and the two surfaces opened with an identical first sentence just 2.51% of the time.

Can a prompt library track AI Overviews and AI Mode citations automatically?

No. A prompt library standardizes and version-controls the questions you ask, but it does not query AI Overviews or AI Mode on a schedule or record which URLs they cite. That requires a separate execution and logging process built alongside the library.

How many times should I run a prompt to trust a citation change?

At least three runs per surface per week, on different days, since Ahrefs found that 45.5% of AI Overview citations change between generations. Act on a band change only after it holds across two consecutive weeks, not after a single run.

Should I monitor more prompts or run each prompt more times?

With a fixed weekly budget, favor more runs per prompt over more prompts. A narrower set tracked three times per surface per week produces a trustworthy stability band; a wider set checked only once per prompt cannot tell a real change from routine variance.