SearchEditorial contact
Data Report

What Makes a Website More Likely to Be Cited by Google AI Mode? What Millions of Citations Can and Can’t Tell Us

Large citation studies reveal recurring source, domain, and technical patterns. They do not reveal a universal ranking formula, and the strongest schema intervention found no reliable AI Mode lift.

By Andrew Ansley · Published Aug 26, 2026 · 18 min read

Halftone editorial illustration of a Google AI Mode answer connected to one cited source while many other source cards recede behind it.

Citation studies show which sources recur, but the visible link is the outcome of a query-specific retrieval process, not a universal checklist.

The short answer

Key takeaways

  • Google says supporting pages must be indexed and snippet-eligible, with no additional technical requirement for AI Mode or AI Overviews.
  • Semrush found about 35% URL overlap between AI Mode citations and Google's top ten for the same original queries, compared with 67% for AI Overviews and 82% for Perplexity.
  • Large observational studies associate citations with established domains, technical clarity, and source types that fit the query, but those patterns do not identify causal ranking weights.
  • In Ahrefs' matched study, adding JSON-LD produced a 2.4% AI Mode citation change that was statistically indistinguishable from zero.
  • Industry, model, country, language, and collection window materially change citation mix, so a global top-domain list cannot substitute for prompt-level measurement.

Millions of citations can make recurring patterns look like a ranking formula. They are not the same thing. A dataset can show which domains, page traits, and source categories appeared together with citations, but it cannot automatically show which trait caused Google AI Mode to select a page.

The useful answer comes from separating three evidence layers: Google's public eligibility guidance, observational citation patterns, and intervention studies that test whether changing one page feature alters citation frequency.

The Short Answer

A website becomes eligible for Google AI Mode citations through ordinary Search fundamentals: Google must be able to crawl and index the page, and the page must be eligible to appear with a snippet. Beyond that threshold, large observational studies repeatedly find citations concentrated among established domains, pages with strong technical foundations, and source types that fit the query. They also show that AI Mode often selects a different URL from the one ranking in Google's top ten.

That does not produce a checklist with known causal weights. Google says there is no special AI Mode markup requirement, and the strongest intervention study in the reviewed evidence found no statistically reliable citation gain after pages added JSON-LD schema. The practical target is therefore not an AI citation formula. It is a discoverable, useful, evidence-rich page inside a credible site, measured against the prompts and source patterns that matter in its own category.

What Google Says Is Required

Google's public guidance establishes an eligibility floor, not a published ranking formula. A supporting page must be indexed and eligible to appear in Search with a snippet. Google says there are no additional technical requirements for AI Overviews or AI Mode.

The retrieval process is broader than one ordinary results page. Google describes retrieval-augmented generation that uses core Search ranking systems and a query fan-out process that issues several related searches across subtopics and data sources. That matters because a page can be relevant to one hidden subquery even when it is not the obvious result for the user's original wording.

Google's July 2026 optimization guide puts the editorial emphasis on non-commodity content: original experience, a useful point of view, clear organization, reliable information, and images or video when those formats help the reader. It explicitly says publishers do not need special AI files, artificial content chunking, AI-specific rewriting, inauthentic mentions, or special schema markup.

The distinction is important. Crawlability, indexation, and snippet eligibility are prerequisites. Helpful content and technical clarity are Google-recommended practices. Neither category discloses the relative weight of a citation-selection system.

AI Mode Does Not Simply Copy the Top Ten

Semrush tested 5,000 randomly selected keywords and collected more than 150,000 citations across Google Search, AI Overviews, AI Mode, ChatGPT, and Perplexity. For each query, it compared cited URLs with Google's top ten organic results.

Google AI Mode had about 35% URL overlap with the top ten. AI Overviews were higher at 67%, and Perplexity reached 82%. At the domain level, AI Mode overlap was higher, about 54%, which suggests that the system often selected a different page from a domain already visible in ordinary Search.

AI Mode citations overlapped less with Google's top ten

URL overlap with Google's top ten organic results for the same query, Semrush sample of 5,000 keywords, 2025

Perplexity had 82% URL overlap, Google AI Overviews 67%, and Google AI Mode about 35%.

Source: Semrush, AI Mode comparison study, July 2025. More than 150,000 citations across platforms. Overlap is sample-specific and does not show that non-overlapping URLs were absent from Google's index or related fan-out results.

The result does not mean that 65% of AI Mode sources were invisible to Google or retrieved independently of Search. It means they were not the same URLs as the ten organic results recorded for the original query in this sample. Query fan-out can retrieve pages for related subqueries, and personalization, location, interface changes, and collection timing can alter the source set.

The more defensible conclusion is narrower: ranking one page for the original query is not sufficient to predict which page AI Mode will cite. Domain-level search visibility still correlated strongly with citations in Semrush's data, but exact URL selection was much less aligned.

Citation Concentration Reveals Source Supply, Not a Universal Preference

Ahrefs' July 2026 United States snapshot shows how concentrated the visible source mix can become. Among the leading domains in its broad AI Mode query set, Reddit accounted for 19.9% of the summed citations to top sources, YouTube 17.0%, Google 11.5%, Facebook 9.3%, Instagram 5.3%, and English Wikipedia 4.9%.

A few social and platform domains dominated the leading sources

Mention share among top Google AI Mode sources, broad United States queries, Ahrefs Brand Radar, July 2026

Reddit led at 19.9%, followed by YouTube at 17.0% and Google at 11.5%.

Source: Ahrefs, 50 Most-Cited Websites in Google AI Mode, July 2026. Mention share uses the summed citations of leading sources, not all web citations, and should not be treated as an industry-specific target.

Those percentages are mention share among the leading sources, not each domain's share of every citation on the web. They also mix very different source functions. A YouTube tutorial, a Reddit experience report, a Google property, a product page, and an encyclopedia entry can answer different subquestions.

Profound's much larger cross-platform corpus reinforces that context dependence. Its 11.84 billion citations covered eight models and 8,061 active categories between April and July 2026. Across all models, company-operated websites supplied roughly 57% of citations, but industry medians varied sharply. Earned media represented a median 59% of citations for pharma and biotech companies and only 11.4% for software-as-a-service companies.

The practical inference is that a publisher should compare itself with the source supply in its category rather than copy the global top-domain list. A medical claim, local recommendation, software configuration question, and product comparison invite different evidence and different source types.

Technical Patterns Are Associations Until an Intervention Tests Them

Semrush's technical study examined five million cited URLs across Google AI Mode and ChatGPT Search. It found that structured data was more common among AI Mode citations than ChatGPT citations. Organization markup appeared on 34% of AI Mode-cited pages, Article on 26%, and Breadcrumb markup on 20%. Other measured types were less common.

Common schema types appeared on a minority of AI Mode-cited pages

Share of cited pages carrying each schema type, Semrush cross-sectional analysis of Google AI Mode citations

Organization schema appeared on 34% of cited pages, Article on 26%, and Breadcrumb on 20%.

Source: Semrush technical SEO study, 2026. Categories overlap. Prevalence among cited pages is an association and does not establish that schema caused selection.

That pattern is useful for describing the cited population. It does not show that adding schema causes a page to be cited. Better maintained sites may be more likely to implement schema, earn links, publish original material, rank across more queries, and satisfy users. Any of those correlated characteristics could explain the difference.

Ahrefs tested that exact causal temptation. It first observed that cited pages were almost three times more likely to contain JSON-LD than non-cited pages. It then tracked 1,885 pages that added JSON-LD between August 2025 and March 2026, matched them against 4,000 control pages, and measured the 30 days before and after the change.

The matched difference-in-differences estimate for Google AI Mode was a 2.4% increase, statistically indistinguishable from zero. ChatGPT's 2.2% estimate was also indistinguishable from zero. AI Overviews showed a 4.6% relative decline, but both treated and control groups had already been falling, and the authors warned against interpreting the result as proof that schema harms citations.

Adding JSON-LD produced no reliable AI Mode citation lift

Matched difference-in-differences estimate after schema adoption, 1,885 treated pages and 4,000 controls, 2025 to 2026

AI Mode increased 2.4% and ChatGPT 2.2%, both statistically indistinguishable from zero; AI Overviews declined 4.6% relative to controls.

Source: Ahrefs schema intervention study, May 2026. The authors caution that the AIO decline occurred amid a broader pre-existing contraction and does not prove schema is harmful.

This is the article's most useful methodological contrast. Cross-sectional data can identify a characteristic that is common among cited pages. A matched intervention can test whether changing that characteristic changes citation frequency. Here, the association survived as a description, but the intervention did not establish schema as an AI Mode citation lever.

The Query Changes the Source Set

Profound analyzed 3.25 billion citations across seven models and fourteen countries in March 2026. Google AI Mode's aggregate social-source rate was 14.5%, close to AI Overviews at 15.3%, but different from ChatGPT at 9.1%, Copilot at 4.3%, Claude at 3.99%, and Gemini at 3.6%.

Social-source reliance varied by model

Aggregate social citation rate, Profound corpus of 3.25 billion citations across 14 countries, March 2026

Google AI Overviews and AI Mode had the highest aggregate social-source rates in this cross-platform sample.

Source: Profound, How query language reshapes AI citations, April 2026. Rates pool countries and categories; the study also found large language-specific shifts, so these are not universal platform constants.

The study's main point was not a global hierarchy of platforms. Query language changed the social-source mix, sometimes sharply. That makes a single English-language citation study a poor basis for universal advice across markets.

Source selection can also change over time. Interfaces, underlying models, index coverage, user preferences, and query fan-out behavior are not fixed. A large historical dataset reduces sampling noise inside its collection window. It does not freeze the product or guarantee that a domain-level pattern applies to one future prompt.

What a Publisher Can Act On

The evidence supports a layered operating model.

  1. Protect eligibility. Keep important pages crawlable, indexable, canonical, and snippet-eligible. Make the main content available without fragile rendering dependencies.
  2. Publish something worth retrieving. Original data, direct experience, primary documentation, transparent methods, and precise answers create information that a synthesis system cannot obtain from a commodity summary alone.
  3. Build coherent site-level coverage. AI Mode's lower URL overlap and higher domain overlap suggest that Google may retrieve a different page from a domain it already understands across the topic. Maintain clear internal relationships among the pages instead of forcing every query variation into one document.
  4. Match the evidence form to the question. Use text, images, video, product feeds, business information, or firsthand discussion only where that form genuinely supplies better evidence.
  5. Treat structured data as clarity infrastructure. Keep accurate markup aligned with visible content and Search eligibility, but do not promise that adding schema will increase citations.
  6. Measure the right prompt set. Track the actual questions, countries, languages, devices, and source types relevant to the organization. Separate mentions, citations, source position, clicks, and business outcomes.

This is compatible with the boundary between GEO and SEO. SEO supplies the discovery foundation. Generative visibility adds a new output and measurement problem. It also complements a broader search measurement stack, where rank, citation, referral, and outcome remain different observations.

Methodology and Limitations

This Data Report synthesizes ten materially useful full-text sources available through August 26, 2026: three official Google Search documents, four disclosed vendor studies, two large Profound observational analyses, and one scholarly preprint about Google AI Overviews.

The studies do not share one population. Semrush's 2025 comparison used 5,000 keywords and more than 150,000 citations. Its technical analysis pooled five million URLs cited by Google AI Mode and ChatGPT Search. Ahrefs' top-domain report is a July 2026 United States snapshot among leading sources. Its schema intervention followed 1,885 treated pages and 4,000 controls. Profound's studies pool billions of citations across several models, categories, countries, and languages. The scholarly study concerns AI Overviews rather than AI Mode.

No figures were pooled into a meta-analysis. Percentages retain each publisher's denominator and measurement label. Vendor studies are treated as disclosed observational or quasi-experimental evidence, not independent audits of Google's systems. Google documentation establishes public eligibility and recommended practice but does not disclose ranking weights. The report does not measure Search Institute's own citation rate or claim that any listed characteristic guarantees inclusion.

All ten sources were acquired through direct MCP Scraper extraction. No search-result snippet or generated AI answer supports a material claim.

Conclusion

Websites cited by Google AI Mode tend to be technically accessible, part of domains with broader search visibility, and represented by pages or source types that fit the question. Large citation datasets make those recurring patterns visible. They also show that AI Mode often cites a different URL than Google's top ten for the original query.

What the datasets cannot provide is a universal recipe. They cannot separate every correlated site characteristic from the underlying quality, authority, content supply, or query mix that produced it. They cannot turn schema prevalence into a causal schema boost, and they cannot guarantee that a global source pattern applies to one industry, language, or future prompt.

The durable strategy is less exotic than the market around it: protect Search eligibility, publish original and useful evidence, organize the site coherently, use accurate technical signals, and measure the prompts that matter. Treat every claimed citation factor according to the evidence design that produced it.

Common questions

Frequently asked questions

Sources

AA

Published by

Andrew Ansley

Andrew Ansley writes about search, information retrieval, AI recommendation systems, and the evidence systems use to form answers.