The short answer
Key takeaways
- Google says supporting pages must be indexed and snippet-eligible, with no additional technical requirement for AI Mode or AI Overviews.
- Semrush found about 35% URL overlap between AI Mode citations and Google's top ten for the same original queries, compared with 67% for AI Overviews and 82% for Perplexity.
- Large observational studies associate citations with established domains, technical clarity, and source types that fit the query, but those patterns do not identify causal ranking weights.
- In Ahrefs' matched study, adding JSON-LD produced a 2.4% AI Mode citation change that was statistically indistinguishable from zero.
- Industry, model, country, language, and collection window materially change citation mix, so a global top-domain list cannot substitute for prompt-level measurement.
Millions of citations can make recurring patterns look like a ranking formula. They are not the same thing. A dataset can show which domains, page traits, and source categories appeared together with citations, but it cannot automatically show which trait caused Google AI Mode to select a page.
The useful answer comes from separating three evidence layers: Google's public eligibility guidance, observational citation patterns, and intervention studies that test whether changing one page feature alters citation frequency.
The Short Answer
A website becomes eligible for Google AI Mode citations through ordinary Search fundamentals: Google must be able to crawl and index the page, and the page must be eligible to appear with a snippet. Beyond that threshold, large observational studies repeatedly find citations concentrated among established domains, pages with strong technical foundations, and source types that fit the query. They also show that AI Mode often selects a different URL from the one ranking in Google's top ten.
That does not produce a checklist with known causal weights. Google says there is no special AI Mode markup requirement, and the strongest intervention study in the reviewed evidence found no statistically reliable citation gain after pages added JSON-LD schema. The practical target is therefore not an AI citation formula. It is a discoverable, useful, evidence-rich page inside a credible site, measured against the prompts and source patterns that matter in its own category.
What Google Says Is Required
Google's public guidance establishes an eligibility floor, not a published ranking formula. A supporting page must be indexed and eligible to appear in Search with a snippet. Google says there are no additional technical requirements for AI Overviews or AI Mode.
The retrieval process is broader than one ordinary results page. Google describes retrieval-augmented generation that uses core Search ranking systems and a query fan-out process that issues several related searches across subtopics and data sources. That matters because a page can be relevant to one hidden subquery even when it is not the obvious result for the user's original wording.
Google's July 2026 optimization guide puts the editorial emphasis on non-commodity content: original experience, a useful point of view, clear organization, reliable information, and images or video when those formats help the reader. It explicitly says publishers do not need special AI files, artificial content chunking, AI-specific rewriting, inauthentic mentions, or special schema markup.
The distinction is important. Crawlability, indexation, and snippet eligibility are prerequisites. Helpful content and technical clarity are Google-recommended practices. Neither category discloses the relative weight of a citation-selection system.
AI Mode Does Not Simply Copy the Top Ten
Semrush tested 5,000 randomly selected keywords and collected more than 150,000 citations across Google Search, AI Overviews, AI Mode, ChatGPT, and Perplexity. For each query, it compared cited URLs with Google's top ten organic results.
Google AI Mode had about 35% URL overlap with the top ten. AI Overviews were higher at 67%, and Perplexity reached 82%. At the domain level, AI Mode overlap was higher, about 54%, which suggests that the system often selected a different page from a domain already visible in ordinary Search.
AI Mode citations overlapped less with Google's top ten
URL overlap with Google's top ten organic results for the same query, Semrush sample of 5,000 keywords, 2025
Perplexity had 82% URL overlap, Google AI Overviews 67%, and Google AI Mode about 35%.
Source: Semrush, AI Mode comparison study, July 2025. More than 150,000 citations across platforms. Overlap is sample-specific and does not show that non-overlapping URLs were absent from Google's index or related fan-out results.
The result does not mean that 65% of AI Mode sources were invisible to Google or retrieved independently of Search. It means they were not the same URLs as the ten organic results recorded for the original query in this sample. Query fan-out can retrieve pages for related subqueries, and personalization, location, interface changes, and collection timing can alter the source set.
The more defensible conclusion is narrower: ranking one page for the original query is not sufficient to predict which page AI Mode will cite. Domain-level search visibility still correlated strongly with citations in Semrush's data, but exact URL selection was much less aligned.
Citation Concentration Reveals Source Supply, Not a Universal Preference
Ahrefs' July 2026 United States snapshot shows how concentrated the visible source mix can become. Among the leading domains in its broad AI Mode query set, Reddit accounted for 19.9% of the summed citations to top sources, YouTube 17.0%, Google 11.5%, Facebook 9.3%, Instagram 5.3%, and English Wikipedia 4.9%.
A few social and platform domains dominated the leading sources
Mention share among top Google AI Mode sources, broad United States queries, Ahrefs Brand Radar, July 2026
Reddit led at 19.9%, followed by YouTube at 17.0% and Google at 11.5%.
Source: Ahrefs, 50 Most-Cited Websites in Google AI Mode, July 2026. Mention share uses the summed citations of leading sources, not all web citations, and should not be treated as an industry-specific target.
Those percentages are mention share among the leading sources, not each domain's share of every citation on the web. They also mix very different source functions. A YouTube tutorial, a Reddit experience report, a Google property, a product page, and an encyclopedia entry can answer different subquestions.
Profound's much larger cross-platform corpus reinforces that context dependence. Its 11.84 billion citations covered eight models and 8,061 active categories between April and July 2026. Across all models, company-operated websites supplied roughly 57% of citations, but industry medians varied sharply. Earned media represented a median 59% of citations for pharma and biotech companies and only 11.4% for software-as-a-service companies.
The practical inference is that a publisher should compare itself with the source supply in its category rather than copy the global top-domain list. A medical claim, local recommendation, software configuration question, and product comparison invite different evidence and different source types.
Technical Patterns Are Associations Until an Intervention Tests Them
Semrush's technical study examined five million cited URLs across Google AI Mode and ChatGPT Search. It found that structured data was more common among AI Mode citations than ChatGPT citations. Organization markup appeared on 34% of AI Mode-cited pages, Article on 26%, and Breadcrumb markup on 20%. Other measured types were less common.
Common schema types appeared on a minority of AI Mode-cited pages
Share of cited pages carrying each schema type, Semrush cross-sectional analysis of Google AI Mode citations
Organization schema appeared on 34% of cited pages, Article on 26%, and Breadcrumb on 20%.
Source: Semrush technical SEO study, 2026. Categories overlap. Prevalence among cited pages is an association and does not establish that schema caused selection.
That pattern is useful for describing the cited population. It does not show that adding schema causes a page to be cited. Better maintained sites may be more likely to implement schema, earn links, publish original material, rank across more queries, and satisfy users. Any of those correlated characteristics could explain the difference.
Ahrefs tested that exact causal temptation. It first observed that cited pages were almost three times more likely to contain JSON-LD than non-cited pages. It then tracked 1,885 pages that added JSON-LD between August 2025 and March 2026, matched them against 4,000 control pages, and measured the 30 days before and after the change.
The matched difference-in-differences estimate for Google AI Mode was a 2.4% increase, statistically indistinguishable from zero. ChatGPT's 2.2% estimate was also indistinguishable from zero. AI Overviews showed a 4.6% relative decline, but both treated and control groups had already been falling, and the authors warned against interpreting the result as proof that schema harms citations.
Adding JSON-LD produced no reliable AI Mode citation lift
Matched difference-in-differences estimate after schema adoption, 1,885 treated pages and 4,000 controls, 2025 to 2026
AI Mode increased 2.4% and ChatGPT 2.2%, both statistically indistinguishable from zero; AI Overviews declined 4.6% relative to controls.
Source: Ahrefs schema intervention study, May 2026. The authors caution that the AIO decline occurred amid a broader pre-existing contraction and does not prove schema is harmful.
This is the article's most useful methodological contrast. Cross-sectional data can identify a characteristic that is common among cited pages. A matched intervention can test whether changing that characteristic changes citation frequency. Here, the association survived as a description, but the intervention did not establish schema as an AI Mode citation lever.
The Query Changes the Source Set
Profound analyzed 3.25 billion citations across seven models and fourteen countries in March 2026. Google AI Mode's aggregate social-source rate was 14.5%, close to AI Overviews at 15.3%, but different from ChatGPT at 9.1%, Copilot at 4.3%, Claude at 3.99%, and Gemini at 3.6%.
Social-source reliance varied by model
Aggregate social citation rate, Profound corpus of 3.25 billion citations across 14 countries, March 2026
Google AI Overviews and AI Mode had the highest aggregate social-source rates in this cross-platform sample.
Source: Profound, How query language reshapes AI citations, April 2026. Rates pool countries and categories; the study also found large language-specific shifts, so these are not universal platform constants.
The study's main point was not a global hierarchy of platforms. Query language changed the social-source mix, sometimes sharply. That makes a single English-language citation study a poor basis for universal advice across markets.
Source selection can also change over time. Interfaces, underlying models, index coverage, user preferences, and query fan-out behavior are not fixed. A large historical dataset reduces sampling noise inside its collection window. It does not freeze the product or guarantee that a domain-level pattern applies to one future prompt.
What a Publisher Can Act On
The evidence supports a layered operating model.
- Protect eligibility. Keep important pages crawlable, indexable, canonical, and snippet-eligible. Make the main content available without fragile rendering dependencies.
- Publish something worth retrieving. Original data, direct experience, primary documentation, transparent methods, and precise answers create information that a synthesis system cannot obtain from a commodity summary alone.
- Build coherent site-level coverage. AI Mode's lower URL overlap and higher domain overlap suggest that Google may retrieve a different page from a domain it already understands across the topic. Maintain clear internal relationships among the pages instead of forcing every query variation into one document.
- Match the evidence form to the question. Use text, images, video, product feeds, business information, or firsthand discussion only where that form genuinely supplies better evidence.
- Treat structured data as clarity infrastructure. Keep accurate markup aligned with visible content and Search eligibility, but do not promise that adding schema will increase citations.
- Measure the right prompt set. Track the actual questions, countries, languages, devices, and source types relevant to the organization. Separate mentions, citations, source position, clicks, and business outcomes.
This is compatible with the boundary between GEO and SEO. SEO supplies the discovery foundation. Generative visibility adds a new output and measurement problem. It also complements a broader search measurement stack, where rank, citation, referral, and outcome remain different observations.
Methodology and Limitations
This Data Report synthesizes ten materially useful full-text sources available through August 26, 2026: three official Google Search documents, four disclosed vendor studies, two large Profound observational analyses, and one scholarly preprint about Google AI Overviews.
The studies do not share one population. Semrush's 2025 comparison used 5,000 keywords and more than 150,000 citations. Its technical analysis pooled five million URLs cited by Google AI Mode and ChatGPT Search. Ahrefs' top-domain report is a July 2026 United States snapshot among leading sources. Its schema intervention followed 1,885 treated pages and 4,000 controls. Profound's studies pool billions of citations across several models, categories, countries, and languages. The scholarly study concerns AI Overviews rather than AI Mode.
No figures were pooled into a meta-analysis. Percentages retain each publisher's denominator and measurement label. Vendor studies are treated as disclosed observational or quasi-experimental evidence, not independent audits of Google's systems. Google documentation establishes public eligibility and recommended practice but does not disclose ranking weights. The report does not measure Search Institute's own citation rate or claim that any listed characteristic guarantees inclusion.
All ten sources were acquired through direct MCP Scraper extraction. No search-result snippet or generated AI answer supports a material claim.
Conclusion
Websites cited by Google AI Mode tend to be technically accessible, part of domains with broader search visibility, and represented by pages or source types that fit the question. Large citation datasets make those recurring patterns visible. They also show that AI Mode often cites a different URL than Google's top ten for the original query.
What the datasets cannot provide is a universal recipe. They cannot separate every correlated site characteristic from the underlying quality, authority, content supply, or query mix that produced it. They cannot turn schema prevalence into a causal schema boost, and they cannot guarantee that a global source pattern applies to one industry, language, or future prompt.
The durable strategy is less exotic than the market around it: protect Search eligibility, publish original and useful evidence, organize the site coherently, use accurate technical signals, and measure the prompts that matter. Treat every claimed citation factor according to the evidence design that produced it.
Frequently asked questions
No. Semrush found about 35% URL overlap between AI Mode citations and Google's top ten for the same original queries. Google also says query fan-out can retrieve pages for related subqueries. Ranking remains relevant to retrieval, but one top-ten URL is not a complete predictor of citation.
The reviewed evidence does not establish that. Schema was common on cited pages, but Ahrefs' matched study found a 2.4% AI Mode change after pages added JSON-LD, statistically indistinguishable from zero. Use accurate schema for its established Search and clarity roles, not as a guaranteed citation lever.
No. They were prominent in broad 2026 samples, but citation mix varied by industry, model, country, and query language. The relevant comparison is the source supply for the prompts and market being measured.
No. Google's July 2026 guidance says Google Search does not use special AI text files and does not require content to be split into tiny chunks for generative features.
Google recommends valuable, non-commodity content, including original experience and useful points of view. That supports publishing original evidence as a sound editorial strategy. The reviewed studies do not provide a universal causal percentage for how much original research increases AI Mode citations.
Track a stable set of relevant prompts and record mentions, citations, cited URLs, source position, country, language, device, and collection time. Keep those observations separate from Search rankings, clicks, conversions, and revenue.
Use them to describe source patterns within a stated platform, sample, category, language, and time window. Do not convert prevalence into causation or a guarantee for one page.
Sources
- “AI features and your website”. Google Search Central; official product documentation; full text accessed August 26, 2026.
- “Optimizing your website for generative AI features on Google Search”. Google Search Central, updated July 10, 2026; official guidance; full text.
- “Top ways to ensure your content performs well in Google's AI experiences on Search”. Google Search Central, May 21, 2025; official guidance; full text.
- “How Google’s AI Mode Compares to Traditional Search and Other LLMs”. Semrush, July 2025; disclosed vendor study of 5,000 keywords and more than 150,000 citations; full text.
- “How Do Technical SEO Factors Impact AI Search?”. Semrush, January 2026; disclosed vendor analysis of five million cited URLs; full text. Correlational design.
- “The 50 Most-Cited Websites in Google AI Mode”. Si Quan Ong, Ahrefs, updated July 21, 2026; vendor-maintained United States snapshot; full text. Mention share is calculated among leading sources.
- “We Tracked 1,885 Pages Adding Schema. AI Citations Barely Moved.”. Louise Linehan, Ahrefs, May 11, 2026; matched difference-in-differences vendor study; full text.
- “Where do AI citations come from?”. Jasman Singh, Profound, August 2026; disclosed vendor analysis of 11.84 billion citations; full text. Cross-platform observational corpus.
- “How query language reshapes AI citations”. Davis McCain, Profound, April 21, 2026; disclosed vendor analysis of 3.25 billion citations; full text. Cross-platform and multilingual corpus.
- “Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact”. Haofei Xu, Umar Iqbal, and Jacob M. Montgomery, submitted May 13, 2026; scholarly preprint; abstract and paper access page reviewed. Concerns AI Overviews rather than AI Mode.
