SearchEditorial contact
Explainer

Is SEO Still About Rankings? Reading the Metrics Marketers Are Using in AI Search

Rank still diagnoses one part of search visibility. A complete 2026 scorecard also needs generative exposure, mentions, citations, referral behavior and business outcomes—without pretending those metrics share one denominator.

By Andrew Ansley · Published Aug 21, 2026 · 13 min read

The short answer

Key takeaways

  • Google still treats average position as useful, but recommends focusing more on impression and click trends than on position alone.
  • Google’s 2026 generative-AI reporting adds a separate exposure layer: impressions plus page, country, device and time breakdowns for a limited set of sites.
  • Mentions, citations, modeled impressions and AI share of voice measure different things and depend on each tool’s prompt corpus and entity definitions.
  • Adobe’s U.S. holiday-retail data shows why downstream metrics matter: AI-referred visits converted 31% more than other traffic while also showing stronger engagement.
  • The defensible replacement for a ranking-only dashboard is a layered scorecard, not one universal AI-visibility number.

Rankings still answer a useful question: when a page appeared in conventional search, where did its top result sit? They do not answer whether an AI system retrieved the page, mentioned the brand, cited the source, sent a visit or influenced revenue.

The evidence supports a layered answer. Keep rank as a diagnostic for classic visibility, then add generative-feature impressions, repeated mention and citation measurements, referral behavior and business outcomes. The difficult part is preserving each metric’s denominator instead of blending unlike signals into one score.

Rankings Still Measure Something Real

Average position is not an obsolete metric. It answers a specific question: when a URL from your site appeared in Google Search, how high was the property’s topmost result, on average, across those impressions?

That is narrower than the everyday phrase “we rank third.” Google’s own methodology says one search-result element can contain several links that all inherit the same position. An AI Overview occupies one position, and every link inside it is assigned that position. A knowledge panel can occupy a numerically lower-looking position while remaining visually prominent. Location, device, history, query mix and result format all affect the average.

This is why position works best as a diagnostic. It can show that a page or query group is moving up or down, help isolate a sudden visibility change and reveal whether technical or content work coincides with broader exposure. It cannot, by itself, show whether the result was noticed, cited in an answer, clicked, trusted or converted.

Google now makes that boundary explicit in its Search Console guidance. The company still documents how to track position history, but recommends focusing more on trends in impressions and clicks than on position alone. That advice predates a complete answer to AI-search measurement, yet it establishes the central point: even within conventional search reporting, rank is one measure rather than the outcome.

AI Search Adds New Observation Points

Google’s AI features do not make conventional SEO irrelevant. Google says pages must still be indexed and eligible to appear with a snippet, and that the same crawlability, internal-linking, textual-content, page-experience and structured-data fundamentals remain useful. AI Overviews and AI Mode can also issue related searches through query fan-out, which expands the set of retrieval decisions behind one visible answer.

What changed is the number of stages between eligibility and a business result.

In June 2026, Google began rolling out dedicated Search Generative AI performance reports to a subset of Search Console properties. The launch documentation lists generative-feature impressions, the pages shown, countries, devices and time breakdowns. Those fields answer a question that an average rank cannot: did a URL from this site appear inside Google’s generative interfaces at all?

The report is important partly because its launch fields stop at exposure. They do not describe the complete path from exposure to value. Google says the data remains included in the overall performance report and that it is considering additional metrics over time. A site can therefore gain generative impressions without knowing, from that dedicated view alone, how much attention, referral traffic or revenue those appearances produced.

That is not a reason to ignore the new report. It is a reason to label the metric correctly. A generative-feature impression is evidence of visible inclusion on a Google AI surface. It is not equivalent to an organic position, a brand mention across multiple answer engines, a citation, a visit or a sale.

A Citation Is Not a Rank—and a Mention Is Not a Click

The measurement problem becomes clearer when the stages are separated.

A July 2026 critical survey of 45 GEO studies describes generative visibility as a partially observable pipeline: search activation, crawling and indexing, retrieval, reranking and context allocation, citation, prominence, factual use, fidelity and user behavior. The paper’s central warning is methodological. Improving one stage does not prove improvement at the next, and a point estimate from one prompt run is not a stable measure of a personalized, changing system.

Commercial AI-visibility tools expose several of those middle stages, but their labels are not interchangeable.

Ahrefs counts a mention when a brand appears at least once in an AI-generated response. It counts a citation when a page or domain appears as a cited source. Its estimated impressions map prompt questions to modeled Google search demand, and its AI share of voice compares that modeled exposure with a selected competitor set. Ahrefs describes those measures as directional indicators rather than actual audience reach.

Semrush uses a different construction. Its AI Visibility Score combines topic coverage with mention consistency across a large prompt-and-response database. The company also reports citations, cited pages, share of voice and sentiment. Semrush explicitly says no platform can provide exact visibility numbers because responses are fast-changing and personalized.

Those caveats do not make the metrics useless. They define their job. A repeated prompt panel can show whether a brand is appearing more consistently, whether competitors occupy topics it misses and which pages are being selected as sources. It cannot turn a modeled prompt corpus into an exact count of people who saw an answer.

Each metric observes a different part of search performance

Measurement layer, unit and strongest interpretation limit as of August 21, 2026

Average position

Observes
Classic Google result placement
Unit or denominator
Topmost property or page position averaged across recorded impressions
Cannot establish
Attention, citation, visit quality or revenue

Generative-feature impressions

Observes
A URL appearing in Google’s generative search features
Unit or denominator
Counted exposure under Google’s surface-specific rules
Cannot establish
Cross-engine visibility, attention or business value

Mentions

Observes
A brand or entity appearing in sampled answers
Unit or denominator
Prompt responses containing the entity at least once
Cannot establish
Source use, citation, reach or favorable context

Citations

Observes
A page or domain visibly selected as a source
Unit or denominator
Cited responses or cited pages in a platform or sampled prompt panel
Cannot establish
How much the source shaped the answer or whether anyone clicked

Referral engagement

Observes
Behavior after an identifiable AI click
Unit or denominator
Sessions, time, depth, bounce or assisted paths
Cannot establish
Zero-click influence or causal credit for the visit

Conversions and revenue

Observes
Observed business outcomes
Unit or denominator
Qualified actions, transactions, pipeline or money
Cannot establish
Which upstream exposure caused the outcome without a valid design

Sources: Google Search Console and Search Central; Martinez’s 2026 GEO review; Ahrefs and Semrush methodologies. Definitions are condensed for comparison; vendor-modeled exposure is not actual audience reach.

The practical error is not choosing one vendor over another. It is placing average position, modeled impressions, mention share, citation counts and conversions in one dashboard as though they share a denominator. They do not. A useful report keeps each metric attached to the stage and population it actually observes.

What the Business Metrics Add

The value of separating stages becomes visible after a click.

Adobe Digital Insights reported that during the 2025 U.S. holiday shopping season, visits from generative-AI referrals behaved differently from other traffic sources. AI-referred shoppers spent 45% more time on site, were 33% less likely to leave immediately, converted at a 31% higher rate and viewed 13% more pages per visit.

AI-referred retail visits showed stronger on-site signals

Relative difference versus other traffic sources, U.S. retail, November–December 2025

Adobe reported that AI-referred retail visits spent 45% more time on site, were 33% less likely to leave immediately, converted 31% more and viewed 13% more pages per visit than other traffic.

Source: Adobe Digital Insights, 2026. Company-reported observations from U.S. retail during the 2025 holiday season; the comparisons do not establish causality or generalize to every industry.

The figures do not prove that an AI citation caused higher purchase intent. They are observational, limited to U.S. retail during an unusually commercial period and published by the analytics provider that measured them. They do show why a ranking-only report can miss the character of the traffic that arrives.

Two pages could hold the same average Google position while producing different generative exposure, citation patterns and visit quality. A page could lose clicks while its brand is mentioned more often inside answers. Another could receive a small volume of AI referrals that convert unusually well. Rank alone cannot adjudicate among those outcomes because it observes none of them.

The Practical Measurement Stack

A defensible 2026 search scorecard needs several layers. The layers should be read together, but not added into one synthetic total.

1. Eligibility and classic search position

Track indexation, crawlability, query-level impressions, average position and result-type changes. Use rank to diagnose where classic visibility moved and whether important pages remain competitive. Pair it with impressions and clicks so an improved position without additional exposure or visits does not become a false victory.

2. Generative-surface exposure

Use Google’s dedicated generative-AI impressions where the pilot is available. Segment by page, country, device and time. This establishes whether URLs entered the visible generative surface, but it should remain separate from conventional impressions when the interfaces and counting rules differ.

3. Mentions, citations and prompt coverage

Run a fixed, disclosed prompt panel across the engines that matter to the audience. Record whether the brand appears, whether the domain is cited, which page is cited, the topic or intent and the response date. Repeat prompts and preserve misses. A one-time screenshot is an example, not a trend.

Competitive share can be useful when the competitor set is stable and the calculation is documented. It becomes fragile when brands, prompts, regions or engines change between reporting periods.

4. Referral behavior

Measure sessions from identifiable AI referrers, landing pages, engagement, assisted paths and conversion rate. Keep in mind that referral analytics only sees users who click and whose source is preserved. It cannot observe people who read an answer, remember a brand and return later through direct or branded search.

5. Business outcomes

Connect search work to qualified leads, transactions, revenue, retention or another outcome the organization actually values. Use experiments, annotations and comparable baselines when claiming that a change caused an outcome. A simultaneous rise in citations and sales is a lead for investigation, not causal proof.

Why No Single AI Visibility Score Is Enough

The pressure to invent one replacement for rank is understandable. Average position gave teams a compact number that could be trended, compared and explained. AI search is harder to summarize because its observable surface is unstable and its internal retrieval process is mostly hidden.

The available evidence argues against solving that complexity with another universal score.

First, tools sample different prompt universes. Ahrefs ties parts of its model to search interest and refreshes question sets on a schedule. Semrush groups prompts into topics and combines coverage with mention consistency. Both approaches can be useful, but their scores answer questions about their own modeled corpora.

Second, the engines differ. A brand can be mentioned without its site being cited. A page can be retrieved without appearing as a visible source. A citation can support only a small part of an answer. The same prompt can produce a different source set across runs, locations, accounts and model updates.

Third, business value sits downstream. The Yext research release reporting that 64% of marketing leaders were unsure how to measure AI-search success is consistent with an immature measurement category, but it does not establish which score should win. The survey was commissioned by a visibility vendor and used a qualified, multi-country consumer sample; the figure should be read as evidence of reported uncertainty, not as a census of the profession.

The better executive summary is therefore layered: classic visibility, generative exposure, mention and citation consistency, attributable visits and business outcomes. Rank remains in the first layer. It does not disappear; it stops pretending to be the whole system.

Methodology and Limitations

This explainer synthesizes four official Google documents, one academic preprint reviewing 45 GEO studies, two AI-visibility vendor methodologies, one first-party analytics analysis and one company-commissioned research release. The research cutoff is August 21, 2026.

The article uses vendor documentation to define the metrics those products expose, not to validate their commercial claims. Ahrefs and Semrush use proprietary prompt corpora, entity resolution and modeling, so their scores are not compared numerically. The Adobe figures are reported observations from U.S. retail visits during the 2025 holiday season and are not generalized to every industry or used as causal evidence. The Yext survey figure is retained with its published sample and sponsorship limitations.

Google’s dedicated generative-AI reports were in a limited rollout at publication. Their availability, fields and counting rules may change. Search Console covers Google surfaces, while cross-engine prompt trackers observe selected samples rather than a complete audience log.

MCP Scraper was attempted first for source discovery and extraction, but its installed app route required reauthentication and its independent route was temporarily unavailable. Full sources were read through current web-search and page-open fallback. No search-result snippet carries a load-bearing claim.

Conclusion: Rankings Are a Diagnostic, Not the Scoreboard

SEO is still about rankings in the same way that commerce is still about store placement: position affects the chance of being found, but it does not describe the whole customer journey.

Average position remains useful for diagnosing classic search visibility and for understanding one part of Google’s result system. Google’s own guidance, however, places impressions and clicks ahead of position alone, and its new generative-AI reporting adds a separate exposure layer. Cross-engine measurement adds mentions and citations. Analytics adds visits, engagement and conversion. Each metric answers a different question.

The defensible replacement for rank tracking is not one new score. It is a measurement stack whose layers keep their own denominators, limitations and owners. Rankings belong in that stack. They are no longer sufficient to stand in for search performance.

Common questions

Frequently asked questions

Sources

Methodology note

Evidence-first explainer using official platform documentation, one academic preprint, vendor metric methodologies, company-reported analytics and a commissioned survey. Research cutoff: August 21, 2026. Unlike denominators are preserved and no visibility scores are pooled.

AA

Published by

Andrew Ansley

Andrew Ansley writes about search, information retrieval, AI recommendation systems, and the evidence systems use to form answers.