Home
Services
AI SEO & Entity Authority AI Content Marketing Link Building & Digital PR AI Search Visibility (AEO) Conversion Rate Optimisation Paid Growth & Performance
Company
About Us Who It's For Careers
More
Blog Get a Quote Book a Free Strategy Call
AI

How to Measure Share of Answer

A page has to clear two gates before it appears as a citation in an AI answer. It has to enter the candidate pool, which is a retrieval problem and largely a search problem. Then it has to be picked out of that pool, which is a selection problem and is not a search problem at all.

The reason this matters commercially is simple. The second gate is where most brands are losing, and almost all the spending goes on the first.

Diagram showing a query fanning out into sub-queries, gate one retrieval assembling a shortlist of five to ten candidate pages, and gate two selection comparing candidates in pairs before citing a subset.

The pipeline. Retrieval assembles a shortlist. A separate process decides which members of that shortlist get named. Production systems typically retrieve five to ten pages per sub-query and cite fewer.

4.89
Average retrieval rank of the references that actually got cited, not rank one
62.5%
Share of cited references that came from the top ten retrieved documents at all
p 0.49
A fourteen-fold authority gap produced no significant difference in output

Gate one decides whether you are in the room

Retrieval assembles a ranked shortlist for each sub-query the model generates. This is the layer that search engine optimisation was designed for, and on this layer it works exactly as advertised.

Everything an index knows is available here. Links, crawl history, site structure, freshness signals, domain-level authority. Google has stated plainly that its AI features sit on top of the existing Search ranking system, which makes ordinary SEO the mechanism at this gate rather than a proxy for it.

One detail is worth more than most content programmes. Published analysis puts roughly 87 percent of ChatGPT-cited pages among Bing's top results. If your sitemap has never been submitted to Bing Webmaster Tools, you are close to invisible on the largest AI search surface in the world, regardless of how you rank in Google.

Gate two decides whether you are the one quoted

By this stage your page has been reduced to an extract. Your backlink profile is not in the extract. Your domain rating is not in the extract. What survives is the text itself.

Agentic retrieval systems do not score candidates in isolation. They run pairwise judgements of the form: given this question, is passage A or passage B more useful? Thousands of those comparisons aggregate into a final ordering. Your page is judged against whatever else happened to be retrieved that day.

Clearing gate one gets you considered. It does not get you quoted, and the evidence that these are separate outcomes is stronger than most of the advice built on top of it.

The cited source is not the top-ranked source

This is the finding that should reshape how the category thinks, and it is barely quoted in the marketing literature. A controlled evaluation of a domain-specific retrieval system examined 277 references that appeared in final answers.

Only 173 of them came from the top ten documents that retrieval had ranked. The cited references sat at a mean rank of 4.89, with a median of 4.67 and a standard deviation of 2.40.

Chart showing that cited references clustered around retrieval rank 4.89 with a standard deviation of 2.40, rather than at rank one where ranking-based SEO assumes citations go.

Where the citations actually came from. If retrieval rank determined citation, cited references would cluster at position one. They cluster around position five. Source: Enhancing Large Language Models with Domain-specific Retrieval Augmented Generation, arXiv 2409.13902.

The generation stage did not reach for whatever retrieval had ranked first. It reached into the middle of the list. The paper states it directly: the top-ranked documents by the retrieval system may not be the ones selected in the final response.

Authority behaves differently at each gate

At gate one, authority is decisive. Domain strength, link profile and topical history are precisely what ranking systems are built to weigh, and a stronger domain enters more candidate pools for more sub-queries. That effect is real and it compounds.

At gate two, it largely stops mattering. A controlled study compared two document sets answering identical clinical questions. One set averaged 303 academic citations. The other averaged 21.

Comparison showing highly cited source documents averaging 303 citations produced answers scoring 4.6 for accuracy, while less cited documents averaging 21 citations scored 4.2, with a p value of 0.49.

A fourteen-fold authority gap, and no reliable advantage. Accuracy, readability and response time all showed no statistically significant difference between the two source sets.

Read that carefully, because it is easy to over-claim from. It does not say authority is worthless. Authority is a large part of what gets documents into a corpus in the first place. What it says is narrower and more useful: once two sources are both in front of the model, being the more authoritative one does not reliably win.

The link between ranking and citation is weakening

For a long time, ranking was a good enough proxy for citation that the distinction did not matter much in practice. That is changing measurably.

Bar chart showing that before 2026 over 92 percent of AI Overview citations came from top ten ranking pages, while after the Gemini 3 upgrade only 38 percent did.

Directional rather than exact. The two figures come from different analysts measuring slightly different units, domains against pages, so read the change as direction rather than a precise delta.

Ranking still helps you into the shortlist. It no longer carries you through the second gate on its own, and the gap between those two statements is where the newer work lives.

Which gate are you failing?

The two failure modes need opposite responses, and telling them apart takes one afternoon rather than a subscription. Server logs are the cheapest diagnostic available and the one most teams have never looked at, because a crawler hit proves you cleared the first gate.

What you observeWhich gateWhat to do about it
No AI crawler hits in your server logsGate oneCrawlability and index coverage. Check robots.txt and submit to Bing.
Crawler hits, but no citations anywhereGate twoEditorial. Your pages are being read and passed over.
Strong rankings, weak AI presenceGate twoThe classic pattern. Rank got you into the pool and nothing carried you further.
Cited on one engine, absent on anotherBothDifferent retrieval corpora. Check index presence per engine before rewriting.
Cited, but for the wrong claimGate twoYour extractable passage is not the one you would have chosen. Rewrite the passage.

What actually moves each gate

The work is different in kind, it sits with different people, and it costs differently.

1

Retrieval

Owned by your SEO function

  • Submit your sitemap to Bing Webmaster Tools
  • Audit robots.txt for GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot by name
  • Render the answer server-side, not after JavaScript runs
  • Read your access logs for crawler activity per URL
  • Keep the authority programme you already run
No new headcount Technical health
2

Selection

Owned by your editors

  • Put a specific number, a named source and a date in the passages you want quoted
  • Quote named experts directly rather than paraphrasing them
  • Answer the question before you contextualise it
  • Make each section stand alone when read out of context
  • Attribute externally, which reads as thoroughness rather than as leakage
Peer-reviewed lift of 30 to 40% Editorial standard

The verdict

The evidence quality across the two gates is lopsided in a way nobody advertises. Gate one rests on Google's own documentation and twenty five years of search practice. Gate two rests on a small number of peer-reviewed studies, all of which point the same way, and none of which are widely quoted.

ClaimStrength of evidence
Retrieval rank does not determine citationMeasured. Cited references sat at mean rank 4.89, not rank one.
Source authority does not decide selectionMeasured. A fourteen-fold citation gap gave no significant difference, p equals 0.49.
Ranking is a weakening proxy for citationVendor-measured, directionally strong. 92 percent falling to 38 percent.
Statistics, sources and quotations lift selectionPeer-reviewed, 30 to 40 percent, from the Princeton GEO study.
You must clear gate one firstConfirmed by Google in its own documentation.
Fund gate one as maintenance and gate two as editorial standards. Gate one is a solved problem you already have staff for, and if it is broken nothing else matters. Gate two is where the differentiated work is, it costs almost nothing in cash, and it is the gate the research actually addresses.

The practical translation is smaller than any agency proposal you will receive. Put a number, a named source and a date in the passages you most want quoted. Answer the question before you contextualise it. Then check your server logs to confirm you are being read at all, because if you are not, none of the rest applies.

The industry has built a large apparatus for measuring gate two and almost no apparatus for improving it. That gap, rather than any single tactic, is the real opportunity.

A note on what this evidence can and cannot support

The two strongest findings here come from domain-specific retrieval evaluations in medicine rather than from open web search. They establish the mechanism rather than the exact magnitudes you would see on your own site. The 92 to 38 percent figures come from two different analysts measuring different units and should be read as direction. Further factorial work on citation preference is in progress, including a 2026 study isolating which content attributes move first-citation preference.

Share this article

Sources and methodology

Where the evidence in this article comes from

  • Enhancing Large Language Models with Domain-specific Retrieval Augmented Generation, arXiv 2409.13902. The source of the 62.5 percent and mean rank 4.89 findings.
  • From Data to Decisions, a controlled retrieval evaluation comparing 30 highly cited papers against 30 less cited papers across 30 clinical questions.
  • What Gets Cited: Competitive GEO in AI Answer Engines, arXiv 2605.25517.
  • Aggarwal et al., GEO: Generative Engine Optimization, ACM SIGKDD 2024. The source of the 30 to 40 percent figures for statistics, citations and quotations.
  • Google Search Central guidance on generative AI features, May 2026.
  • seoClarity and Ahrefs analyses of AI Overview citation sources, before and after the January 2026 Gemini 3 upgrade.

Where a figure is vendor-published rather than peer-reviewed, the article says so at the point of use. Figures reflect July 2026 and this area changes quickly.

Find your gate

Which gate is costing you citations?

We audit both. Crawler access and index presence for gate one, extractability and citation discipline for gate two, with the server-log evidence to tell them apart.