Key takeaways
- A cited URL may be selected without materially supporting the generated answer.
- Track the pipeline from search activation and retrieval through citation, claim support, referral, and commercial outcome.
- Sampling must repeat prompts, platforms, locations, and dates because generative results are not a fixed ranking list.
- Human claim-level review is still necessary when the question is whether your evidence was represented faithfully.
01
Separate being listed from being used
A source can appear beside an answer without supplying the sentence a reader remembers. It may support a minor detail, sit behind an expandable panel, or be one of several retrieved pages the model barely used. Research on citation generation repeatedly treats citation correctness and completeness as separate problems because a plausible-looking reference does not automatically support the adjacent claim.
A 2026 GEO measurement paper proposes distinguishing citation selection from citation absorption: whether a page is chosen as a source, and whether its language, facts, evidence, or structure actually contributes to the answer. That distinction gives marketing teams a more honest unit of analysis than “we appeared.”
02
Build a funnel that keeps unlike signals apart
Record whether the experience triggered search, whether your domain was retrieved or cited, where the citation appeared, which claim it supported, whether the representation was accurate, whether someone clicked, and what happened after the visit. These stages influence one another, but they are not interchangeable.
Microsoft’s AI Performance preview provides citation totals, average cited pages, grounding-query samples, page-level activity, and trends. Microsoft cautions that citation count does not indicate ranking, authority, placement, or the role of a page in an answer. Treat the dashboard as visibility evidence, then add the layers it cannot observe.
Swipe to compare every column
| Stage | Evidence to retain | Do not conclude |
|---|---|---|
| Activation | Did the interface search the web? | Every prompt has the same opportunity |
| Selection | URL cited, platform, prompt, time | The page shaped the central answer |
| Absorption | Claim-to-source support review | A citation proves faithful use |
| Referral | Landing URL and tagged referrer where available | No click means no influence |
| Outcome | Qualified action, pipeline, or task completion | A visit caused the outcome by itself |
03
Sample a changing system without pretending it is fixed
Run a defined prompt set across the same platform, market, language, device context, and logged-in state where those factors are relevant. Repeat it on several dates and retain screenshots or structured observations. Report the share of runs with a citation and the distribution, not a single “position.”
Keep navigational, informational, comparison, and high-stakes prompts in separate groups. A brand query and an open-ended category comparison exercise different retrieval behavior. When the sample is small, show the denominator. “Cited in 3 of 10 observed runs” is useful; “30% AI visibility” without the method is theatre.
04
Review the claim that matters, not every generated word
Choose a few commercially or reputationally important claims: price boundaries, capabilities, locations served, product differences, research findings, and safety constraints. For each observed answer, ask whether the cited page entails the claim, merely relates to it, conflicts with it, or leaves it unsupported.
Use the findings to improve the source page when the limitation is yours: unclear scope, stale facts, missing evidence, or a comparison that buries its criteria. When the system misreads a clear source, document the failure rather than rewriting the page into unnatural fragments. Measurement should protect the truth of the content, not reward whatever wording a model happens to copy.
Primary sources and further reading
Use the source material to validate details against your own context and current platform configuration.
- Microsoft Bing: AI Performance in Webmaster Tools
- From Citation Selection to Citation Absorption: a GEO measurement framework
- On the Capacity of Citation Generation by Large Language Models
- VeriCite: rigorous verification for citations in retrieval-augmented generation
This field note follows the XenGrowth editorial policy: primary sources where available, visible limitations, material review dates, and no invented first-hand experience.
Stay with the problem



