Skip to content
GEO Measurement

A GEO Test You Can Defend in the Monday Meeting

Run an AI-search visibility test that another person can repeat. Discovery, retrieval, citation, answer use, traffic, and business outcomes stay separate instead of becoming one convenient score.

Research notes and a measurement worksheet for an AI search experiment

Field note

By XenGrowth EditorialPublished Reviewed 11 min read

Key takeaways

  • GEO is a partially observable pipeline; retrieval, citation, answer use, traffic, and revenue are separate outcomes.
  • Repeat prompts across paraphrases, sessions, dates, and engines because a single generated answer is not a stable rank check.
  • Predefine the edit, comparison set, observation window, primary measure, and conditions that would weaken the conclusion.
  • Keep a human fidelity review: being cited is not success when the answer misstates the source or drops an important condition.

01

Stop treating a generated answer like position three

A conventional rank tracker observes a relatively legible output. Generative discovery adds more hidden stages: whether search activates, which queries fan out, what is indexed, what is retrieved, how context is allocated, which sources are cited, which claims absorb source material, and whether a person acts. A recent critical survey of GEO research argues that these stages should not be collapsed into one visibility score.

Choose one primary outcome for the test. It might be source selection for a stable question panel, accurate use of a newly published dataset, or qualified visits to a decision guide. Record the other stages as diagnostics. This prevents a citation increase from quietly becoming a traffic or revenue claim.

02

Build a question panel with controlled variation

Start with ten to twenty questions tied to real reader jobs: direct explanation, comparison, constraint, local or industry context, implementation, and follow-up. Create two or three natural paraphrases for each. Do not stuff brand names into every prompt; include branded questions only when the business genuinely wants to measure them.

Run each variant repeatedly across a fixed engine and interface, then repeat on a separate date. Record location, language, account state where relevant, model or product label, timestamp, and whether web search was visibly used. Keep exploratory prompts outside the stable panel so curiosity does not corrupt the baseline.

  • Stable core question
  • Natural paraphrases
  • Engine and interface
  • Date, language, and location
  • Search activation
  • Sources, citation placement, and answer fidelity

03

Change one material thing and write down the counterfactual

A useful content test might add a dated method, publish an original comparison, clarify a buried definition, or repair crawlability. Avoid changing the title, structure, evidence, links, and distribution at the same time when the goal is learning. If multiple changes are operationally necessary, call the work a release rather than pretending it isolated a factor.

Choose comparison pages or questions that did not receive the edit. State what you would expect to observe if the change had no effect. Competition and platform updates can move every source at once; controls make that visible even when they cannot create a laboratory-perfect counterfactual.

04

Report a visibility vector, not a victory number

Summarize discovery eligibility, observed retrieval or citation, prominence, factual use, fidelity, identifiable visits, and business quality separately. Add sample size and uncertainty. Research published in 2026 emphasizes that citation breadth and citation absorption can diverge: a page may appear in the source list without materially supporting the answer, or strongly influence a narrow answer with few visible citations.

End with the next editorial decision. Strengthen the evidence, fix a misleading passage, improve retrieval context, wait for more observations, or stop. A defensible test is valuable even when the result is null because it prevents the team from scaling a superstition.

Swipe to compare every column

StageEvidence to recordFailure to watch
DiscoveryIndexing, crawl access, observable search activationTesting a page the engine cannot access
SelectionRetrieved or cited URL across repeated runsOne screenshot presented as stable visibility
AbsorptionClaims or evidence accurately used in the answerCitation present but source meaning absent
OutcomeVisits, qualified actions, commercial contextAttributing all change to the content edit

Primary sources and further reading

Use the source material to validate details against your own context and current platform configuration.

This field note follows the XenGrowth editorial policy: primary sources where available, visible limitations, material review dates, and no invented first-hand experience.

Stay with the problem

Explore AI search & GEO