Guide · Updated 10 August 2026 · 12 min
How to measure whether AI search recommends your company
A repeatable test for brand mentions, citations and recommendations in AI answers, without treating a variable answer like a search ranking.
First-pass estimate: 90 minutes for a first baseline

One encouraging screenshot can conceal several problems. An AI answer may cite a company page without naming the company. It may recommend the company using a third-party source. It may get the reason to buy wrong.
Measure those outcomes separately. Start with 12 to 20 questions drawn from real buying decisions, run them more than once in the products your buyers use and keep the full answers. A spreadsheet is enough for the first baseline.
What an AI search audit actually measures
Here, AI search means a product or search feature that retrieves current web pages and uses them in a generated answer. ChatGPT Search, Google AI Mode, Google AI Overviews and Microsoft Copilot qualify. They use different systems and present results differently.
Traditional search usually gives the user a ranked set of links. AI search can retrieve several sources, combine their information and present a conclusion before the user visits any page. That creates several points at which a company may appear or disappear.
- 01Buyer question
The problem, company context and constraints the buyer describes.
- 02Retrieval
The searches the product runs and the pages it considers relevant.
- 03Citation
The sources it displays as support for part of the answer.
- 04Representation
What the answer says your company does, suits and proves.
- 05Recommendation
Whether the company is presented as a sensible option for this buyer.
The chain is a way to read the output, not a claim about the hidden architecture of every product. A retrieved page may never appear as a citation. A citation may support a general fact without earning the company a mention. A mention can be neutral or even cautionary.
Record four outcomes:
- Mention: does the answer name the company at all?
- Recommendation: does it present the company as relevant to the buyer’s stated situation?
- Citation: does it link to a page owned by the company, or to a third party discussing it?
- Accuracy: do the cited pages actually support what the answer says about the company?
Track traffic beside them. A recommendation can influence a shortlist without a click, while a cited page can receive visits even when the company was never recommended. Referral data cannot substitute for reading the answer.
What the audit can and cannot tell you
A controlled audit shows how selected products answered a defined question set under recorded conditions. Repeated runs can expose recurring competitors, inaccurate descriptions and sources that shape the answer.
It says nothing about prompts outside the sample or the total number of buyers who saw the company. It cannot reveal why a system selected a source, and tomorrow’s answer may differ. Report the result plainly: the company appeared in this share of observations for these questions, products and conditions.
Build the question set around buying decisions
Skip the thirty variations of “best software.” Mine sales calls, discovery notes, support questions, proposal objections and internal site search for decisions buyers are actually trying to make.
Organise the questions into four groups:
- Understanding the problem. The buyer is trying to explain what is going wrong and what kinds of solution might exist. For example: “How can a regulated B2B team reduce the time spent checking campaign claims?”
- Choosing an approach. The buyer understands the problem and is comparing categories, operating models or ways to solve it. For example: “Should we buy a monitoring platform or build a reviewed workflow in our existing tools?”
- Building a shortlist. The buyer wants suitable providers for a defined situation. For example: “Which providers help a European software company build a market intelligence workflow with analyst review?”
- Qualifying a choice. The buyer is checking fit, risk, evidence, implementation requirements or alternatives. For example: “What should I compare when choosing between two AI content workflow providers?”
A company can look highly visible for its own name and disappear while a buyer is still defining the problem. Reporting by buying stage exposes that gap. It also shows when a page is being used as a source while the company itself is absent from the shortlist.
Keep two prompt sets separate. Use unbranded questions to test discovery and recommendation. Use branded questions such as “What does [company] provide, and who is it suitable for?” to test whether the company is described accurately. A branded prompt cannot demonstrate unprompted recommendation because the company was supplied in the question.
There is no correct question count for every business. Twelve well-chosen questions make a manageable first pass. Add segments, countries or prompt variations only when they represent a real commercial difference.
Test products under recorded conditions
Choose products your buyers are likely to use and name the exact surface. “Google” is too vague: an ordinary result page, an AI Overview and AI Mode are different experiences. An AI Overview may not appear for a query at all. Google explains that AI Overviews appear only when its systems decide they add value, while AI Mode and AI Overviews may use different models and techniques.
For every observation, record:
- Product and surface
- Whether web search or retrieval was visibly used
- Exact prompt
- Country and language
- Signed-in or signed-out state
- Memory or personalisation setting where the product exposes one
- Date and run number
Use a new or temporary conversation for each run so earlier prompts do not shape later answers. The tests still will not be identical. OpenAI documents that ChatGPT Search may rewrite a prompt into several searches and can use general location and relevant memory. Record those conditions.
Repeat the important questions. A single answer is one observation. A 2026 study comparing Google, OpenAI and Perplexity systems found substantial differences in retrieval, source diversity and stability, including variation across executions and time. The research was published in Findings of ACL 2026.
Three questions at each of four buying stages.
The products most relevant to the market being studied.
Fresh observations under the same recorded conditions.
A practical first sample, not a universal research standard.
Two runs expose immediate instability. Add a third when an important question produces conflicting answers, and retain every result. Stop there rather than rerunning until a preferred answer appears.
Read the answer and its evidence separately
Read each response, then open every source attached to a claim about the company. Record the URL, its publisher and whether the page supports the statement.
A citation icon shows that a source was associated with an answer; it does not guarantee support for every nearby sentence. Research on generative search has found gaps between generated claims and cited evidence. The systems studied in 2023 have since changed, so its percentages should not be treated as current performance. The underlying verification problem remains worth testing.
A recommendation can be clear while its evidence is weak
A fictional provider is recommended for a regulated European company because the answer says it offers a particular data-residency option.
The linked pricing page describes plans and features but does not state where customer data is hosted.
- Company mentioned
- Yes
- Presented as suitable
- Yes
- Company page cited
- Yes
- Reason verified
- No
The company received visibility while the decisive claim remained unsupported. That result belongs in the correction queue, even though a simple monitoring tool might colour it green.
Alongside the raw response, record the competitors named and the role each company receives. Use plain categories such as lead recommendation, shortlist option, incidental mention and not mentioned. Generated answers do not always behave like ranked lists, so “named first” should not automatically be treated as position one.
Calculate a baseline without inventing one score
Keep the denominator visible and report each measure by product and buying stage.
- Mention rate: unbranded observations that name the company divided by all valid unbranded observations.
- Recommendation rate: unbranded observations that present the company as suitable divided by all valid unbranded observations.
- Owned-source citation rate: web-grounded observations that cite at least one company-owned page divided by all valid web-grounded observations.
- Claim accuracy: checked claims about the company that are supported by an authoritative source divided by all company claims checked.
Do not average them into a proprietary “AI visibility score.” A combined number can rise while factual accuracy falls. It can also drop simply because the latest audit added harder, more valuable prompts.
Keep referral sessions, conversions and sales evidence beside the audit rather than inside the score. They show what happened after some users followed a link or disclosed their research path. They do not reveal every answer that influenced a buyer without a click.
Diagnose the gap before choosing the fix
Connect recurring observations with a likely problem and a specific check. Otherwise the audit ends as a generic list of AI-search tactics.
| What you observe | What may be missing | What to check next |
|---|---|---|
| The company and its pages never appear. | Eligibility, crawl access, relevance or sufficient public information. | Indexing, robots controls, CDN access, canonical URLs, internal links and whether a useful page answers the question. |
| A company page is cited but the company is not discussed. | The page supplies a fact but does not establish the company as a suitable provider. | Whether the page clearly connects the evidence to the offer, intended customer and relevant use case. |
| The company is mentioned inaccurately. | Outdated, ambiguous or contradictory information across sources. | Product pages, documentation, profiles, review pages and the exact sources used in the answer. |
| The company is accurate but rarely recommended. | Weak fit for the tested situation, unclear differentiation or insufficient proof. | Whether the question matches the actual ideal customer and whether the public evidence supports a reason to choose the company. |
| Recommendations rely only on third-party sources. | The company site may not contain the useful comparison, evidence or detail those sources provide. | Which facts third parties contribute and whether the company can publish a clearer authoritative source without copying their opinion. |
| Citations rise but useful visits do not. | The source may answer the question completely, attract low-intent research or appear without a prominent recommendation. | Landing pages, referral sessions, buyer stage, calls to action and any sales conversations that mention AI-assisted research. |
The middle column contains hypotheses, not diagnoses. Confirm the issue before changing the site. A company that is correctly excluded from an unsuitable shortlist does not have a visibility problem.
Combine manual observations with platform data
Some platform data is now available, although each source records a different slice.
Google now reports AI feature traffic within Search Console, and in June 2026 it began rolling out dedicated generative AI performance reports showing impressions, pages, countries, devices and dates for participating sites. This is useful visibility data, but it does not reproduce every answer or explain whether a company was recommended.
Bing Webmaster Tools introduced an AI Performance report covering citations across Microsoft Copilot, AI summaries in Bing and selected partner integrations. Bing explicitly notes that total citations do not indicate placement or presentation inside a particular answer.
OpenAI advises publishers to allow OAI-SearchBot if they want pages to be eligible for summaries and citations in ChatGPT Search. The same guidance explains that referral links include utm_source=chatgpt.com, which makes resulting visits identifiable in analytics. OAI-SearchBot concerns search visibility, while GPTBot is the separate control described for potential model training.
Platform reports show impressions, citations or pages over time. Web analytics records identifiable visits. Controlled prompts preserve the answer and its competitive context, while sales and customer research can reveal AI-assisted discovery in a real purchase. Keep the sources separate so each claim retains its proper denominator.
Improve the specific weakness the audit found
If important pages are inaccessible, fix indexing, crawl access, canonical URLs and internal links. Put the important information in text. When answers describe the company incorrectly, create one current source for the disputed fact and remove contradictions from pages you control.
If the company is understood but rarely recommended, inspect fit and evidence before rewriting the whole site. The answer may be a sharper audience definition, clear implementation requirements, an honest comparison or first-party evidence for the decisive claim. Research earns attention by answering a question people genuinely have, not by carrying an “original” label.
When a third-party page repeatedly shapes the answer, read it closely. Correct factual errors through the publisher’s legitimate process and improve the authoritative information on your own site. Do not manufacture reviews, community posts or artificial citations.
Do not make llms.txt a prerequisite. Google states that sites do not need new machine-readable AI files or special schema to appear in its AI features. Keep structured data accurate and consistent with the visible page, but do not present it as a guaranteed AI ranking lever. OpenAI’s documented search requirement is access for OAI-SearchBot, not an llms.txt file.
Rerun without erasing the baseline
Keep the core questions, wording and conditions stable so later observations remain comparable. Save the full answers and source URLs, not only the score, because a citation can change while the company remains mentioned.
At the same time, a frozen question set will gradually stop representing the market. Reserve part of the audit for new questions from sales conversations, product changes and emerging alternatives. Keep stable and exploratory results separate.
A fast-moving category may justify monthly checks; a stable one may need fewer. Rerun after a material change to the offer, position or public evidence.
Questions worth answering
- What does AI search visibility mean?
- It describes whether a company, product or source appears when an AI system answers a relevant buyer question. Measure mentions, source citations, recommendations and referral visits separately; they are not the same result.
- How do you measure AI search visibility?
- Choose a fixed set of real buyer questions and run each one more than once under documented conditions. Save the full answers and cited URLs. Then compare mention, citation and recommendation rates by question type and system, alongside first-party analytics and available platform reports.
- Does structured data guarantee inclusion in AI answers?
- No. Valid structured data can clarify page content and entities, but it cannot guarantee a citation or recommendation. The markup also needs to match the visible page.
- Do you need an llms.txt file for AI search visibility?
- No major answer system documents llms.txt as a requirement for ranking or inclusion. It is an optional experiment, not a replacement for crawlable pages, clear entity information, useful evidence and established search fundamentals.