Skip to content

Method · Updated 2 August 2026 · 15 min

How to measure whether AI search recommends your company

A repeatable method for measuring brand mentions, citations and recommendations across AI answer systems without pretending they behave like rankings.

A hand-drawn magnifying glass inspecting whether an AI answer cites a company website

An AI answer can cite your website without naming your company, recommend your company using a third-party source, or mention you while getting the reason to buy wrong. Those are different outcomes, so a useful audit measures them separately.

Start with 12 to 20 questions that represent real buying decisions, test them in the AI search products relevant to your market, repeat each run and record four things: whether your company is mentioned, whether it is presented as a suitable choice, which sources support the answer and whether the claims are accurate. The result is a controlled baseline. It is not a measure of every buyer conversation or your share of an entire market.

This article explains how to build that baseline, what the numbers mean and how to decide what needs fixing. You can run the first version yourself with a spreadsheet. The important part is not the tool. It is separating the stages of visibility so that one encouraging screenshot does not become a strategy.

What an AI search audit actually measures

In this article, AI search means a product or search feature that can retrieve current web pages and use them to generate an answer. ChatGPT Search, Google AI Mode, Google AI Overviews and Microsoft Copilot are examples, but they do not work as one shared index or present results in one standard format.

Traditional search usually gives the user a ranked set of links. AI search can retrieve several sources, combine their information and present a conclusion before the user visits any page. That creates several points at which a company may appear or disappear.

  1. 01Buyer question

    The problem, company context and constraints the buyer describes.

  2. 02Retrieval

    The searches the product runs and the pages it considers relevant.

  3. 03Citation

    The sources it displays as support for part of the answer.

  4. 04Representation

    What the answer says your company does, suits and proves.

  5. 05Recommendation

    Whether the company is presented as a sensible option for this buyer.

The chain is a simplified model, not a claim that every product follows five visible steps. Its purpose is to show why one score is inadequate. Retrieval is not the same as citation, citation is not the same as a company mention, and a mention is not necessarily a recommendation.

A useful audit therefore records four separate outcomes:

  • Mention: does the answer name the company at all?
  • Recommendation: does it present the company as relevant to the buyer’s stated situation?
  • Citation: does it link to a page owned by the company, or to a third party discussing it?
  • Accuracy: do the cited pages actually support what the answer says about the company?

Traffic is a fifth, downstream outcome. A recommendation may influence a shortlist without creating a click, while a cited page may receive visits from people who never saw the company as a recommendation. Referral traffic matters, but it should not be used as a substitute for examining the answer.

What the audit can and cannot tell you

A controlled audit can show how selected products answered a defined set of questions under recorded conditions. It can reveal recurring competitors, inaccurate descriptions, important sources and gaps that appear across multiple runs. It can also give you a baseline that can be repeated after a product launch, positioning change or substantial content update.

It cannot tell you how every buyer prompts an AI system, estimate the total number of buyers who saw your company, prove why a system selected a source or promise that the same answer will appear tomorrow. There is no universal AI search ranking shared by all products.

That limitation is not a reason to avoid measurement. It is a reason to describe the result correctly: this company appeared in this proportion of recorded observations for these questions, products and conditions.

Build the question set around buying decisions

Do not begin with a list of keywords or thirty variations of “best software.” Start with decisions a real buyer is trying to make. Sales calls, discovery notes, support questions, proposal objections and internal site-search data are useful sources because they preserve the language and constraints buyers actually use.

Organise the questions into four groups:

  1. Understanding the problem. The buyer is trying to explain what is going wrong and what kinds of solution might exist. For example: “How can a regulated B2B team reduce the time spent checking campaign claims?”
  2. Choosing an approach. The buyer understands the problem and is comparing categories, operating models or ways to solve it. For example: “Should we buy a monitoring platform or build a reviewed workflow in our existing tools?”
  3. Building a shortlist. The buyer wants suitable providers for a defined situation. For example: “Which providers help a European software company build a market intelligence workflow with analyst review?”
  4. Qualifying a choice. The buyer is checking fit, risk, evidence, implementation requirements or alternatives. For example: “What should I compare when choosing between two AI content workflow providers?”

These groups matter because a company can be highly visible for its own name and absent when a buyer is still defining the problem. It can also be cited as a useful source without being treated as a provider. Reporting by buying stage makes that difference visible.

Keep two prompt sets separate. Use unbranded questions to test discovery and recommendation. Use branded questions such as “What does [company] provide, and who is it suitable for?” to test whether the company is described accurately. A branded prompt cannot demonstrate unprompted recommendation because the company was supplied in the question.

There is no scientifically correct number of questions for every business. Twelve well-chosen questions can support a manageable first pass. A formal programme may need more segments, countries and prompt variations. Add complexity when it represents a real commercial distinction, not because a larger spreadsheet looks more authoritative.

Test products under recorded conditions

Choose the products your buyers are reasonably likely to use and name the exact surface you tested. “Google” is not precise enough because an ordinary result page, an AI Overview and AI Mode are different experiences. An AI Overview may not appear for a particular query at all. Google explains that AI Overviews appear only when its systems decide they add value, while AI Mode and AI Overviews may use different models and techniques.

For every observation, record:

  • Product and surface
  • Whether web search or retrieval was visibly used
  • Exact prompt
  • Country and language
  • Signed-in or signed-out state
  • Memory or personalisation setting where the product exposes one
  • Date and run number

Use a new or temporary conversation for each run so that earlier prompts do not shape later answers. This reduces one source of variation, but it does not make every test identical. OpenAI documents that ChatGPT Search may rewrite a prompt into several searches and can use general location and relevant memory. Record those conditions instead of assuming they have disappeared.

Most importantly, repeat the important questions. A single answer is an observation, not a stable ranking. A 2026 study comparing Google, OpenAI and Perplexity systems found substantial differences in retrieval, source diversity and stability, including variation across executions and time. The research was published in Findings of ACL 2026.

Question set12

Three questions at each of four buying stages.

Selected surfaces3

The products most relevant to the market being studied.

Separate runs2

Fresh observations under the same recorded conditions.

Observations72

A practical first sample, not a universal research standard.

Two runs will not remove uncertainty, but they immediately expose answers that change. If a question is commercially important or the two runs disagree, add a third run and retain all results. Do not keep rerunning until the answer you prefer appears.

Read the answer and its evidence separately

For each response, first read what the answer says. Then open every source attached to a claim about your company. Record the cited URL, who controls it and whether it supports the statement.

This verification is not pedantry. A citation icon shows that a source was associated with an answer, not that every nearby sentence is supported by that page. Research on generative search has found meaningful gaps between generated claims and cited evidence. The systems studied in 2023 have since changed, so the old percentages should not be treated as current performance, but the underlying verification problem remains useful to test.

Illustrative observation

A recommendation can be clear while its evidence is weak

Generated answer

A fictional provider is recommended for a regulated European company because the answer says it offers a particular data-residency option.

Cited page

The linked pricing page describes plans and features but does not state where customer data is hosted.

Company mentioned
Yes
Presented as suitable
Yes
Company page cited
Yes
Reason verified
No

The company received visibility, but the most important claim still needs correction. That may require a clearer authoritative page, correction of an external source, or both. Counting the observation as an uncomplicated success would hide the risk.

Alongside the raw response, record the competitors named and the role each company receives. Use plain categories such as lead recommendation, shortlist option, incidental mention and not mentioned. Generated answers do not always behave like ranked lists, so “named first” should not automatically be treated as position one.

Calculate a baseline without inventing one score

Calculate each measure with a visible denominator and report it by product and buying stage.

  • Mention rate: unbranded observations that name the company divided by all valid unbranded observations.
  • Recommendation rate: unbranded observations that present the company as suitable divided by all valid unbranded observations.
  • Owned-source citation rate: web-grounded observations that cite at least one company-owned page divided by all valid web-grounded observations.
  • Claim accuracy: checked claims about the company that are supported by an authoritative source divided by all company claims checked.

These measures answer different questions. Do not average them into a proprietary “AI visibility score.” A combined number can improve while factual accuracy falls, or decline because you added more difficult and commercially valuable prompts.

Keep referral sessions, conversions and sales evidence beside the audit rather than inside the score. They show what happened after some users followed a link or disclosed their research path. They do not reveal every answer that influenced a buyer without a click.

Diagnose the gap before choosing the fix

The audit becomes useful when it changes a decision. Instead of producing a generic list of AI search tactics, connect each recurring observation with a likely problem and a specific check.

How to investigate common AI search visibility gaps
What you observeWhat may be missingWhat to check next
The company and its pages never appear.Eligibility, crawl access, relevance or sufficient public information.Indexing, robots controls, CDN access, canonical URLs, internal links and whether a useful page answers the question.
A company page is cited but the company is not discussed.The page supplies a fact but does not establish the company as a suitable provider.Whether the page clearly connects the evidence to the offer, intended customer and relevant use case.
The company is mentioned inaccurately.Outdated, ambiguous or contradictory information across sources.Product pages, documentation, profiles, review pages and the exact sources used in the answer.
The company is accurate but rarely recommended.Weak fit for the tested situation, unclear differentiation or insufficient proof.Whether the question matches the actual ideal customer and whether the public evidence supports a reason to choose the company.
Recommendations rely only on third-party sources.The company site may not contain the useful comparison, evidence or detail those sources provide.Which facts third parties contribute and whether the company can publish a clearer authoritative source without copying their opinion.
Citations rise but useful visits do not.The source may answer the question completely, attract low-intent research or appear without a prominent recommendation.Landing pages, referral sessions, buyer stage, calls to action and any sales conversations that mention AI-assisted research.

The middle column contains hypotheses, not diagnoses. Confirm the issue before changing the site. A company that is correctly excluded from an unsuitable shortlist does not have a visibility problem.

Combine manual observations with platform data

The claim that no analytics property can reveal AI visibility is no longer accurate.

Google now reports AI feature traffic within Search Console, and in June 2026 it began rolling out dedicated generative AI performance reports showing impressions, pages, countries, devices and dates for participating sites. This is useful visibility data, but it does not reproduce every answer or explain whether a company was recommended.

Bing Webmaster Tools introduced an AI Performance report covering citations across Microsoft Copilot, AI summaries in Bing and selected partner integrations. Bing explicitly notes that total citations do not indicate placement or presentation inside a particular answer.

OpenAI advises publishers to allow OAI-SearchBot if they want pages to be eligible for summaries and citations in ChatGPT Search. The same guidance explains that referral links include utm_source=chatgpt.com, which makes resulting visits identifiable in analytics. OAI-SearchBot concerns search visibility, while GPTBot is the separate control described for potential model training.

These sources complement the manual audit:

  • Platform reports show broader impressions, citations or pages over time.
  • Web analytics shows identifiable visits and what some visitors do next.
  • Controlled prompts show the answer, competitive context and factual quality.
  • Sales and customer research show whether AI-assisted discovery influenced a real buying decision.

No single layer can replace the others.

Improve the specific weakness the audit found

If important pages are not eligible or accessible, begin with the ordinary search foundation: indexing, crawl access, canonical URLs, internal links and important information available as text. If answers describe the company incorrectly, create one clear, current source for the disputed fact and correct contradictions across pages you control.

If the company is understood but not recommended, inspect fit and evidence before rewriting everything. The right response may be a more precise audience definition, a documented methodology, a useful comparison, clear implementation requirements, transparent limitations or first-party evidence that supports a meaningful claim. Original research is valuable when it answers a question buyers and other authors genuinely need answered. Its value does not come from being labelled “original.”

When a third-party page repeatedly shapes the answer, read it closely. Correct factual errors through the publisher’s legitimate process and improve the authoritative information on your own site. Do not manufacture reviews, community posts or artificial citations.

Do not make llms.txt a prerequisite. Google states that sites do not need new machine-readable AI files or special schema to appear in its AI features. Keep structured data accurate and consistent with the visible page, but do not present it as a guaranteed AI ranking lever. OpenAI’s documented search requirement is access for OAI-SearchBot, not an llms.txt file.

Rerun without erasing the baseline

Keep the core questions, wording and conditions stable so later observations remain comparable. Save the full answers and source URLs, not only the score, because a citation can change while the company remains mentioned.

At the same time, a frozen question set will gradually stop representing the market. Reserve part of the audit for new questions from sales conversations, product changes and emerging alternatives. Keep stable and exploratory results separate.

The right cadence depends on how quickly the category changes and how important the channel is. A fast-moving market may justify monthly checks, while a stable category may need less frequent review. Always rerun after a material change to the offer, company position or public evidence.

The goal is not to prove that an AI system “likes” the company. It is to understand where a buyer can find it, what they are told, which evidence shapes the answer and whether the result deserves to influence a real decision. That is a much more useful baseline than one citation count.

Questions B2B teams ask

What does AI search visibility mean?
AI search visibility describes how a company, product or source appears when an AI answer system responds to relevant buyer questions. A useful measurement separates being mentioned, being cited as a source, being recommended for a use case and receiving a visit from the answer.
How do you measure AI search visibility?
Create a fixed portfolio of real buyer questions, run each question repeatedly under documented conditions, save the full answers and cited URLs, then compare mention, citation and recommendation rates by question type and system. Combine these observations with first-party analytics and the platform reports available to you.
Does structured data guarantee inclusion in AI answers?
No. Valid structured data can help systems understand page content and entities, but it does not guarantee a citation or recommendation. The markup must also match what a visitor can see on the page.
Do you need an llms.txt file for AI search visibility?
No major answer system documents llms.txt as a ranking or inclusion requirement. Treat it as an optional experiment, not a substitute for crawlable pages, clear entity information, useful evidence and established search fundamentals.

Read next

Why B2B marketing attribution is degrading, and what to measure instead

A decision-focused measurement model for B2B teams facing consent gaps, private buyer journeys and AI-assisted research.

If any of this describes your team

Describe the problem and the teams involved. I will tell you whether I can help and suggest a sensible next step.