A technical reference on AI search visibility. This site sells nothing, takes no engagements and endorses no products. Consulting enquiries are handled separately at hartzer.com.

Hartzer.it.com logoHartzer.it.comAI search visibility reference
Abstract horizontal band illustration representing Building an AI visibility report that survives scrutiny
Measuring AI Search Visibility

Building an AI visibility report that survives scrutiny

Four instruments, none complete, one of them a survey. A defensible report says which number came from which, how many samples it rests on, and what it cannot see.

ObservedEvery input is a sample, a floor or a subset. A report can still be honest, provided it says which of the three each number is.

What the report is for

An AI visibility report exists to support a decision: whether to keep investing in a body of content, where the exposure is concentrated, and whether something changed that anyone should act on. It is not a scoreboard, and treating it as one is what produces the monthly ritual of explaining a number that moved for reasons nobody can name.

Most of these reports fail for a reason that has nothing to do with data quality. They stack four different kinds of number in one chart — a count, a sample, a floor and a subset — and then draw a line through them. The count comes from a platform reporting its own output. The sample comes from a tool running prompts. The floor comes from your analytics, which sees only the visits that kept a referrer. The subset comes from Search Console's generative AI report, whose impressions already sit inside its Web total. Each is legitimate. Averaging them is not.

The work of building a defensible report is mostly the work of keeping them apart, labeling each one, and saying out loud what the report cannot see. That is less satisfying than a single index and it is the version that holds up when somebody arrives with a vendor chart that disagrees.

Keep the four numbers apart on the page

Citation, mention, impression and referral are four measurements, not four views of one. They are produced by different instruments, they move independently, and they come apart in ways that are named and documented.

  • Cited but not mentioned. Seer Interactive named this the ghost citation in March 2026: the answer uses your page as a source without ever speaking your brand name.
  • Mentioned but not cited. The answer recommends you and links somewhere else, which produces no citation, no referral, and possibly a customer.
  • Impression without a click. The ordinary case, and the reason an impressions-only report cannot be read as a performance report.
  • Referral without any of the above being visible to you. Clicks from Google's AI surfaces arrive with a google.com referrer that your analytics cannot distinguish from organic.

Semrush's June 2026 visibility index keeps mentions and citations deliberately separate, defining mentions as how often a company appears in an answer and citations as which domains and pages the platforms use as evidence, and reports the two sets diverging sharply on Gemini. Any report that blends them into one figure has thrown away the only information that tells a reader what to do next: a mention problem lives largely off your own site, and a citation problem lives on it.

A sampling design that can carry a claim

This is the section most reports skip and the one that determines whether anything else in them means anything. The object being measured is generated rather than looked up, and the instability has been quantified by two vendors with no reason to advertise it.

SE Ranking parsed the same 10,000 keywords three times on a single day in AI Mode, with data collected 20 June 2025 and published 29 August 2025: 9.2% of cited URLs appeared in all three runs, domain-level overlap across all three was 14.7%, and any two runs overlapped at roughly 18.5% to 19% on URLs. Ahrefs, on 540,000 query pairs of September 2025 US data published 15 December 2025, found that 45% of AI Overview citations change between generations (Ahrefs). Both were funded by the vendors that ran them; neither has been contradicted.

A single measurement per prompt is therefore one draw from a distribution whose variance you did not measure, and any vendor reporting a change from single-sample data is making a category error rather than a rounding error. A design that survives that looks like this: freeze the prompt or keyword set for the whole reporting period; run each prompt several times inside every collection window; report the share of runs in which you appeared rather than a yes or no; hold country, device and account state constant and name them; collect on a stated schedule rather than opportunistically; and print the run count beside every rate. No platform publishes variance figures for its own outputs, which means your confidence interval can only come from your own repeated runs. That is the argument for running them.

Say which numbers are floors

Some figures in the report are complete for what they cover, and some are lower bounds of unknown tightness. A reader cannot tell them apart from the chart, so the report has to say.

Your AI referral segment is a floor for three separate reasons, and all three belong in a footnote wherever the number appears. It cannot see the most common outcome of an AI answer, the brand named without a link, because there is no referral event to capture. It loses visits from native apps that send no referrer, from restrictive referrer policies, and from copied and pasted URLs. And it excludes Google's own AI surfaces entirely, because their clicks carry a google.com referrer, which for most sites means the segment omits the majority of the AI exposure it claims to measure. It will also break without warning: every host in it except ChatGPT's UTM parameter is an undocumented convention that the operator never promised.

Search Console's impressions are a different animal — a count of Google's own counting, for your property, not a sample. They still under-register in one specific way. AI Overview impressions register when a citation is scrolled or expanded into view, so citations buried behind expansion controls are systematically undercounted, and a third-party tracker that expands the block will always find more citations than Google records impressions. That divergence is expected and should be explained once in the report rather than investigated every month.

The comparisons that are not comparisons

Most of the damage in visibility reporting is done by charts that look like comparisons and are not. These are the ones to refuse by policy, so that the argument happens once rather than every quarter.

  • One vendor's score against another's. Different prompt sets, surfaces, sampling depths, numerators and denominators. There is no conversion factor between them.
  • This month against last month on a changed prompt set. The most common way a visibility chart misleads without anyone intending it to. Ask in writing whether the prompt set changed between the two dates being plotted.
  • Generative AI impressions against Web impressions as an AI share. The first is a subset already inside the second.
  • A tracker's citation count against Search Console impressions. Different events by construction, for the scroll-and-expand reason above.
  • Your presence rate against a published one. Semrush's ten-million-keyword panel recorded AI Overview presence at 6.49% in January 2025, 24.61% in July 2025 and 15.69% in November 2025. One panel, one year, three numbers that have been used to argue opposite conclusions. A presence rate without its month, panel and country is not a benchmark.

The methodology box

Every recurring report should carry a short, fixed block stating how the numbers were made. It takes a paragraph, it shortens the rest of the report, and it is the only defense available when someone brings a figure that disagrees with yours.

It should name the instrument behind each metric, the collection window, the size of the prompt or keyword set and whether it changed since the last edition, the number of runs per prompt, the surfaces covered and whether the tool queried a consumer product or an API, the country and device profile, and the known blind spot for each number. Where a figure comes from a study rather than from your own collection, name the study, its date, and who funded it — Pew Research Center's July 2025 click study is the only non-vendor-funded piece of evidence in this whole area, and every other number in circulation was produced by a company selling something to the people quoting it. That is not disqualifying and it is disclosable.

A methodology box also does something less obvious. It makes the report falsifiable, which is what turns it from an assertion into evidence, and it forces the person assembling it to notice when a number cannot be sourced. Numbers that cannot survive being described usually should not have been on the page.

What to leave off

Several standard report elements are worse than nothing, because they carry an authority the underlying data cannot support.

  • A single composite visibility score. No platform defines one; vendors define their own from different inputs at different sampling rates, and the index is not comparable to any other or to itself across a prompt-set change.
  • Citation position, or any share weighted by it. Google states that all links in an AI Overview are assigned the same position; Microsoft states that its citation data does not indicate ranking, authority, or the role of any page in an answer. A position-weighted score contradicts both.
  • An AI Overview click-through rate. Impressions with no clicks on one side, and clicks pooled into an unfilterable Web total on the other.
  • A growth multiple without its base. Sixteenfold growth in AI referral traffic and a 0.32% share of all website traffic describe the same finding (SE Ranking, 101,574 sites, published 18 June 2026). Either half alone misleads, in opposite directions.
  • A value multiplier presented as a conversion rate. Semrush's figure of 4.4 times the value of an organic visit is a projection over 500-plus digital marketing and SEO topics, and Semrush's own note calls its traffic and value projections extrapolations of historical data and adoption rates.
  • A week-on-week change without an error bar. In this subject that is not conservative reporting; it is a claim the instrument cannot support.

What the first page should say

Strip the report to the statements the evidence can actually carry, and there are about five of them. They are less exciting than an index and they are defensible.

Say whether you appeared at all, and in what share of sampled runs, on a frozen prompt set with the run count printed. Say whether the direction has held across enough consecutive windows to be distinguishable from regeneration noise, and say plainly when it has not. Say where the exposure is concentrated, remembering that AI referral traffic lands on homepages in about 60% of cases against 17% for organic search, so page-level analysis of arrivals will tell you far less than page-level analysis of citations. Say what was never fetched, which is the one diagnosis that separates a retrieval failure from a selection failure and the one that produces immediate work. And say what would have to change for the conclusion to change — a rate that moves outside the range your own repeated runs produce, a prompt set that has to be revised, a platform that starts reporting clicks.

That last sentence is the one almost no report contains, and it is the one that makes the rest of it credible. A report that states in advance what would falsify it is doing measurement. A report that only ever confirms the direction of the previous one is doing decoration.

Frequently asked questions

How many runs per prompt is enough?

There is no published standard to copy, because no platform publishes variance figures for its own outputs. The workable test is empirical: run enough times that your reported rate is stable across consecutive collection windows on an unchanged prompt set, then hold that depth constant and print it. Given 9.2% three-way URL reproducibility on identical same-day queries, one run per prompt is not a measurement at any frequency.

Can the report include a share-of-voice number?

Only with its denominator named. Microsoft's Citation Share is the one platform-defined version, computed as your citations divided by all citations shown for the same grounding query, and Microsoft states it exposes no competitor domains and is not a competitive scoreboard. Every other share figure on the market is a share of a prompt list the vendor wrote, which is a legitimate thing to report as long as the report says so.

Should AI referral traffic be reported as revenue?

It can be, provided it is labeled a floor. The engagement direction is reasonably well supported — AI-referred visitors spent 67.7% longer on site than organic visitors in the largest cross-site dataset available — but the referral count itself misses mention-without-link entirely and excludes Google's AI surfaces by construction, and no study has ever measured the size of that gap.

A vendor chart disagrees with mine. Who is wrong?

Quite possibly neither. Ask four questions: which prompt set, how many runs per prompt, which surfaces, and which dates. Two tools can be internally consistent and differ by a large multiple on the same brand in the same month, because surface selection alone moves the number — AI Overviews and AI Mode share only about 13.7% of their citations.

How often should this report run?

Monthly, on a frozen prompt set, is the shortest interval that usually produces signal. Weekly reporting in this subject mostly reproduces regeneration noise, and it trains an audience to react to movement that no action caused. Where a platform report gives you a count rather than a sample, a shorter cadence is defensible because the variance being reported is Google's, not your sampler's.

What should the report do when a platform changes its reporting?

Break the series and say so. Google's generative AI performance report began collecting on 18 May 2026 and rolled out to a subset of properties first, so early data is not comparable across sites or backdated. A visible break with a dated note is honest; a continuous line drawn across a definitional change is not.

Top