AI visibility tracking is now a standing agenda item in most B2B marketing reviews we join, and it usually starts from the same uncomfortable place. The company publishes expert content, invests in technical SEO, and still cannot answer a simple question from the board: when a buyer asks an assistant who can build their B2B portal, does our name appear at all?
The gap shows up in a very specific way. A sales manager writes in the CRM that the prospect "found us somewhere in ChatGPT", web analytics reports the session as Direct, and the marketing dashboard shows nothing that connects the two. Nobody is lying. The data simply was never designed to carry that link.
This article is the measurement system we use in practice with mid-market B2B clients: working definitions for each metric, formulas with explicit denominators, a registry of buyer questions, an observation protocol that survives platform changes, the real boundaries of GA4 and Search Console, CRM attribution rules, and a first-month pilot plan.
One frame first, because it saves a lot of arguments later. This is a measurement system, not a growth guarantee. Monitoring tells you what happened on a sample of questions you chose, on platforms you checked, in markets you defined. It does not measure total brand exposure, it does not prove causality, and it does not promise that your company will be recommended.
What AI Visibility Means for a B2B Company
AI visibility is the observed presence of your company in answers to a defined set of buyer questions, on defined platforms, in a defined market and language, during a defined period. Every word in that definition is load-bearing: remove the sample, the platform or the market, and the number stops meaning anything you can act on.
That definition rules out the phrasing most vendors and most internal decks use. "Our AI visibility is 34%" is not a statement about your brand. It is a statement about 34% of some sample, and unless the sample is published next to the number, the reader has no way to judge it.
It also rules out the reverse mistake. If a check could not be completed, or the AI Overviews block did not appear, that is an unavailable observation. It is not evidence that the brand is absent. We treat those two states as different data categories from the very first spreadsheet, because merging them corrupts every period comparison that follows.
There is a research background to this. The GEO paper by Aggarwal and co-authors, accepted to KDD 2024, formalized visibility in generative answers as something you can define and measure rather than guess at, and reported improvements of up to 40% in visibility metrics under their experimental setup (GEO: Generative Engine Optimization{rel="nofollow noopener noreferrer"}). That is an academic result in a controlled setting. It is not a forecast for your pipeline, and we never present it as one to a client.
Brand Mentions, Recommendations and Website Citations
Three events get collapsed into one word in most reports, and they carry completely different commercial weight.
- Mention. The brand name appears in the answer in any role, including a neutral alphabetical list of vendors or a passing reference in a paragraph about the market.
- Recommendation. The model explicitly proposes the company as a fit for the buyer's stated task: "for a manufacturing customer portal with ERP integration, consider X".
- Citation. The answer links to a domain you track. Split this into your own domains and third-party pages that discuss you, because they are earned in completely different ways.
The practical consequence is the one that surprises digital marketing teams most. A company can be cited without being recommended, and recommended without being cited.
We see this pattern in our own content regularly. An article we published on connecting a corporate site to CRM and ERP gets used as the source for the explanatory part of an answer, and then the same answer lists three other vendors as candidates for the work. Our page supplied the reasoning; someone else got the shortlist slot. If the report shows only "we were present", that answer looks like a win. It is not. It is a citation with a missed recommendation, and the fix is a different piece of work entirely.
Which Buyer Decisions the Measurement Should Support
Metrics earn their place by changing a decision. In B2B, three decisions actually matter, and every observation you record should map to one of them.
| Buyer decision | What the buyer is trying to settle | Example question we track |
|---|---|---|
| Understand the problem | Whether there is a problem worth solving and what it is called | "How do we connect a corporate site to CRM and ERP?" |
| Compare approaches | Which architecture or delivery model fits | "Build vs buy for a B2B customer portal?" |
| Shortlist a supplier | Who can credibly do the work | "How do I compare B2B portal development contractors?" |
The buying committee complicates this in a useful way. A CTO, a head of operations and a CFO ask different questions about the same purchase, in different vocabulary, and an assistant answers each of them separately. A registry that contains only the CTO's questions measures one third of the decision.
Our rule of thumb: if a metric cannot change a marketing action or a sales action, it does not belong in the report. Share of voice on questions nobody asks before buying is a decoration, not a KPI.
Which AI Visibility Metrics to Track
Five metrics cover what a B2B company actually needs: mention rate, domain citation rate, recommendation rate, AI share of voice and description accuracy. Each one needs an explicit denominator printed next to it, and each one has a specific way of misleading you when the denominator is hidden.
Fix the notation before writing a single formula. These are this article's working definitions, not an industry standard, and no industry standard currently exists.
N = number of valid answers in a defined slice
M = number of valid answers containing the brand
C = number of valid answers citing a tracked domain
R = number of valid answers with an explicit recommendation of the company
S = number of valid answers to supplier-selection questions
T = total brand presences across the fixed competitive set in the slice
A "slice" means one combination of platform, interface, market, language and period. Aggregate across slices only when you have decided, in writing, that the slices are comparable.
Mention Rate, Citation Rate and Recommendation Rate
Three ratios, three different denominators, and the third one is where most reports go wrong.
| Metric | Formula | Denominator | What it answers | How it misleads |
|---|---|---|---|---|
| Mention rate | M / N x 100% |
Valid answers in the slice | How often the brand appears at all | Counts neutral list entries as success |
| Domain citation rate | C / N x 100% |
Valid answers in the slice | How often a tracked domain is used as a source | Mixes own domains with third-party pages if not split |
| Recommendation rate | R / S x 100% |
Valid answers to supplier-selection questions only | How often the brand is proposed as a fit | Collapses toward zero if computed over all questions |
| Check success rate | N / A x 100% |
All attempted checks A | Data quality of the slice | Ignored, which hides a broken sample |
Recommendation rate uses S, not N, on purpose. A recommendation is only a meaningful event where the buyer asked for a supplier. Dividing recommendations by all answers, including "what is a customer portal", produces a number that drifts every time you add explanatory questions to the registry.
Failed checks never disappear. They are excluded from valid answers and reported separately as a data-quality line. A slice with a 96% check success rate and a slice with a 61% check success rate are not comparable, no matter how similar their mention rates look.
AI Overviews needs three denominators of its own, because the block itself is a variable:
- Block appearance share = queries with a block / successfully checked queries
- Brand share within appearing blocks = blocks containing the brand / blocks that appeared
- Brand presence across the sample = queries with both block and brand / successfully checked queries
If no block appears in a slice, the second metric is undefined. Reporting it as 0% is a factual error that has cost more than one team a pointless quarter of "recovery" work.
AI Share of Voice and Description Accuracy
AI share of voice compares you to a fixed competitive set, and it is only readable when that set is frozen and published.
AI Share of Voice = M / T x 100%
Rules that make the number honest: each brand is counted once per answer regardless of how many times it is named, your own brand is part of the competitive set, and when T = 0 the metric is undefined rather than zero. Changing the competitor list mid-quarter changes T and therefore changes your share of voice without anything happening in the market.
Hypothetical example (teaching numbers, not research data). Take a slice of 120 valid answers. The brand appears in 24 of them, so the mention rate is 24 / 120 = 20%. Across the same 120 answers, all brands in the fixed competitive set together produce 80 presences, counted once per brand per answer. AI share of voice is then 24 / 80 = 30%. These figures illustrate the arithmetic and the denominators only. They are not a benchmark, not a target, and not derived from any study.
Description accuracy is the qualitative metric, and in vendor-selection contexts it can matter more than frequency. Score each branded answer on a short fixed scale, for example:
- Accurate - services, markets, industries and evidence of experience are stated correctly.
- Partially accurate - correct core, with one material error or omission.
- Inaccurate - wrong services, wrong market, wrong company, or invented claims.
An inaccurate mention can be worse than no mention. If an answer positions an ERP and B2B platform company as a small web development studio, that answer actively removes you from a shortlist you would otherwise have joined. Frequency metrics cannot see this. A scored sample of branded answers can.
Why Visibility, Traffic and Revenue Need Separate Measures
Three populations, three data sources, three reports: sampled answers from monitoring, sessions from web analytics, records from CRM. They describe different things and are collected under different rules.
The most common arithmetic error we are asked to review is a homemade CTR: site visits divided by mentions in test answers. Those numbers come from different populations. Test answers are a sample you generated; visits come from real users you did not sample. The ratio has no defined meaning, and it always looks plausible enough to reach a board slide.
There is also no universal benchmark for "good visibility". Nobody can tell you what mention rate a mid-market B2B software company should have in the German market on solution-comparison questions, because no comparable public dataset exists. The only honest comparisons are against your own baseline and against your fixed competitive set, under an unchanged protocol.
Methodology drift is a first-class risk and belongs in the risk register, not in a footnote. Changing questions, competitors, platforms, models or repeat counts breaks period comparison. The report header must carry the registry version, and any change gets a dated note.
How to Build a Representative Set of Buyer Questions
The question registry is the measurement instrument. If it is weak, every downstream number is decoration, no matter how sophisticated the tool that produced it.
Build the registry from what buyers actually say. Keyword tools are a supplement here, not the source, because people phrase things differently to an assistant than they type into a search box. "b2b portal development company" is a search query. "We manufacture industrial fittings and our dealers order by email, what should we build?" is a prompt. The split is the same in any vertical: "real estate website development" is a query, while "our agents lose listings because the site cannot filter by district" is a prompt.
Freeze the registry as a baseline, version it, and mark any change in the report. Keep a shared cross-market core so markets remain comparable, plus a local extension per market for questions that only exist there.
Sources of Questions: Sales, CRM and Search Data
Sales conversations are the primary source of real wording, and they are usually sitting unused.
- Sales call notes and objection lists. The exact sentence a prospect used to describe their problem is worth more than ten keyword variants.
- CRM records. First-touch descriptions, RFP texts, qualification questions and lost-deal reasons show what buyers were comparing when they chose someone else.
- Search data. Search Console queries and keyword candidates, each one checked against a simple test: would a human actually phrase it this way to an assistant?
- Support and delivery teams. The questions clients ask in month three often predict what the next prospect asks in week one.
Our practical filter is one line: a question enters the registry only if a real prospect could plausibly ask it before choosing a vendor. Questions that pass this filter tend to be longer, more situational and much less elegant than keyword research output. That is the point.
Problem Research, Solution Comparison and Supplier Selection
Balance the three stages deliberately. A registry made only of supplier-selection questions overstates commercial visibility, because those are the questions where brand names naturally appear. A registry made only of explanatory questions understates it.
Stage 1, problem research. "How do we connect a corporate site to CRM and ERP?" "When does a manufacturer need a customer portal?" "What breaks when dealer orders are handled by email?"
Stage 2, solution comparison. "Build vs buy for a B2B customer portal." "What integration approach fits a legacy ERP with no modern API?" "Headless commerce or a custom portal for dealer ordering?"
Stage 3, supplier selection. "How do I compare B2B portal development contractors?" "What should a mid-market company ask a software vendor before signing?" "Does company X have experience integrating our system?"
The mix should reflect where your commercial risk sits. If your problem is that buyers never reach the shortlist stage aware of you, stage 1 and stage 2 questions carry the diagnostic value.
Branded Questions, Non-Branded Questions and Wording Variants
Branded questions test description accuracy. Non-branded questions test discoverability. You need both, and you should never blend them into a single headline metric.
Wording variants matter more than teams expect. A preprint on paraphrase brittleness in retrieval-augmented commercial recommendation reports that recommendation outputs shifted noticeably when the same request was paraphrased, within the conditions the authors studied (Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation{rel="nofollow noopener noreferrer"}). It is a 2026 preprint by authors affiliated with Unusual, and we cite it narrowly: it supports testing several phrasings per question, and it does not generalize to every platform or interface.
Keep the variant count per question consistent. If one question is checked with five phrasings and another with one, the heavily checked question silently gains weight in every aggregate.
Here is the registry shape we use. Publish it, version it, and keep it in a place sales can read.
| ID | Buyer task | Role | Stage | Market | Language | Question | Variants | Priority |
|---|---|---|---|---|---|---|---|---|
| Q-014 | Connect website to CRM and ERP | Head of IT | Problem research | US | EN | "How do we connect a corporate website to our CRM and ERP?" | 3 | High |
| Q-021 | Decide on a customer portal | COO | Problem research | DE | DE | "Wann braucht ein Mittelstaendler ein Kundenportal?" | 3 | High |
| Q-033 | Compare delivery approaches | CTO | Solution comparison | US | EN | "Build vs buy for a B2B dealer portal?" | 3 | Medium |
| Q-047 | Shortlist a contractor | CEO | Supplier selection | US | EN | "How do I compare B2B portal development contractors?" | 3 | High |
| Q-052 | Verify a specific vendor | Head of IT | Supplier selection | DE | DE | "Welche Erfahrung hat Webdelo mit ERP-Integrationen?" | 2 | Medium |
How to Run Comparable Checks Across Platforms and Markets
Comparability is a protocol problem, not a tooling problem. Two teams using the same monitoring service and different protocols will produce numbers that cannot be compared, and the tool will not warn them.
The rule that saves the most rework: never merge API results with user-interface results into one metric. They are different products with different retrieval behaviour, different defaults and different availability. Track them as separate slices, always.
Verify interface and regional availability before launch, not after the first anomaly. A market where a given mode is unavailable produces a stream of unsupported conditions, and if you discover that in week six you will have to discard weeks one to five.
Recording Interface, Language, Market, Date and Conversation Context
Log every observation with the same fields, every time. The list below is the minimum we implement for clients.
| Field | Why it exists |
|---|---|
| Question ID | Links the observation to the registry version |
| Exact wording used | The variant actually submitted, not the canonical question |
| Date and time | Answers change; undated observations are unusable |
| Platform | ChatGPT, Gemini, Google AI Overviews, others |
| Interface | Web UI, mobile app, API |
| Mode or model | Named model or mode where visible |
| Language | Language of the prompt |
| Regional settings available | What location or region controls were actually applied |
| Conversation context | Fresh conversation by default; otherwise describe the carried context |
| Answer text | Full raw answer, stored verbatim |
| Links | Every cited URL, with domain classification |
| Check status | Valid / block absent / check failed / condition unsupported |
Store the raw answer, not only the parsed verdict. When you refine the definition of "recommendation" in month four, raw answers let you recompute the whole history. Parsed verdicts do not.
Keep market, language and user location as three distinct dimensions. A German-language check is not automatically a German-market check, and a check run from one location is not a statement about buyers elsewhere. Collapsing these three into "DE" is the single most common reason cross-market reports stop making sense. Location matters most for services bought locally: a dental SEO company is recommended in one city and absent in the next, and a report without the location field will call that noise.
Fresh conversation by default. If context is deliberately carried over, that is a separate experimental condition and it gets its own slice.
Repeat Checks, Prompt Variants and a Fixed Baseline
Fix the repeat count per question and keep it stable across periods. Repeats reduce part of the random variation between runs. They do not make your sample representative of the whole audience, and no amount of repetition will.
Aggregate by question first, then across questions. Compute a per-question mention rate over its repeats and variants, then average across questions. Pooling all answers into one bucket lets frequently checked questions dominate the result for reasons that have nothing to do with the market.
Define the baseline period explicitly and freeze four things with it: the question registry version, the competitor set, the platform and interface list, and the repeat count. Write these into the report header so that anyone reading month six can see whether they are looking at the same instrument as month one.
Any protocol change gets a dated note in the report header. One deliberate change per period is the discipline that makes the next comparison readable.
Missing AI Overviews, Failed Checks and Unsupported Conditions
Three statuses, three different meanings, and none of them is zero visibility.
- Block absent. The check ran successfully and no AI Overviews block appeared for that query. This is a valid observation about the SERP. It enters the block appearance share denominator and leaves the "brand share within blocks" metric untouched.
- Check failed. A collection error: timeout, rate limit, parsing failure, session problem. This is not an observation at all. It is excluded from valid answers and reported in the check success rate.
- Condition unsupported. The platform, mode or regional setting is not available in that market or interface. This is a scope limitation. Report the slice as not covered, and say so explicitly rather than leaving a blank cell that readers will interpret as zero.
Publish the check success rate next to the visibility numbers as a first-class data-quality metric. Set a minimum threshold below which the slice is reported as insufficient data rather than as a number. Where that threshold sits is a judgement call, but it must be decided before you see the results, not after.
We have watched a "sharp visibility drop" turn out to be a collection failure three times in the last year. Every time, the fix was in the log, and every time the team had already spent a week discussing content strategy.
Comparing Russian, German and English Results
Literal translation of questions is not enough. Match the buyer task first, then adapt the phrasing so it sounds like something a native speaker would actually type.
German market notes. Use the terms buyers use: KI-Sichtbarkeit for AI visibility, Nachvollziehbarkeit for the traceability of reporting, Mittelstand for the mid-market segment, qualifizierte Anfragen for qualified inquiries. German B2B buyers ask more procedural questions earlier in the cycle, and the registry should carry that.
US market notes. Vendor evaluation, qualified leads, pipeline, buying committee. Supplier-selection questions arrive earlier and more directly.
Russian-language notes. English terms often need explanation in the answer itself, which changes the shape of what the model returns. Keep the reader's language, the location where the check was run and the target market of the buying company as three separate fields.
Report each market separately. A blended cross-market average hides exactly the decisions the report is supposed to support, and there is no reasonable way to weight markets that differ in deal size, cycle length and platform availability.
How to Measure AI Traffic in GA4 and Search Console
Traffic measurement moves you from sampled answers to observed visits, which is a completely different population with completely different limits. Monitoring told you what a model said to you; analytics tells you what real users did, minus everything the browser and the assistant app did not pass along.
Google Analytics documents an AI Assistant channel in its default channel group definitions (Google Analytics: Default channel group{rel="nofollow noopener noreferrer"}). Traffic arriving from AI Overviews and AI Mode inside Google Search is classified under Organic Search, because it is search traffic: it lands in the same reports as your SEO traffic.
On the Search Console side, Google states that results from AI features are included in the general Web search type report (Google Search Central: AI features and your website{rel="nofollow noopener noreferrer"}). There is no separate AI Overviews traffic report to open. Any dashboard that claims to show your exact AI Overviews click volume as a distinct line is presenting an estimate, and it should be labelled as one.
Validate with a controlled test visit and a controlled form submission before trusting anything. Every time we skip this step, we find something broken later, and it is always more expensive to find later.
The AI Assistant Channel and Session Sources
Check the metric scope every time you read a number: user, session or event. A "sessions" chart and a "users" chart of the same channel tell different stories, and reports that mix them silently are common.
Read session source and medium alongside the channel grouping rather than trusting the channel label alone. Channel groupings are rule-based classifications that evolve; the underlying source and medium values are what you can audit and, if necessary, reclassify.
Referrer loss is normal in assistant apps. Depending on the client, the platform and how the link is opened, part of genuine AI-originated traffic arrives with no referrer and lands in Direct. Expect this, document it, and stop treating Direct as a residual bucket you can ignore.
The corresponding discipline: do not attribute all Direct traffic and all branded-search growth to AI visibility work. Both move for many reasons, including offline activity, events, PR and simple seasonality. If a report needs a causal claim, it needs an experiment, not a channel chart.
AI Overviews and the Limits of Organic Search Reporting
State the boundary from Google's own documentation and stop there. AI feature results are folded into the Web search type report; impressions and clicks from those features are counted within it rather than broken out as a separate AI surface.
What you can observe:
- Landing-page level performance over time
- Query-level trends inside the Web report
- Impression and click shifts on pages you deliberately changed
- Position and coverage changes for the queries in your registry
What you cannot observe:
- The exact user prompt behind a single visit
- Whether a specific session came from an AI Overviews click or a classic blue link
- Whether a user read an AI answer about you and never clicked at all
The stance we take with clients is deliberately narrow: track direction and landing-page patterns, and refuse to fabricate per-visit precision. A marketing lead who reports "we cannot separate this, here is what we can see instead" keeps credibility. A dashboard that invents the separation loses it at the first audit.
Landing Pages, Key Events and Lost Attribution
Landing-page reporting is the most stable signal available, because the landing page is recorded on your own property regardless of what the referrer did.
Instrument these key events at minimum:
- Form submission, with form identity and page context
- Document download, such as a capability deck or technical brief
- Meeting booking, including the scheduler flow if it sits on another domain
- Pricing or engagement-model page depth, as an intent signal
- Contact reveal actions, such as email or phone clicks
The engineering part is where this usually fails in practice, and it is the part that gets skipped. In the audits we run, and in the tracking section of every SEO site audit, the recurring defects are the same short list: event parameters that were renamed but not updated in the reporting layer, consent configuration that drops events for entire regions, cross-domain forms and schedulers that start a new session and lose the original source, and redirect chains that strip the referrer before the first page view fires.
Run a data validation checklist before the first management report:
- Complete a control visit from an assistant interface and find it in the raw event stream.
- Submit a control form and confirm the key event fires with the expected parameters.
- Confirm the landing page, source, medium and campaign values recorded on that session.
- Verify cross-domain continuity end to end, including any external scheduler.
- Check consent behaviour for each market you report on.
- Confirm the CRM record created by the control submission carries the acquisition fields.
If any step fails, fix it before publishing numbers. Publishing first and correcting later teaches management to distrust the dashboard permanently.
How to Connect AI Interactions to B2B Leads and Deals
Three classes of evidence must stay separate in the CRM: an observed visit, a known assisted interaction, and a customer self-report. They differ in strength, they are collected differently, and merging them produces a number that cannot be defended in front of a commercial director.
Keep them separate and the report becomes readable: "requests with an observed AI-originated visit", "requests with a documented AI touch reported by sales", "requests where the client self-reported an AI assistant". Merge them and you get "AI leads", which nobody can audit.
Two hard rules before any of this is built. Never push customer contact data into web analytics to stitch reports together; join records through internal identifiers instead. And never add open pipeline value to won revenue in a single figure - they are different states of a different thing.
Passing Available Acquisition Data to CRM
Pass only what is genuinely available, and pass it raw.
| CRM field | Source | Notes |
|---|---|---|
| Request ID | Web form / CRM | Primary key for the inquiry |
| Account | CRM | Company-level record, not the person |
| Channel | Analytics layer | Stored as captured, normalized later |
| Source and medium | Analytics layer | Raw values, including empty |
| Landing page | Web form hidden field | First page of the session |
| First-touch date | Analytics layer | Where a first-touch value exists |
| Campaign | UTM parameters | Present only when tagged |
| Touch history available | Analytics layer | Explicit yes / no / partial |
| Self-reported discovery | Form field | Free text or fixed list plus "other" |
| Qualification status | Sales | Against written criteria |
| Deal | CRM | Linked opportunity, if any |
Store raw values and normalize in the reporting layer, not at capture time. Normalizing at capture destroys the ability to re-classify later, and classification rules always change.
Handle the offline path explicitly. Phone calls, email replies and referrals still need a source field, and that field needs an explicit "unknown" option that sales is allowed to choose. An "unknown" rate of 30% is information. A forced guess is noise that looks like information.
This is where the engineering work actually earns its place, and it is unglamorous: passing hidden fields through cross-domain forms, normalizing values, verifying that events fire with the right parameters, deduplicating records, wiring CRM data into the reporting layer, and keeping all of it working after the next website release. Most measurement projects we are called into do not fail on strategy. They fail because the form on the pricing page never passed the landing page to the CRM.
Observed Visits, Assisted Interactions and Self-Reported Discovery
Add one well-worded open question to the inquiry form: "How did you first hear about us?" Treat the answer as evidence, not proof.
Self-reports skew toward the last memorable touch. A buyer who read three of your articles over four months and then asked an assistant a final comparison question will often name the assistant, because that is what they remember. Report self-reports as their own metric with their own denominator, and never upgrade a self-report into an observed visit in aggregate reporting.
Assisted interactions are the middle class of evidence: a documented AI touch mentioned during a call and logged by sales. "They said our name came up when they asked Gemini about portal vendors" is a real data point if it is recorded in a structured field with a date. It is not a session, and it should never appear in a sessions chart.
Give sales a two-second way to log it. A required picklist on the first call - observed / mentioned by client / unknown - produces more usable data than a free-text field nobody fills in.
Account-Level Deduplication and Long Sales Cycles
Deduplicate at three levels: lead, account and deal. In B2B, several members of the buying committee arrive separately, often weeks apart, often through different channels, and each of them creates a record.
Without account-level deduplication, a single opportunity with four stakeholders appears as four leads. Multiply that across a quarter and the lead count is inflated by a factor nobody can reconstruct afterwards. Repeat inbound requests from the same account have the same effect and must be collapsed under a documented rule.
Long cycles change what month one can deliver. If the average cycle from first touch to signature is five to nine months, the first month of monitoring gives you a baseline and a working process. It does not give deal outcomes, and promising deal outcomes at that horizon is how measurement programmes get cancelled in month three.
Choose a reporting horizon that matches the actual sales cycle, and state that horizon in the report itself. Visibility and traffic can be reported monthly. Commercial outcomes need a window long enough to contain a real deal.
Attributed Pipeline, Won Revenue and Causal Limits
Report attributed pipeline and won revenue separately, each tagged with its evidence class. Open pipeline is a forecast about the future; won revenue is a fact about the past. Adding them produces a number that means nothing and always flatters the programme.
Attribution shows association under known observation rules. It does not prove incremental sales. A deal where an AI touch was observed might have closed anyway through a referral you did not record.
The phrasing we recommend to management is deliberately plain: "requests where an AI touch was observed or reported", not "leads generated by AI". The first sentence is defensible in any audit. The second one is not, and the moment a CFO tests it, the whole report loses standing.
When a genuine causal claim is required - for a budget decision, for example - it requires an experiment design: a controlled change, a comparison group or a pre-registered before-and-after with everything else held constant. That is a different project from attribution reporting, and it should be scoped as one.
How to Choose an AI Visibility Monitoring Tool
Two different products are sold under one name, and choosing the wrong one wastes a year. The first is a monitor for your own buyer questions, matching the registry approach in this article. The second is a broad visibility index that estimates presence across a large general question set.
Before comparing features, compare definitions. Ask the vendor what they call a region, a citation and visibility. In our experience these three words mean different things in almost every product, and a feature comparison built on mismatched definitions is worthless.
Ask specifically what happens to failed checks and missing AI Overviews inside the vendor's math. If a failed check is counted as a non-mention, the tool will understate your visibility every time its collection degrades, and you will read that as a market change.
Keep a manual control sample even after buying a tool. Twenty questions checked by hand each month, under your own protocol, tells you whether the vendor's numbers still behave.
Manual Monitoring, Custom Buyer Questions and Broad Visibility Indexes
- Manual checks. Cheapest start, full control of conditions, complete access to raw answers, poor scalability. Correct choice for a pilot and for a permanent control sample.
- Custom question monitoring. Automates your own registry across platforms and markets. Matches the method described here and produces numbers you can defend.
- Broad visibility indexes. Useful for market context and competitor movement at a category level. Weak for your specific buying decisions, because the underlying question set is not yours.
For terminology and product framing, Ahrefs Brand Radar, Peec AI and SEOWORK are reasonable reference points to look at while you build your requirements. We name them as examples of how the category is being described, explicitly not as verified market leaders or as a ranking, and we have not run a comparative evaluation of their methodologies.
Collection Methods, Market Coverage and Metric Definitions
The checklist below is what we hand to clients before demo calls.
- Collection method. API or user interface? Both? Are they reported separately, or blended into one number?
- Market and language coverage. Can the vendor actually reproduce the market you sell into, including regional settings, or only the language?
- Model and mode coverage. Which platforms, which models, which modes, and how quickly do they adapt when an interface changes?
- Denominators. Are N, the check success rate and the sample composition exposed in the product, or only the final percentage?
- Failed-check handling. How are errors, absent blocks and unsupported conditions treated in the math?
- Repeat policy. Fixed repeats per question? Configurable? Consistent across periods?
- Competitive set control. Can you define and freeze your own competitor list?
A vendor that cannot answer the denominator question in a demo will not answer it in production either.
Raw Answers, Historical Data, Export and Integrations
Access to raw answers is the single most important requirement on the list. Without stored raw answers you cannot recompute anything when your definition of "recommendation" or "citation" changes, and it will change within the first two quarters.
- Historical retention. How long is history kept, and can it be recalculated after a definition change?
- Export. CSV and API access to observations, not just to charts.
- Integrations. Can the data reach your analytics and CRM reporting layer without manual copying?
- Ownership. If you leave the vendor, do you keep your baseline and your history?
That last question is the one most teams forget to ask, and it determines whether your first year of measurement is an asset or a rental.
How to Read the Report and Decide What to Change
One page for management, three blocks: visibility, traffic, requests and deals. Every number carries its denominator, its period, the registry version and the check success rate, or it does not go on the page.
Diagnostics are hypotheses to test, never proven causes. The report says what was observed; the discussion after it decides what to try next.
The first question after any sharp change is always the same: did the measurement change? Protocol edits, interface updates and collection failures produce movements that look exactly like performance movements, and they are far more common.
A Dashboard for Marketing and Sales
Header block. Registry version, competitor set, platforms and interfaces, repeat count per question, reporting period, check success rate per slice, change log entries for the period.
Marketing block. Mention rate, domain citation rate split by own and third-party domains, recommendation rate, AI share of voice, description accuracy distribution. Reported per market, per platform, never blended.
Traffic block. AI Assistant channel sessions, organic sessions on registry-relevant landing pages, landing-page table, key events, and a data-quality note listing known gaps such as consent effects or referrer loss.
Commercial block. Requests with an observed AI-originated visit, requests with a reported or assisted AI touch, qualified requests against written criteria, distinct accounts, attributed pipeline, won revenue. Each line tagged with its evidence class, pipeline and revenue never summed.
From Observations to Testable Hypotheses
| Observation | Hypothesis to test | First check |
|---|---|---|
| Few mentions across the registry | Pages do not answer the buyer tasks in the registry | Compare page content against the exact question wording; review which sources the answers do draw on |
| Citations without recommendations | Service descriptions and evidence of relevant experience are too thin to support a recommendation | Review service pages, case evidence, industry and technology specifics |
| Recommendations without visits | The answer satisfies the buyer without a click, or the brand is named without a link | Check landing-page trends and branded query trends rather than assuming loss |
| Visits without requests | Landing page or form is the bottleneck | Run the control submission; review page relevance to the originating question |
| Requests that do not qualify | Visibility grew on questions outside the real purchase decision | Review registry stage balance and qualification criteria |
| Sharp spike or drop | Measurement changed, not the market | Check protocol changes, interface availability, collection errors, check success rate |
Note the third row. Rising mentions with flat visits is not automatically a failure. It may mean answers now resolve the question inside the assistant. That is a real outcome with real limits on measurability, and the honest report says so.
Distinguishing Business Changes from Measurement Changes
A change log is a mandatory report component, not an optional appendix. Every registry edit, competitor set change, platform addition, repeat count change and tooling switch gets a date and an owner.
The operating rule is one deliberate change per period. Two changes in the same period make attribution of the result impossible, and the team ends up arguing from preference rather than data.
Separate three sources of movement before discussing performance: seasonality in your market, interface and model updates on the platform side, and your own registry or protocol edits. Only what remains after those three is a candidate for a real performance change.
When the sample is broken, escalate to "insufficient data" rather than reporting a confident number. It is an unpopular slide once and a credibility asset permanently.
How to Launch a Monitoring Pilot in the First Month
The goal of month one is a working process and a comparison point, not deal results. Anyone promising commercial outcomes in four weeks on a B2B cycle is describing a different business than yours.
Assign owners on day one: marketing owns the question registry, sales owns qualification and the assisted-touch field, and the IT partner owns data flow, integrations and validation. Unowned steps are the ones that quietly do not happen.
Keep the pilot small enough to actually finish. Two markets, two platforms, thirty to fifty questions and a fixed repeat count beats an ambitious plan that stalls in week three. Define what "done" means before starting. The same scope rule applies to a local business: an appliance repair SEO pilot runs on one city and twenty questions, and the discipline is identical, only the scale is smaller.
| Week | Focus | Output |
|---|---|---|
| 1 | Goals, markets, registry, competitor set, platform list | Versioned registry v1.0, frozen competitor set, documented protocol |
| 2 | Baseline observations and analytics verification | Baseline dataset with check success rates, validated GA4 configuration |
| 3 | CRM integration, self-report field, deduplication | Acquisition fields in CRM, control submission verified end to end |
| 4 | First report, hypotheses, next experiment | Management dashboard v1, prioritized hypothesis list, one planned change |
This is an example plan. The real timeline depends on the readiness of your analytics, your CRM and your web forms, and week three expands considerably when the forms are older than the CRM.
Goals, Question Registry and Baseline
Week 1 settles the scope: business goals for the programme, target markets, the buyer questions themselves, the competitor set and the platform list. Write down what decision each metric is supposed to inform, and drop the metrics that cannot answer that question.
Week 2 produces the baseline observations and verifies analytics in parallel. Run the full registry under the protocol, log every field, and record the check success rate per slice before looking at any visibility number.
Freeze the baseline and version it. Everything you report for the next two quarters compares against this, so the freeze needs a date, an owner and a stored copy of the registry.
Decide the reporting cadence and the audience for each report at the same time. Marketing may want it monthly; a commercial director usually wants a quarterly view aligned to the sales cycle.
Analytics, CRM Integration and Data Validation
Week 3 is the engineering week. Pass acquisition data to the CRM, add the self-report field to the inquiry form, and implement deduplication rules at lead and account level.
Run a control visit and a control form submission end to end, from an assistant interface through to the CRM record. Confirm that the landing page, source, medium and first-touch date survive every hop, including cross-domain schedulers.
Validate event parameters, consent handling and cross-domain flows for each market in scope. Consent behaviour differs enough between markets that a single validation pass is not sufficient.
Document known gaps instead of quietly filling them with assumptions. "Referrer is lost for assistant app traffic on iOS in this configuration" is a useful line in the report. A silently imputed value is not.
First Review, Ownership and Next Experiments
Week 4 produces the first report, a prioritized hypothesis list and exactly one deliberate change for the next period. Resist the urge to change five things because the first report looked uncomfortable.
Define who acts on each diagnostic signal. Low recommendation rate on supplier-selection questions is a content and evidence problem for marketing. Visits without requests is a landing page and form problem shared between marketing, engineering and the web design agency that owns the page. Off-target requests is a qualification conversation with sales.
Set the horizon for commercial evaluation according to the real sales cycle, and put that horizon in writing so nobody asks for deal attribution in week six.
Plan the next experiment as a testable statement: "Adding a technical integration evidence section to the customer portal page will raise the recommendation rate on Q-047 and Q-052 over the next two measurement cycles." That is testable. "Improve our GEO" is not.
Frequently Asked Questions
Can AI Visibility Be Measured for Free?
Yes, manually. A registry of twenty to forty buyer questions, a spreadsheet log with the fields listed above, a fixed protocol and a consistent repeat count will give you a defensible baseline without any paid tool. The output is genuinely comparable over time as long as the protocol does not change.
Cost appears with scale. Multiple markets, multiple languages, several platforms and interfaces, repeated checks and consistent scheduling turn into many hours per cycle very quickly. Most teams start manually, prove the method matters, and automate once the registry is stable and someone has to run it every week.
How Many Questions and Repeat Checks Are Needed?
There is no universal minimum, and any specific number offered without reference to your sample is guesswork. The honest answer is that you need enough questions and repeats to make your denominators stable, so that a small change in one answer does not swing the reported percentage.
Practically: keep the repeat count identical for every question, keep the registry frozen between comparisons, and publish N with every metric. If a slice has fewer than a couple of dozen valid answers, report it as directional rather than as a percentage. Growing the sample later is fine; changing it silently is not.
Why Do Monitoring Tools Disagree?
Because they measure different things. Collection method, regional configuration, model and mode versions, check dates, repeat counts, and the definitions of "citation" and "visibility" all differ between products, and each of those differences alone can move a headline number by a wide margin.
Disagreement between tools is expected and is not evidence that one of them is broken. What is comparable is a within-tool trend measured under a fixed protocol. Comparing your Tool A number against a competitor's Tool B number, or against a public index, produces a conversation with no defensible conclusion.
Can Every AI Overviews Visit Be Identified?
No. Google includes results from AI features in the general Web search type report in Search Console, so there is no separate AI Overviews traffic report to isolate (Google Search Central: AI features and your website{rel="nofollow noopener noreferrer"}). In GA4, traffic from AI Overviews and AI Mode falls under Organic Search, while a separate AI Assistant channel exists for assistant platforms.
Track landing-page level patterns, query-level trends and key events instead, and accept that part of genuine AI-originated traffic will arrive without a referrer and land in Direct. Report the boundary explicitly rather than presenting an estimate as a measurement.
What If Mentions Increase but Leads Do Not?
Check the chain one step at a time rather than concluding that visibility does not work. Are the questions where mentions grew actually part of a purchase decision, or are they explanatory questions? Is the brand being mentioned or recommended, since only recommendations move shortlists? Does the landing page answer the originating question, and does the form work in every market?
Then check qualification. If requests grew but qualified requests did not, the registry may be drawing an audience outside your ICP. Each of these is a separate fix, and testing them one per period is the only way to learn which one mattered.
Where to Take This Next
The system is short to state and slow to build: a defined set of buyer questions, a fixed observation protocol, metrics with honest denominators, verified web analytics, deduplicated CRM data, and a report that separates what was observed from what was reported and from what was assumed.
Most of the difficulty is engineering, not strategy. In the projects we run, the work that actually makes the numbers trustworthy is integration between the website, analytics and CRM, normalization of acquisition data, verification that key events fire with the right parameters, deduplication at lead and account level, and wiring all of it into a report that a commercial director can read without a translator.
Webdelo has been building B2B platforms, ERP and CRM integrations and high-load services since 2006, with more than 200 delivered projects and teams working with mid-market companies in the US, Germany and Eastern Europe. We are an official resident of Moldova IT Park. We do not promise that a measurement system will increase your visibility, because measurement and growth are different projects with different levers.
If you are setting up AI visibility tracking now, or you already have numbers you do not fully trust, we are glad to look at your current setup with you: the question registry, the analytics configuration, the CRM fields and what your report can and cannot legitimately claim. Book a working session with our team and bring the dashboard you have.
Frequently Asked Questions
Can AI Visibility Be Measured for Free?
Yes, manually. A registry of twenty to forty buyer questions, a spreadsheet log, a fixed protocol and a consistent repeat count give you a defensible baseline without any paid tool, and the output stays comparable over time as long as the protocol does not change. Cost appears with scale: multiple markets, languages, platforms and scheduled repeat checks turn into many hours per cycle. Most teams start manual AI visibility tracking, prove the method matters, and automate once the registry is stable.
How Many Questions and Repeat Checks Are Needed?
There is no universal minimum, and any specific number offered without reference to your sample is guesswork. You need enough questions and repeats to make your denominators stable, so that a small change in one answer does not swing the reported percentage. Practically: keep the repeat count identical for every question, keep the registry frozen between comparisons, and publish N with every metric. If a slice has fewer than a couple of dozen valid answers, report it as directional rather than as a percentage.
Why Do Monitoring Tools Disagree?
Because they measure different things. Collection method, regional configuration, model and mode versions, check dates, repeat counts and the definitions of citation and visibility all differ between products, and each of those differences alone can move a headline number by a wide margin. Disagreement is expected and is not evidence that one tool is broken. What is comparable is a within-tool trend under a fixed protocol; comparing your Tool A number against a competitor's Tool B number produces no defensible conclusion.
Can Every AI Overviews Visit Be Identified?
No. Google folds results from AI features into the general Web search type report in Search Console, so there is no separate AI Overviews traffic report to isolate. In GA4, traffic from AI Overviews and AI Mode falls under Organic Search, while a separate AI Assistant channel exists for assistant platforms. Track landing-page patterns, query-level trends and key events instead, and accept that part of genuine AI-originated traffic arrives without a referrer and lands in Direct. Report the boundary explicitly rather than presenting an estimate as a measurement.
What If Mentions Increase but Leads Do Not?
Check the chain one step at a time rather than concluding that visibility does not work. Are the questions where mentions grew actually part of a purchase decision, or are they explanatory questions? Is the brand being mentioned or recommended, since only recommendations move shortlists? Does the landing page answer the originating question, and does the form work in every market? Then check qualification: if requests grew but qualified requests did not, the registry may be drawing an audience outside your ICP. Test one fix per period.