"Just check whether we come up in ChatGPT." It sounds like a reasonable request from a board, and it produces a reasonable-looking answer: you type the prompt, screenshot the reply, and report a number. The trouble is that the number is confidence dressed as fact. Ask the same question twice and the model can hand you two different answers, because a large language model doesn't return a ranking. It returns a probability. Bolt a rank-tracking mindset onto a system that has no rank to hold, and you get a figure that collapses under the first informed follow-up. This piece is about the other kind of number: one built on a method you can put in front of a board and defend, line by line.

Why the spot-check falls apart in the room

Let's start with why the question is being asked, because the pressure is real. Business buyers have moved their research into AI, though the scale depends on how you count it. In Forrester's 2025 Buyers' Journey Survey, 94% of business buyers reported using AI somewhere in their buying process (up from 89% a year earlier), and twice as many named generative AI or conversational search as a more meaningful information source than any other, ahead of vendor websites and sales teams (Forrester). Gartner measures something narrower and lands lower: across two surveys at its 2026 CSO & Sales Leader Conference, 45% of B2B buyers said they used generative AI primarily to gather vendor and product information during a recent purchase, with 69% still validating those answers with a human sales rep (Demand Gen Report). The two aren't directly comparable. Forrester counts any AI use, while Gartner counts AI used specifically for vendor research, so treat it as a range rather than a settled stat. Either way, the answer machine is in the room when your buyers decide.

The naive number doesn't just wobble. It misreads what's happening. [AU] An Australian study of 115 businesses tracked over a year found AI-referred traffic up a median 1,200% year on year, with ChatGPT driving 90.2% of identifiable AI visits (SmartCompany). The authors' own caveat is the line for a board: measuring in-answer brand visibility is an analytics challenge, because direct referrals stay modest even as your brand's appearances inside the answers climb. A spot-check catches almost none of that. As NP Digital puts it, applying rank-tracking logic to a probabilistic system produces "a fundamentally different kind of measurement entirely," because there is "no rank one to hold" (NP Digital).

There's a second, quieter flaw. Generic prompts like "best CRM in 2026" describe an abstract user with no history, constraints or real intent. Ask one and you're measuring how the model answers a person who rarely exists. That is like polling a single voter, then reporting the election result.

What a number you can defend actually measures

So what's the alternative? Not more prompts. The instinct, when one generic prompt feels thin, is to run a thousand variations, but NP Digital calls this the "scaling trap": every topic branches into phrasings, intents and personas until you're running tens of thousands of prompts and still haven't fixed the representativeness problem underneath. More volume just makes a flawed measurement more expensive.

The fix is to improve the quality of the input rather than the quantity. NP Digital's approach (which it labels SPIV) injects persona and intent variables into each prompt so it models a specific person, in a specific situation, trying to reach a specific outcome. A focused set of 15 to 30 prompts mapped to your key buyer personas and intent stages, the team argues, produces more actionable signal than hundreds of generic variations (NP Digital). For a comms audience, that is the whole difference: a small, representative, documented prompt set is something you can hand an auditor. A brute-forced pile of queries is not.

A defensible method also spans models rather than a single chatbot. Visibility that holds on Gemini but vanishes on ChatGPT is a finding, not a rounding error. The market is already pricing this in. [UK] London-based Searchable raised £10.3m in May 2026 for a platform that tracks brand visibility across ten AI engines and folds in Search Console and analytics data; founder Chris Donnelly argues customers arriving from an AI recommendation "convert at 3x higher" and that being absent from those answers cedes ground to competitors daily (UKTN). UK trade press now frames this as "the AI visibility war," with agencies launching their own tracking tools (UKTN). One note of discipline: the field goes by GEO or AEO in vendor decks, but Google has pushed back that "good SEO is good GEO," so treat it as a contested practice rather than settled doctrine.

The finding that changes the conversation

A proper method earns its fee at exactly this point. The single most consequential thing it surfaces is the gap between where you look strong and where you actually get chosen. In one NP Digital audit, a brand showed visibility above 65% for general licensing topics across platforms, the broad category terms that make standard tracking look healthy. On the compliance and banking topics tied directly to buyer decisions, its visibility was zero across ChatGPT, Google AI Overviews and Perplexity (NP Digital).

Sit with that for a moment, because it is the sentence that reframes the boardroom conversation. The headline number said "we're visible." The truthful number said "we are invisible at the exact moment a buyer decides." Unlike a low Google ranking, which is visible and actionable, AI invisibility is silent. You won't know it is happening unless you measure for it deliberately. A category-term spot-check will report the flattering 65% every time and never once mention the zero.

Making it board-ready: the definition before the deck

Turning this into reporting an executive can read comes down to two layers. The primary layer is the number a board wants: how often you appear, and your share of voice against competitors. Beneath it sits a second layer that spot-checks skip: how stable that appearance is across repeated runs, how crowded the topic is with rivals, and whether your presence holds from one AI engine to the next or collapses when a buyer switches from ChatGPT to Gemini. The technical names (run length, entropy, divergence between models) matter far less to a board than the plain question they answer. Is this figure solid, or a fluke?

Which brings us to the word that gets abused most: share of voice. It only means something if everyone can see the denominator. Share of what, across which prompts, on which engines, over what window? A percentage with no defined bottom is a batting average with nobody keeping score, impressive until someone asks how you got it. Define the denominator before the number goes in the deck.

This is where governance stops being abstract. [AU] In a paper on director responsibility in the era of generative AI, NSW Chief Justice Andrew Bell warned that Australia's technology-neutral Corporations Act duties already apply to AI: directors must keep critical oversight of AI outputs, treat the tool as "an aide... rather than as a substitute" for their own judgement, and cannot "shut their eyes... and then claim that because they did not see the misconduct, they did not have a duty to look." He calls AI-related litigation "inevitable" (ARN). A director signing off on a visibility figure they can't interrogate is the exact posture that warning describes. An auditable method isn't polish; it's cover.

There's a data angle too. [AU] When you stand up recurring AI-visibility reporting (pulling prompts, personas, sometimes client data into a tracking pipeline), that pipeline sits inside Australia's privacy regime. The OAIC opened its first privacy compliance sweep in early 2026, checking whether privacy policies clearly disclose how data is collected and used, with penalties of up to $66,000 under the 2024 Privacy Act amendments (OAIC). For a firm running white-labelled reporting under NDA, the method has to be privacy-clean, not just accurate.

Key takeaways

  • A single spot-check isn't a measurement. LLMs are probabilistic, with no rank to hold, so one prompt produces a number that changes on the next run and can't survive scrutiny.
  • Structure beats scale. A documented set of 15 to 30 persona-and-intent prompts, run across multiple AI engines, tells you more than tens of thousands of generic queries (NP Digital).
  • Look for the decision-stage zero. Strong visibility on broad category terms routinely hides near-zero presence on the high-intent topics where buyers actually choose. That gap is the finding worth reporting.
  • Define share of voice before you present it. No number goes in a board pack without a visible denominator: share of what, across which prompts and engines, over what window.
  • Build it audit- and privacy-clean. Director-duty expectations and active OAIC privacy enforcement mean the method behind the figure matters as much as the figure, especially for white-labelled, NDA-bound reporting.

Frequently asked questions

What is AI visibility measurement?

AI visibility measurement is how reliably and how favourably your brand appears in the answers generated by tools like ChatGPT, Perplexity and Google AI Overviews. Done properly it treats visibility as a probability distribution across specific buyer contexts rather than a single score, because the same prompt can return different answers each time (NP Digital).

Is asking ChatGPT whether my brand comes up a reliable check?

No. A one-off prompt models a hypothetical user and returns a probabilistic answer that can change on the next run, so it produces a confident-looking figure with no defensible method behind it. It also misses in-answer brand appearances that never show up as direct referral traffic, something Australian buyer-behaviour data has flagged as a genuine analytics challenge (SmartCompany).

What is AI share of voice, and can it go in a board report?

AI share of voice is how often your brand appears in AI answers relative to competitors, and it belongs in a board report only once the denominator is defined. Share of what, across which prompts, on which engines, over what window? Without that, the percentage is unauditable and shouldn't be presented as a metric.

How many prompts do you need to measure AI visibility properly?

Fewer than most teams assume, if they're the right ones. NP Digital argues a focused set of 15 to 30 prompts mapped to real buyer personas and intent stages gives more actionable signal than hundreds or thousands of generic variations, because representativeness matters more than volume.

What should a board-ready AI visibility report include?

It should include the headline visibility and share-of-voice numbers, a clearly stated denominator and prompt method, results across multiple AI engines rather than one, and some measure of how stable those numbers are. For Australian and UK organisations, it should also be built to withstand director-duty and privacy-governance scrutiny, so the method is defensible if questioned.

Ready to give your board a number that holds?

That's the gap M2.0's AI Visibility Check is built to close: a defined prompt set mapped to your buyers' real decision questions, measured across the major AI engines, and reported in a format your board can actually read, with the denominator, the method and the confidence range on the page rather than buried in a tool. If you need it white-labelled and NDA-safe for a client's board pack, that is how we deliver it by default. Ask us for a board-deck-ready sample.

Amina
Editorial Team