A single blended AI visibility percentage averages four different measurements into one that answers nothing. The engines share a training recipe but diverge at retrieval, the layer that decides who gets quoted, and Australian evidence now shows both the engine mix and the citations behind it moving faster than a quarterly reporting cycle. Measure and report per engine instead.
Ask most agencies whether a client is showing up in AI and one number comes back. A blended share-of-voice percentage, averaged across the assistants, tidy enough for a slide. It is a comforting number, and it is an average of four separate measurements that are visibly pulling apart.
The engines run a shared recipe right up to the moment that decides who gets quoted, then they split. That split has now been measured, including here in Australia, and it is wide. The fix sits in method rather than in a better KPI: which numbers you separate, what you can honestly promise about each, and how to add a per-engine line to reporting you already deliver.
Your client has already read the headline
Clients have stopped asking whether AI visibility matters. Australian trade press settled that this week, when iTnews carried who owns your brand's AI reputation, which is partner content rather than newsroom reporting and worth reading as such, but which lands in front of an Australian business-IT audience all the same. It names ChatGPT, Gemini, Perplexity and Google's AI Overviews as the surfaces buyers now use instead of visiting a brand site, and it walks a reader through the falling click-through trajectory as settled context. Assume your client has already been given that trajectory. The piece also does the thing this argument works against: it names four engines and then treats them as one undifferentiated surface.
The market is already pricing the granular version. Digiday's briefing on creator pitches has Deloitte Digital's Jenny Kelly, head of content, creator and AI, coaching creator-representation companies on that skill, while Tinuiti's Crystal Duncan, evp of brand engagement, describes pitches moving from "I had my last pasta video go viral" to "my video about PFAS in your water shows up when you search it". The same piece names the limit: a creator saying "I got cited" is closer to an anecdote than a stat. Your clients will hit that wall next, and the reputational half of it gets fuller treatment in how AI describes your brand.
The engines share a recipe right up to the point that matters
Every major assistant runs a version of the same build. A per-engine GEO breakdown from Neil Patel's team sets out four stages: pretraining, instruction tuning, preference optimisation to shape the model's habits, and product-level decisions about what to retrieve and rank. The first three are broadly shared. The fourth decides whose name appears. Think of four researchers given the same education, then sent to four different libraries with different rules about which shelves they may use. The training explains why the answers sound alike. The library explains why the citations do not match.
Some engines fetch live sources the moment you ask, which the industry calls retrieval-augmented generation. Others lean on patterns already baked into the model's weights. Google's AI Mode shows the local version: its Australian launch ran on customised Gemini models using query fan-out, which splits one prompt into multiple sub-searches.
Two surfaces can agree on what to say and still disagree on who gets credit. Ahrefs' comparison of AI Mode and AI Overviews, published in December 2025 across 730,000 response pairs, found average semantic similarity of 86%, yet the same URLs appeared in both only 13.7% of the time. Ahrefs sells SEO tooling and has a commercial interest in per-engine tracking. The test case is still the cleanest available, because both surfaces come from one company, one index and the same underlying models.
The pattern holds beyond Google. An academic study of 55,936 queries across six LLM-based search engines, including ChatGPT, Gemini, Perplexity and Grok, reports that "Even for services provided by the same company (e.g., Google, AI Mode, and Gemini), the sourced domains differ substantially (≥30%)." It is a preprint with no pairwise overlap table, so it establishes direction rather than a ChatGPT-versus-Perplexity figure. Australia has its own version: a study of 11,500 queries comparing Google Search, AI Overviews and Gemini found similarity scores of just 11% to 18%. Three surfaces in one company's stack, agreeing on sources less than a fifth of the time. [AU]
Four engines, four selection jobs
Each engine bends the recipe toward a different job. Perplexity is built around source-based retrieval, so it reads like a research assistant: citation-heavy, grounded in what is live on the web. Claude is tuned to go deep, holding long context across multi-step, agent-style work. ChatGPT covers the widest surface and the largest install base. Gemini's edge appears the moment a task touches Google's own ecosystem, from Docs and Sheets to YouTube.
That decides who gets cited. Profound's analysis of 680 million citations from August 2024 to June 2025 found ChatGPT's most-cited domain was Wikipedia at 7.8% of its citations, while Perplexity leaned to Reddit at 6.6% and YouTube at 2.0%. Profound sells AI-visibility software and the dataset predates AI Mode, so year-stamp it in any client deck. An encyclopedic lean on one engine and a community lean on another means a client can win one surface and stay invisible on the next without changing a word. Earning mentions in the source classes an engine favours is digital PR work aimed at AI citations. Australia has a worked example: the same Perth research firm found the Starworks business directory appeared 199 times across 2,400 AI responses in its September 2026 study, third as a recommendation source behind Google and Reddit. [AU]
Preference optimisation deserves one plain sentence, because it is the least mystical part of this: human graders pick which of two answers they prefer, and the model learns to produce more of that kind. Anthropic's 2022 hh-rlhf release carries 169,000 human preference rows and Nvidia's 2024 HelpSteer2 grades on five axes, and an independent analysis found list-shaped answers scored higher in a majority of the pairs examined, framed by its author as an informed inference rather than settled fact.
Which engines you report on should follow your client's market, not an imported US list. In Australia, ChatGPT leads and the tail is thinner than most agencies assume. Roy Morgan's AI tool tracking put ChatGPT at 10.5 million Australians against Gemini's 5 million and Claude's 777,000, from 14,646 respondents surveyed January to March 2026, with Perplexity not reported at all. Telsyte's 2026 Australian study, from 2,023 respondents, puts ChatGPT higher at 13.8 million and leaves Perplexity outside the top ten. Two credible sources, two sizes for the same lead, so present usage as a range. Statcounter's AI chatbot traffic share is the only current measure carrying all four engines, recording ChatGPT at 72.32% and Perplexity at 3.4% for Australia in August 2026 against 77.44% and 2.81% for the United Kingdom, though its methodology is undocumented, so treat it as a traffic proxy rather than a survey of people. [AU] [UK]
One local quirk belongs in the client conversation too. Perplexity is bundled into Optus mobile plans here, so its Australian install base reflects a telco deal as much as user preference.
What nobody can tell you, and why an honest method says so
Start with the ground rule that engine comparison states up front: nobody outside OpenAI, Anthropic, Google and Perplexity knows the current, exact selection formula, and those systems shift on a rolling basis. Anyone selling you a formula is selling a guess.
A single engine also disagrees with itself. A 2026 preprint on measuring AI search visibility ran eight prompts per campaign across four verticals on ChatGPT, Gemini, Google AI Mode and Perplexity between January and March 2026, and found "only about 35% of the cited sources overlap between two consecutive days". Simultaneous runs inside 24 hours were no better, so the instability is stochastic rather than time-driven, and rank-biased overlap ran from 0.088 for ChatGPT to 0.254 for Google AI Mode. Engines differ in how unstable they are, so one blended figure mixes signals of very different reliability.
The mix moves too. Similarweb's 2026 AI search figures estimate that over the twelve months to May 2026 ChatGPT's share of worldwide generative-AI web traffic fell from roughly 76% to about 53% while Gemini rose from under 9% to around 27%. Those are panel-based global estimates from a commercial measurement vendor, so use them for direction only. Even regulators publish at different resolutions: Ofcom's Online Nation 2025 found 54% of UK adults using generative AI tools, up from 31% in 2024, and gave absolute visit counts for ChatGPT while reporting Gemini, Claude and Perplexity as growth rates with no base underneath. [UK]
In Australia the honesty requirement has a legal shape. The ACCC's guidance on false or misleading claims is blunt: a business must be able to prove any claim it advertises, claims must be accurate, truthful and based on reasonable grounds, and for claims about future matters the business carries the burden. A blended score with an undisclosed method is harder to substantiate than four measured numbers carrying dates and sample sizes. To be accurate about the regulator, its 2026-27 enforcement priorities name manipulative and false practices in digital markets and contain no AI-specific item. The general substantiation rules already apply. GEO, AEO and "AI search visibility" remain useful shorthand and contested doctrine at once, so any method you sell should hold up if the label disappears. [AU]
The per-engine reporting line you can add this quarter
Start where the traffic lands, then check where it converts. A study of 115 Australian businesses covering September 2024 to September 2025 found ChatGPT drove 90.2% of identifiable AI referral traffic in Australia, Perplexity 4.25% and Gemini 2.5%. The part that makes your case sits under the headline: Perplexity and Gemini users converted three to six times higher than ChatGPT users. That window closed in September 2025, so present it as the latest Australian baseline rather than a current-quarter figure. A blended percentage weighted by volume would bury the highest-converting traffic in that dataset almost entirely. [AU]
So let's make this a line item rather than a project. Add an AI share-of-voice line, per engine, to the influencer briefs and post-campaign reports you already deliver, covering the prompts that genuinely precede a purchase. Attach the prompt set, the date and the number of runs, and report a range rather than a single figure. That is also the practical bridge from being mentioned to being the source AI recommends.
A competitor will quote the blended score. The per-engine breakdown survives scrutiny in the room, which makes "we measure these separately, and this is why" a positioning asset rather than a caveat, and defensible measures a board will trust are built from the same parts. The UK side is open ground: UKTN's coverage of an AI visibility tool for B2B companies built by Clarity Global shows a funded category with vendor noise but no published per-engine methodology standard. [UK] Sized properly, this is an add-on rather than a new budget line: per-engine AI visibility measurement, white-labelled, delivered against a retainer you already hold.
Key takeaways
- Treat AI visibility as four measurements. Two surfaces from the same company agreed on the answer 86% of the time and on the cited URLs only 13.7% of the time.
- Report each engine with its prompt set, date and run count, and give a range. A 2026 preprint measured only about 35% overlap in cited sources between consecutive days.
- Pick engines by market. Australian surveys put ChatGPT well ahead of Gemini and Claude, with Perplexity unreported or outside the top ten, and the UK orders similarly.
- Do not weight by volume alone: one Australian study found ChatGPT carried 90.2% of AI referral traffic while Perplexity and Gemini users converted three to six times higher.
- Selling an AI visibility number in Australia means meeting the ACCC's substantiation standard: accurate, truthful, based on reasonable grounds, with the method disclosed.
Frequently asked questions
Does one AI visibility score cover ChatGPT, Gemini and Perplexity?
No. The engines share a training and tuning recipe but make different product-level retrieval decisions, and that is the layer deciding who gets cited. Ahrefs found Google's own AI Mode and AI Overviews reached 86% semantic similarity yet cited the same URLs only 13.7% of the time, so averaging separate engines hides the difference a client is paying you to move.
Is using ChatGPT for SEO the same job as optimising for Perplexity?
Not in practice. Perplexity is built around source-based retrieval of what is live on the web, while ChatGPT covers a wider surface with a larger install base. Analysis of 680 million citations from August 2024 to June 2025 found ChatGPT's most-cited domain was Wikipedia at 7.8% of citations against Perplexity's lean toward Reddit at 6.6%. That data comes from an AI-visibility vendor and predates AI Mode, so read the shape rather than the decimal.
Which AI engine should an Australian client be measured on first?
ChatGPT, on current Australian evidence, then Gemini. Roy Morgan put ChatGPT at 10.5 million Australian users against Gemini's 5 million from 14,646 respondents surveyed in early 2026, and Telsyte reported the same ordering from 2,023 respondents. One Australian study also found Perplexity and Gemini users converted three to six times higher, so volume alone is the wrong sole criterion.
How often should AI visibility be measured to be reliable?
Repeatedly, and reported as a range. A 2026 preprint running four engines across four verticals measured only about 35% of cited sources overlapping between consecutive days, with simultaneous runs inside 24 hours no better, so a single measurement is one draw from a distribution rather than a stable number.
Start with the engine your buyers actually open
Amina helps professional-services firms get cited in AI answers, not just ranked in Google. We measure share of AI visibility engine by engine and build the content and structure that earn the citation. See Amina's AI SEO and visibility service.


