skipToContent
Back to all posts

Measuring AI Visibility: Figures Instead of Gut Feeling

August 10, 2026
By Michael Kaiser
GEOAI VisibilityMonitoringChatGPTSME
Four text cards on the left whose connecting lines lead to a bar chart with turquoise error bars

The board wants figures before it releases budget: how present is our company in AI answers, and where does the competition stand? So someone opens ChatGPT, types the obvious question and gets a list of five providers. Their own company is not on it. Screenshot, slide, alarm. Two days later a colleague asks the same question and the company sits in second place.

Both measurements are worthless. Not because anyone made a mistake, but because a single query carries no information with these systems.

Key takeaway: AI systems answer the same question differently, and this is not random noise but measurable and systematic. A Conductor study covering 14,000 queries shows that for purchase-intent questions, only 40 percent of the brands named overlap between two runs of the same question. A defensible measurement therefore needs three things: a fixed question list, several repetitions per question, and an analysis that states its uncertainty. Anyone deriving a percentage from five queries is reporting randomness with decimal places.

Why does the same question produce a different answer every time?

Because language models do not look answers up. They calculate probabilities word by word. Tiny numerical differences are enough for a different word to win at a fork, and from there the answer runs in another direction. With provider lists, that is exactly what decides who gets named.

What is notable is that the variation persists even when you set the model temperature to zero. Thinking Machines Lab showed why in September 2025: the compute kernels on the graphics card produce different results depending on how many requests are being processed at the same time. Your request does not change, but its neighbourhood in the batch does, and with it the rounding in the last decimal place. Only after the team made the affected kernels independent of batch size were 1,000 consecutive runs bit-for-bit identical.

The practical consequence: this variation is not a quirk of individual vendors, and it does not disappear once you find the right switch. It belongs to the method.

Conductor measured its magnitude systematically in July 2026: 70 questions, 50 independent runs each, four systems (ChatGPT, Perplexity, Claude and Gemini), ten industries, 14,000 queries in total. The result contradicts the common assumption that the answers are simply arbitrary.

Type of questionOverlap of named brandsStability of the first mention
Purchase intent40 %59 %
Support47 %72 %
Navigation47 %71 %
Recommendation49 %72 %
Pricing58 %59 %
Educational60 %30 %
Comparison63 %91 %

How to read this: for purchase-intent questions, two answers to the same question overlap in only 40 percent of the brands they name. Precisely where it matters most commercially, the spread is widest. Comparison questions are the most stable at 63 percent, and the brand named first stays the same in 91 percent of runs. Educational questions show a pattern of their own: the set of brands is fairly stable at 60 percent, but which of them comes first changes in 70 percent of runs.

A second finding from the same study: Gemini names 9.2 brands per answer on average, ChatGPT only 5. Anyone merging both systems into one metric is comparing unlike things.

Context worth knowing: Conductor sells software for search and AI visibility and therefore has an interest in the topic being seen as demanding. That said, the methodology is disclosed and reproducible, and the order of magnitude matches what we see in our own measurement series. The figures are usable, but their origin should be known.

Which metrics actually exist?

AI visibility is not a single number but a bundle of six. They answer different questions, and most internal debates derail because two people are talking about different ones.

MetricWhat it answersHow it is produced
Mention rateDo we appear at all?Share of runs containing a mention
Share of mentionsHow big is our slice?Own mentions divided by all brand mentions
Position in the answerAre we named first or fifth?Rank per answer, averaged across runs
Citation rateIs our website linked as a source?Count of source references
Factual accuracyIs what is said about us correct?Manual comparison against our own master data
SentimentAre we recommended or named with caveats?Rating against a fixed scheme

The most important of the six is factual accuracy, and it is the only one that cannot be automated. A mention containing false claims can be commercially worse than no mention at all. How widespread this is was shown by a survey from the agency maxonline in March and April 2026: across 150 mid-sized companies in eleven industries in the German-speaking region, tested with more than 450 questions, ChatGPT returned false or invented details for 96 percent of queries about managing directors. The same caveat applies here: the study comes from an agency selling GEO services. The finding does match what anyone can verify by querying their own management team.

What you can do about false claims is covered in a separate article on false AI claims about your company. In short: there is no reporting route, only the detour via the sources.

How many repetitions does a defensible figure need?

More than most people expect. If you ask a question five times and are named twice, that is 40 percent on paper. Statistically, the true value sits somewhere between 12 and 77 percent. That range is too wide to support a decision.

The table below shows how uncertainty develops with the number of repetitions, in each case at an observed share of 40 percent. The intervals are calculated using the Wilson method, which is more reliable for small samples than the widely used approximation.

Repetitions per questionObserved95 percent intervalWidth
52 of 512 % to 77 %65 points
104 of 1017 % to 69 %52 points
3012 of 3025 % to 58 %33 points
5020 of 5028 % to 54 %26 points
10040 of 10031 % to 50 %19 points

Anyone wanting a mention rate accurate to plus or minus 10 percentage points needs around 80 repetitions of the same question. No monitoring vendor says this gladly, because it makes the effort visible.

From this follows a practical rule that limits effort without becoming dishonest: with small samples, report categories rather than percentages. Four levels are enough for any board presentation.

  1. Never named: in not a single run
  2. Occasionally: in fewer than half the runs
  3. Regularly: in more than half
  4. Always: in every run

That keeps the statement defensible, and nobody has to pretend that a move from 38 to 42 percent is a result. For metrics you do want to report as percentages, the floor is 30 observations per cell, as is standard in market research.

How do you assemble the question list?

With 25 to 50 questions your customers actually ask, and then those questions stay unchanged. Word for word, over months. The moment you adjust the wording, you are measuring the wording rather than the development of your visibility.

Below 25 questions, individual cases dominate the result; much above 50 and the list stops being maintained without automation. In practice the useful range for a mid-sized company is 30 to 40 questions.

The composition should reflect the intent types that performed with differing stability in the Conductor study. A workable split:

ShareQuestion typeExample from our own list
40 %Purchase intentWhich agency in Baden-Württemberg builds AI phone assistants for medical practices?
25 %ComparisonWhat is the difference between a chatbot and an AI agent?
15 %PricingWhat does an AI chatbot cost for a mid-sized company?
10 %EducationalHow does an AI phone assistant work in the trades?
10 %Brand directWhat is ArkeonTech?

The most common mistake here: almost only brand questions get asked. If you want to know what a prospect experiences who does not yet know your company, you have to ask the problem question, not the company name. Brand questions measure whether the system knows you. Problem questions measure whether it recommends you. For the pipeline, the second one counts.

How do you calculate your share against competitors?

Through your share of all brand mentions. You count every provider named in each answer, sum across all runs, and divide your own mentions by the total. With 40 own mentions among 500 brand mentions overall, your share is 8 percent.

Two rules make the comparison fair. First, evaluate each system separately. Because Gemini names 9.2 brands on average and ChatGPT names 5, a combined figure dilutes the statement. Second, the competitive set is defined in advance and not adjusted afterwards. Otherwise the temptation arises to redefine the field until the numbers look acceptable.

For a board presentation, share of mentions is the most intuitive figure because it reads like market share. It should never stand alone, though, but always alongside the number of observations it rests on.

What Search Console shows that no monitoring tool sees

There is one measurement channel that costs nothing and appears in no product demo: your own Google Search Console. Search queries now show up there that no human types.

In our own data for arkeontech.de, several fully written-out prompts appeared among the normal search terms in July and August 2026. One of them ran to more than 300 characters and ended with an instruction to search online for specialised providers and name specific ones. Strings like this arise when an AI system runs a web search on behalf of a user and the query reaches Google.

The value of these finds is considerable, for two reasons. First, they contain the actual wording of the customer problem, in more detail than any keyword tool provides. Second, next to it stands the position your page was served at. That is a genuine visibility measurement in a generative context, free of charge, in your own account.

How to find them: in the Search Console performance report, scan the queries by length. The prompts sit alphabetically among the short keywords and are easy to miss. Look for queries longer than roughly 80 characters and for full sentences, often with instructions such as research online or name specific providers.

One caveat belongs here: these queries are rare, and they capture only the part of AI usage that runs through a Google search. As a complete measurement it does not qualify. As an early warning system and a source for your own question list, it is excellent.

What mistakes do most measurements make?

Six patterns come up repeatedly, and five of them cost nothing but discipline.

Measuring in a logged-in account. Anyone testing in their own ChatGPT account is measuring what the system has stored about them. Saved memories and earlier conversations feed into the answer. Measure logged out or in a temporary chat without history, otherwise you are measuring your own usage history.

Confusing being mentioned with being described correctly. Your company gets named, everyone is pleased, and nobody checks the details in the subordinate clause. Every mention needs a fact check attached.

Phrasing the question slightly differently each time. Even a changed word order can tip the result. The question list gets frozen and versioned.

Deriving a trend from one measurement. A trend needs at least three measurement points at equal intervals and the same methodology across all three.

Measuring through the API and talking about the product. Querying the API measures the model. Using the web interface measures the product, including the web search in front of it and the vendor's system instructions. Both are legitimate, but they are different things. For what a prospect actually sees, the product is what counts.

Keeping no raw data. Without stored answer texts there is no way to establish later why a value jumped. Save every answer in full, with date, system and question.

Measure yourself, buy a tool, or hire an agency?

All three work. They differ in effort and in how much control you keep over the methodology.

RouteFits whenLimits
Yourself, with a spreadsheetYou are starting out and 30 questions sufficeTime cost rises steeply with each repetition
Monitoring softwareYou measure regularly across several systemsMethodology is fixed, comparability between vendors is low
Agency or consultancyYou need interpretation and measures, not just figuresOnly worthwhile if the methodology is disclosed

The monitoring software market settled in 2026; published prices range from around 29 US dollars a month for simple checks to 295 or 399 US dollars for packages covering several systems and competitive comparison, with enterprise plans quoted individually.

Whichever route you take, one test question applies: ask to be shown how many runs sit behind each reported percentage. Anyone with no answer to that, or who deflects, is not delivering a measurement but a snapshot with decimal places. We have collected more questions of this kind in our checklist for the first meeting with an AI agency.

How often should you measure?

Monthly for ongoing observation, weekly only while a campaign is running. Measuring more often produces movement in the chart that is mostly spread rather than development.

The reason lies in how slowly these systems move. If a system answers from its training state, the statement changes at the earliest with the next training cycle, and that takes months. If it searches live, changes to your content take effect considerably faster, but even then weeks pass before this shows up in a mention rate. Why systems react so differently comes down to where they draw their knowledge from, which we described in more detail using the example of an AI agent's knowledge base.

A realistic cycle for a mid-sized company: 30 questions, three systems, five runs per question. That is 450 queries a month and roughly half a working day once the analysis is set up. For the ten most important questions it is worth going to 30 runs quarterly, so that at least those metrics hold up as percentages.

Frequently asked questions

Is it enough to simply ask ChatGPT about my company? No, for two reasons. First, the answer varies considerably between runs; for purchase-intent questions only around 40 percent of the named brands overlap. Second, asking about the company name measures awareness, not recommendation. What matters is the question asked by a prospect who does not yet know you.

How many questions do I need to start? 25 to 50, with a sensible focus at 30 to 40. Below 25, individual cases shape the picture; above 50 the list stops being maintained without automation. More important than the count is that the wording stays unchanged.

Do I have to measure through the API, or is the web interface enough? It depends on the question you want answered. The API delivers the model without a web search in front of it and without the vendor's system instructions; the web interface delivers what a customer actually sees. For visibility in customer contact, the web interface is the right measurement point, even though it is harder to automate.

How do I recognise in Search Console that a query came from an AI system? By its length and sentence structure. Queries over roughly 80 characters, written in full sentences and often carrying instructions such as research online or name specific providers, practically never come from a person typing into a search box. They arise when an AI system searches on behalf of a user.

What is a good value for the mention rate? There is no general answer, because it depends on the size of the competitive field. What carries meaning is the comparison with the three to five providers you actually compete against, and the development across several measurement cycles. An absolute target without that context is a number without meaning.

How long until measures show up in the measurement? With systems that search the web live, first changes are possible after a few weeks. With answers from the training state it takes until the next training cycle, so months. Plan for an observation period of at least a quarter before evaluating any measure.

Conclusion

The board's question is legitimate: figures belong on the table before a budget release, not gut feeling. But the obvious answer, asking ChatGPT once, is not a measurement. It is a sample of size one in a system demonstrably known to vary.

A defensible measurement is no great art, but it has requirements: fixed questions, several runs, separate analysis per system, stored raw data, and an honest statement of how certain a figure is. Meet those and you can say after three cycles whether anything is moving. Skip them and you are producing charts.

Getting started costs less than the numbers suggest: write down 30 questions, query three systems while logged out, five times per question, save the answers with dates. Half a day. After that you know where you stand, and the second cycle next month shows a direction for the first time.

This article reflects the state of knowledge as of 10 August 2026.

Sources

  • Conductor, AI Brand Recommendation Study (17 July 2026): 14,000 queries, 70 questions, 50 runs per cell, four systems
  • Thinking Machines Lab, Defeating Nondeterminism in LLM Inference (September 2025)
  • maxonline Marketing hfw GmbH, AI hallucination study for the German-speaking region (March to April 2026): 150 companies, eleven industries, more than 450 questions
  • Own calculation of confidence intervals using the Wilson method
  • Google Search Console, performance data for arkeontech.de (July and August 2026)
Free tool

Can ChatGPT find your website at all?

The AI visibility check tests in under a minute whether AI assistants can fetch, understand and cite your page. Eight checks, a concrete fix per finding. No signup.

Run the free check
Matching ArkeonTech service

AI chatbot for sales & support

Answers customer enquiries in seconds, qualifies leads and hands over to your team - live in 2-4 weeks.