Buying AI Visibility: What a Provider Must Prove Since the IAB Standard

Anyone holding a quote for AI visibility work had, until three weeks ago, no way to check whether the numbers in it were worth anything. Since 3 August 2026 there is a yardstick, and it comes from the industry body rather than from a vendor.
In brief: On 3 August 2026 the IAB published "Measuring Visibility in the AI Era", the first measurement standard for visibility inside AI answers. It sorts every metric into four categories and separates two quality tiers: directional measurement, which shows trends but is explicitly insufficient for budget decisions, and decision-grade measurement, which meets nine requirements covering scope, repeatability and documentation. Disclosure obligations for measurement vendors come on top. The trigger is a confused market: more than 20 firms sell measurement tools whose differing methods produce different results for the same brand, and only 16 percent of brands track their AI visibility systematically at all. For companies evaluating a quote, the standard is above all one thing: a list of questions a serious provider can answer.
What did the IAB publish, and why now?
The Interactive Advertising Bureau released the framework "Measuring Visibility in the AI Era" in New York on 3 August 2026. It came out of a working group chaired by Caroline Gigerech, the body's VP of AI, with contributions from Walmart, EMARKETER, Acxiom, WPP Media, Microsoft Clarity, PMG, the Alliance for Audited Media and Tinuiti, among others.
The trigger is a market failure. More than 20 firms now sell tools for measuring AI visibility, each with its own methodology, and the same brand receives different numbers from different vendors. Without shared definitions, a reliable signal cannot be told apart from noise.
Two figures from the framework itself show how much faster demand is growing than measurement quality: ChatGPT has over 900 million weekly active users, and Google's AI Overviews reach over 2.5 billion users a month. Meanwhile only 16 percent of brands track their visibility there systematically.
Which four categories does the standard define?
The framework sorts every metric into four groups, the "4 P's" in the original. They build on one another: first whether a mention happens at all, then how prominent it is, then in what light, then to what effect.
| Category | Question | Metrics |
|---|---|---|
| Presence | Does the brand appear in the answer? | mention rate, citation rate, share of category mentions, visibility momentum |
| Prominence | Where in the answer? | position in the text, rank within lists |
| Portrayal | In what context and how accurately? | sentiment, framing, hallucination rate, factual inaccuracy rate |
| Persuasion | Does the mention lead anywhere? | recommendation strength, post-citation click-through rate |
Two distinctions inside this matter more in practice than they first sound.
First, the standard separates mention from citation. A brand can be named constantly in answers and still almost never be linked as a source. These are two different metrics with two different levers, and conflating them means measuring past your own intervention.
Second, it separates hallucination rate from factual inaccuracy. In the first case the model invents a connection with no source behind it. In the second there is a real source, but the detail drawn from it is wrong. The first is a model problem; the second is one you can fix through source work.
What separates a dependable number from a snapshot?
This is the most useful part of the framework. It defines two quality tiers and names nine criteria that distinguish them.
Directional means the measurement shows whether things are moving up or down. It serves internal monitoring and competitive awareness. For a budget decision it is explicitly not enough.
Decision-grade means the number can carry an investment decision. That requires nine conditions.
| Criterion | Requirement for decision-grade |
|---|---|
| Query volume | at least 50 queries |
| Sample size | sufficient for the precision reported |
| Query type coverage | all four intent types, not just one kind of question |
| Testing frequency | weekly or more often |
| Repeatability | a defined tolerance band within which values may move |
| Data validation | a documented verification procedure |
| Methodology documentation | complete and traceable |
| Platform coverage | several AI systems |
| Cross-platform aggregation | disclosed method of combination |
The figure of 50 is a floor, not a target. How many repetitions a specific claim actually needs depends on the value being measured; the arithmetic is in our article on the methodology of visibility measurement.
The repeatability criterion is the notable one. It does not demand that a measurement be exactly reproducible, which is impossible with language models. It demands that the vendor states in advance how much the values are allowed to move. Without a tolerance band, there is no way to tell an effect from noise.
What must a provider disclose?
The framework sets out disclosure obligations in three groups. They are the most practical checklist currently available for evaluating a quote, because every item is a question you can ask in a meeting.
| Group | What must be disclosed |
|---|---|
| What is measured | which platforms at which model version, how the prompt library was built, volume and distribution of queries, origin of the queries (synthetic or from real user behaviour) |
| How it is measured | collection architecture (active simulation, passive observation, platform data or hybrid), capture methodology, panel validity, procedure for classifying accuracy |
| How data is handled | history and versioning, re-baselining after platform changes with documentation |
Query origin is where quotes differ most visibly. A prompt library invented at a desk measures something other than one derived from actual search behaviour. Both are legitimate, but it has to be stated.
Re-baselining sounds technical and has an immediate consequence for reporting: if a vendor changes its model or its prompt library, the numbers before and after are no longer comparable. If that is not in the report, a decline looks like a failure when in fact only the instrument changed.
Which promises do not hold up technically?
There is no position one in ChatGPT. Language models do not return the same answer to the same question, and this is not a teething problem but part of how they work. Anyone guaranteeing a fixed placement is selling something the technology cannot deliver.
The Semrush Visibility Index 2026 shows how large the variation is: over 1,200 brands across 22 industries were examined from January to April 2026 using 126 million queries in ChatGPT, Gemini, Google AI Mode and AI Overviews. Only 36 of those brands appeared among the hundred most-mentioned on every platform in every month. If brands of that size cannot hold their visibility steady, a guarantee to a mid-sized company is not credible.
Three further formulations deserve scrutiny:
"We will get you into the AI training data." Nobody can commit to that. Which content enters a training run is decided by the model provider, and the next run is months away. What can actually be influenced is retrieval at runtime, and that takes effect in weeks.
"Visibility score 78 out of 100." A composite score without a disclosed formula is packaging, not a metric. Ask which of the four categories it draws on and how they are weighted.
"We measure daily across all AI systems." Possible, but expensive. Ask for the query volume per run. Daily measurement with five questions is worth less than weekly measurement with fifty.
How do you evaluate a specific quote?
Six questions you can ask in a first meeting, each with an answer that can be held against the standard.
First: is the measurement supplied directional or decision-grade? A provider who knows the difference and states their own tier has read the framework. Both answers are legitimate; the first is simply cheaper and carries no budget decision.
Second: how many queries sit behind each percentage in the report? Below 50, the number is not decision-grade under the standard.
Third: what tolerance band applies? Without one, effect cannot be separated from fluctuation.
Fourth: where do the questions come from? Invented or drawn from real search behaviour, and if real, from which source.
Fifth: what happens when a platform changes? A provider with a procedure re-baselines and says so.
Sixth: are mention and citation reported separately? If not, there is no basis for deciding whether to work on source presence or on discoverability.
Further questions, aimed less at measurement than at the working relationship, are in our checklist for the first meeting.
What does this mean for a mid-sized company?
The honest answer: not every company needs decision-grade measurement. It exists for budget decisions, and anyone without a large sum to allocate gets a long way on directional measurement, provided it is labelled as such.
What every company does need is the ability to judge a quote. The standard supplies a shared vocabulary for that for the first time. When two vendors report different numbers for the same firm, that is not proof that one is lying; it usually points to different prompt libraries, platforms or collection methods. Which is precisely why the disclosure obligations are the more valuable half of the framework.
Two practical consequences follow. Ask to see the prompt library before signing, not after. And insist that the first report establishes a baseline against which later reports are measured, rather than merely describing a state.
For assessing your own position, the order from the visibility cluster still applies: first establish whether the company is mentioned at all, then whether it is described accurately, and only then measure how often.
Frequently asked questions
What is the IAB measurement standard for AI visibility? "Measuring Visibility in the AI Era" is a framework published by the Interactive Advertising Bureau on 3 August 2026. It defines a shared vocabulary for visibility metrics in AI answers, sorts them into four categories, distinguishes two quality tiers and sets disclosure obligations for measurement vendors.
What does decision-grade measurement mean? A measurement meeting nine requirements: at least 50 queries, a sufficient sample, coverage of all four query types, weekly or more frequent collection, a defined tolerance band, documented data validation, complete methodology documentation, several platforms and a disclosed aggregation method. It is the precondition for basing a budget decision on a number.
Is directional measurement worthless? No. It shows movement and suits internal monitoring and competitive comparison. The standard merely records that it is insufficient for budget decisions. It becomes a problem only when a vendor presents it as decision-grade.
Can an agency guarantee a mention in ChatGPT? No. Language models are not deterministic and there is no fixed placement. The Semrush Index 2026 shows that of more than 1,200 brands examined, only 36 appeared among the hundred most-mentioned on every platform in every month. A guarantee promises something the technology cannot deliver.
What is the difference between mention and citation? A mention means the brand name appears in the answer. A citation means your own page is given as a source and linked. Either can occur without the other, the standard tracks them as separate metrics, and they call for different measures.
Why do two vendors report different numbers for the same company? Because more than 20 vendors measure with different methods: different prompt libraries, different platforms and model versions, different collection architectures. The standard does not resolve this, but it makes the differences visible once both vendors disclose their methodology.
Does a mid-sized company need decision-grade measurement? Only if a sizeable investment decision depends on it. For ongoing monitoring, directional measurement suffices, provided it is honestly labelled. More important than the tier is having the prompt library, the scope and the tolerance band in writing.
Conclusion
The IAB framework changes nothing about the technology behind AI answers. It changes something about the negotiating position: for the first time there is a yardstick that does not come from a vendor, and with it a basis for comparing quotes that were previously incomparable.
For buyers it reduces to one sentence. A provider does not have to deliver the best number, but they do have to be able to say how it came about. Anyone who discloses scope, query origin and tolerance band is working transparently, even if the measurement is only directional. Anyone offering a score without a formula and a guarantee without a technical basis is not.
Sources
- Interactive Advertising Bureau, "Measuring Visibility in the AI Era", published 3 August 2026, working group chaired by Caroline Gigerech (VP AI, IAB), with contributions from Walmart, EMARKETER, Acxiom, WPP Media, Microsoft Clarity, PMG, emberos, Alliance for Audited Media, Tinuiti and IQRush.ai
- Semrush AI Visibility Index 2026, 126 million US queries from January to April 2026, more than 1,200 brands across 22 industries, four platforms (ChatGPT, Gemini, Google AI Mode, AI Overviews)
Can ChatGPT find your website at all?
The AI visibility check tests in under a minute whether AI assistants can fetch, understand and cite your page. Eight checks, a concrete fix per finding. No signup.
Run the free checkAI visibility: analysis and correction
What ChatGPT, Perplexity and Gemini say about your company, which sources it comes from and where competitors stand - with an action list instead of a score.
More Articles

AI Agents in Customer Service: Revolutionizing Digital Communication in 2025
Discover how AI agents are revolutionizing customer service in 2025, enabling 24/7 intelligent communication and boosting business efficiency.

AI Automation for SMEs: A Guide to Boosting Efficiency in 2025
Our 2025 guide for SMEs on using AI agents to increase efficiency, reduce operational costs, and overcome labor shortages. Calculate your ROI.
