AI Visibility Measurement: The New IAB Standard
The IAB has standardised AI visibility measurement. Only 16 percent of brands measure it.
On 3 August 2026, the Interactive Advertising Bureau published Measuring Visibility in the AI Era, the first standardised set of measurement guidelines for tracking how brands and publishers appear inside AI-generated answers. Buried in the framework is the number that should concern every marketing leader: only 16 percent of brands systematically track AI visibility today.
The reason is a familiar one. Over 20 companies now sell AI visibility measurement tools, each using a different methodology, and they can produce different answers for the same brand. Without a shared standard, a chief executive asking how the company shows up in ChatGPT gets a number nobody can validate.
Third Hemisphere is an Australian communications agency working with B2B companies across climate, technology, and finance in Asia-Pacific. Its practice for this layer is The Fourth Hemisphere: Marketing for AI, which measures how five major AI engines represent a brand and directs the work needed to change it. The IAB framework gives that work an industry vocabulary, which makes this a useful moment to explain what should actually be on the dashboard.
What is AI visibility measurement?
AI visibility measurement is the practice of sampling AI-generated answers to buyer questions at scale, then recording whether a brand appears, where it appears, how accurately it is described, and whether the answer prompts action. It replaces the ranking position as the unit of measurement, because an AI answer has no ranked list of links to occupy.
What are the 4 P's of AI Visibility?
The IAB organises metrics into a causal hierarchy of four categories.
Presence. Does the brand appear in an AI response at all? Measured through mention rate, citation rate, share of voice, and visibility momentum.
Prominence. Where and how prominently does it appear? This covers placement, ranking order, and whether content is drawn on substantively or cited in passing.
Portrayal. In what context, and with what accuracy? Metrics include sentiment, framing, hallucination rate, and factual inaccuracy rate.
Persuasion. Does the visibility drive action? Metrics include recommendation strength and post-citation click-through rate.
The hierarchy is the useful part. Presence without accurate portrayal is a liability. A company mentioned in 60 percent of category answers, described with the wrong product scope or the wrong geography, is worse off than a company mentioned in 20 percent and described correctly.
Caroline Giegerich, VP of AI at the IAB, leads the working group behind the framework, a cross-industry team drawn from brands, agencies, publishers, and measurement providers. The framework's own diagnosis is that the market lacks definitional agreement: no common definition of a mention, no standard for what counts as a citation, and no shared method for judging whether a tool's output is rigorous enough to inform a decision.
What separates directional data from decision-grade data?
This is the part worth taking to a budget meeting. The IAB introduces a two-tier quality classification.
Directional measurement identifies patterns and signals trends. It supports early detection and competitive awareness. The framework states plainly that it is insufficient for budget allocation or executive strategy decisions.
Decision-grade measurement meets a higher standard across six dimensions: query volume, sample size, prompt type coverage, testing cadence, reproducibility, and platform coverage. The framework treats it as the required standard before making any budget or strategy call.
The practical test is sample size and repetition. A single prompt run once in one engine tells you almost nothing, because AI answers vary between runs. Anything below a meaningful query volume sits in exploratory territory, which means most of the ad hoc checks marketing teams currently run would fail the standard they are being used to justify.
How does this map to what Third Hemisphere already measures?
The Fourth Hemisphere Audit runs 50 buyer questions across five AI engines, then scores the results on four dimensions: Presence, Authority, Consensus, and Truth. The overlap with the IAB hierarchy is close enough to be useful and different enough to be worth explaining.
Presence maps directly. Authority sits close to the IAB's Prominence, asking whether the brand is drawn on as a source rather than mentioned in passing. Truth maps to Portrayal, covering factual accuracy and the specific risk of an engine asserting something about the company that is wrong. Consensus is the dimension the IAB framework distributes across categories rather than naming: whether independent sources describe the company the same way, which is what allows an engine to state something with confidence.
The scale carries weight for the same reason the IAB says it does. Fifty questions across five engines produces enough observations to distinguish a real position from a single lucky answer. One question asked once produces an anecdote.
What the largest dataset available says about where the opportunity sits
The 2026 AI Visibility Index analysed 126 million United States AI search prompts between January and April 2026, establishing benchmarks across 22 industries. Four findings are directly relevant to Australian B2B companies.
Visibility concentration varies sharply by category. In News and Media, the three most visible brands accounted for 82.9 percent of total category visibility. In Consumer Electronics the top three held 76.9 percent. In Finance, the top three accounted for 41.4 percent, and in Industrial 42.2 percent. Less concentrated categories leave room for a company to build a position, which is a reasonable description of the sectors Third Hemisphere works in.
Engines behave differently. ChatGPT cites an average of 15 sources per response and leans on community and reference platforms including Reddit and Wikipedia. Gemini cites an average of three sources, drawing on a smaller pool. A company can perform strongly in one engine and be absent from another, which is why platform coverage sits inside the decision-grade standard.
Being mentioned and being cited are separate outcomes. On Gemini, the overlap between mentioned brands and cited domains can be as low as 30 percent. A company can be recommended in an answer without its own website appearing as the evidence, which means the material doing the persuading belongs to somebody else.
Integration beats separation. Among organisations that combined search and AI visibility work into a single workflow, 81 percent reported increased traffic or leads from AI platforms. Among those managing the two separately, 36 percent reported the same.
Why Portrayal is the metric with legal consequences
Most attention goes to Presence, because it is the easiest thing to sell and the easiest to celebrate. Portrayal carries the risk.
An engine that describes an Australian company as operating in markets it has left, holding certifications it does not hold, or offering a product still in development creates an exposure that has nothing to do with marketing performance. In regulated sectors, that includes financial services, healthcare, and climate disclosure, an inaccurate machine-generated description of a company's status can travel further than a corrected media report.
Hallucination rate and factual inaccuracy rate belong in the same review cycle as any other claim the business is responsible for. Checking them requires asking the questions on a schedule and logging the answers, which is the same discipline as media monitoring, applied to a different surface.
What to do first
Four steps produce a defensible baseline without a large budget.
Write the question set before choosing a tool. Fifty buyer and investor questions, covering category, comparison, risk, and pricing. The question set is the asset, because it defines what you are measuring against.
Run it across all five major engines, and run it repeatedly. Reproducibility is part of the standard. Two runs a fortnight apart will show you how stable your position actually is.
Separate accuracy findings from visibility findings. Wrong descriptions go to whoever owns claims and compliance. Absent descriptions go to content and media.
Ask any vendor which tier their data meets. The IAB framework exists partly so buyers can ask that question. A provider who cannot describe their query volume, sample size, cadence, and platform coverage is selling directional data at decision-grade prices.
Two questions for the next board or leadership meeting
Is the AI visibility number we report decision-grade or directional? The IAB framework defines the difference across query volume, sample size, prompt type coverage, testing cadence, reproducibility, and platform coverage. Directional data supports awareness. It does not support budget allocation, and a number presented without its tier disclosed should be treated as directional.
Do we know our factual inaccuracy rate in AI answers? Portrayal covers sentiment, framing, hallucination rate, and factual inaccuracy rate. An incorrect machine-generated description of a company's status, market, or product is a compliance issue rather than a marketing one, and the only way to find it is to ask the questions and read the answers.
The takeaway
The IAB has given the market a shared vocabulary and a quality bar for AI visibility, which means the excuse for guessing has gone. Build a question set, run it across every major engine on a schedule, and treat accuracy as seriously as presence.
Third Hemisphere runs this as the Fourth Hemisphere Audit, covering 50 buyer questions across five engines with PACT scoring across Presence, Authority, Consensus, and Truth. Jeremy Liddle (LinkedIn) and Hannah Moreno (LinkedIn) lead the practice. To see your current baseline, book a consultation, or read more in Third Hemisphere's insights.