AI Visibility12 min read·2,200 words

KPIs for AEO and GEO: The 7 Numbers That Actually Measure AI Visibility

Faisal Zaman
Published July 27, 2026
Part of the AI Visibility & GEO cluster
Read the pillar guide →

There are seven KPIs worth tracking for AEO and GEO, and none of them is a keyword ranking. The most important is AI share of voice: the percentage of relevant AI answers that name your brand, but on its own it can’t tell you why the number moves. The other six close that gap.

This is the full measurement framework underneath the AEO + GEO playbook, built from Princeton’s GEO study, Pew Research’s click data, and our own citation research across more than 175 live searches.

Why Rankings Stopped Being Enough

Pew Research tracked the real browsing behavior of 900 US adults and found click-through drops from 15% to 8% when a Google AI Overview appears on the results page. You can hold position one and still lose the visit, because the answer above you resolved the query before anyone scrolled down to click.

Share of searches where a user clicked a resultPew Research Center · 900 US adults · March 2025No AI summary shown15%AI Overview shown8%Link inside the AI Overview1%An AI Overview cuts click-through by nearly half.
Fig. 1When an AI Overview appears, click-through nearly halves. Source: Pew Research Center, 2025.

A ranking report can’t see any of this. It tells you where your page sits in a list of blue links that a shrinking share of searchers ever reach. What replaces it is a small set of KPIs that measure the thing that actually happened: were you named, was the claim believed, and did anyone act on it.

The 7 KPIs, at a Glance

Each one answers a different question. Together they explain the whole funnel.

  • AI share of voice: are you in the answer at all?
  • Citation / adoption rate: when you make a claim, does the engine believe it?
  • Consensus coverage: do independent sources back you up, so the engine has a reason to?
  • Passage liftability: is your content even structurally eligible to be quoted?
  • AI referral traffic: is any of this turning into visits?
  • AI referral conversion: are those visits worth more than an average click?
  • Click displacement: how much organic click-through are you losing to the answer itself?

Start with the one that’s easiest to run without any tooling.

1. AI Share of Voice

AI share of voice is the percentage of relevant AI-generated answers that mention your brand compared with your competitors, across assistants like ChatGPT, Perplexity, and Gemini. If ten buying-intent questions in your category produce answers that name a competitor eight times and you twice, your share of voice is 20%.

It is the direct descendant of share of voice in traditional media and paid search, moved to the surface where research now happens: the answer box and the chat window. No ratings panel reports it for you, so you measure it yourself.

Influence can grow while traffic stays flat. If you only measure sessions, you will miss the entire shift.

Building the query set that matters

The exercise is only as good as the questions, and it stays fixed once you start.

  • Category queries: “best CRM for a small team,” “top project management tools.” Test whether you appear in the consideration set at all.
  • Recommendation queries: “what should I use to send invoices.” Test whether the model actively names you.
  • Direct brand queries: “is [brand] any good.” Test accuracy, not just presence.
  • Comparison queries: “[you] vs [competitor].” Show exactly where you win and lose head-to-head.

Fifteen to thirty questions is plenty. Fewer and one odd answer swings the whole number; many more and the monthly run becomes a chore you skip.

The free method, step by step

  1. 1Lock the query set from the categories above, and write it down so it never drifts.
  2. 2Run every query on each assistant. ChatGPT, Claude, Perplexity, and Gemini, in a clean session each time.
  3. 3Log the outcome per query, per assistant. Mentioned? Which competitors appeared? Accurate? What sentiment?
  4. 4Note the sources. On Perplexity especially, record which pages it cited: that tells you where the answer is coming from.
  5. 5Re-run the identical set next month and compare the percentages.
Run in a private or logged-out session where possible. Personalization and chat memory can quietly bias an assistant toward brands you’ve mentioned before, flattering your numbers and ruining the comparison.

A mention that describes you wrong is not a win. Track accuracy and sentiment alongside the raw mention rate, or the headline number will mislead you.

2. Citation / Adoption Rate

Share of voice tells you if you were named. It doesn’t tell you whether a specific claim you make about yourself actually gets believed and repeated. That’s a separate KPI, and it’s measurable: run the exact commercial query your own “best X” or comparison content makes a claim on, and check whether the AI Overview’s generated answer adopts your framing, or just cites your page as one source among several while naming someone else.

We tested this directly across 129 live searcheswhere a vendor ranked its own product #1 on its own blog. The self-ranking got adopted into Google’s actual generated answer…

60%
of the time, but almost always only when independent sources already named the same vendor
0
times did a bare 'best overall' claim win against an entrenched category leader with no outside backing

Full breakdown, including the case where the winning line was sourced to a direct competitor’s page, not the vendor’s own, is in our citation study.

Citation rate is the KPI that catches the gap between “we said it” and “the model believed it.” A claim with zero adoption is a content problem or a consensus problem, not a coincidence.

3. Consensus Coverage

This is the leading indicator behind KPI 2. Consensus coverage is the percentage of a cited AI answer’s own sources that independently name you: not sources you control, sources the engine chose on its own.

71%
of an AI Overview's own cited sources named the vendor it ultimately adopted
10%
named the vendor it ultimately ignored; five of seven scored exactly zero

That’s roughly a 7× gap, and it’s the cleanest split we’ve found in any of our research. To measure it yourself: pull the source list from an AI Overview or Perplexity answer in your category, open each one, and count how many mention you by name. Below roughly 30%, expect to be filed as a source but not named as a pick.

4. Passage Liftability

This KPI runs on your own site, not on search results. It asks: of the passages on your key pages, how many are structurally eligible to be quoted at all? A correct fact written the wrong way is invisible to a retrieval system no matter how authoritative your domain is.

We audited this across 49 live searches spanning health, finance, food, travel, and B2B topics, comparing the exact sentence an AI Overview echoed against a nearby, ignored paragraph on the same page.

100%
of lifted passages opened with the answer, vs. ~18% of the ignored ones nearby
98%
were self-contained: no dangling pronoun, no 'as mentioned above', vs. ~37% of the ignored ones

To score your own pages: pick your 10 most important passages and check each against those two traits. A page can pass this audit with zero new content: often it’s a restructuring exercise, not a writing one. Full method and examples in the playbook.

5. AI Referral Traffic (Volume & Growth)

The first four KPIs measure visibility. This one measures whether any of it converts into a visit. Semrush analyzed more than 1 billion lines of US clickstream data and found ChatGPT outbound referral traffic grew 206% year over year between January 2025 and January 2026, reaching roughly 170,000 unique referred domains per month.

  • Segment it in GA4 now. If AI referrals sit inside a generic “referral” or “direct” bucket, you can’t prove any of this internally later.
  • Track the trend, not the absolute number. Volume is still small for most sites; the growth rate is the signal.

6. AI Referral Conversion Rate

Volume without conversion is a vanity number. Seer Interactive’s case-study data puts ChatGPT-referred traffic at around 16% conversion against roughly 1.8%for Google organic. Similarweb’s 2026 clickstream figures are more conservative at 7.1%, still second only to paid search. Both point the same direction.

Someone acting on an AI recommendation has already had their objections handled by the model. At a 7–16% conversion rate, a few hundred AI-referred visits can outperform several thousand organic ones.

7. Click Displacement

The last KPI is the one everyone feels but few formally track: how much of your organic click-through is being absorbed by the answer itself before anyone reaches your page. Pew found users ended their browsing session 26% of the time after seeing an AI summary, versus 16% without one.

Track it by comparing click-through rate on queries where an AI Overview appears against queries in the same category where it doesn’t. A widening gap tells you displacement is accelerating in that category faster than your other KPIs are compensating for it.

Two KPIs to Stop Chasing

Not everything that’s easy to count is worth counting. Two metrics get tracked constantly and move almost nothing.

Raw citation count

Being cited as a source is not the same as being named as the pick. In the same research behind KPI 2, most of the vendors whose self-ranking got ignored were still cited as a source in the very answer that passed over them: five of seven losing vendors scored a source citation with zero adoption. If your GEO tool reports “number of citations,” ask whether it’s counting that, or counting adoption. They are very different numbers wearing the same label.

Schema markup coverage

We tested this directly across 70 pages, rendered in a real browser, JSON-LD read after JavaScript ran: 92% of cited pages carried schema markup, and so did 97% of the ones that weren’t cited.Having it told us nothing. The types most associated with “optimizing for search” (AggregateRating, SoftwareApplication) actually skewed toward the pages that got ignored. Track schema for hygiene, not as a KPI you expect to move citation.

Reading the Numbers: What Good Looks Like

Every KPI here means nothing in isolation; it only means something against a baseline and a set of competitors. The first month you measure is not a grade, it is a starting line.

As a rough orientation, treat AI share of voice under roughly ten percent as an entity or consensus problem: the model barely knows you exist for these queries. A middle band means you’re in the consideration set but not the default answer, usually a clarity and coverage issue. A high mention rate with a low citation/adoption rate on your own comparison claims points at a consensus gap specifically: read KPI 3 before assuming it’s a content problem.

Segment share of voice by query type, too. It’s common to score well on direct brand questions, because your own pages answer those, while scoring poorly on category and recommendation questions, which depend on third-party consensus you haven’t built yet. That split tells you exactly where to spend.

Watch accuracy as its own line inside share of voice. A brand can climb from twenty to forty percent mention rate while its descriptions stay wrong, which looks like progress but converts badly. If that’s happening, fix how AI describes your brand before pushing for more mentions.

Mistakes That Ruin the Data

Monthly is the right cadence for most of these KPIs, with a deeper look each quarter; see how often to run an AI visibility audit. A few habits quietly destroy the whole exercise:

  • Changing the query set. The single biggest mistake. Rewrite the questions and you lose comparability. Freeze the queries, vary only the date.
  • Running in a personalized session that biases the model toward you.
  • Tracking only mention rate and missing that the mentions are inaccurate, uncited as a source, or never adopted as a claim.
  • Measuring once and treating a single snapshot as a trend. The value is entirely in the movement over time.
  • Confusing citation with adoption, and schema coverage with citation likelihood. Both are covered above, and neither is a reliable KPI on its own.
The tooling in this space is maturing fast, but don’t wait for it. The brands that started tracking these seven numbers a year early are the ones who can now point at a trend line instead of a hunch.

Frequently asked questions

Seven: AI share of voice (the % of relevant AI answers naming you), citation/adoption rate (whether your specific claims get adopted, not just cited), consensus coverage (how many independent sources back you), passage liftability (whether your content is structured to be quoted), AI referral traffic, AI referral conversion rate, and click displacement. Raw citation count and schema coverage are commonly tracked but don't reliably predict AI visibility.
AI share of voice is the percentage of relevant AI answers that mention your brand versus your competitors, across assistants like ChatGPT, Perplexity, and Gemini. It is the AI-era equivalent of share of voice in traditional media, and it's the headline KPI in this framework, though it needs the other six to explain why it moves.
Not reliably. Across 70 pages tested in a rendered browser, schema markup was nearly universal on both cited (92%) and non-cited (97%) pages, and the types most associated with 'optimizing for search' (AggregateRating, SoftwareApplication) actually skewed toward pages that weren't cited. Track schema for hygiene, not as a citation KPI.
Most of them, yes. AI share of voice just needs a fixed query set run across assistants monthly and logged in a spreadsheet. Citation rate and consensus coverage need you to manually check AI Overview sources against your claims. Referral traffic and conversion need GA4 segmented correctly. None require paid tooling to start.

Want this run on your brand?

A 30 minute call on your business, then a free opportunity check. If I can't move your numbers, I'll tell you and you'll have paid nothing.

Book a Free Fit Call
FreeNo obligationHonest answer either way