There are seven KPIs worth tracking for AEO and GEO, and none of them is a keyword ranking. The most important is AI share of voice: the percentage of relevant AI answers that name your brand, but on its own it can’t tell you why the number moves. The other six close that gap.
This is the full measurement framework underneath the AEO + GEO playbook, built from Princeton’s GEO study, Pew Research’s click data, and our own citation research across more than 175 live searches.
Why Rankings Stopped Being Enough
Pew Research tracked the real browsing behavior of 900 US adults and found click-through drops from 15% to 8% when a Google AI Overview appears on the results page. You can hold position one and still lose the visit, because the answer above you resolved the query before anyone scrolled down to click.
A ranking report can’t see any of this. It tells you where your page sits in a list of blue links that a shrinking share of searchers ever reach. What replaces it is a small set of KPIs that measure the thing that actually happened: were you named, was the claim believed, and did anyone act on it.
The 7 KPIs, at a Glance
Each one answers a different question. Together they explain the whole funnel.
- AI share of voice: are you in the answer at all?
- Citation / adoption rate: when you make a claim, does the engine believe it?
- Consensus coverage: do independent sources back you up, so the engine has a reason to?
- Passage liftability: is your content even structurally eligible to be quoted?
- AI referral traffic: is any of this turning into visits?
- AI referral conversion: are those visits worth more than an average click?
- Click displacement: how much organic click-through are you losing to the answer itself?
Start with the one that’s easiest to run without any tooling.
1. AI Share of Voice
AI share of voice is the percentage of relevant AI-generated answers that mention your brand compared with your competitors, across assistants like ChatGPT, Perplexity, and Gemini. If ten buying-intent questions in your category produce answers that name a competitor eight times and you twice, your share of voice is 20%.
It is the direct descendant of share of voice in traditional media and paid search, moved to the surface where research now happens: the answer box and the chat window. No ratings panel reports it for you, so you measure it yourself.
Influence can grow while traffic stays flat. If you only measure sessions, you will miss the entire shift.
Building the query set that matters
The exercise is only as good as the questions, and it stays fixed once you start.
- Category queries: “best CRM for a small team,” “top project management tools.” Test whether you appear in the consideration set at all.
- Recommendation queries: “what should I use to send invoices.” Test whether the model actively names you.
- Direct brand queries: “is [brand] any good.” Test accuracy, not just presence.
- Comparison queries: “[you] vs [competitor].” Show exactly where you win and lose head-to-head.
Fifteen to thirty questions is plenty. Fewer and one odd answer swings the whole number; many more and the monthly run becomes a chore you skip.
The free method, step by step
- 1Lock the query set from the categories above, and write it down so it never drifts.
- 2Run every query on each assistant. ChatGPT, Claude, Perplexity, and Gemini, in a clean session each time.
- 3Log the outcome per query, per assistant. Mentioned? Which competitors appeared? Accurate? What sentiment?
- 4Note the sources. On Perplexity especially, record which pages it cited: that tells you where the answer is coming from.
- 5Re-run the identical set next month and compare the percentages.
A mention that describes you wrong is not a win. Track accuracy and sentiment alongside the raw mention rate, or the headline number will mislead you.
2. Citation / Adoption Rate
Share of voice tells you if you were named. It doesn’t tell you whether a specific claim you make about yourself actually gets believed and repeated. That’s a separate KPI, and it’s measurable: run the exact commercial query your own “best X” or comparison content makes a claim on, and check whether the AI Overview’s generated answer adopts your framing, or just cites your page as one source among several while naming someone else.
We tested this directly across 129 live searcheswhere a vendor ranked its own product #1 on its own blog. The self-ranking got adopted into Google’s actual generated answer…
Full breakdown, including the case where the winning line was sourced to a direct competitor’s page, not the vendor’s own, is in our citation study.
Citation rate is the KPI that catches the gap between “we said it” and “the model believed it.” A claim with zero adoption is a content problem or a consensus problem, not a coincidence.
3. Consensus Coverage
This is the leading indicator behind KPI 2. Consensus coverage is the percentage of a cited AI answer’s own sources that independently name you: not sources you control, sources the engine chose on its own.
That’s roughly a 7× gap, and it’s the cleanest split we’ve found in any of our research. To measure it yourself: pull the source list from an AI Overview or Perplexity answer in your category, open each one, and count how many mention you by name. Below roughly 30%, expect to be filed as a source but not named as a pick.
4. Passage Liftability
This KPI runs on your own site, not on search results. It asks: of the passages on your key pages, how many are structurally eligible to be quoted at all? A correct fact written the wrong way is invisible to a retrieval system no matter how authoritative your domain is.
We audited this across 49 live searches spanning health, finance, food, travel, and B2B topics, comparing the exact sentence an AI Overview echoed against a nearby, ignored paragraph on the same page.
To score your own pages: pick your 10 most important passages and check each against those two traits. A page can pass this audit with zero new content: often it’s a restructuring exercise, not a writing one. Full method and examples in the playbook.
5. AI Referral Traffic (Volume & Growth)
The first four KPIs measure visibility. This one measures whether any of it converts into a visit. Semrush analyzed more than 1 billion lines of US clickstream data and found ChatGPT outbound referral traffic grew 206% year over year between January 2025 and January 2026, reaching roughly 170,000 unique referred domains per month.
- Segment it in GA4 now. If AI referrals sit inside a generic “referral” or “direct” bucket, you can’t prove any of this internally later.
- Track the trend, not the absolute number. Volume is still small for most sites; the growth rate is the signal.
6. AI Referral Conversion Rate
Volume without conversion is a vanity number. Seer Interactive’s case-study data puts ChatGPT-referred traffic at around 16% conversion against roughly 1.8%for Google organic. Similarweb’s 2026 clickstream figures are more conservative at 7.1%, still second only to paid search. Both point the same direction.
Someone acting on an AI recommendation has already had their objections handled by the model. At a 7–16% conversion rate, a few hundred AI-referred visits can outperform several thousand organic ones.
7. Click Displacement
The last KPI is the one everyone feels but few formally track: how much of your organic click-through is being absorbed by the answer itself before anyone reaches your page. Pew found users ended their browsing session 26% of the time after seeing an AI summary, versus 16% without one.
Track it by comparing click-through rate on queries where an AI Overview appears against queries in the same category where it doesn’t. A widening gap tells you displacement is accelerating in that category faster than your other KPIs are compensating for it.
Two KPIs to Stop Chasing
Not everything that’s easy to count is worth counting. Two metrics get tracked constantly and move almost nothing.
Raw citation count
Being cited as a source is not the same as being named as the pick. In the same research behind KPI 2, most of the vendors whose self-ranking got ignored were still cited as a source in the very answer that passed over them: five of seven losing vendors scored a source citation with zero adoption. If your GEO tool reports “number of citations,” ask whether it’s counting that, or counting adoption. They are very different numbers wearing the same label.
Schema markup coverage
We tested this directly across 70 pages, rendered in a real browser, JSON-LD read after JavaScript ran: 92% of cited pages carried schema markup, and so did 97% of the ones that weren’t cited.Having it told us nothing. The types most associated with “optimizing for search” (AggregateRating, SoftwareApplication) actually skewed toward the pages that got ignored. Track schema for hygiene, not as a KPI you expect to move citation.
Reading the Numbers: What Good Looks Like
Every KPI here means nothing in isolation; it only means something against a baseline and a set of competitors. The first month you measure is not a grade, it is a starting line.
As a rough orientation, treat AI share of voice under roughly ten percent as an entity or consensus problem: the model barely knows you exist for these queries. A middle band means you’re in the consideration set but not the default answer, usually a clarity and coverage issue. A high mention rate with a low citation/adoption rate on your own comparison claims points at a consensus gap specifically: read KPI 3 before assuming it’s a content problem.
Segment share of voice by query type, too. It’s common to score well on direct brand questions, because your own pages answer those, while scoring poorly on category and recommendation questions, which depend on third-party consensus you haven’t built yet. That split tells you exactly where to spend.
Watch accuracy as its own line inside share of voice. A brand can climb from twenty to forty percent mention rate while its descriptions stay wrong, which looks like progress but converts badly. If that’s happening, fix how AI describes your brand before pushing for more mentions.
Mistakes That Ruin the Data
Monthly is the right cadence for most of these KPIs, with a deeper look each quarter; see how often to run an AI visibility audit. A few habits quietly destroy the whole exercise:
- Changing the query set. The single biggest mistake. Rewrite the questions and you lose comparability. Freeze the queries, vary only the date.
- Running in a personalized session that biases the model toward you.
- Tracking only mention rate and missing that the mentions are inaccurate, uncited as a source, or never adopted as a claim.
- Measuring once and treating a single snapshot as a trend. The value is entirely in the movement over time.
- Confusing citation with adoption, and schema coverage with citation likelihood. Both are covered above, and neither is a reliable KPI on its own.