AI Visibility20 min read·5,417 words

The AEO + GEO Playbook: How to Get Cited by AI Answer Engines, Not Just Ranked on Google

Faisal Zaman
Published July 20, 2026

For two decades, SEO had one job: get the page to position one. You researched keywords, built the best page on the internet for them, earned links, and waited for the click.

Today, that no longer works. Google now resolves a large share of searches on the results page itself, and a growing volume of research never touches a search engine at all: it happens inside ChatGPT, Claude, Perplexity, and Gemini. The page still matters. The click is disappearing.

Which forces a different unit of optimization. Answer engines cite passages, not pages. They break the web into paragraph-sized chunks, match a question against those chunks, and rebuild an answer from the handful they trust. Your 3,000-word guide is not the thing being ranked. One paragraph inside it is.

Everything that follows reduces to three moves:

  • Find the gap. Locate the questions where no strong answer exists yet.
  • Answer it in one liftable block. Self-contained, answer-first, quotable without context.
  • Prove it. Statistics, named sources, and claims that agree with the rest of the web.
SEARCH ENGINE (SEO)ANSWER ENGINE (AEO / GEO)UNIT RANKEDA whole pageUNIT RANKEDA single passageWON BYBacklinks + authorityWON BYConsensus + clarityDELIVERED ASA list of blue linksDELIVERED ASOne synthesized answerSUCCESS ISClicks & trafficSUCCESS ISCitations & visibility
Fig. 1The unit ranked, the signal that wins, and the definition of success all change at once.

1. AI Overviews Cut Click-Through Rates Nearly in Half

The clearest evidence that the click is collapsing comes from outside the SEO industry. In July 2025, Pew Research Center tracked the actual browsing behaviour of 900 US adults who agreed to share their activity.

When an AI summary appeared, users clicked a traditional result 8% of the time. When no summary appeared, they clicked 15% of the time. Clicks on a link inside the AI Overview itself: 1%.

Share of searches where a user clicked a resultPew Research Center · 900 US adults · March 2025No AI summary shown15%AI Overview shown8%Link inside the AI Overview1%An AI Overview cuts click-through by nearly half.
Fig. 2Click-through with and without an AI Overview. Source: Pew Research Center, July 2025, based on 900 US adults’ browsing data from March 2025.

Three things follow from that gap.

  • Rankings decouple from traffic. You can hold position one and watch sessions fall, because the answer above you resolved the query first.
  • Sessions end sooner. Pew found users ended their browsing session 26% of the time after seeing an AI summary, versus 16% without one.
  • Being named is the new ranking. If the summary resolves the question, appearing inside it is the only visibility left on that search.

That shift only makes sense once you understand what the engine is actually assembling an answer from.

2. Answer Engines Retrieve Passages, Not Pages

Retrieval systems do not evaluate your page as a whole. They split content into small units, convert each into a vector, and match the query against those vectors. The chunk is the atom of AEO, not the URL.

The Princeton and IIT Delhi team behind the first formal study of this behaviour built their benchmark, GEO-bench, around roughly 10,000 queries across nine datasets, and measured visibility using a position-adjusted word count of how much of your content survived into the generated answer. Not whether you ranked. How much of you got quoted.

The metric that matters in generative search is not position. It is how many of your words the model chose to reuse. Aggarwal et al., “GEO: Generative Engine Optimization,” KDD 2024
  • Length stops being authority. A tight paragraph that resolves the question outranks a 2,500-word guide burying the answer in section seven.
  • Structure becomes a ranking factor. Question-named headings and answer-first paragraphs make a passage mechanically easier to lift.
  • Context-dependent prose gets skipped.A paragraph opening “as mentioned above” cannot travel, so it never gets retrieved.

We tested this ourselves rather than taking it on faith. Across 49 live Google searches spanning health, finance, travel, food, legal, and B2B topics, we found the exact sentence each AI Overview echoed on the cited page, then compared it against a nearby paragraph on the same page that was ignored. The split was almost total:

100%
of lifted passages opened with the answer, vs. ~18% of the ignored ones nearby
98%
were self-contained: no dangling pronoun, no “as above,” vs. ~37% of the ignored ones

The clearest way to see this is to watch an engine assemble an answer in public.

3. One Synthesized Answer Can Outrank Ten Blue Links

We searched “what is answer engine optimization” on Google in July 2026. The AI Overview returned a complete definition above every organic result, and cited a Reddit thread from r/localseo as a primary source, alongside a vendor glossary page.

A forum comment outranked the SEO industry’s own publishers on its own terminology. That is worth sitting with.

AI Overview
Answer Engine Optimization is the practice of structuring content so AI tools like ChatGPT, Perplexity, and Google AI Overviews can easily understand, trust, and cite it.
source +1source +2Reddit threadr/localseoVendor glossaryseobility.netOne synthesized answer. Two named sources. Zero clicks required.Illustration of an observed AI Overview for “what is answer engine optimization”
Fig. 3Illustration of the AI Overview observed for “what is answer engine optimization,” Google, July 2026. Two named sources, one synthesized answer, no click required.
  • Clarity beats domain authority. The Reddit answer won because it defined the term plainly in its opening sentence, not because the domain outranked anyone.
  • Community content is training data. Reddit, G2, and YouTube comments are heavily represented in what these systems retrieve.
  • Citations are winnable. If a forum post can take the slot, a properly structured page can take it back.

Underneath that behaviour sits a three-stage pipeline worth knowing precisely.

01Query Processing
Strips the question to key terms, then expands it with assumed intent.
02Knowledge Retrieval
Converts text to vectors and matches the query against paragraph-level chunks.
03Answer Generation
Merges the top chunks into prose, token by token, citing what it trusts.
Fig. 4Query in, answer out. Each stage is a place you can intervene.

Query processing strips the question to key terms, then expands it with assumed context: Claude branches into beginner and expert readings, Perplexity infers location. Knowledge retrieval matches the expanded query against stored chunks. Answer generation merges the winners into prose. Knowing how a model expands a query tells you which angles to cover.

4. Adding Statistics and Quotations Lifts AI Visibility by up to 37%

Most AEO advice is assertion. The Princeton GEO study is the rare piece of controlled evidence: nine content-modification strategies, tested across ~10,000 queries, each measured against an unoptimized baseline.

Statistics Addition improved visibility by 37%. Quotation Addition improved position-adjusted visibility by 22%. Citing sources landed in the same 30–40% band. The three strongest methods all clustered there.

Change in AI visibility vs unoptimized baselineAggarwal et al., “GEO: Generative Engine Optimization” · KDD 2024 · ~10,000 queriesStatistics Addition+37%Quotation Addition+22%Cite Sources+30%Keyword Stuffing-10%BASELINEStuffing performs worse than doing nothing
Fig. 5Change in AI visibility versus unoptimized baseline. Source: Aggarwal et al., KDD 2024 (arXiv 2311.09735).
  • Numbers are extraction handles. A model reaching for a defensible claim takes the sentence carrying a figure over the one carrying an adjective.
  • Attribution transfers trust.Naming your source lets the model inherit that source’s credibility instead of gambling on yours.
  • Fluency is measurable. Clear, well-formed prose scored materially better than dense or hedged writing across the same queries.
This article is built to its own spec. Every section names a source and carries an exact figure, because the study says that is what gets quoted.

The same experiment also found a tactic that actively backfires.

5. Keyword Stuffing Now Performs Worse Than Doing Nothing

Of the nine strategies tested, keyword stuffing was the only one to score below the control. It came in 10% under the unoptimized baseline. Adding keywords made pages measurably less likely to be cited than leaving them alone.

  • Density signals nothing. Embeddings encode meaning, so repetition adds no semantic information the model can use.
  • Stuffing degrades fluency. It damages the readability that the same study found improves citation rates.
  • The old playbook has negative value. Not neutral. Negative. Legacy SEO artifacts are now a liability worth auditing out.

Keyword stuffing is the only tested tactic that performs worse than publishing nothing at all.

If old tactics lose ground, the question becomes where the genuinely winnable topics are.

6. Hallucinations Mark the Gaps Worth Writing

A hallucination is not only a model failure. It is a signal that no strong, consistent source exists for that question, so the model guessed. Every guess marks an unclaimed citation.

A useful test: run your target query through a model yourself and read the answer for hedges, vague generalities, or an outright wrong specific. Each one is a marker that no strong source exists yet for that exact question, which is exactly the gap a well-sourced, specific answer can fill.

  • Specificity beats volume. Narrow, multi-condition questions have far fewer good answers online than broad ones, so the bar to become the best available answer is lower.
  • Regulated niches are still winnable.YMYL categories (health, finance, legal) reward sourcing and precision; they don’t categorically punish smaller sites that get the specifics right.
  • Coverage precedes citation. A model can only cite a passage that exists. Depth on a topic has to come before AI visibility on it can follow.

To find those gaps at scale, ask a model to decompose a query the way it decomposes it internally.

7. Hub-and-Spoke Content Compounds Faster Than Isolated Posts

A broad query is never one question. Behind “best CRM for small business” a model silently generates six to eight sub-questions and answers them before responding. Mapping that fan-out is the highest-leverage thing you can do before writing.

BROAD QUERY“best crm for small business”Which CRM is cheapest to start?Best CRM for a 5-person team?Does it integrate with Gmail?How long does setup take?Free plan or free trial?Best for non-technical founders?1 QUERY TYPED6–8 SUB-QUESTIONS THE MODEL ANSWERS SILENTLY
Fig. 6One typed query fans out into the sub-questions a model resolves silently.

Independent research on content clustering points the same direction. According to a 2026 review of topic-cluster case studies compiled by Digital Applied, successful clusters typically run 8–12 spokes per pillar, and clustered content reportedly earns 3.2× more AI citations than standalone posts covering the same topics individually: the mechanism being that a fully-covered topic gives a retrieval system more distinct, matchable chunks to draw an answer from.

  • Clusters match more embeddings. Answering every sub-question makes one asset eligible across a whole family of related retrievals.
  • Hubs concentrate authority. Internal links from spokes back to the hub make the hierarchy legible to a retrieval system.
  • Breadth buys credibility. Comprehensive coverage of a topic earns more trust on the commercial query at the bottom of the funnel.

Coverage gets you retrieved. Three other filters decide whether you get used.

8. Consistency, Consensus, and Context Decide Who Gets Cited

Once several candidate passages are retrieved, the model has to choose. Three filters do most of that work, and they are cumulative: failing one is enough to be passed over.

1Consistency
Your own claims match everywhere
Same facts across site, docs, and profiles
2Consensus
Third parties agree with you
Reviews, press, and forums echo the story
3Context
Written for a specific reader
Answers a real segment, not an average
Fig. 7The three filters applied before a passage is committed to an answer.

We tested the consensus half of this directly. Across 129 live searches where a vendor ranked its own product #1 on its own blog, Google adopted that self-ranking into its generated answer 60% of the time, but almost always only when independent sources (review sites, Reddit, competitor pages) already named the same vendor. In one case, the winning line was sourced to a direct competitor’s page, not the vendor’s own. Full breakdown in our citation study.

  • Consistency is cheapest. Reconciling contradictory claims across your own pages costs nothing and removes the fastest reason to distrust you.
  • Consensus replaces backlinks. Independent third parties echoing your facts is the modern equivalent of an earned link.
  • Context beats comprehensiveness. Content written for a named segment matches intent better than content written to offend nobody.

Passing those filters still depends on how the passage itself is built.

9. Self-Contained Chunks Beat Comprehensive Pages

The test for any paragraph is simple: would it still make sense as the only thing someone read? If it opens with “this is why it matters,” it fails: it is tethered to its neighbours and cannot be lifted.

✗ BURIED ANSWERAnswer stranded on line 4 → weak match✓ SELF-CONTAINED CHUNK
The average CRM takes 2–4 weeks to fully set up for a small team, covering data import, pipeline config, and integrations.
Question + full answer in one block →clean, liftable, citable
Fig. 8Left: the answer depends on the lines around it. Right: question and answer live in one block.

This follows directly from how retrieval works, described above: a system matching query vectors against passage vectors can only surface a passage that is itself a complete, self-contained match. A paragraph that depends on the one before it to make sense is structurally invisible to that process, no matter how good the page around it is.

The cleanest proof we found sits on one page. Microsoft’s security glossary defines ransomware twice: once in the flowing body copy, once in a boxed “Key takeaways” callout a few lines down. Same fact, same author, same page: Google’s AI Overview lifted only the boxed version.

  • Front-load the answer. Lead with the conclusion, then add nuance, so the first sentence can stand alone.
  • One idea per block. A chunk resolving a single sub-question retrieves far better than one covering three.
  • Fix structure before adding content. A page with broken chunking gets no benefit from more words; restructuring the existing content usually matters more than expanding it.

Two more from the same 49-search sample, both lifted almost verbatim because the source sentence needed no context to stand alone:

Google AI Overview for 'what is api rate limiting' with the lifted sentence from Postman's blog highlighted
Fig. 9Postman’s one-sentence definition of API rate limiting, lifted almost word-for-word.
Google AI Overview for 'how to remove a tick' with the lifted CDC sentence highlighted
Fig. 10The CDC’s tick-removal steps, compressed into one self-contained sentence and lifted whole.

One content type is naturally shaped this way, and it is badly underused, almost.

10. A Labeled Answer Block Beats a Rhetorical Question

The common advice is “add an FAQ section.” Our data says that’s half right. Literal question-and-answer shaping helped: it showed up in roughly 30% of the lifted passages in our 49-search sample and in none of the ignored controls, but a rhetoricalquestion (“Ever wonder why...”) was a negative signal every time we saw one. The engine wants a direct answer, not a hook.

The strongest version of this we found wasn’t FAQ markup at all. It was a plainly labeled summary blocknear the top of the page (“Key Takeaways,” “Quick Answer,” “The Short Answer”), and that exact boxed line getting lifted whole. We saw it on Fidelity (“Key takeaways”), Healthline (“Key takeaways”), the CDC (“Key Points”), and a Walser auto-service page that literally headed its own answer “The Short Answer.”

Q: How long does it take to rank a new page?1. Direct answer“Most pages take 3–6 months to rank.”2. Context sentence“It depends on competition and domain age.”3. Brand mention“We track this in a monthly ranking report.”4. Guardrail“Low-competition terms can move in weeks.”
Fig. 11The four layers of an FAQ answer built to be quoted.

Compare two answers to the same question.

Weak: “Setup time varies depending on a number of factors and your needs.” Unliftable: it answers nothing.

Citable: “Most teams are fully set up in two to four weeks. That covers data import, pipeline configuration, and integrations; a solo user with clean data can be live in a day. If you are migrating from another CRM, add a week for mapping custom fields.” Any sentence can be quoted and remain correct.

Google AI Overview for 'what is mindfulness' with the lifted answer-first sentence from Mindful.org highlighted
Fig. 12Mindful.org’s opening definition, answer-first and self-contained, lifted whole.
  • Lead with the unqualified answer. No throat-clearing before the first useful sentence.
  • Label the summary if you can.A visibly boxed “Key Takeaways”/“Quick Answer” line outperformed the identical fact left unboxed in body copy, on the same page, every time we found the pair.
  • Skip the rhetorical hook. A question posed for effect, rather than answered immediately, was never the passage that got lifted.

All of which raises the fair question of whether traffic from these surfaces is worth anything.

11. AI Referrals Convert Better Than Organic Search

The volume is small. The quality is not. Semrush analysed more than 1 billion lines of US clickstream data and found ChatGPT outbound referral traffic grew 206% year over year between January 2025 and January 2026, reaching roughly 170,000 unique referred domains per month.

On conversion, Seer Interactive’s case-study data puts ChatGPT-referred traffic at around 16% against roughly 1.8%for Google organic. Similarweb’s 2026 clickstream figures are more conservative at 7.1%, still second only to paid search. Both point the same direction.

  • Intent arrives pre-qualified. Someone acting on a recommendation has already had their objections handled by the model.
  • Small numbers still pay. At a 7–16% conversion rate, a few hundred visits can outperform several thousand organic ones.
  • Segment it now. If AI referrals sit inside a generic bucket in GA4, you cannot prove any of this internally.

12. Entity Consistency Is the Cheapest Trust Win Available

Underneath consensus sits a concept that separates people who dabble in GEO from people who move the number: the entity. To a model your brand is not a website. It is a node in a knowledge graph, with attributes attached: what you do, who you serve, where you operate, whether you are credible.

Every mention of you anywhere is training data for that node. When your G2 profile, LinkedIn, Crunchbase entry, and your own homepage describe you identically, the model forms a confident picture. When they disagree, the node stays fuzzy, and fuzzy entities do not get recommended.

This is why running the same audit across many sites tends to surface the same root cause repeatedly: not a shortage of good content, but contradictory titles, headers, and descriptions accumulated across pages that were each written in isolation. Reconciling that is a QA problem, not a creative one, and it’s usually the highest-leverage fix available before writing anything new.

  • Audit before you publish. Reconciling your existing footprint costs a day and removes the fastest reason a model has to distrust you.
  • Claim the structured records. Wikidata entries and knowledge panels feed directly into how both search engines and models resolve who you are.
  • Standardise the description. One sentence describing your business, repeated verbatim everywhere, beats ten well-written variations.

Consistency across properties tells a model who you are. Internal links tell it what you consider important.

Internal linking is the most underrated lever in AEO because it is the clearest structural signal you fully control. It defines hierarchy, concentrates authority, and makes topical relationships explicit rather than inferred.

A common failure pattern: the product and the content are both strong, but authority pools in the wrong places because supporting pages never link back to the pages meant to convert or be cited. Rebuilding the link graph around a clear hub-and-spoke hierarchy, every spoke pointing back to its hub, consolidates relevance signals onto the pages that actually matter, instead of scattering them evenly across the site.

  • Link from spokes to hubs. Authority should flow toward the page you actually want retrieved, not scatter evenly across the site.
  • Use descriptive anchors.Anchor text is a labelled edge in the graph; “click here” labels nothing.
  • Audit orphans first. A page nothing links to reads as a page nothing vouches for.

Structure you build carefully can still be destroyed in a single afternoon by a platform change.

14. Migrations Are Where SEO Equity Quietly Dies

Nothing erases years of accumulated authority faster than an unmanaged replatform. Broken URLs, missing redirects, lost metadata, and altered architecture can vaporise rankings in days, and the damage usually surfaces weeks later when the traffic report lands.

Protecting against it during a platform migration (Magento to Shopify, a domain change, a CMS switch) is unglamorous but entirely preventable: map and implement URL redirects, preserve metadata and content structure, then audit post-migration crawlability and indexation before declaring anything finished. Done properly, nothing dramatic happens during the switch, and that absence of drama is the actual deliverable.

  • Map redirects before launch, not after. A redirect written in response to a traffic drop has already cost you the traffic.
  • Preserve the chunk, not just the URL. If the answer paragraph gets reformatted into a widget, you keep the ranking and lose the citation.
  • Re-audit indexation post-launch. Confirm the new structure is actually being crawled rather than assuming the redirect map covered it.

Once the structure survives, you can make it explicit rather than leaving it to inference.

15. Schema Is Table Stakes, Not a Lever

This is the one place our own research overturned the standard advice, including an earlier version of this section. The usual claim is that FAQPage and similar schema help a passage get cited. We tested it properly and it doesn’t hold up.

We loaded 70 vendor “best X software” pages in a real, rendered browser (37 that ended up cited in Google’s AI Overview, 33 that didn’t), and pulled every schema type present after JavaScript ran. Schema itself is nearly universal on both sides: 92% of cited pages carried it, and so did 97% of the ones that weren’t cited. Having it tells you nothing.

What we didn’t expect: the specific types most associated with “optimizing for search” skewed toward the pages that weren’t cited. AggregateRating appeared on 24% of the ignored pages and just 3% of the cited ones. SoftwareApplication ran 30% versus 5%. FAQPage itself was flat: 39% versus 35%, no real difference either way. The only type that leaned toward the cited pages was plain Article.

The likely explanation isn’t that schema causes exclusion. It’s that heavy rating/product markup is a tell for a page built to sell, and that’s the kind of page an answer engine skips over regardless of what’s declared in the JSON-LD.

Full breakdown, including the matched-pair example where the more heavily schema’d page lost to a plainer one on the identical query, is in our citation study.

  • Don’t chase schema as a citation lever.Add Organization and Article for hygiene; it won’t move whether you get cited.
  • Mark up what you already wrote.Structure the prose first: schema describes it, it doesn’t fix it.
  • Spend the time on consensus instead.Section 8 has the number: it separates winners from losers by a 7× margin. Schema doesn’t move the needle at all.

Markup is mechanical, and now measured. The research feeding it is where quality is actually won or lost.

16. Prompt Quality Caps Research Quality

If you use language models to map fan-outs, analyse competitors, and draft briefs, your prompting discipline sets a ceiling on your output. A model has no context by default, and vague inputs produce the generic content that answer engines pass over.

Every serious prompt carries three elements: context (who the audience is and what the goal is), constraints (what to leave out, not just what to include), and format (the shape of the output). Two amplifiers help further: assigning a role, and stating the task explicitly rather than letting it be inferred.

Assembled, that looks like: “Act as a skeptical buyer researching project management tools for a six-person remote design studio on a tight budget. You already use Slack and Figma. List the eight questions you would need answered before choosing, in order of importance, then flag which ones most vendors fail to answer clearly. Numbered list, one line each, no preamble.” That returns a sharper brief than most humans write from a keyword.

  • Specify what to exclude. Constraints do more work than instructions in practice.
  • Clear the thread when it drifts. Iterative correction pollutes context; the more you patch a bad answer in place, the less consistent it becomes.
  • Treat prompts as briefs. If it would be too thin to hand a freelancer, it is too thin to hand a model.

Better research produces better content. Proving it worked requires a different dashboard entirely.

17. Measurement Has to Change With the Channel

A discipline you cannot measure is one you cannot defend at budget time. The awkward reality of AEO and GEO is that influence can grow while sessions stay flat, which, given Pew’s finding that click-through nearly halves under an AI Overview, is the expected outcome rather than a failure.

Four layers are worth instrumenting.

  1. 1AI share of voice. A fixed query set run on a schedule across ChatGPT, Claude, Perplexity, and Gemini, tracking what share of answers mention you versus each competitor.
  2. 2Citation accuracy and sentiment. Appearing is not enough if the description is wrong. Log how accurately and how favourably you are characterised.
  3. 3SERP feature ownership. Which target queries trigger AI Overviews and snippets, and who currently holds them. These shift faster than classic rankings.
  4. 4Segmented AI referrals. Split ChatGPT and Perplexity traffic out of your generic referral bucket, and watch branded search as a secondary signal.

Whatever you track, freeze the query set. The temptation is to keep refining the questions, but the moment they change you lose month-over-month comparability. Lock the queries, vary only the date.

18. Funnel Stage Determines the Winning Format

Gap-hunting produces a pile of questions. Funnel mapping tells you which to write first, and in what shape. Sort any query set by intent and a grammar emerges.

How, what, and why questions are awareness, and they want educational prose. “Best X” is consideration, and it wants a comparison hub. “X vs Y”is evaluation, and it wants a table. Branded and “near me” queries are decision, and they want a landing page that converts.

Comparison tables punch far above their weight in AEO specifically. They are dense, structured, and unambiguous: exactly the format a model can lift wholesale without having to summarise anything. A well-built comparison table is often the single most-cited element on a page.

The stage almost everyone skips is pre-awareness, where the reader does not yet know they have the problem you solve. They are searching symptoms, not solutions: “why does my knee ache after running” rather than the name of any product category. That gap is usually wide open precisely because it looks unrelated to the sale.

  • Write decision-stage content first. It converts immediately and funds the rest of the programme.
  • Fill earlier stages for authority. Breadth across the cluster is what makes the commercial pages credible to a model.
  • Match format to stage.A comparison query answered in flowing prose loses to a competitor’s table every time.

You rarely need to invent these topics from scratch. Someone is usually already winning them.

19. Displacement Beats Invention

If a competitor already holds the citation you want, the concept is validated: you only need to out-answer them. That is a substantially cheaper bet than creating demand for a topic nobody has proven yet.

The loop we run is four steps. Research who currently wins the answer. Replicate the structure readers and models clearly reward. Reciprocate by matching their coverage and adding what they omit. Replace them by resolving the question more directly than they do.

Choosing the right target matters more than the technique. Sort competitors into tiers: the genuinely strong sources worth learning from, the mid-tier, and the ones ranking on thin, uncited content in a low-competition space. That third group is where the fast wins live.

When analysing a competitor with a language model, always use the exact URL rather than the brand name. Ambiguous company names are a leading cause of hallucinated competitor analysis, and a confidently wrong summary is worse than no summary. Ask what their primary USPs are, what subject areas they consistently cover, and how credible the source actually is.

  • Grade before you commit. Check whether their page is genuinely AEO-built or just an old SEO artifact ranking on domain strength.
  • Look for unsupported claims. Assertions without citations read as low-confidence to a model, and low confidence is an opening.
  • Check for self-contradiction. Competitors who state different facts on different pages are the easiest to displace.

One caveat applies to all of this: there is no single “the AI” to optimise for.

20. Different Models Retrieve, and Fail, Differently

Treating ChatGPT, Claude, Perplexity, and Gemini as one audience is the most common measurement mistake in GEO. They expand queries differently, weight sources differently, and hallucinate in different directions.

We’ve seen this show up directly in our own AI Overview research: the same underlying page can win a citation on one commercial query and be completely passed over on a near-identical one, purely because of which sources each surface’s retrieval step happened to weight that time. Measuring a single assistant and generalizing from it will systematically misread your actual visibility.

The behavioural differences are consistent enough to plan around. ChatGPT tends to accept a user’s framing and follow it. Claude pushes back and wants evidence before committing to a recommendation. Perplexity leans hardest on live retrieval and surfaces citations most aggressively, which makes it the best early-warning system for whether your content is actually indexed and quotable.

  • Track each surface separately. A gap in one is not a gap in all, and averaging them hides the actionable signal.
  • Use Perplexity as a canary. If it will not cite your page, the passage is probably not cleanly extractable yet.
  • Write for the sceptical reader. Content with evidence attached satisfies the models that demand it and loses nothing with the models that do not.

So What Should You Actually Do?

Some of this is genuinely unsettled. Citation behaviour differs between models, retrieval changes with every release, and nobody, including anyone selling you a guarantee, can promise a specific model cites you on a specific query.

But the direction is not ambiguous. Pew has quantified the click loss. Princeton has quantified which content changes earn citations. Our own client data shows the same tactics producing +318% keyword growth and 128 AI mentions from a standing start. The mechanics are knowable, and most competitors are still optimizing for a results page that increasingly answers on its own.

Start here.

  1. 1Baseline your AI share of voice this week. Fix a query set across ChatGPT, Claude, Perplexity, and Gemini, log who gets cited, and never change the questions: only the date.
  2. 2Audit your own contradictions before writing anything new. Reconcile conflicting claims across your homepage, docs, pricing, and third-party profiles.
  3. 3Add a statistic and a named source to your ten most valuable pages. The Princeton data says this is the single highest-return edit available.
  4. 4Rewrite one FAQ hub to the four-layer template. Direct answer, scope, optional brand mention, guardrail: then mark it up with FAQ schema.

None of this is gaming a model. It is being the clearest, best-evidenced answer to the exact questions your customers ask, structured so a machine can lift it without friction. That advantage does not expire with the next model update.

Frequently asked questions

AEO (Answer Engine Optimization) is about winning the answer inside a search results page: featured snippets, People Also Ask, and Google and Bing AI Overviews. GEO (Generative Engine Optimization) is about being cited or recommended inside standalone chat assistants like ChatGPT, Claude, Perplexity, and Gemini. Same mechanics, different surfaces.
Passages. Retrieval systems convert content into embeddings at the paragraph level and match a query against those chunks, not against whole pages. A self-contained paragraph that answers one question completely outperforms a long page where the answer is scattered.
Hunt hallucinations. When you ask a model a specific, multi-condition question and it guesses or gets it wrong, you have found a content gap where no strong source exists. Publishing the direct, well-attributed answer to that exact question is the most reliable way to earn a citation.
No. Nobody can guarantee that any specific model cites you on any specific query. What you can do is put yourself in the strongest realistic position: structurally clear, well-attributed, factually consistent content that answers the exact questions people ask.

The AI Visibility & GEO cluster

13 deep-dives that expand on the sections above.

The 10 Best AEO and GEO Agencies for SaaS in 2026 (Ranked, With Receipts)

I ranked the 10 best AEO and GEO agencies for SaaS with a disclosed methodology and real screenshots, including my own agency at #1. Here's why, and where it doesn't win.

Content That Gets Cited by AI: What 20 Real Searches Reveal

We ran 20 Google searches and read every cited source. Real examples of the content that gets cited by AI, broken down by the structure that won each slot.

KPIs for AEO and GEO: The 7 Numbers That Actually Measure AI Visibility

A rankings report can't see AI visibility. Here are the 7 KPIs that can: AI share of voice, citation/adoption rate, consensus coverage, and more, backed by 175+ live searches and Princeton's GEO study.

Why ChatGPT Recommends Your Competitor Instead of You

The specific reasons ChatGPT names a competitor when someone asks for a tool like yours, and the fixes that put you in the recommendation set.

How to Get Cited by Perplexity: A Practical Guide

Perplexity cites its sources on every answer, which makes it the best place to earn and verify AI citations. Here is exactly how to become one of those sources.

How to Win a Google AI Overview Citation

Google's AI Overview lifts answers from pages it trusts. Here is how to structure content so yours is the passage it quotes, and the citation it links.

AI Mode vs AI Overviews: What's Different for Brands

Google's AI Mode and AI Overviews are different surfaces with different rules. Here is what separates them and what each means for your visibility strategy.

Entity SEO: Making AI Understand Who You Are

Before AI can recommend you, it has to understand you as an entity. Entity SEO is how you make the web tell one clear, consistent story about your brand.

Wikidata & Knowledge Panels for Brand Entities

Wikidata and the Google knowledge panel feed directly into how AI understands your brand. Here is how to claim and shape both, the right way.

How Often Should You Re-Run an AI Visibility Audit?

AI answers shift faster than search rankings. Here is how often to re-run an AI visibility audit, and what to check each time, without wasting effort.

Why AI Describes Your Brand Wrong (and How to Fix It)

When ChatGPT or Gemini gets your brand wrong, it is a content and consistency gap, not bad luck. Here is how to diagnose the cause and correct it at the source.

GEO vs SEO: What Actually Changes in Your Workflow

GEO and SEO share most of their mechanics but differ where it counts. Here is what actually changes in your day-to-day workflow when you optimize for both.

How to Get Cited in Google’s AI Overview for “Best Software” Searches (New Data)

I ran 70 real 'best software' searches to reverse-engineer what Google's AI Overview actually cites. The consensus rule, the narrow-slot rule, and 5 plays to get your product named.

Want this run on your brand?

A 30 minute call on your business, then a free opportunity check. If I can't move your numbers, I'll tell you and you'll have paid nothing.

Book a Free Fit Call
FreeNo obligationHonest answer either way