For two decades, SEO had one job: get the page to position one. You researched keywords, built the best page on the internet for them, earned links, and waited for the click.
Today, that no longer works. Google now resolves a large share of searches on the results page itself, and a growing volume of research never touches a search engine at all: it happens inside ChatGPT, Claude, Perplexity, and Gemini. The page still matters. The click is disappearing.
Which forces a different unit of optimization. Answer engines cite passages, not pages. They break the web into paragraph-sized chunks, match a question against those chunks, and rebuild an answer from the handful they trust. Your 3,000-word guide is not the thing being ranked. One paragraph inside it is.
Everything that follows reduces to three moves:
- Find the gap. Locate the questions where no strong answer exists yet.
- Answer it in one liftable block. Self-contained, answer-first, quotable without context.
- Prove it. Statistics, named sources, and claims that agree with the rest of the web.
1. AI Overviews Cut Click-Through Rates Nearly in Half
The clearest evidence that the click is collapsing comes from outside the SEO industry. In July 2025, Pew Research Center tracked the actual browsing behaviour of 900 US adults who agreed to share their activity.
When an AI summary appeared, users clicked a traditional result 8% of the time. When no summary appeared, they clicked 15% of the time. Clicks on a link inside the AI Overview itself: 1%.
Three things follow from that gap.
- Rankings decouple from traffic. You can hold position one and watch sessions fall, because the answer above you resolved the query first.
- Sessions end sooner. Pew found users ended their browsing session 26% of the time after seeing an AI summary, versus 16% without one.
- Being named is the new ranking. If the summary resolves the question, appearing inside it is the only visibility left on that search.
That shift only makes sense once you understand what the engine is actually assembling an answer from.
2. Answer Engines Retrieve Passages, Not Pages
Retrieval systems do not evaluate your page as a whole. They split content into small units, convert each into a vector, and match the query against those vectors. The chunk is the atom of AEO, not the URL.
The Princeton and IIT Delhi team behind the first formal study of this behaviour built their benchmark, GEO-bench, around roughly 10,000 queries across nine datasets, and measured visibility using a position-adjusted word count of how much of your content survived into the generated answer. Not whether you ranked. How much of you got quoted.
The metric that matters in generative search is not position. It is how many of your words the model chose to reuse. Aggarwal et al., “GEO: Generative Engine Optimization,” KDD 2024
- Length stops being authority. A tight paragraph that resolves the question outranks a 2,500-word guide burying the answer in section seven.
- Structure becomes a ranking factor. Question-named headings and answer-first paragraphs make a passage mechanically easier to lift.
- Context-dependent prose gets skipped.A paragraph opening “as mentioned above” cannot travel, so it never gets retrieved.
We tested this ourselves rather than taking it on faith. Across 49 live Google searches spanning health, finance, travel, food, legal, and B2B topics, we found the exact sentence each AI Overview echoed on the cited page, then compared it against a nearby paragraph on the same page that was ignored. The split was almost total:
The clearest way to see this is to watch an engine assemble an answer in public.
3. One Synthesized Answer Can Outrank Ten Blue Links
We searched “what is answer engine optimization” on Google in July 2026. The AI Overview returned a complete definition above every organic result, and cited a Reddit thread from r/localseo as a primary source, alongside a vendor glossary page.
A forum comment outranked the SEO industry’s own publishers on its own terminology. That is worth sitting with.
- Clarity beats domain authority. The Reddit answer won because it defined the term plainly in its opening sentence, not because the domain outranked anyone.
- Community content is training data. Reddit, G2, and YouTube comments are heavily represented in what these systems retrieve.
- Citations are winnable. If a forum post can take the slot, a properly structured page can take it back.
Underneath that behaviour sits a three-stage pipeline worth knowing precisely.
Query processing strips the question to key terms, then expands it with assumed context: Claude branches into beginner and expert readings, Perplexity infers location. Knowledge retrieval matches the expanded query against stored chunks. Answer generation merges the winners into prose. Knowing how a model expands a query tells you which angles to cover.
4. Adding Statistics and Quotations Lifts AI Visibility by up to 37%
Most AEO advice is assertion. The Princeton GEO study is the rare piece of controlled evidence: nine content-modification strategies, tested across ~10,000 queries, each measured against an unoptimized baseline.
Statistics Addition improved visibility by 37%. Quotation Addition improved position-adjusted visibility by 22%. Citing sources landed in the same 30–40% band. The three strongest methods all clustered there.
- Numbers are extraction handles. A model reaching for a defensible claim takes the sentence carrying a figure over the one carrying an adjective.
- Attribution transfers trust.Naming your source lets the model inherit that source’s credibility instead of gambling on yours.
- Fluency is measurable. Clear, well-formed prose scored materially better than dense or hedged writing across the same queries.
The same experiment also found a tactic that actively backfires.
5. Keyword Stuffing Now Performs Worse Than Doing Nothing
Of the nine strategies tested, keyword stuffing was the only one to score below the control. It came in 10% under the unoptimized baseline. Adding keywords made pages measurably less likely to be cited than leaving them alone.
- Density signals nothing. Embeddings encode meaning, so repetition adds no semantic information the model can use.
- Stuffing degrades fluency. It damages the readability that the same study found improves citation rates.
- The old playbook has negative value. Not neutral. Negative. Legacy SEO artifacts are now a liability worth auditing out.
Keyword stuffing is the only tested tactic that performs worse than publishing nothing at all.
If old tactics lose ground, the question becomes where the genuinely winnable topics are.
6. Hallucinations Mark the Gaps Worth Writing
A hallucination is not only a model failure. It is a signal that no strong, consistent source exists for that question, so the model guessed. Every guess marks an unclaimed citation.
A useful test: run your target query through a model yourself and read the answer for hedges, vague generalities, or an outright wrong specific. Each one is a marker that no strong source exists yet for that exact question, which is exactly the gap a well-sourced, specific answer can fill.
- Specificity beats volume. Narrow, multi-condition questions have far fewer good answers online than broad ones, so the bar to become the best available answer is lower.
- Regulated niches are still winnable.YMYL categories (health, finance, legal) reward sourcing and precision; they don’t categorically punish smaller sites that get the specifics right.
- Coverage precedes citation. A model can only cite a passage that exists. Depth on a topic has to come before AI visibility on it can follow.
To find those gaps at scale, ask a model to decompose a query the way it decomposes it internally.
7. Hub-and-Spoke Content Compounds Faster Than Isolated Posts
A broad query is never one question. Behind “best CRM for small business” a model silently generates six to eight sub-questions and answers them before responding. Mapping that fan-out is the highest-leverage thing you can do before writing.
Independent research on content clustering points the same direction. According to a 2026 review of topic-cluster case studies compiled by Digital Applied, successful clusters typically run 8–12 spokes per pillar, and clustered content reportedly earns 3.2× more AI citations than standalone posts covering the same topics individually: the mechanism being that a fully-covered topic gives a retrieval system more distinct, matchable chunks to draw an answer from.
- Clusters match more embeddings. Answering every sub-question makes one asset eligible across a whole family of related retrievals.
- Hubs concentrate authority. Internal links from spokes back to the hub make the hierarchy legible to a retrieval system.
- Breadth buys credibility. Comprehensive coverage of a topic earns more trust on the commercial query at the bottom of the funnel.
Coverage gets you retrieved. Three other filters decide whether you get used.
8. Consistency, Consensus, and Context Decide Who Gets Cited
Once several candidate passages are retrieved, the model has to choose. Three filters do most of that work, and they are cumulative: failing one is enough to be passed over.
We tested the consensus half of this directly. Across 129 live searches where a vendor ranked its own product #1 on its own blog, Google adopted that self-ranking into its generated answer 60% of the time, but almost always only when independent sources (review sites, Reddit, competitor pages) already named the same vendor. In one case, the winning line was sourced to a direct competitor’s page, not the vendor’s own. Full breakdown in our citation study.
- Consistency is cheapest. Reconciling contradictory claims across your own pages costs nothing and removes the fastest reason to distrust you.
- Consensus replaces backlinks. Independent third parties echoing your facts is the modern equivalent of an earned link.
- Context beats comprehensiveness. Content written for a named segment matches intent better than content written to offend nobody.
Passing those filters still depends on how the passage itself is built.
9. Self-Contained Chunks Beat Comprehensive Pages
The test for any paragraph is simple: would it still make sense as the only thing someone read? If it opens with “this is why it matters,” it fails: it is tethered to its neighbours and cannot be lifted.
This follows directly from how retrieval works, described above: a system matching query vectors against passage vectors can only surface a passage that is itself a complete, self-contained match. A paragraph that depends on the one before it to make sense is structurally invisible to that process, no matter how good the page around it is.
The cleanest proof we found sits on one page. Microsoft’s security glossary defines ransomware twice: once in the flowing body copy, once in a boxed “Key takeaways” callout a few lines down. Same fact, same author, same page: Google’s AI Overview lifted only the boxed version.
- Front-load the answer. Lead with the conclusion, then add nuance, so the first sentence can stand alone.
- One idea per block. A chunk resolving a single sub-question retrieves far better than one covering three.
- Fix structure before adding content. A page with broken chunking gets no benefit from more words; restructuring the existing content usually matters more than expanding it.
Two more from the same 49-search sample, both lifted almost verbatim because the source sentence needed no context to stand alone:


One content type is naturally shaped this way, and it is badly underused, almost.
10. A Labeled Answer Block Beats a Rhetorical Question
The common advice is “add an FAQ section.” Our data says that’s half right. Literal question-and-answer shaping helped: it showed up in roughly 30% of the lifted passages in our 49-search sample and in none of the ignored controls, but a rhetoricalquestion (“Ever wonder why...”) was a negative signal every time we saw one. The engine wants a direct answer, not a hook.
The strongest version of this we found wasn’t FAQ markup at all. It was a plainly labeled summary blocknear the top of the page (“Key Takeaways,” “Quick Answer,” “The Short Answer”), and that exact boxed line getting lifted whole. We saw it on Fidelity (“Key takeaways”), Healthline (“Key takeaways”), the CDC (“Key Points”), and a Walser auto-service page that literally headed its own answer “The Short Answer.”
Compare two answers to the same question.
Weak: “Setup time varies depending on a number of factors and your needs.” Unliftable: it answers nothing.
Citable: “Most teams are fully set up in two to four weeks. That covers data import, pipeline configuration, and integrations; a solo user with clean data can be live in a day. If you are migrating from another CRM, add a week for mapping custom fields.” Any sentence can be quoted and remain correct.

- Lead with the unqualified answer. No throat-clearing before the first useful sentence.
- Label the summary if you can.A visibly boxed “Key Takeaways”/“Quick Answer” line outperformed the identical fact left unboxed in body copy, on the same page, every time we found the pair.
- Skip the rhetorical hook. A question posed for effect, rather than answered immediately, was never the passage that got lifted.
All of which raises the fair question of whether traffic from these surfaces is worth anything.
11. AI Referrals Convert Better Than Organic Search
The volume is small. The quality is not. Semrush analysed more than 1 billion lines of US clickstream data and found ChatGPT outbound referral traffic grew 206% year over year between January 2025 and January 2026, reaching roughly 170,000 unique referred domains per month.
On conversion, Seer Interactive’s case-study data puts ChatGPT-referred traffic at around 16% against roughly 1.8%for Google organic. Similarweb’s 2026 clickstream figures are more conservative at 7.1%, still second only to paid search. Both point the same direction.
- Intent arrives pre-qualified. Someone acting on a recommendation has already had their objections handled by the model.
- Small numbers still pay. At a 7–16% conversion rate, a few hundred visits can outperform several thousand organic ones.
- Segment it now. If AI referrals sit inside a generic bucket in GA4, you cannot prove any of this internally.
12. Entity Consistency Is the Cheapest Trust Win Available
Underneath consensus sits a concept that separates people who dabble in GEO from people who move the number: the entity. To a model your brand is not a website. It is a node in a knowledge graph, with attributes attached: what you do, who you serve, where you operate, whether you are credible.
Every mention of you anywhere is training data for that node. When your G2 profile, LinkedIn, Crunchbase entry, and your own homepage describe you identically, the model forms a confident picture. When they disagree, the node stays fuzzy, and fuzzy entities do not get recommended.
This is why running the same audit across many sites tends to surface the same root cause repeatedly: not a shortage of good content, but contradictory titles, headers, and descriptions accumulated across pages that were each written in isolation. Reconciling that is a QA problem, not a creative one, and it’s usually the highest-leverage fix available before writing anything new.
- Audit before you publish. Reconciling your existing footprint costs a day and removes the fastest reason a model has to distrust you.
- Claim the structured records. Wikidata entries and knowledge panels feed directly into how both search engines and models resolve who you are.
- Standardise the description. One sentence describing your business, repeated verbatim everywhere, beats ten well-written variations.
Consistency across properties tells a model who you are. Internal links tell it what you consider important.
13. Internal Links Tell a Retrieval System What Matters
Internal linking is the most underrated lever in AEO because it is the clearest structural signal you fully control. It defines hierarchy, concentrates authority, and makes topical relationships explicit rather than inferred.
A common failure pattern: the product and the content are both strong, but authority pools in the wrong places because supporting pages never link back to the pages meant to convert or be cited. Rebuilding the link graph around a clear hub-and-spoke hierarchy, every spoke pointing back to its hub, consolidates relevance signals onto the pages that actually matter, instead of scattering them evenly across the site.
- Link from spokes to hubs. Authority should flow toward the page you actually want retrieved, not scatter evenly across the site.
- Use descriptive anchors.Anchor text is a labelled edge in the graph; “click here” labels nothing.
- Audit orphans first. A page nothing links to reads as a page nothing vouches for.
Structure you build carefully can still be destroyed in a single afternoon by a platform change.
14. Migrations Are Where SEO Equity Quietly Dies
Nothing erases years of accumulated authority faster than an unmanaged replatform. Broken URLs, missing redirects, lost metadata, and altered architecture can vaporise rankings in days, and the damage usually surfaces weeks later when the traffic report lands.
Protecting against it during a platform migration (Magento to Shopify, a domain change, a CMS switch) is unglamorous but entirely preventable: map and implement URL redirects, preserve metadata and content structure, then audit post-migration crawlability and indexation before declaring anything finished. Done properly, nothing dramatic happens during the switch, and that absence of drama is the actual deliverable.
- Map redirects before launch, not after. A redirect written in response to a traffic drop has already cost you the traffic.
- Preserve the chunk, not just the URL. If the answer paragraph gets reformatted into a widget, you keep the ranking and lose the citation.
- Re-audit indexation post-launch. Confirm the new structure is actually being crawled rather than assuming the redirect map covered it.
Once the structure survives, you can make it explicit rather than leaving it to inference.
15. Schema Is Table Stakes, Not a Lever
This is the one place our own research overturned the standard advice, including an earlier version of this section. The usual claim is that FAQPage and similar schema help a passage get cited. We tested it properly and it doesn’t hold up.
We loaded 70 vendor “best X software” pages in a real, rendered browser (37 that ended up cited in Google’s AI Overview, 33 that didn’t), and pulled every schema type present after JavaScript ran. Schema itself is nearly universal on both sides: 92% of cited pages carried it, and so did 97% of the ones that weren’t cited. Having it tells you nothing.
What we didn’t expect: the specific types most associated with “optimizing for search” skewed toward the pages that weren’t cited. AggregateRating appeared on 24% of the ignored pages and just 3% of the cited ones. SoftwareApplication ran 30% versus 5%. FAQPage itself was flat: 39% versus 35%, no real difference either way. The only type that leaned toward the cited pages was plain Article.
The likely explanation isn’t that schema causes exclusion. It’s that heavy rating/product markup is a tell for a page built to sell, and that’s the kind of page an answer engine skips over regardless of what’s declared in the JSON-LD.
Full breakdown, including the matched-pair example where the more heavily schema’d page lost to a plainer one on the identical query, is in our citation study.
- Don’t chase schema as a citation lever.Add Organization and Article for hygiene; it won’t move whether you get cited.
- Mark up what you already wrote.Structure the prose first: schema describes it, it doesn’t fix it.
- Spend the time on consensus instead.Section 8 has the number: it separates winners from losers by a 7× margin. Schema doesn’t move the needle at all.
Markup is mechanical, and now measured. The research feeding it is where quality is actually won or lost.
16. Prompt Quality Caps Research Quality
If you use language models to map fan-outs, analyse competitors, and draft briefs, your prompting discipline sets a ceiling on your output. A model has no context by default, and vague inputs produce the generic content that answer engines pass over.
Every serious prompt carries three elements: context (who the audience is and what the goal is), constraints (what to leave out, not just what to include), and format (the shape of the output). Two amplifiers help further: assigning a role, and stating the task explicitly rather than letting it be inferred.
Assembled, that looks like: “Act as a skeptical buyer researching project management tools for a six-person remote design studio on a tight budget. You already use Slack and Figma. List the eight questions you would need answered before choosing, in order of importance, then flag which ones most vendors fail to answer clearly. Numbered list, one line each, no preamble.” That returns a sharper brief than most humans write from a keyword.
- Specify what to exclude. Constraints do more work than instructions in practice.
- Clear the thread when it drifts. Iterative correction pollutes context; the more you patch a bad answer in place, the less consistent it becomes.
- Treat prompts as briefs. If it would be too thin to hand a freelancer, it is too thin to hand a model.
Better research produces better content. Proving it worked requires a different dashboard entirely.
17. Measurement Has to Change With the Channel
A discipline you cannot measure is one you cannot defend at budget time. The awkward reality of AEO and GEO is that influence can grow while sessions stay flat, which, given Pew’s finding that click-through nearly halves under an AI Overview, is the expected outcome rather than a failure.
Four layers are worth instrumenting.
- 1AI share of voice. A fixed query set run on a schedule across ChatGPT, Claude, Perplexity, and Gemini, tracking what share of answers mention you versus each competitor.
- 2Citation accuracy and sentiment. Appearing is not enough if the description is wrong. Log how accurately and how favourably you are characterised.
- 3SERP feature ownership. Which target queries trigger AI Overviews and snippets, and who currently holds them. These shift faster than classic rankings.
- 4Segmented AI referrals. Split ChatGPT and Perplexity traffic out of your generic referral bucket, and watch branded search as a secondary signal.
Whatever you track, freeze the query set. The temptation is to keep refining the questions, but the moment they change you lose month-over-month comparability. Lock the queries, vary only the date.
18. Funnel Stage Determines the Winning Format
Gap-hunting produces a pile of questions. Funnel mapping tells you which to write first, and in what shape. Sort any query set by intent and a grammar emerges.
How, what, and why questions are awareness, and they want educational prose. “Best X” is consideration, and it wants a comparison hub. “X vs Y”is evaluation, and it wants a table. Branded and “near me” queries are decision, and they want a landing page that converts.
Comparison tables punch far above their weight in AEO specifically. They are dense, structured, and unambiguous: exactly the format a model can lift wholesale without having to summarise anything. A well-built comparison table is often the single most-cited element on a page.
The stage almost everyone skips is pre-awareness, where the reader does not yet know they have the problem you solve. They are searching symptoms, not solutions: “why does my knee ache after running” rather than the name of any product category. That gap is usually wide open precisely because it looks unrelated to the sale.
- Write decision-stage content first. It converts immediately and funds the rest of the programme.
- Fill earlier stages for authority. Breadth across the cluster is what makes the commercial pages credible to a model.
- Match format to stage.A comparison query answered in flowing prose loses to a competitor’s table every time.
You rarely need to invent these topics from scratch. Someone is usually already winning them.
19. Displacement Beats Invention
If a competitor already holds the citation you want, the concept is validated: you only need to out-answer them. That is a substantially cheaper bet than creating demand for a topic nobody has proven yet.
The loop we run is four steps. Research who currently wins the answer. Replicate the structure readers and models clearly reward. Reciprocate by matching their coverage and adding what they omit. Replace them by resolving the question more directly than they do.
Choosing the right target matters more than the technique. Sort competitors into tiers: the genuinely strong sources worth learning from, the mid-tier, and the ones ranking on thin, uncited content in a low-competition space. That third group is where the fast wins live.
When analysing a competitor with a language model, always use the exact URL rather than the brand name. Ambiguous company names are a leading cause of hallucinated competitor analysis, and a confidently wrong summary is worse than no summary. Ask what their primary USPs are, what subject areas they consistently cover, and how credible the source actually is.
- Grade before you commit. Check whether their page is genuinely AEO-built or just an old SEO artifact ranking on domain strength.
- Look for unsupported claims. Assertions without citations read as low-confidence to a model, and low confidence is an opening.
- Check for self-contradiction. Competitors who state different facts on different pages are the easiest to displace.
One caveat applies to all of this: there is no single “the AI” to optimise for.
20. Different Models Retrieve, and Fail, Differently
Treating ChatGPT, Claude, Perplexity, and Gemini as one audience is the most common measurement mistake in GEO. They expand queries differently, weight sources differently, and hallucinate in different directions.
We’ve seen this show up directly in our own AI Overview research: the same underlying page can win a citation on one commercial query and be completely passed over on a near-identical one, purely because of which sources each surface’s retrieval step happened to weight that time. Measuring a single assistant and generalizing from it will systematically misread your actual visibility.
The behavioural differences are consistent enough to plan around. ChatGPT tends to accept a user’s framing and follow it. Claude pushes back and wants evidence before committing to a recommendation. Perplexity leans hardest on live retrieval and surfaces citations most aggressively, which makes it the best early-warning system for whether your content is actually indexed and quotable.
- Track each surface separately. A gap in one is not a gap in all, and averaging them hides the actionable signal.
- Use Perplexity as a canary. If it will not cite your page, the passage is probably not cleanly extractable yet.
- Write for the sceptical reader. Content with evidence attached satisfies the models that demand it and loses nothing with the models that do not.
So What Should You Actually Do?
Some of this is genuinely unsettled. Citation behaviour differs between models, retrieval changes with every release, and nobody, including anyone selling you a guarantee, can promise a specific model cites you on a specific query.
But the direction is not ambiguous. Pew has quantified the click loss. Princeton has quantified which content changes earn citations. Our own client data shows the same tactics producing +318% keyword growth and 128 AI mentions from a standing start. The mechanics are knowable, and most competitors are still optimizing for a results page that increasingly answers on its own.
Start here.
- 1Baseline your AI share of voice this week. Fix a query set across ChatGPT, Claude, Perplexity, and Gemini, log who gets cited, and never change the questions: only the date.
- 2Audit your own contradictions before writing anything new. Reconcile conflicting claims across your homepage, docs, pricing, and third-party profiles.
- 3Add a statistic and a named source to your ten most valuable pages. The Princeton data says this is the single highest-return edit available.
- 4Rewrite one FAQ hub to the four-layer template. Direct answer, scope, optional brand mention, guardrail: then mark it up with FAQ schema.
None of this is gaming a model. It is being the clearest, best-evidenced answer to the exact questions your customers ask, structured so a machine can lift it without friction. That advantage does not expire with the next model update.