Generative Engine Optimization: What Is Supported by Evidence and What Remains Unclear

Cover image placeholder

GEO Optimization: Between Facts and Hype

You optimize your content for Google. But your audience is increasingly asking ChatGPT, Perplexity, and Gemini the same questions instead. And you have no idea whether you even show up there.

The scale of this is measurable now: BrightEdge finds AI Overviews present in roughly 48 percent of tracked search queries in 2026 — up 58 percent year over year. In parallel, a whole market has grown up around Generative Engine Optimization within two years, complete with checklists, special files, and tactics sold as secret levers. Some of it is well studied, some is a plausible guess, and some has no evidence behind it at all.

This article keeps three things strictly separate: what’s measured, what Google explicitly says is unnecessary, and what simply has no evidence either way. No blanket debunking. No number without a source, a sample size, and a date.

The short version

  • GEO isn’t a break from SEO. Google said so itself on May 15, 2026: from Google Search’s perspective, optimizing for generative AI search is optimizing for the search experience — and is therefore still SEO.
  • Not every AI system can be reached. Systems with live web access respond to SEO work as soon as your content is indexed and ranked. Pure language models can only be reached through the training-data path — and that’s tied to the model’s release cycle, not to how fast you work.
  • The link between Google rankings and AI citations is real, but platform-dependent. It’s strong for Google’s own AI formats, much weaker for ChatGPT. A one-size-fits-all strategy falls short.
  • The weighting is shifting. Brand mentions correlate markedly more strongly with AI visibility than plain backlink metrics. Topical breadth beats a single primary keyword.
  • AI visibility has been partly measurable natively since June 3, 2026. Search Console now shows impressions from AI Overviews and AI Mode separately — without clicks or search queries.

What is Generative Engine Optimization?

Generative Engine Optimization means optimizing web content to appear as a source in the answers of generative AI systems — not as a link in a results list, but as the basis for the answer itself.

The term comes from a paper by researchers at Princeton, Georgia Tech, the Allen Institute for AI, and IIT Delhi. It appeared as a preprint in November 2023 and was presented at the ACM SIGKDD conference in 2024. The authors group systems that assemble an answer from multiple sources and phrase it with a language model under the label “generative engines” — and describe GEO as a way to increase the visibility of your own content within those answers. Their explicit starting point is the position of content creators, who have almost no influence over whether and how they appear once an answer is synthesized.

In the DACH market, GEO is the most commonly used umbrella term. A handful of other acronyms circulate alongside it, and no consensus definition exists — not in the industry, and not in the academic literature either. In practice, the terms are used largely interchangeably.

TermAngleDistinction, where it can be substantiated
GEO – Generative Engine OptimizationAppearing as a source in a synthesized answerThe only term with an academic origin; names the actual mechanism (retrieval and synthesis) and is the leading term in the DACH region
LLMO – Large Language Model OptimizationThe language model itself as the targetAn industry coinage; foregrounds the model layer and the training-data path
AEO – Answer Engine OptimizationBeing the answer, in whatever formatThe oldest of the terms, dating to the era of featured snippets, position zero, and voice assistants; today usually read as a subset of GEO
AI SEOColloquial umbrella termNo methodology of its own, but the phrase that lands most reliably in a client conversation

You’ll also occasionally run into GAIO, AIO, or AISO. They describe the same field and never caught on; AIO is additionally ambiguous, since the same abbreviation also stands for “All-in-One.”

You won’t find precise overlap percentages between these terms here. Claims like “these disciplines overlap 90 percent” have no reliable source behind them. What can be substantiated is the qualitative point: the levers overlap heavily, and all of them assume a working SEO foundation. Whichever acronym you use, the work itself doesn’t change.

Which AI systems can you actually reach?

This distinction determines whether your work pays off in weeks or in years. Most GEO guides skip right past it. Behind it sit three technically very different system types.

Classic language models (LLMs)

They’re built for dialogue with you and formulate coherent, context-aware answers from what they learned during training. Their knowledge is frozen: every model has a fixed knowledge cutoff. Whatever happened after that, it doesn’t know — new content only enters the model through a new training run or a successor model.

Knowledge intake here is static. A chatbot in default mode, with web search not enabled, is one example.

Classic search engines (information retrieval)

Information-retrieval systems draw on content from external document collections, databases, or knowledge stores — static or continuously updated. Unlike a language model, which draws on training data, they search existing collections on demand and return what they found: an ordered list of documents, results, or text passages, usually with a source, a date, and a relevance score.

What they don’t do: formulate a coherent answer from multiple findings. Evaluating and synthesizing the results stays your job.

RAG systems (Retrieval-Augmented Generation)

These hybrids combine both. The approach traces back to a 2020 paper by Lewis et al., which describes the knowledge stored in the model as parametric memory and the retrieved documents as non-parametric memory — exactly the dividing line between the first two system types. A RAG system pulls current information through a search function and has a language model turn it into a coherent answer. Because they draw on both static training and dynamic sources, their answers are more current and stay closer to verifiable facts.

This is where Google AI Overviews and AI Mode belong, along with Perplexity, ChatGPT Search, Microsoft Copilot, and Gemini with search enabled. Google names the mechanism explicitly in its own documentation: retrieval-augmented generation and query fan-out — breaking your query into several sub-queries, each with its own sources retrieved.

Many assistants switch between modes depending on the query, by the way: for current or very specific questions, search kicks in; for general knowledge, the model answers on its own.

Why only retrieval-based systems can be optimized

Here’s the key insight: you can only influence what actively looks something up. RAG systems and assistants with web access pull in external content — and that’s exactly the content you control. In a pure language model, by contrast, the knowledge is already baked in and can no longer be changed after the fact.

That produces two completely different time horizons. Caution is warranted here, because a lot of numbers circulate for which no actual measurement exists.

Via retrieval, your work pays off as fast as your content gets indexed and ranked — it’s the same content classic search already draws on. No study exists yet on how long it takes from publication to a first AI citation. From project work, we know it’s weeks to a few months; that’s a field observation, not a measurement, and this article doesn’t treat it as one.

Via the training-data path, it structurally takes longer, and at least part of that can be quantified. Historically, the gap between a model’s knowledge cutoff and its release has typically run six to eighteen months. For current flagship models, that gap has shrunk noticeably, while older models still running inside integrations lag years behind — GPT-4o, for instance, with a cutoff in October 2023.

The limits of that number matter: it describes the model’s release cycle, not your brand. No provider publishes whether or when content actually gets pulled into a training corpus. How long it takes for a brand to show up in training-based answers is therefore not known — only that it can’t happen faster than the cycle of whichever model you’re up against.

That’s good news for anyone already doing SEO: practically every AI system currently gaining meaningful market share runs on retrieval — and draws on exactly the content that’s already classically indexed.

And the honest limitation: for the training-data path, there’s no shortcut. No tool and no technical fix gets your brand into a model on a short timeline. Anyone who promises you that is selling something they can’t deliver.

Does your SEO work pay off in AI visibility?

Yes — but how much depends on the system. And this is exactly where most articles get sloppy, picking a single overlap number and generalizing it across the board.

The research, laid out:

StudyFindingScope and date
AhrefsGoogle AI Overviews (running an adapted Gemini 2.5 model at the time): 76% of citations came from Google’s top 10, 9.5% from positions 11–100, 14.4% not in the top 100 at all1.9M citations from 1M AI Overviews, July 2025
AhrefsGoogle AI Overviews (running Gemini 3 globally since late January 2026): down to just 38% from the top 10; 31.2% from positions 11–100, 31.0% beyond position 100863,000 keywords, 4M AIO URLs, February 2026
BrightEdgeGoogle AI Overviews: share of citations from the organic top 10 held steady at roughly 17% across the entire measurement periodGenerative Parser, 9 industries, February 2025 to February 2026
BrightEdgeGoogle AI Overviews / SGE (period spans Gemini 1.5 through 2.5): overlap with organic top-100 rankings rose from 32% to 54%; healthcare 75.3%, education 72.6%, insurance 68.6%, e-commerce 22.9%9 industries, May 2024 to September 2025
BrightEdge (via Search Engine Land)Perplexity (model not disclosed): 60% overlap with Google’s top 10; healthcare 82%, restaurants 27%Generative Parser, sample size not published, April 2024
AhrefsPerplexity, ChatGPT, Gemini, and Copilot (model versions not disclosed): Perplexity roughly 29%; ChatGPT, Gemini, and Copilot combined roughly 12%15,000 long-tail prompts, August 2025
Seer InteractiveSearchGPT / ChatGPT with web search (model not disclosed): 87% match with Bing’s top 20; only 56% against Google’s top 20Just over 500 citations, February 2025
SE RankingChatGPT (model version and mode not disclosed): number of referring domains is the single strongest factor for citations129,000 domains, 216,524 pages, 20 niches, Nov. 2025 / March 2026

These numbers only look like they contradict each other. They’re measuring different things: sometimes Google’s own formats, sometimes third-party systems, sometimes against Google, sometimes against Bing. The two BrightEdge rows aren’t measuring the same thing either — the 17 percent refers to the top 10, the 54 percent to organic rankings overall, including positions 11 through 100. And none of the eight studies discloses which model version it measured against. “AI Overviews in July 2025” and “AI Overviews in February 2026” share a name but ran on different models: in late January 2026, Google made Gemini 3 the global default for AI Overviews.

Three numbers deserve an explicit caveat. Seer Interactive’s 87 percent rests on just over 500 citations — a small sample, and it measures against Bing’s top 20, not Google’s. The same study also checked against Google and landed at 56 percent; that’s the figure comparable to the other rows. There are also indications ChatGPT has since shifted its retrieval sources. The 60 percent for Perplexity dates from 2024, and the more recent Ahrefs measurement sits markedly lower, at roughly 29 percent — though Ahrefs tested exclusively long-tail prompts, so part of that gap is methodological, not just a matter of time. And BrightEdge’s 17 percent isn’t a counterweight to the Ahrefs decline: BrightEdge describes that share as essentially unchanged across its entire measurement window, so it isn’t seeing a drop at all, just a consistently low baseline.

The defensible summary, then: for Google’s AI formats, your organic visibility remains the strongest lever among the known ones — but a weaker one than in 2025, since depending on the measurement, a top-10 position now explains only 17 to 38 percent of citations. For Perplexity, the link sits in the middle; for ChatGPT, it’s weak relative to Google — there, the path runs more through Bing, referring domains, and brand presence. Anyone who applies a single percentage across every system is measuring wrong.

What’s shifting right now

The top-10 overlap is falling

From 76 percent in July 2025 to 38 percent in February 2026, measured by Ahrefs across 863,000 keywords and 4 million AIO URLs. What’s more revealing than the headline is the shift underneath it: in 2025, only 9.5 percent of citations came from positions 11 through 100, and 14.4 percent from pages outside the top 100. In 2026, the rest splits almost evenly — 31.2 percent from positions 11 through 100, 31.0 percent from pages that didn’t rank in the top hundred results at all.

One caveat belongs here, and it comes from Ahrefs itself: part of the decline traces back to improved citation detection in its own tooling. To some extent, the second survey is measuring differently from the first, not just measuring something different. And the second available data series doesn’t support the decline: BrightEdge puts top-10 overlap at roughly 17 percent and describes that share as essentially unchanged across its entire measurement window. Two vendors, two methodologies — a decline at Ahrefs, a consistently low baseline at BrightEdge. What they agree on is the outcome: a top-10 position now explains only a fraction of citations.

As an explanation, Ahrefs points to query fan-out alongside the measurement change: because the system asks several semantically related sub-questions per query, it pulls in sources that don’t even rank on page 1 for the original query. Coverage has also tied this to timing — in late January 2026, Google made Gemini 3 the default model for AI Overviews worldwide. Google hasn’t confirmed any change to fan-out behavior in connection with that. What’s established so far is the shift itself, not its cause.

An AirOps analysis shows how strong the fan-out effect is — though it looks at ChatGPT, not Google’s AI Overviews. Its 15,000 prompts expanded into 43,233 search queries; 89.6 percent of queries triggered at least two sub-queries, and among cited pages that showed up in a top-20 SERP at all, 32.9 percent were findable exclusively through a sub-query. The same analysis also provides the flip side of the 12 percent figure above: measured not just against the original query but against all sub-queries combined, 55.8 percent of the pages ChatGPT cited rank in Google’s top 20. The link between ranking and citation is larger than the overlap numbers suggest — it just no longer hangs on the primary keyword.

In May 2025, Ahrefs analyzed 75,000 brands — exclusively ones with a Domain Rating above 40, which removes small brands from the sample. Brand mentions there correlate with AI Overview visibility at a Spearman coefficient of 0.664, classic backlinks at 0.218. Two further brand signals sit in between: branded anchor text at 0.527 and search volume for the brand name at 0.392. An expanded analysis from December 2025 confirms the pattern across all three major platforms — 0.664 for AI Overviews, 0.709 for Google AI Mode, 0.656 for ChatGPT. The strongest single correlation shows up for brand mentions on YouTube, at roughly 0.737.

What this gap is not: proof of causation. Ahrefs itself points out these are correlations, and that the values are moderate on the Spearman scale. Both quantities are also tied to a shared third factor — brand awareness. Big brands naturally have both more mentions and more visibility.

In practice, it still means something. Backlinks haven’t become worthless; they keep working, but mostly indirectly, through rankings. Brand mentions are also the only lever that also feeds the training-data path.

Platforms follow different rules

The 12 percent figure above isn’t an outlier — it’s the norm outside of Google’s own formats. Systems like ChatGPT frequently favor topically deep specialist publications over large generalist domains.

The takeaway: a strategy meant to apply equally across every platform isn’t a strategy.

What demonstrably works

Five levers with solid evidence behind them — each with the caveat that belongs to it:

TacticFindingSource and context
Answer-first — the direct answer in the first one to two sentences of a section44.2% of all citations come from the first 30% of a text, 31.1% from the middle, 24.7% from the last thirdKevin Indig, February 2026: 18,012 verified citations, isolated from roughly 1.2M ChatGPT answers
Adding numbers, quotes, and sourcesThe three most effective of nine tested methods achieved 30–40% relative improvement on the position-adjusted word-count metric; keyword stuffing had no effectPrinceton GEO paper, KDD 2024: 10,000 queries, but against a simulated engine built on GPT-3.5 — the magnitude transfers, the exact figures don’t
Building topical authority57% faster visibility growth, 62% higher probability of traffic in the first weekGraphite: 332 URLs, 12 domains — measures organic visibility, not AI citations; the link to GEO is indirect, via rankings
Third-party editorial coverage84% of all citations come from earned media, paid and advertorial content account for 0.3%; stable between 82 and 89% across three study editions since July 2025Muck Rack, May 2026: 25M cited links across 17 industries
Publisher authorityCoverage from publishers in the DR-81 cluster got cited in 43% of cases, in the DR-62 cluster in just 2%Stacker, April 2026: 215 stories, 75 brands — the reported r = 0.99 correlation is based on four cluster averages, not individual stories. The direction holds up; don’t over-read the strength

What stands out: none of these levers are new. They’re the same principles that already work in classic SEO — substance, structure, authority.

What Google doesn’t require — and where evidence is simply missing

There’s a distinction worth making here that almost never comes up in GEO discussions. There are three different levels of evidence, and they are not equivalent.

TacticEvidence levelWhat can be said
llms.txtMeasuredOf 137,210 domains Ahrefs checked against its own log data in June 2026, roughly 28 percent had an llms.txt — 97% of those files were never fetched even once in May 2026. Ahrefs itself frames the adoption rate as an upper bound, since the sample overrepresents technically sophisticated sites. Notably: on domains without the file, no AI bot requested it at all. SE Ranking checked roughly 300,000 domains and found no measurable effect on citation likelihood — the model was actually more accurate without the file. Google also lists it as unnecessary.
Breaking content into micro-chunksGoogle statement plus contradicting dataNamed as unnecessary in Google’s May 15, 2026 mythbusting section. SE Ranking measures the opposite: sections under 50 words got 2.7 citations, sections of 120–180 words got 4.6, longer ones got 5.7.
Rewriting content “for AI readability”Google statementAlso named as unnecessary
Buying staged brand mentionsGoogle statementAlso named as unnecessary; consistent with Muck Rack’s finding of 0.3% paid content across all citations
Extensive schema markup as a GEO leverGoogle statement, data inconclusiveNot required to appear in generative AI features. A Search Engine Land experiment with three test pages points the other way — at that sample size, a hint, not proof. SE Ranking found the opposite: pages without FAQ schema got slightly more ChatGPT citations (4.2 versus 3.6). Still worthwhile for rich results in classic search — that’s a separate question.
Prompt-optimized phrasing in body textNo evidence of effectNeither proven nor disproven. No study exists at all.
Special meta tags for AI crawlersNo evidence of effectNo documented tags that AI systems preferentially read

Why this distinction matters. “Not required” is not the same as “doesn’t work.” And “there’s no evidence of an effect” is not the same as “it’s proven not to work.” Mixing up these three levels makes you sound more certain — and leaves you exposed the moment someone brings a counterexample.

The defensible position: put your effort where the evidence is. For everything else, you now know exactly where the evidence stands, and you can decide for yourself.

The foundation: no findability, no citation

Before any fine-tuning matters at all, an AI system has to be able to find your content in the first place. Three areas decide that, and none of them are new.

Technical. Crawlability and indexability need to be consistent — robots.txt, meta robots, canonicals, status codes. The Otterly AI Citations Report from February 2026 analyzed over a million citations and found technical obstacles for AI crawlers on 73 percent of the sites examined. Behind that are usually robots.txt rules, CDN settings, or content that only renders via JavaScript — almost never deliberately aimed at AI crawlers, just historically grown.

Particularly relevant: the main content needs to be in the HTML that actually gets served. Vercel and MERJ jointly studied how AI crawlers behave in December 2024 and found no evidence that any of them execute JavaScript. They do download it — GPTBot on 11.50 percent of requests, ClaudeBot on 23.84 percent — but neither one executes it. The only major crawler that actually renders is Googlebot. The simplest test: open the page, view source, search for a sentence from the main content. If it’s not there, these systems don’t see it either.

Load time shows up in the data too: fast pages with a First Contentful Paint under 0.4 seconds averaged 6.7 citations at SE Ranking, while pages beyond 1.1 seconds averaged 2.1.

Content. What matters is the question behind the search query — and whether you cover a topic across several connected pages instead of cramming everything onto one page stuffed with keywords.

Authority. At SE Ranking, no other single metric explains ChatGPT citations as well as the number of referring domains. Below 2,500 such domains, sites sit at 1.6 to 1.8 citations; beyond 350,000, it’s 8.4. There’s a notable threshold around 32,000, where the value nearly doubles from 2.9 to 5.6. And rankings remain the gatekeeper: sites averaging position 1 to 45 got 5 citations, sites averaging position 64 to 75 got only 3.1.

Fine-tuning: what else the data shows

The findings below come mostly from SE Ranking’s analysis of 129,000 domains. They show correlations, not causation — and SE Ranking itself stresses that the factors aren’t independent of each other: maximizing one at the expense of the rest ends up producing a worse overall picture.

  • Data density. Pages packing in 19 or more data points averaged 5.4 citations; pages with barely any numbers averaged 2.8.
  • Expert quotes. Citing a named expert lifts the average from 2.4 to 4.1 citations.
  • Readability. Cited content scored markedly better on readability measures than lower-performing content. Shorter sentences beat dense technical prose.
  • Don’t over-optimize titles. Perhaps the study’s most surprising finding: pages whose title stuck very tightly to the target keyword averaged 2.8 citations — pages with more broadly worded titles averaged 5.9. The same pattern held for URLs.
  • Substance beats brevity. Beyond 2,900 words, the average was 5.1 citations; below 800 words, 3.2.
  • Freshness beats age. Content updated within the last three months averaged 6.0 citations, stale content 3.6. Raw age, by contrast, barely mattered: brand-new content 3.6, content one to five years old 3.1.
  • Presence beyond your own website. Maintaining a profile on Trustpilot, G2, Capterra, Sitejabber, or Yelp brought 4.6 to 6.3 citations — without such profiles, it was 1.8. For Quora and Reddit, the range stretches from 1.7 to 7.0, though the top end assumes millions of mentions. That’s out of reach for most; the direction still holds.

How to measure AI visibility

Until recently, this was the weakest point in every GEO discussion. Plenty of advice, no data.

Since June 3, 2026, there’s a native report. Google launched a dedicated view for generative AI features under “Performance” in Search Console. It shows impressions from AI Overviews, AI Mode, and generative features in Discover, separately from regular organic data, broken down by page, country, device, and time period.

What it doesn’t show: clicks, click-through rate, position, or search queries. So you can see that you’re showing up — but not what it’s actually doing for you. The rollout is happening gradually and started with a subset of sites — market reports point to a beta launch in the UK. If the section is still missing for you, that’s not a setup problem on your end.

Google has also introduced a toggle that lets you control whether your content can appear in AI Overviews, AI Mode, and Discover. Inclusion is the default. Google has explicitly stated that this toggle is not a ranking signal for classic search — but opting out does cost you traffic and impressions from the generative features.

What this report leaves open. It measures exclusively Google’s own surfaces. ChatGPT, Perplexity, and Gemini outside of Google Search remain invisible. And because search queries are missing, you can’t hold AI visibility directly against your organic performance per query.

For everything outside of Google, you therefore need a second data source. That’s precisely the gap searchsquare’s Performance module closes: visibility in ChatGPT, Perplexity, and Gemini sits alongside position, real clicks, search volume, and competition — per keyword, in one view.

That doesn’t solve GEO. It just answers the question every further decision depends on: where you actually stand right now.

Your first steps

  1. Clarify which systems actually matter for you. Grounding or training data — that determines whether you’re thinking in weeks or in years.
  2. Check the foundation before you fine-tune. Is the main content in the HTML that actually gets served? Do AI crawlers get through with a 200 status code? Without that, no further measure gets you anywhere.
  3. Check Search Console for the new generative AI features section. If it’s not there yet, keep checking back.
  4. Rebuild for answer-first. A direct answer in the first one to two sentences of every section, then context and evidence — but no micro-sections.
  5. Add numbers, sources, and expert quotes wherever you’re currently just asserting. It’s one of the few effects that’s been cleanly measured.
  6. Think in topic clusters, not single keywords. Query fan-out rewards coverage, not one top position. What Google explicitly does not want in the same guidance: spinning up a separate thin page for every conceivable sub-question — that falls under the policy against scaled content abuse.
  7. Work on mentions beyond your own website — editorial, in industry communities, on review platforms.
  8. Cut what has no evidence behind it, and put that time into substance instead.

Check where you actually stand

The Search Console report shows you Google’s own surfaces. For everything beyond that, you need a second source.

In searchsquare, position, real clicks, search volume, and competition sit side by side per keyword — and right next to it, whether your content is showing up in ChatGPT, Perplexity, and Gemini. Instead of five or six separate tools, you only need one.

Try it free. Ranking, search volume, and competitor data for your domain are ready instantly. Your own click and conversion data joins in as soon as you connect Search Console and GA4. Onboarding by real SEOs is included.

Add your domain and get started

Bottom line

Generative Engine Optimization extends your SEO work to an additional visibility layer — it doesn’t replace it. Google says as much itself now.

What’s changing is the weighting: topical breadth over single keywords, brand mentions alongside backlinks, platform-specific approaches instead of a universal recipe. And what’s changing is measurability — slowly, but visibly.

What isn’t changing: content without substance doesn’t get ranked or cited. The shortest path into AI answers runs through good work, not through a file in your root directory.

FAQ

Is GEO the same as SEO?

Not the same, but not a separate discipline either. On May 15, 2026, Google stated that optimizing for generative AI search is, from Google Search's perspective, still SEO. What's shifting is the weighting of individual levers, not their fundamental relevance.

What's the difference between GEO, LLMO, and AEO?

All three terms describe the same underlying goal from slightly different angles, and no consensus definition exists. GEO targets citation in generative answers, LLMO additionally emphasizes the training-data path, and AEO also covers non-generative answer formats like featured snippets. In practice, they're used largely interchangeably.

Which AI systems can I actually reach with SEO?

Anything that looks things up live on the web: Google's AI Overviews and AI Mode, Gemini with search enabled, Perplexity, ChatGPT Search, and Microsoft Copilot. You can only reach pure language models without web access over the long term, through brand building. How long that takes isn't documented: the only known figure is the gap between a model's knowledge cutoff and its release, which historically ran six to eighteen months.

Do I need structured data to appear in AI Overviews?

No. Google explicitly states that additional markup is not required to appear in generative AI features. The data here is inconclusive: a Search Engine Land experiment with three test pages points the other way but is far too small to prove anything, and SE Ranking found slightly more ChatGPT citations on pages without FAQ schema across 129,000 domains. Schema markup still makes sense for rich results in classic search — but that's a separate question.

Does an llms.txt actually help?

Not based on current evidence. Ahrefs looked at 137,210 domains in June 2026: roughly 28 percent had set up the file, and 97 percent of those were never fetched at all in May 2026. SE Ranking found no measurable effect on citation likelihood. Google also lists the file as unnecessary. Of all the GEO tactics in circulation, this one has the clearest evidence stacked against it.

How many AI citations come from Google's top 10?

It depends on the system — and for Google specifically, it even depends on who's measuring. For Google's own AI Overviews, Ahrefs puts it at roughly 38 percent in February 2026, down from 76 percent in July 2025, though part of that decline traces back to improved measurement. BrightEdge, using a different methodology, measures only around 17 percent and finds that figure stable across its entire tracking period. For Perplexity, it's roughly 29 percent; for ChatGPT, Gemini, and Copilot combined, around 12 percent on average. One thing worth flagging: these numbers measure against the original user query. Factor in the sub-queries from query fan-out, and AirOps finds that 55.8 percent of the pages ChatGPT cites rank in Google's top 20. There's no single number that applies across every platform.

Are backlinks still relevant for AI visibility?

Yes, but mostly indirectly: they drive rankings, and rankings are the entry ticket for grounding systems. In SE Ranking's predictive model, the number of referring domains is the single most important factor for ChatGPT citations. Brand mentions, on top of that, correlate noticeably more strongly with AI visibility than classic backlinks do — 0.664 versus 0.218 for AI Overviews, and 0.656 for ChatGPT, measured across 75,000 brands (Ahrefs). Two different measurement approaches, pointing the same way: the two belong together.

What is query fan-out?

The mechanism by which an AI search system breaks a single user query into several semantically related sub-queries and pulls sources for each one. Google names it explicitly in its own documentation. An AirOps analysis of ChatGPT shows how significant this second path to citation is: among cited pages that showed up in a top-20 SERP at all, 32.9 percent were findable exclusively through a sub-query. The practical consequence: your page can get cited without ranking on page 1 for the original query.

How much do AI Overviews cut organic click-through rate?

Substantially, though the numbers vary. Pew Research observed click-through rate falling from 15 to 8 percent on queries with an AI Overview present. Seer Interactive, across 5.47 million queries, found a 61 percent drop; Ahrefs found 58 percent for content in position 1. When a page does get cited, it captures roughly 120 percent more organic clicks per impression than uncited pages on the same results page — but that still sits about 38 percent below what it would earn on a results page with no AI Overview at all. Getting cited is a clear advantage in the new environment, not a replacement for the old one.

Can I see in Search Console whether I'm showing up in AI answers?

Partly, since June 3, 2026. The new generative AI features section shows impressions from AI Overviews, AI Mode, and generative Discover features, broken down by page, country, device, and time period — but no clicks, no position, and no search queries. The rollout is happening gradually, initially for a subset of sites.

How do I measure visibility in ChatGPT, Perplexity, and Gemini?

The Search Console report covers exclusively Google's own surfaces. For the other systems, you need a separate data source that captures visibility there and sets it against your organic performance. It's also worth a second look at your server logs — that's where you can see which AI crawler is requesting which page, and what it gets back.

Sources

Every number cited in this article can be verified through the sources below.

Google (primary sources)

02Google Search Central Blog, June 3, 2026 — Introducing Search Generative AI performance reports in Search Console
04Google Search CentralAI Features and Your Website

Studies and analyses

01Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, DeshpandeGEO: Generative Engine Optimization (preprint November 2023, presented at ACM SIGKDD 2024)
02WikipediaGenerative engine optimization (term overview incl. AEO and AIO)
03Lewis et al. (2020) — Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (parametric vs. non-parametric memory)
04Otterly.AILLM Knowledge Cutoff Dates (gap between knowledge cutoff and release)
05Temso AI, July 2026 — AI Knowledge Cutoff Dates: Every Major LLM (current cutoffs compared)
06Ahrefs, July 2025 — 76% of AI Overview Citations Pull From the Top 10 (1.9M citations)
07Ahrefs, February 2026 — AI Overview Citations From Top-Ranking Pages Drop Sharply (863,000 keywords, 76% → 38%)
08Ahrefs, August 2025 — AI Search Overlap (15,000 prompts, 12%, Perplexity ~29%)
09AhrefsAI Overview Brand Correlation (75,000 brands, 0.664 vs. 0.218)
11Ahrefs, June 2026 — llms.txt Study (137,210 domains, 97% with zero fetches)
12Ahrefs, February 2026 — AI Overviews Reduce Clicks Update (58% CTR decline, position 1)
13SE Ranking, November 2025 / March 2026 — How to Optimize for ChatGPT (129,000 domains, 216,524 pages)
15BrightEdgeAI Overviews at the One-Year Mark (48% coverage, 58% growth)
16BrightEdgeRank Overlap After 16 Months of AIO (32% → 54%, YMYL figures)
18Seer Interactive, February 2025 — 87% of SearchGPT Citations Match Bing's Top Results (just over 500 citations)
19Seer InteractiveAI Overview CTR Impact Study (5.47M queries)
21Kevin Indig, February 2026 — The Science of How AI Pays Attention (18,012 verified citations)
23Muck Rack, May 2026 — What Is AI Reading? (25M links, 84% earned media)
24Stacker, April 2026 — Pickup Quality: The X-Factor for LLM Visibility (215 stories, 75 brands)
25Otterly.AI, February 2026 — The AI Citations Report 2026 (over 1M citations)
26Vercel & MERJ, December 2024 — The Rise of the AI Crawler (JavaScript rendering by AI crawlers)
27GraphiteTopical Authority Whitepaper (332 URLs, 12 domains)
Raul Dahm Cardo

About the author

Raul Dahm Cardo

SEO Consultant & Co-Founder, searchsquare

Six years at SISTRIX, many of them as a seminar leader for several hundred trained companies — alongside freelance consulting and hands-on implementation for companies he helps become visible. Raul knows SEO from both sides: consulting and execution. What interests him is the point where data turns into a decision — and that's exactly where searchsquare comes in. His articles grow out of ongoing SEO projects.