Generative Engine Optimization (GEO)
Generative Engine Optimization is the practice of structuring content so AI answer engines — ChatGPT, Perplexity, Google AI Overviews, Gemini and Claude — cite it as a source. Unlike SEO, which competes for a page ranking, GEO competes at the passage level inside a retrieval system.
This page states what the evidence actually supports, and what it does not. Every figure names its study. Where the honest answer is “no measurable effect”, it says so — including for two things much of the industry sells.
How do AI engines choose what to cite?
AI engines choose passages, not pages. A prompt is fanned out into multiple sub-queries — around 8.5 for a reasoning-mode model — each retrieved separately. Retrieved pages are split into short passages, every passage is embedded, and the one that best matches a sub-query gets quoted and attributed.
The consequence is the whole of GEO: a 2,000-word article is roughly 25 competing passages, and a passage that only makes sense in context is worthless — because being lifted out of that context is exactly what happens to it.
Why does GEO matter now?
Search increasingly ends without a click. SparkToro and Similarweb clickstream data put US Google searches ending without a click at 68.01% between January and April 2026, up from 60.45% in 2024. Where an AI Overview appears, click-through rate drops by roughly 60%. The answer is increasingly the destination.
Being cited is still worth chasing on its own terms: Seer Interactive, analysing 5.47 million queries across 53 brands, found pages cited in AI Overviews received about 120% more organic clicks per impression than uncited ones. Citation and traffic are complements, not substitutes.
What do the numbers actually say?
The figures below cover the most-replicated findings in generative search research as of August 2026, with the study behind each one.
| Finding | Value | Source |
|---|---|---|
| Citations from the top 30% of a page | 44–55% | Passionfruit; Indig (1.2M citations) |
| Sub-queries per prompt (reasoning mode) | ~8.5 | Passionfruit |
| Citation lift from data-rich vs thin pages | 4.31× | Yext (17.2M citations) |
| Sources cited by all four major engines | 3.8% | Writesonic (161,286 prompts) |
| Cited URLs still cited after 28 days | 10.6% | Digital Authority Partners + Profound |
| Effect of adding schema on AI Overviews | −4.6% | Ahrefs (1,885 pages vs 4,000 controls) |
| Effect of llms.txt on citation frequency | none measured | SE Ranking (~300,000 domains) |
GEO vs SEO: what is the actual difference?
GEO and SEO share most of their foundations and differ in their unit of competition. SEO optimises a document to rank; GEO optimises a passage to be extracted. The practical differences are narrow but they matter, and in one case the two disciplines pull in opposite directions.
| SEO | GEO | |
|---|---|---|
| Unit of competition | The page | The passage |
| Winning looks like | A high ranking | Being quoted in an answer |
| Keyword density | Helps | Measured at −10% (KDD 2024) |
| Structure | Aids scanning | Decides extractability |
| Freshness | Modest factor | Modest factor (~26% newer) |
| Stability | Rankings persist | ~10.6% persist 28 days |
That third row is the one to notice. Keyword density is a classical SEO signal, and the peer-reviewed KDD study measured keyword stuffing at −10% for generative citation. Optimising hard for one can cost you the other, which is why a single blended “AI score” hides more than it shows.
What does the industry get wrong about GEO?
The most common mistake is treating GEO as a set of new files and tags to add. The two most-sold additions have both been tested and neither survived.
SE Ranking tested ~300,000 domains and found no relationship with citation frequency; removing the variable made their model more accurate. Google compared it to the meta keywords tag. It is genuinely useful for developer documentation, where agents ingest a markdown bundle — not for a marketing site.
Ahrefs tracked 1,885 pages that added JSON-LD against 4,000 controls: −4.6% in AI Overviews, nothing meaningful elsewhere. A separate experiment found engines extract visible HTML during retrieval and ignore JSON-LD. Emit schema for entity disambiguation and agent readability, not for citations.
Brand mentions do correlate with citations more strongly than backlinks, but the correlation is confounded by brand size — big brands have both. Google's own 2026 guidance says seeking inauthentic mentions "isn't as helpful as it might seem", and platform detection makes it a bad trade.
Only about 10.6% of cited URLs persisted across 28 days. Any dashboard reporting week-over-week movement as progress is mostly reporting noise.
What actually works?
Five things carry most of the evidence. They are unglamorous, which is probably why they are undersold.
- 1Be crawlable by the right bots
Crawler accessibility scores 9.5/10 for evidence quality — the highest of 23 factors, ahead of search rank. Blocking a training bot costs nothing; blocking a retrieval bot removes you entirely.
- 2Answer first
44–55% of citations come from the top 30% of a page, in two independent analyses. Lead every section with its answer.
- 3Write self-contained passages
A section opening with "It also supports…" is uninterpretable once lifted out. Name the subject in the first sentence, every time.
- 4Be specific
Data-rich pages are cited 4.31× more per URL than thin ones. Concrete figures, named entities, real prices.
- 5Cite your own sources visibly
The largest single effect measured in the peer-reviewed literature. Treat the magnitude as a lab artifact and the direction as real.
Where to go next
Free, no signup. Tells you whether OAI-SearchBot, PerplexityBot and Claude-SearchBot can fetch a page at all — the first item on the list above, and the one with the strongest evidence behind it.
GEO templates →What a passage contract actually looks like: what each retrievable passage must open with, what it must contain, and how many concrete figures it needs.
How does ActiveGeo write for GEO?
ActiveGeo writes to a passage contract rather than a page outline. Each template declares what every retrievable passage must open with and contain, the draft is scored against 15 deterministic checks, and anything below the bar gets one targeted revision. Every check shows its rule and the study behind it — you can audit the score rather than trust it.
One thing it will not do: write you onto Wikipedia, Reddit or G2. For commercially valuable queries the large majority of citations come from third-party media, and no content tool changes that. What it can do is make your own pages structurally worth citing, and tell you when an engine cannot even reach them.
Frequently asked questions
Is GEO the same as SEO?
GEO is not the same as SEO. SEO competes for a page ranking; GEO competes at the passage level inside a retrieval system, because answer engines split a page into short passages and quote only the best match. In practice GEO is roughly 70% classical technical SEO, 20% third-party brand presence, and 10% AI-specific structuring.
Does llms.txt work?
No. SE Ranking tested around 300,000 domains and found no meaningful relationship between llms.txt and AI citation frequency — their model became more accurate when the variable was removed. Zyppy scored it 2.0 out of 10, last of 23 factors. No AI provider has confirmed reading it.
Does schema markup improve AI citations?
Barely. Ahrefs tracked 1,885 pages that added JSON-LD against 4,000 controls and measured −4.6% in AI Overviews and no meaningful change in AI Mode or ChatGPT. Schema is still worth adding for entity disambiguation and agent readability — just not as a citation lever.
How long is a page cited for?
Not long. Across 1,127 tracked URLs, only about 10.6% persisted as citations over 28 days, ranging from roughly 11% on Gemini to 44% on Perplexity. AI visibility behaves like a high-variance position, not like a stable ranking.
Can a small brand win AI visibility?
Only in the long tail. Analysis of 100,000+ prompt responses found household-name brands visible about 73% of the time, mid-market 44%, and niche brands 11%. A small brand can realistically be the cited answer for narrow, high-intent questions — not for broad category queries.
Do AI crawlers run JavaScript?
No major AI crawler executes JavaScript. Vercel and MERJ analysed over 500 million GPTBot fetches and found zero evidence of JS execution. Client-rendered content does not rank poorly in AI search — it does not exist there.
Sources
- Aggarwal et al., GEO: Generative Engine Optimization, ACM SIGKDD 2024
- Zyppy — 23 factors that get content cited (54 studies)
- Ahrefs — we tracked 1,885 pages adding schema
- SE Ranking — llms.txt shows no clear effect across 300K domains
- Vercel — the rise of the AI crawler
- Writesonic — cross-engine citation overlap (161,286 prompts)
- SparkToro — Google zero-click searches reach 68%
- Google Search Central — optimizing for AI features
Last updated 1 August 2026