Generative Engine Optimization (GEO): The 2026 Guide
The complete guide to generative engine optimization (GEO) in 2026: how AI search changes CRO and SEO, what research shows works, and how to test it.

GEO (generative engine optimization) is the practice of structuring content so that generative systems cite your page when they compose an answer. In 2026 it stopped being a niche topic for a simple reason: a meaningful share of the questions that used to end in a click are now answered before the click, inside an AI summary or a conversation with an assistant. This guide covers what actually changed (with attributed data, not panic), how a generative engine decides what to cite, what the academic research found to work, what all of it means for conversion work, and the part almost nobody handles honestly: how to test a GEO change with statistical rigor when the AI-sourced traffic segment is far too small to reach sample size. It includes live calculators and an end-to-end worked example.
What generative engine optimization is, and where it diverges from SEO
GEO is optimizing to be extracted, not to be listed. Classic SEO competes for a position in a list of blue links, and the click is the final product. GEO competes for the citation inside an answer synthesized by a model, and the final product may be a click, a brand mention with no click, or both.
The term comes from an academic paper, “GEO: Generative Engine Optimization”, by Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, presented at ACM SIGKDD 2024, which defined the problem and measured solutions under controlled conditions (arXiv 2311.09735).
| Dimension | Classic SEO | GEO |
|---|---|---|
| Target | A position in a list of results | Being cited or paraphrased inside a generated answer |
| Competing unit | The whole page | The extractable passage: a claim, a number, a definition |
| Success signal | Organic click | Citation, brand mention, and the click when it happens |
| Structure that helps | Title, headings, internal links, speed | All of that, plus atomic claims, sourced data and schema |
| How it is measured | Impressions, average position, CTR | Indirect metrics and manual checks; measurement still imperfect |
The point that causes the most confusion: GEO does not replace SEO, it depends on SEO. In the architecture of today’s generative engines, the page has to be retrievable (indexed, crawlable, relevant to the query) before it is a candidate for citation. A site invisible to search stays invisible to the assistant.
What actually changed in search
Three shifts overlap, and it is worth separating what has measured data behind it from what is market speculation.
The first is the summary generated on the results page itself. A Pew Research Center analysis, built from real browsing data of 900 adults in the United States during March 2025, found that users clicked a traditional result on 8% of the searches where an AI summary appeared, versus 15% of the searches without one, and that only 1% clicked a link inside the summary itself (Pew Research Center). That is the most concrete evidence available on the effect, and it has limits: one sample, one country, one month.
The second is the migration of the start of research toward conversational assistants. Many questions that used to begin in a search box now begin in a chat, which removes volume from search without producing a single row in your Search Console.
The third is a composition effect on the traffic that remains. When simple questions get answered before the click, the people who do click tend to be the ones who still need something the answer did not give them: a comparison, a price, a technical detail, proof. Traffic shrinks in volume and changes in profile.
One reading caveat applies to this entire guide: those numbers describe a market average, measured in a specific context. The real effect on your site depends on how much of your traffic comes from broad informational queries versus transactional and branded ones. Before reacting, measure your own composition.
How a generative engine decides what to cite
Simplifying just enough to be useful without being false, current systems work in three stages:
- Retrieval. The system fetches documents relevant to the question, usually through a search index or its own corpus.
- Synthesis. The model reads the retrieved passages and writes an answer in natural language, combining information from several sources.
- Attribution. The system attaches citations to the parts of the answer that came from specific sources.
The GEO paper reproduced that pipeline in the lab, with the retrieval stage fed by a real search engine and the synthesis performed by a language model, so it could measure the effect of content changes on final visibility in the answer.
One practical consequence of that pipeline usually goes unnoticed: the system rarely reads your whole page when composing the answer. It works with chunks, and the unit competing for the citation is the passage, not the article. An excellent piece whose central idea only holds up after three paragraphs of setup is structurally disadvantaged against a piece that states the conclusion first and explains afterward. That is not an aesthetic preference, it is a consequence of how extraction works: if the passage needs the previous paragraph to make sense, it tends to be discarded as a candidate.
That gives us the golden rule of GEO, and it is short: every section should contain at least one sentence that stays true and comprehensible if it is ripped out of the page and pasted somewhere else on its own. If no sentence in your section passes that test, it probably will not be cited, no matter how good the surrounding content is.
GEO by query type: where the effect is big and where it is small
Treating “AI search” as one block leads to bad decisions, because the effect varies enormously by question type. The classic intent split still holds, with distinct behavior in each band:
| Query type | Example | Effect of the generated answer | What to do |
|---|---|---|---|
| Broad informational | “what is a/b testing” | High: the answer often suffices and the click drops | Write to be cited, and accept the no-click mention as part of the return |
| Specific informational | “how to calculate sample size for a 2.5% conversion rate” | Medium: the answer gives context, but the user wants a tool or a detail | Offer what the answer cannot deliver: a calculator, a template, a worked example |
| Comparative | “tool x versus tool y” | Medium to low: the user wants to check the source before deciding | Keep comparisons neutral, current, and explicit about criteria |
| Transactional | “sign up”, “pricing” | Low: the decision requires the site | Prioritize price clarity, friction and proof, which is classic CRO work |
| Branded | “donnu pricing” | Low, but sensitive: the answer may describe your brand wrong | Keep official pages about product, price and policy clear and current |
The strategic reading of that table: the content most affected is exactly the informational top of funnel, which historically produces the most volume and the least direct conversion. That means the loss of sessions can be large while the loss of revenue is small, and looking only at the sessions chart leads to conclusions that are far too dramatic. Before reacting to a drop, split traffic by query type and find out where it actually happened.
What the research shows works
The central finding of the GEO paper is that optimization methods aimed at generative engines can increase visibility in the answers by up to 40%, measured on the large-scale benchmark the authors built for the study (Aggarwal et al., GEO, KDD 2024). The detail most useful to anyone writing content is which kind of change carried that gain: the tactics that add verifiable evidence to the text, meaning quoting credible sources, adding statistics and citing sources, moved visibility the most, while stuffing the target keyword landed below the unoptimized baseline. One result is worth recording because it is easy to read backwards: the paper also tested a method it named “Authoritative”, which rewrites the text in a more persuasive and authoritative tone without adding any evidence, and found no significant improvement from it. Adding a source moved the needle; sounding like an authority did not.
| Tactic | What it means in practice | Why it tends to work |
|---|---|---|
| Statistics with an explicit number | Replacing “many users abandon the cart” with the measured number, sourced | A passage with a number is easier to cite as a complete unit of information |
| Source citation | Attributing every claim to whoever produced it, with a link | Raises the odds the model treats the passage as a verifiable claim |
| Quoting a credible source | Bringing in the original sentence from a recognized source | Gives the system a self-contained passage, ready to reproduce |
| Clear, direct language | Writing the answer before the explanation | Makes the claim extractable without depending on surrounding context |
| Keyword repetition | Stuffing the target term in over and over | Weak performance in the study: the text is still hard to extract |
An important interpretation caveat. Those results came from a pipeline reproduced in the lab, with models and indexes from a specific generation, not from a test inside today’s commercial products. The direction is solid and consistent with practical experience; the exact number should not be treated as a promise.
GEO in practice: the page optimization checklist
Translating the research and the available documentation into concrete editorial decisions:
| Element | What to do | Why |
|---|---|---|
| Opening of each section | Start with an atomic claim that answers the subheading | It is the passage a system can extract without needing the full paragraph |
| Data | Always with number, unit, period and attributed source | Without a source, the passage tends to be dropped or rewritten without credit |
| FAQ | A question-and-answer block with FAQPage markup | Explicitly pairs question and answer, the shape a generated response needs |
| Visible freshness | A visible updated date and references from the current year | Signals the page is a safe citation |
| Heading structure | Hierarchy without skipped levels, a single H1 | Delimits the units of meaning that will be extracted |
| Tables and lists | Comparisons in tables, steps in numbered lists | Formats that survive extraction without losing meaning |
| Original content | Your own data, your own calculation, a worked example | What exists nowhere else cannot be synthesized from another source |
FAQ markup deserves an honesty note. Google documents the FAQPage schema and the conditions under which it is used, and the FAQ rich result itself stopped appearing broadly in search (Google Search Central, FAQPage structured data). In other words: the reason to keep the schema today is not to win an ornament in the result, it is to make explicit, to any automated reader, which question each passage answers. The specification lives at schema.org/FAQPage.
Notice that almost everything on that list was already good writing practice before generative AI existed. That is the best news about GEO: most of the work is the same old discipline, applied with more rigor.
A four-step plan for the content you already have
Almost nobody starts from zero. The fastest route to a GEO result runs through rewriting what already has traction, not through publishing more:
- Pick the 20 pages that receive the most informational traffic. They have the most to lose to the generated answer and the most to gain from the citation. Set transactional and branded pages aside for now.
- Rewrite the opening of every section. One direct answer sentence before the explanation, in every H2. It is the highest-return change per hour of work, because it creates the extractable passages that were missing.
- Put a source on every number. Every unattributed statistic becomes a binary decision: find the real source and cite it, or remove the claim. A number without a source is a liability, not an asset, and it is the first thing that undermines a page’s credibility with any reader, human or not.
- Add a question-and-answer block with FAQPage markup. Four to six real questions, answered completely and briefly, covering what people actually ask about that topic.
Once that is done on the 20 main pages, measure for a quarter before expanding. That is where most GEO programs get lost: they start producing new volume before knowing whether the rewrite worked.
A mention with no click is a result, and it belongs in the accounting
An uncomfortable part of GEO is that a good share of the return never shows up in analytics. When an assistant answers a question by citing your brand and the user does not click, something real happened: the brand was presented as a trustworthy source, in a decision context, with no media cost. That produces no session, no event, no line in the acquisition report.
Ignoring that effect leads to underinvesting in the channel. Overstating it leads to justifying anything with a number nobody can measure. The honest middle path is to treat the mention as a brand signal, measured with the same imperfect proxies brand has always used: the trend in searches for the company name, direct traffic, and the “how did you hear about us” question on the signup form, which remains one of the most useful and most underrated instruments there is.
What not to do is convert mentions into estimated revenue using an invented multiplier. An invented number in a spreadsheet becomes a target in three months and an excuse in six.
What this changes in CRO
Here is the part almost no GEO guide connects, and the part that changes what a conversion team should be doing.
If the traffic reaching the site shrinks in volume and arrives more decided, conversion work shifts. Fewer people need to be convinced the problem exists, because the generated answer already explained it. More people arrive wanting to confirm one specific detail and move on.
The mistake would be to turn that hypothesis into an immediate site redesign. The correct way to handle it is the usual one: turn it into a hypothesis, size it, test it. The complete conversion rate optimization guide covers the process, and the next block deals with the specific problem AI search creates for that test.
Blocking or allowing the AI crawlers
A decision that comes up early in any GEO discussion: is it worth blocking generative system crawlers in robots.txt? Both sides have a legitimate argument, and the answer depends on how your business makes money.
A business that lives on advertising and page volume takes a direct loss when its content is synthesized without a click, and blocking reduces that extraction, at the cost of disappearing from the answers. A business that lives on a product, such as a SaaS, is usually on the other side of that ledger: appearing as a cited source in an answer about the problem the product solves is cheap distribution, and blocking means giving it up.
There is also a technical detail that confuses a lot of people: the training crawler and the live search crawler are not always the same agent. Blocking the first does not necessarily remove the page from answers generated with live retrieval, and blocking the second has an immediate effect on citation. Before touching robots.txt, separate the two cases and decide on each, instead of applying a blanket block that produces a different effect from the one intended.
Publishing more AI-generated content is not a GEO strategy
Worth saying plainly, because it is the most tempting shortcut of the moment: raising publication volume with automatically generated text does not improve the odds of being cited, and often makes them worse.
The reason is structural. What the research identified as the determinant of visibility in generative answers was the presence of statistics, sources and authoritative quotes, which is exactly what mass-produced generic text does not have. An article that reorganizes what already exists in ten other articles offers nothing the model cannot synthesize from those ten other sources, so there is no reason for it to be picked.
What survives that logic is content carrying something that exists nowhere else: your own calculation, an end-to-end worked example, data you collected, a comparison nobody had made with that criterion. It costs more per piece and it is the only thing that keeps working when the marginal cost of producing average text tends toward zero.
How to measure GEO without fooling yourself
There is no reliable “AI citations” dashboard today. The set below is what you can assemble with common tools, with the limitations stated:
| Signal | How to collect it | The honest limitation |
|---|---|---|
| AI crawler hits | Server logs filtered by crawler user agent | Crawling is not citation: being read does not guarantee being used |
| Referral traffic from assistants | Source reports in analytics | Many assistants send little or no source information |
| Manual checks on target queries | Running your main questions through the assistants and recording who got cited | Answers vary by session, region and model version |
| Branded search | Search Console, trend of queries containing the brand name | Slow effect, subject to other causes |
| Conversion of the traffic that arrives | An A/B test on the site, with the business metric | Measures the funnel effect, not the citation itself |
The practical rule: use the first four as a direction thermometer and only the last one as a basis for decisions about site changes. A thermometer does not decide a test.
Testing a GEO change with rigor
Suppose a site with 120,000 visitors a month, of whom 3% (3,600 a month, roughly 840 a week) arrive through AI assistants. That segment converts better: 4.2%, against the site-wide average of 2.6%. The team wants to test a new comparison page, designed for the intent of someone who arrives already informed, and wants to know whether it converts better.
The temptation is to run the test on the AI segment only, since the page was designed for it. The arithmetic shows why that does not work:
- Detecting a 15% relative gain on a 4.2% baseline requires 17,050 visitors per variation. At 840 visitors a week in that segment, the test would take 285 days.
- Lowering the ambition to a 30% relative gain drops the sample to 4,544 per variation, and the duration is still 76 days.
Two-proportion normal approximation, 2 variations (50/50). Tweak the inputs and watch it update live.
Now the alternative: running the same test across the whole site, with a 2.6% baseline and 28,000 visitors a week. Detecting the same 15% relative gain requires 28,039 visitors per variation, and the duration drops to 15 days.
Two-proportion normal approximation, traffic split evenly across variations. The date uses your timezone and updates live.
The correct reading of that result is subtle and important. The whole-site test answers a different question from the segment test: it measures the average effect across everyone, not the effect on people who arrived through AI. If the new page works very well for the already-informed visitor and badly for the cold visitor, the average effect can land near zero and hide both.
The honest path, then, is this: decide on the whole-site test, and read the AI segment as a secondary signal, knowing it is exploratory. A difference observed inside a small segment is there to generate the next hypothesis, never to declare victory. It is the same discipline as any post-hoc segment analysis: the more slices you look at after the fact, the higher the chance of finding a pattern that is only noise.
| Test design | Question it answers | Duration | Good enough to decide? |
|---|---|---|---|
| AI segment only, MDE 15% | Effect on people arriving through AI | 285 days | No, not feasible |
| AI segment only, MDE 30% | A large effect on people arriving through AI | 76 days | Maybe, if the expected effect really is large |
| Whole site, MDE 15% | Average effect across all traffic | 15 days | Yes |
| Whole site with a secondary segment read | Average effect, plus a clue about the segment | 15 days | Yes, decide on the overall result; the segment generates hypotheses |
Common GEO mistakes
| Mistake | Warning sign | Fix |
|---|---|---|
| Treating GEO as a separate discipline | An “AI SEO” team duplicating the content work | GEO is the same good writing, with stricter structure and sourcing |
| Redesigning the whole site because of AI traffic | A full rebuild for a channel worth a few percent | Measure your real composition before reallocating effort |
| Assuming you lost traffic to AI without checking | A sessions drop attributed to AI with no query-type analysis | Split informational from transactional and branded queries before concluding |
| Stuffing the page with the keyword | The target term repeated in every paragraph | Cosmetic tactics performed poorly in the research; prefer data and sources |
| Publishing a number without a source | “Studies show that 70% of users…” | Without a verifiable source, remove or soften the claim |
| Deciding a site change on the AI segment | Victory declared on a few hundred visitors | Decide on the sized test; a small segment is exploratory |
| Confusing crawling with citation | “The crawler came through, so we are being cited” | Crawling is a precondition, not a result; check citation on your target queries |
Make this automatic with Donnu
This whole guide converges on one point: changing content with AI in mind is cheap, but knowing whether the change improved the result requires the same statistics as always. One primary metric chosen before anyone looks at the dashboard, and the discipline not to turn a small segment into a victory. Donnu covers the part that is easiest to get wrong: you define the hypothesis and the metric, Donnu collects with a snippet that does not block page load, and returns an honest verdict with the 95% confidence interval in front, holding back the winner call until the variation has accumulated at least 200 visitors and 7 days on air, whether the visitor came from a traditional search, an AI summary or an assistant.
Start a 14-day free trial and test your next content change with rigor instead of intuition. For the method behind it, see statistical significance in A/B testing and how to run an A/B test.
References
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K. and Deshpande, A. GEO: Generative Engine Optimization. ACM SIGKDD 2024. arxiv.org/abs/2311.09735.
- Pew Research Center. Google users are less likely to click on links when an AI summary appears in the results. July 2025. pewresearch.org.
- Google Search Central. FAQPage structured data. developers.google.com/search/docs/appearance/structured-data/faqpage.
- Schema.org. FAQPage specification. schema.org/FAQPage.
Read next
Frequently asked questions
- What is GEO (generative engine optimization)?
- GEO is the practice of structuring content so that generative systems (Google AI Overviews, ChatGPT, Perplexity, Gemini and similar assistants) are more likely to cite or paraphrase your page when they answer a question. The term comes from an academic paper presented at ACM SIGKDD 2024, which measured, under controlled conditions, which content characteristics increase visibility inside AI-generated answers.
- Does GEO replace SEO?
- No. GEO depends on SEO. In most current generative systems the page has to be retrievable by a search index before it can become a citation candidate. What changes is the target. Classic SEO optimizes for a position in a list of links, GEO optimizes for being the source extracted inside a synthesized answer. In practice the two disciplines share most of the technical work and diverge on how the text itself is structured.
- Does AI search really reduce clicks to my site?
- On informational queries, the available data says yes. A Pew Research Center analysis using real browsing data from 900 US adults in March 2025 found that users clicked a traditional result on 8% of searches where an AI summary appeared, versus 15% of searches without one, and that only 1% clicked a link inside the summary itself. That is a meaningful drop, measured in a specific sample over a specific period, and it does not apply equally to transactional or comparison queries.
- What does the research show actually works in GEO?
- The paper that coined the term tested several tactics on a large-scale benchmark and reported that GEO methods can boost visibility in generative responses by up to 40%. The most effective tactics in the study were the ones that add verifiable evidence to the text: quoting credible sources, adding statistics, and citing sources. Repeating the target keyword landed below the unoptimized baseline, and rewriting the text in a more authoritative tone without adding evidence produced no significant improvement.
- How do I measure whether my GEO work is paying off?
- No single metric settles it, and it is honest to say the measurement is still imperfect. The most useful set today combines four signals: AI crawler hits in your server logs, referral traffic from generative assistants in analytics, periodic manual citation checks on your target queries, and the trend in branded search. None of them proves causality alone, but together they show direction.
- Can you A/B test a GEO change?
- You can test the effect of the change on site conversion, and that is usually what matters to the business. What you almost never can do is isolate the AI-sourced traffic segment as the decision metric, because it is normally far too small to reach sample size in a reasonable window. The practical answer is to run the test across the whole site and treat the AI segment as a secondary, exploratory read.