Generative Engine Optimization: What Actually Gets You Cited
Last updated · published
Generative Engine Optimization is the practice of shaping content, markup and measurement so that AI assistants surface and cite your pages. The term comes from a 2023 Princeton-led paper, “GEO: Generative Engine Optimization” (Aggarwal et al.), which tested whether rewriting a source changes how often a generative engine quotes it.
Separate what that research showed from what the industry has since asserted, because the gap is wide and most GEO advice lives on the wrong side of it.
The paper tested text rewrites. Adding citations, quoting authorities and adding statistics all improved visibility on its benchmark. It did not test schema markup, FAQ blocks or TL;DR summaries.
TL;DR
- The Princeton paper reported gains of up to 40 percent across its best rewriting methods, and 115 percent for a source ranked fifth, measured on its own benchmark.
- Those methods were about the words: cite sources, quote authorities, add statistics. Schema and page furniture were not part of the study.
- Google states plainly that no special structured data is needed for AI features, so treat schema as good practice rather than a citation lever.
- Blocking a training crawler is not the same as blocking a retrieval crawler. Google-Extended does not affect Search inclusion.
- Do not invent an
ai_referralsource. Tag placements for what they are, and measure assistants through the referrer and GA4’s AI Assistant channel.
What the Research Actually Supports
The paper introduced both the term and GEO-bench, a benchmark for testing how rewriting a source changes its visibility inside a generated answer. Its headline result is that the best methods lifted visibility by up to 40 percent, with Cite Sources reaching 115 percent for a source that started ranked fifth.
Read those numbers as what they are: one benchmark, one moment in the technology, and a set of text-level interventions.
Three things the study supports directly:
Cite your own sources. Content that attributes well gets quoted more. Name the study, the year and the metric, and link the primary document.
Quote authorities. Direct quotation from a credible source improved visibility across the tested engines.
Third, add real statistics. Specific numbers with provenance beat general claims, which is the same discipline as the first two wearing different clothes.
What the study did not test, but the industry now asserts confidently: FAQ blocks, TL;DR summaries, heading density and JSON-LD. Those may well help retrieval systems chunk a page. That is a reasonable hypothesis, and it is how we structure our own content. It is not a research finding, and Google’s own guidance says there is no special structured data you need to add for AI features.
Two more factors are real but unglamorous. Freshness matters for time-sensitive queries, and a visible date beats a header field nobody can see. Source trust matters everywhere, which is the part of GEO that is just SEO with a new name.
The Vendor Layer
A market formed around this fast. Profound, AthenaHQ, Otterly.AI, BrandRank.AI, Peec AI, Goodie AI and Daydream all monitor how often models mention or cite a brand, by running fixed prompt sets against the major engines and parsing what comes back.
Some call it AI search optimisation, some AI brand visibility, some Answer Engine Optimization. The distinctions are mostly positioning. GEO is the broader term and the one with a paper behind it.
What these tools give you: prompt-level visibility, share of voice inside answers, competitor comparison, trend lines. What they cannot give you: clicks. For that you need your own logs.
Know If You Are Being Cited
Three signals, in increasing order of reliability.
Monitoring tools tell you when a model mentions you. Useful, indirect, and priced in a way that moves often enough that you should check the vendor page rather than trust a number in an article.
Referrers tell you when someone clicked. Assistant hostnames to expect: chatgpt.com, perplexity.ai, claude.ai, copilot.microsoft.com, gemini.google.com, chat.deepseek.com, grok.com, meta.ai, you.com. ChatGPT also appends utm_source=chatgpt.com to citation links, so some of this traffic arrives tagged. GA4 classifies several of these into a native AI Assistant channel on the medium ai-assistant, and the ones it misses need a custom channel group. The mechanics are in the AI traffic channel guide.
Be honest about the gap: app and in-window experiences often strip the referrer, so this signal is a floor.
Server logs tell you who crawled. A weekly grep over access logs gives a crawl baseline, and most CDNs now ship an AI crawler dashboard.
Crawl is not citation. A busy crawler does not mean you are being quoted. A silent one does mean you are invisible to that system, which is worth investigating.
Tag What You Control. Measure the Rest by Referrer.
Here is where a lot of GEO advice, including the earlier version of this page, goes wrong.
It is tempting to reserve a tidy pair of values, utm_source=ai_referral and utm_medium=ai_search, and put them on any link you think an assistant might cite. Do not do this. It breaks in two directions at once.
UTM parameters beat the referrer in every analytics tool. So a guest post tagged ai_referral reports as AI traffic for every human who reads that guest post, whether an assistant was involved or not. Meanwhile the one visitor who genuinely arrived through a ChatGPT citation of that article carries the partner site as the referrer, not chatgpt.com, because the click happened inside the article.
You end up mislabelling ordinary referral traffic as AI, and missing the AI traffic you were trying to count.
The rule is simpler. Tag a placement for what it is: utm_source=<partner>, utm_medium=referral, and a campaign name that identifies the piece. Then measure assistants where they are actually visible, in the referrer and in GA4’s AI Assistant channel.
That still needs a controlled vocabulary, because the failure mode is the usual one. When half the team writes partner-acme and the other half writes Acme Partners, one placement becomes three rows. Terminus, the marketing taxonomy governance platform, validates those values when the link is created. For the wider convention, see the 2026 UTM tagging guide.
What this buys you is a clean split: placements you tagged, assistants you can see by referrer, and an honest unknown for the rest. That is a measurable GEO programme. A single invented source value is not.
Crawlers: Training Is Not Retrieval
This is the distinction that decides whether a robots.txt edit costs you citations.
Retrieval and user-fetch agents fetch pages to answer a live question. Block those and you lose in-answer citation. OpenAI’s OAI-SearchBot and ChatGPT-User sit here, as do Claude-User and Claude-SearchBot, and Perplexity-User.
Training agents collect data to train models. GPTBot, ClaudeBot, CCBot, Google-Extended and Applebot-Extended sit here. Blocking them is a licensing and policy decision, not a visibility one. Google says directly that Google-Extended does not affect inclusion in Google Search.
Anthropic’s agent names changed, and old articles still list Claude-Web, which is retired. The current three are ClaudeBot, Claude-User and Claude-SearchBot. Check each vendor’s documentation before editing robots.txt, because this inventory has churned repeatedly.
As for llms.txt, the proposed standard published at llmstxt.org in late 2024: no major vendor has committed to honouring it. It costs almost nothing to publish and promises nothing in return.
The consent and copyright questions behind all of this remain unsettled. Decide your stance with legal, not with a blog post.
A 90-Day Plan
Days 1 to 30, measure what is true today. Inventory the pages that already rank for your priority queries. Baseline 30 days of server logs for the crawler agents above, split into training and retrieval. Run a citation baseline, either with a vendor tool or a fixed set of 30 to 50 prompts you run weekly by hand. Check which assistants your GA4 AI Assistant channel already catches, and build a custom channel group for the rest.
Days 31 to 60, fix the words first. On your top 20 pages, apply what the research supports: cite primary sources, quote authorities by name, and replace vague claims with specific figures that carry a year and a link. Then apply what we believe helps retrieval: a self-contained answer in the first 80 words, a short TL;DR, question-shaped FAQ headings, and a visible last-updated date. Add Article, Organization and FAQPage markup because it is good practice, not because it buys citations.
Set the tagging convention while you are here: real sources and mediums for every placement you control, and no invented AI source.
Days 61 to 90, close the loop. Report AI-referred sessions and conversions weekly next to the crawl baseline. Find queries where competitors are cited and you are not, and classify each gap as missing content, weak sourcing or stale dates. Rewrite the worst ten. Re-run the citation baseline and compare.
Expect partial movement inside a month and more over a quarter or two. Then do it again, because a fixed playbook and a moving set of engines is not a durable combination.
FAQ
What is GEO?
Designing content, markup and measurement so generative AI systems surface and cite your pages. The term comes from the 2023 Princeton paper “GEO: Generative Engine Optimization” by Aggarwal et al.
How does GEO differ from SEO?
SEO competes for a clickable link on a results page. GEO competes to be quoted inside a generated answer. They overlap heavily, because both reward trusted, well-sourced, well-structured content. GEO puts more weight on the sourcing inside your text, because that is what the research measured.
Does schema markup get me cited?
There is no evidence that it does. Google says no special structured data is required for AI features, and the Princeton study tested text rewrites rather than markup. Publish schema because it helps search features and keeps your metadata honest.
What were the actual numbers in the paper?
Up to 40 percent improvement in visibility for the strongest methods, and 115 percent for a source that began ranked fifth, measured on the paper’s own benchmark with its own metric. Treat them as directional evidence for the approach rather than a forecast for your site.
Which AI crawlers should I allow?
Allow the retrieval agents if you want citations: OAI-SearchBot, ChatGPT-User, Claude-User, Claude-SearchBot, Perplexity-User. Training agents such as GPTBot, ClaudeBot, CCBot, Google-Extended and Applebot-Extended are a policy choice, and blocking Google-Extended does not affect Google Search inclusion.
Is llms.txt a real standard?
It is a real proposal from late 2024 with no major vendor commitment behind it. Publishing one is cheap insurance, not a strategy.
Should I tag links with an AI source value?
No. UTMs override the referrer, so an invented ai_referral source relabels every ordinary visitor to that page as AI traffic, while the real assistant click arrives carrying the citing site as its referrer. Tag placements for what they are and measure assistants through the referrer and GA4’s AI Assistant channel.
Can I pay to be cited?
Not for the organic citation list. Advertising has arrived around these products (ChatGPT began testing ads in February 2026, and AI Overviews carry ads), and vendors describe those placements as separate from the citations models produce when answering. Assume that separation holds only as long as it is stated.
How long does GEO take?
Partial movement within 30 days for content changes, and more over 90 to 180 days. Trust signals move slowest, because they are earned rather than configured.
Does GEO replace SEO?
No. Many users still click search results, and several assistants draw on the same underlying index. What changes is that thin, click-optimised pages perform badly when the unit of value is a quotable, self-contained passage.
Every account starts with a 21-day trial, no credit card required.