Generative Engine Optimization: How to Get Cited by AI in 2026

Imdad Khan Author
9 min read
Share
Generative Engine Optimization: How to Get Cited by AI in 2026

Most guides to optimising for AI search are written as if the whole field were settled. It is not. Even the name is contested — you will see GEO, GAIO, AEO and AIO used for roughly the same idea, and the acronym you pick changes nothing about what actually works. This guide uses GEO, generative engine optimisation, because that is the term the industry has largely settled on.

What is worth your attention is the mechanism underneath. When someone asks ChatGPT, Perplexity or Claude a question, the model does not return ten links. It returns an answer, and cites a handful of sources — typically two to seven domains. Getting into that handful is a different job from ranking on page one, and the gap between the two is widening.

The shift is measurable, not theoretical

What changed on Google specifically is covered in Google AI Mode vs traditional search.

The clearest number available: research from Brandlight found the overlap between the top links on Google and the sources AI engines actually cite has fallen from around 70% to below 20%.

Read that carefully, because it cuts both ways. It means ranking well on Google no longer guarantees you get cited. It also means sites that never cracked the top three on Google can still be cited, because the selection criteria are different. For a small site, that second half is the interesting part.

The crawler layer nobody checks

This is the most common and most expensive mistake in the whole field, and it takes ten minutes to check.

AI companies run two separate kinds of bot, and people routinely block both when they meant to block one:

BotWhat it is forBlock it and you lose
GPTBotTraining OpenAI modelsNothing citation-related
ClaudeBotTraining Anthropic modelsNothing citation-related
Google-ExtendedTraining Google modelsNothing citation-related
OAI-SearchBotChatGPT search and citationsCitations in ChatGPT
ChatGPT-UserFetches a page a user asked aboutLive lookups of your pages
Claude-SearchBotClaude search and citationsCitations in Claude
PerplexityBotPerplexity index and citationsCitations in Perplexity

The practical consequence: you can decline to have your work used as training data while staying fully eligible to be cited. Those are separate decisions, and a lot of sites have accidentally made both at once by pasting in a block-list they found on a forum.

If you want that split, it looks like this:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

Accidentally blocking OAI-SearchBot is the single highest-impact error in this area. It is worth opening your own robots file right now and reading it before you go any further.

What our own robots file looks like, and why

For transparency, this site names no AI bots at all. The file blocks some WooCommerce log directories and the admin area, and nothing else. That means every crawler in the table above is allowed by default.

That is a deliberate choice rather than an oversight: the tools and guides here are free and the point is for people to find them, so being quoted by an AI assistant is the same win as being found in search. If you sell the content itself, you would reasonably decide differently.

What actually gets quoted

Once the crawlers can reach you, the question is whether a model finds your page worth pulling a sentence from. The pattern across the research is consistent, and it runs against a decade of SEO habit.

Density beats length

A 2,000-word article that takes four paragraphs to reach the point is bad raw material for a model. The engine is looking for a passage it can lift that answers the question on its own. Short, self-contained, factual passages get quoted. Long wind-ups do not.

The practical test: take any section heading on your page and read only the first two sentences underneath it. Do they answer the heading? If they need the next three paragraphs for context, a model has nothing clean to take.

Specifics beat adjectives

“Fast charging” is unquotable. “140W output, enough to bring a MacBook Pro to 50% in about 40 minutes” is quotable, because it contains something a model can attribute to you. Numbers, dates, model names, measured results and named trade-offs all raise the odds. Marketing adjectives lower them.

Sources beat assertions

Content that cites where its facts came from is treated more favourably than content that simply asserts. This is the same instinct behind E-E-A-T, and it has carried over almost intact.

Structure beats prose walls

Tables, clear headings that match real questions, and short definitional paragraphs give an engine obvious extraction points. A comparison table is close to ideal: it is dense, factual, and unambiguous about what belongs to which item.

Entities, and why a single Wikipedia line outperforms a month of blogging

Models do not only read your site. They build a picture of who you are from everywhere you appear, and they weight independent mentions more heavily than anything you publish about yourself.

That is why entity work pays disproportionately. A mention on a source the models already trust can drive citations across several engines for years, while a self-published claim on your own domain carries much less weight. The practical version for a small site: make sure the basic facts about you are consistent everywhere, and that your own site declares them in structured data.

The concrete version of that is the sameAs property on your Organization schema, listing every profile you control. It is the difference between an engine guessing that six accounts belong to you and being told so directly.

A worked example from this site

Our Organization schema had no sameAs property at all until recently. Six social profiles existed, all genuinely ours, and nothing in the markup connected them to the domain. To a model assembling an entity, they were six unrelated accounts with a similar name.

Adding the property took one filter and a list of URLs. What it does not do is work on its own: sameAs is a claim you make about yourself, and the confirming signal is each profile linking back to the site. One direction without the other is half a handshake.

The wider lesson is that most GEO work is not new work. It is finishing the technical jobs that were always on the list — schema, consistent naming, crawlable pages, honest sourcing — and which mattered less when the only reader was a ranking algorithm.

How to tell whether any of it is working

This is the genuinely awkward part, and most guides skate over it. There is no Search Console for AI answers. Impressions and clicks do not exist in the way you are used to.

What you can do:

  • Ask the engines directly. Put the questions your content answers into ChatGPT, Perplexity and Claude, and record whether you are cited. Tedious, but it is a real measurement, and repeating it monthly gives you a trend.
  • Watch referral traffic by source. Perplexity and ChatGPT pass referrers. Small numbers, but the direction is informative.
  • Track branded search. AI citations often produce a branded search rather than a click, so a rise in people searching your name can be the visible edge of citations you never saw.

Expect three to six months of consistent work before any of these move. Anyone promising faster is selling something.

What we would not bother with

Two things get pushed hard and appear to be mostly noise.

Stuffing content with “according to experts” phrasing. The research supports citing real sources, not imitating the surface features of cited text. Fake authority reads as fake to a model and to a person.

Publishing volume for its own sake. The advice to increase output assumes quality holds. Fifty thin pages give an engine fifty weak candidates. Ten dense ones give it ten strong ones, and the dense pages also stand a chance in ordinary search.

A checklist you can run this week

  1. Open your robots file and confirm OAI-SearchBot, Claude-SearchBot and PerplexityBot are not blocked.
  2. Decide separately whether you want to allow the training bots. There is no wrong answer, only an uninformed one.
  3. Check your Organization schema has a sameAs list covering every profile you control.
  4. Make each profile link back to your domain, so the claim is confirmed from both ends.
  5. Take your five most important pages and rewrite the first two sentences under each heading so they answer the heading on their own.
  6. Replace three adjectives with three numbers.
  7. Add a source for any statistic you state.
  8. Ask each engine five questions your content should answer, and write down whether you are cited. That is your baseline.

The schema and entity work described here overlaps heavily with ordinary technical SEO — how to do an SEO audit covers the wider check. If the sourcing and experience side is what you need, what E-E-A-T actually means is the companion piece, and what Rank Math got wrong on this site is a concrete example of schema going astray.

Frequently asked questions

Is GEO different from SEO, or just a rebrand?
Mostly the same foundations with a different output. Crawlability, schema, clear structure and honest sourcing serve both. What changes is the target: you are optimising to be quoted rather than clicked, which rewards density over length.

Should I block AI bots from training on my content?
That depends on your business, and it is genuinely separate from citation eligibility. If your content is the product, blocking training is reasonable. If your content exists to be found, blocking it costs you little either way. Just do not block the search bots by accident while doing it.

Does GEO work for a small site?
Better than you might expect. The overlap between Google’s top results and AI citations is now under 20%, so the incumbents’ advantage transfers less than it used to. Being dense and specific is available to anyone.

How long before I see anything?
Three to six months of consistent work is the realistic range. The crawler fix is the exception — if you were accidentally blocking a search bot, unblocking it can show up much faster.

Do I need to publish more to get cited?
No. You need quotable passages. One dense page beats five padded ones, and padding actively hurts because it buries the extractable parts.

Tooling for this: our Surfer SEO review covers content optimisation in practice.

Reference. Google’s own position on content quality is set out in Google Search Central’s helpful content documentation.

Was this article helpful?
Scroll to Top