What is LLM SEO?
LLM SEO is search engine optimisation aimed at large language models rather than at a page of blue links. The work is to make your content easy for a model to retrieve, easy to lift a clean answer from, and safe to attribute, so your brand is part of the answer when someone asks ChatGPT, Perplexity, Gemini or Google's AI Overviews about your category.
The naming is a mess, so clear it up once. LLM SEO, LLM optimisation, LLMO and SEO for LLMs all describe the same discipline. Generative engine optimisation names the same work from the engine's side. Answer engine optimisation is the older, broader term, covering snippets and voice answers too. Pick one label internally and get on with the work.
LLM SEO
Optimising for retrieval and citation by large language models. Also written LLM optimisation or SEO for LLMs.
LLMO
Large language model optimisation. Same discipline, shorter label.
GEO
Generative engine optimisation. The same work named after the engine rather than the model.
AEO
Answer engine optimisation. The broader parent term, covering snippets and voice answers too.
The distinction that matters is not between acronyms. It is between the two things a model can do with your content: recall it from training, or retrieve it live when someone asks. Different mechanisms, different levers, and most practical decisions follow from which one you are influencing.
How LLMs retrieve and cite: the two pathways
An LLM can reach your content two ways: from training data, absorbed during training and recalled from memory, or through live retrieval, where the model searches at the moment of the question and reads the results before answering. Almost all citation happens through live retrieval, because that is the only pathway that produces a link.
Training data is slow, opaque and outside your control. Content published today may appear in a model trained next year, you cannot check, and recall rarely produces a clickable source. Live retrieval is where the addressable work sits, and its pipeline is predictable.
- 01
The question is rewritten.
What the model does
turns a conversational question into several search queries, often phrased unlike anything a human would type.
What you can influence
cover the natural-language phrasings of your topic, not just the head keyword.
- 02
A search runs against an index.
What the model does
queries an underlying search index. ChatGPT leans on Bing, Perplexity on its own crawl, AI Overviews and AI Mode on Google.
What you can influence
be indexed everywhere that matters. Absent from Bing means largely absent from ChatGPT.
- 03
Candidate pages are fetched.
What the model does
fetches a shortlist of results and pulls the rendered text.
What you can influence
crawler access and server-rendered content. If your answer only exists after client-side JavaScript runs, assume it is never read.
- 04
Pages are split into passages.
What the model does
chunks each page and scores the passages against the rewritten query. It is not judging your page as a whole; it is hunting the block that answers the question.
What you can influence
passage design. Sections that stand without the paragraph before them survive chunking.
- 05
Passages are ranked and reconciled.
What the model does
compares passages across sources, favours claims several sources agree on, and drops sources that contradict the consensus or themselves.
What you can influence
consistency. One figure published three ways is the fastest route to being dropped.
- 06
An answer is composed and attributed.
What the model does
writes an answer from the surviving passages and cites the sources it leaned on hardest.
What you can influence
liftability. The quoted passage is usually the one that already reads like a finished answer.
How each assistant sources its answers
Each assistant answers from a different index, which is why the same question returns different sources in each. Optimising for "AI" as one destination is a category error: there are four or five destinations with different entry requirements.
ChatGPT and ChatGPT Search
Draws on the Bing index for live retrieval. Crawlers: OAI-SearchBot for the search index, ChatGPT-User for live fetches, GPTBot for training. Bing Webmaster Tools is the highest-leverage account in the discipline and almost nobody in the UK mid-market has it.
Perplexity
Runs its own crawl and index rather than licensing one. Crawler: PerplexityBot. It cites more sources per answer than any other assistant, which makes it the easiest first citation to earn.
Google AI Overviews
Composed over Google's index on the classic results page, governed by Googlebot, with Google-Extended controlling training use. Ranking in the top ten correlates strongly with citation but is neither necessary nor sufficient.
Google AI Mode
Same index, different surface, different traffic problem: a conversational session that may never return a page of links. Report it separately from AI Overviews.
Gemini
Grounds answers in Google Search results. Work that earns AI Overview citations tends to surface here too.
Claude and Microsoft Copilot
Claude retrieves through a web search tool as ClaudeBot; Copilot sits on Bing. Both reward the same substrate: clean indexation, clear passages, consistent entity data.
Overlap between these citation sets is lower than most people expect. A page cited by Perplexity is not automatically cited by ChatGPT. Our brand tracker, covering 8,400 prompts across ChatGPT, Perplexity, Gemini and Claude, breaks cross-engine agreement down in the AI search visibility statistics study. Measure each engine separately.
What is the difference between traditional SEO and LLM SEO?
Roughly eighty per cent of the work is identical. Crawlability, indexation, speed, information architecture, internal linking and earned links all still apply, because live retrieval runs on a search index and search indexes work as they always did. What changes is the unit of competition, the shape of a winning answer, and how you know it worked.
| Signal | Traditional SEO | LLM SEO |
|---|---|---|
| Unit of competition | The page | The passage |
| Query input | One typed query | Several machine-rewritten queries |
| Winning shape | A page satisfying intent overall | A self-contained, finished answer block |
| Index that matters | Google, Bing, Perplexity's own crawl | |
| Role of position | Position one wins the click | Making the fetched shortlist is what counts |
| Role of links | A primary ranking signal | Still a signal, plus unlinked mentions |
| Consistency | Contradictions cost little | Contradiction gets the source dropped |
| Feedback loop | Days to weeks | Weeks to months, and noisier |
| Primary metric | Rankings, impressions, clicks | Citations, mentions, sessions, enquiries |
| Reporting source | Search Console | GA4 channels, trackers, prompt panels |
The row that trips people up is the last one. Traditional SEO gives you a daily rank check. LLM SEO gives you a citation that appears and vanishes depending on phrasing. That is a reason to measure differently, not to skip measurement.
What actually influences whether an LLM cites you
Content structure is the strongest lever you directly control. Across 12,400 AI Overview queries, a definitive opening sentence under each H2 correlated with citation at 0.49, and citation density, meaning how often the page cites named checkable sources, correlated at 0.42. Both figures and the full factor set are in the AEO statistics study.
Ranked by what we see moving the needle on client accounts:
- 1
Definitive opening sentences.
A question-shaped heading answered in the first sentence beneath it, in a form that lifts whole. Correlation 0.49.
- 2
Citation density.
Pages that name their own sources get quoted more than pages that assert. Correlation 0.42.
- 3
Indexation in Bing.
Binary, and frequently the whole explanation for absence from ChatGPT.
- 4
Passage self-containment.
Sections that survive being read without the paragraph above them.
- 5
Entity consistency.
The same organisation name, address, founder and description everywhere a model can find you.
- 6
Off-site brand mentions.
Including unlinked ones, on forums, review sites, listicles and press.
- 7
Original data.
Models quote numbers, and prefer them with a named source and a stated method.
- 8
Freshness signals.
Visible publish and update dates, honestly maintained.
- 9
Schema markup.
Organization, Article, FAQPage and sameAs, used to disambiguate rather than decorate.
- 10
Classic authority.
Domain strength helps but is not the gate. On this topic's own results, DR 30 and DR 40 sites rank beside DR 91, and our DR 56 site is cited in AI Overviews on commercial terms.
How to structure content for retrieval
Write so that any single section can be lifted out of the page and still make sense. That is the whole rule, and everything below follows from it.
Do
Don't
Open every H2 with a one-sentence answer to that heading
Open with context that sets up the answer later
Phrase headings as the questions people ask
Use clever, abstract or brand-flavoured headings
Repeat the subject noun rather than saying “it”
Lean on the previous paragraph for meaning
Put comparisons in a real HTML table with a caption
Describe a comparison across several paragraphs
Name your sources, with figure and method attached
Assert numbers with no attribution
State one canonical version of each figure site-wide
Let three pages quote three versions of one metric
None of this requires shorter content. Long pages are cited constantly. What gets cited is a well-shaped passage inside a long page, which is a different thing from a short page.
What content formats get cited most by LLMs?
Definitions, question-and-answer blocks, comparison tables, step lists and original data with a stated method. Each survives chunking intact because each is already a complete unit of meaning. Narrative case studies, opinion essays and pages whose argument only lands cumulatively get read and rarely get quoted.
Schema, entities and brand consistency
Structured data does not make a model cite you, but it removes the ambiguity that stops it. Its job in LLM SEO is identity: proving that this organisation, this author and this article are the same entities the model met elsewhere, so it can attribute a claim to you with confidence.
Four types carry almost all the value. Organization schema, with a stable identifier, legal name, address and a sameAs array pointing at every profile you control. Article schema on every post, with a named author and publish and modified dates that are true. FAQPage schema on question blocks, which puts your answers into the machine-readable layer in the shape a retrieval system wants. And Person schema for the author, so expertise attaches to a name.
Entity consistency is the unglamorous half. If your company name is written four ways across your site, your public company record, your LinkedIn page and your review profiles, you are asking a model to reconcile four entities into one and hoping it guesses right. It often does not, and the failure is silent: you are simply not the brand it names. Fix the trading name, address format and founder's name everywhere before spending on content. This is ordinary technical SEO agency work and the cheapest win in the discipline.
- Organization schema with stable identifier and full sameAs array
- Article schema with named author and real publish and modified dates
- FAQPage schema matching the visible copy exactly
- Person schema for the author, linked from every article
- Identical organisation name and address in every structured data block
- Directories and review profiles aligned to that same name and address
- One canonical business description, reused verbatim off-site
Crawler access, Bing indexing and llms.txt
You cannot be cited by a model that was never allowed to read you, and you cannot be retrieved from an index you are not in. Two of the three items below are essential. The third is optional and probably does nothing yet.
01 · Allow the crawlers you want, knowingly.
OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot and Googlebot handle retrieval and citation. GPTBot, Google-Extended and CCBot relate to training collection. Blocking training crawlers while allowing retrieval crawlers is a defensible publisher position. Blocking everything, which plenty of sites did in 2023 and never revisited, removes you from the answer set. Check robots.txt today rather than assuming.
02 · Get into Bing, properly.
Verify the site in Bing Webmaster Tools, submit the sitemap, enable IndexNow, and compare the indexed page count against your sitemap count. This is the most commonly skipped step in the discipline, and for ChatGPT visibility it is closer to a prerequisite than a tactic.
03 · llms.txt: honest verdict.
llms.txt is a proposed plain-text file at your site root listing your key pages for language models, like a curated contents page. It has real search demand and plenty of advocacy. There is no published confirmation from OpenAI, Anthropic, Google or Perplexity that any of them read it, and we have seen no citation change on accounts that added one. Ship it for the option value; do not let anyone bill you for it as a strategy.
How do you measure LLM SEO results?
You measure LLM SEO with four things: an AI channel in your analytics, a citation tracker or manual prompt panel, branded search volume in Search Console, and attributed enquiries wherever you record leads. Rankings do not apply and impressions do not exist. What you have instead is a referral trail, more traceable than this discipline's reputation suggests.
We know because we have traced it. In July 2026 two client enquiries arrived through our own AI Assistant channel and were attributable to ChatGPT. Not impressions, not a share-of-voice score: two named enquiries, from a site at DR 56. Here is how that was set up and followed.
--
AI Assistant sessions
monthly, GA4
--
Cited pages
URLs appearing in AI answers
--
Prompt panel mentions
out of 20 prompts
--
AI-attributed enquiries
the commercial number
How do I track AI referral traffic?
Set up an AI channel in GA4 and stop relying on default reports. GA4 now surfaces an AI Assistant channel of its own, which is where we first saw this traffic appear as a named line rather than scattered referrals. Treat it as the headline and build a custom channel group underneath as a cross-check, because the default grouping is not retroactive.
To build it: in GA4 open Admin, then Data display, then Channel groups, and create a custom channel group. Add a channel named AI search, defined as session source matching a regular expression covering the assistant domains: chatgpt.com|chat.openai.com|perplexity.ai|gemini.google.com|copilot.microsoft.com|claude.ai|bing.com/chat|edgeservices.bing.com|you.com. Order it above Referral and Organic Search so it claims those sessions first. Add a second condition on session campaign or manual source containing chatgpt.com, because some links arrive tagged rather than referred.
How do I monitor AI search visibility?
Watch the referral pattern, not just the volume. AI-sourced sessions look distinctly different from organic ones, and the differences are how you sanity-check the channel.
- They land on deep pages, not the homepage. The model cites a specific answer, so the visitor arrives at it.
- Session counts are small and engagement is high. This traffic reads.
- Some of it is invisible. In-app browsers, and anyone copying a link out of a chat, land in Direct with no referrer. Assume your AI channel undercounts.
- Referrer strings differ between assistants and change without notice. Review the expression quarterly.
- There is no keyword. Search Console will not help. The nearest proxy is branded query volume, which is why it sits in the monthly set below.
How do I audit my site for AI search visibility?
Run a manual prompt panel monthly, at zero cost, before you buy a tool. Write twenty prompts a real buyer would type: category, comparison, problem-shaped, and two or three direct brand questions. Run each in ChatGPT, Perplexity, Gemini and Claude, logged out, with memory and personalisation off. Record: were you mentioned, were you cited with a link, which page, which competitors appeared. Repeat the identical panel monthly.
Twenty prompts across four assistants takes about ninety minutes. It gives you a baseline, tells you which pages models reach for, and means that when you buy a tracker you can check whether it detects the citations you already know exist.
The attribution chain: how we traced two enquiries to ChatGPT
- 1
Channel. The session lands in the AI Assistant channel in GA4, with the assistant domain as the source. That is the weakest link, and alone it proves nothing.
- 2
Landing page. Record which URL the session entered on. This tells you which page the model cited, the most actionable fact in the chain.
- 3
Session path. Follow the pages viewed. AI-sourced visitors typically read the cited page, then a service page, then the contact page, in one session.
- 4
Form capture. The enquiry form stores the session's attribution with the submission, so the source travels with the lead rather than being reconstructed later. Without this the chain breaks.
- 5
Enquiry record. The stored source is written to the enquiry record, so the lead carries the assistant domain as its origin from the moment it exists.
- 6
Confirmation. Ask, in the first reply: how did you come across us? Two July 2026 enquiries came back naming ChatGPT, matching the channel data. A self-report is not evidence alone, but agreement between report and channel is as close to proof as this discipline gets.
Two enquiries is a small number and we are not presenting it as a benchmark. It is a method: six steps that turn an anecdote into something a finance director will accept. The wider picture, across 14.7 million AI-attributed sessions, is in our ChatGPT and AI search referral statistics study.
What to track monthly
| Metric | Where it lives | What it tells you |
|---|---|---|
| AI channel sessions | GA4 custom channel group | Direction of travel, never a full count |
| Landing pages from AI sessions | GA4 landing page report | Which pages models actually cite |
| Prompt panel mentions | Your own 20-prompt panel | Visibility across engines, like for like |
| Cited URLs and terms | Ahrefs Brand Radar or equivalent | Where citations sit, and who shares them |
| Branded search volume | Search Console | Lagging proof exposure creates demand |
| Bing indexed pages | Bing Webmaster Tools | Whether you are eligible for ChatGPT |
| AI-attributed enquiries | Enquiry records | The only number that survives a board |
Review monthly, judge quarterly. Citations move week to week for reasons unconnected to your site, and reacting to one month's dip is how good pages get rewritten for nothing.
What you cannot measure
Four things, and anyone claiming otherwise is selling. You cannot see how many people saw your brand named without clicking, so there is no impressions equivalent. You cannot recover the prompt that produced a citation. You cannot separate AI-influenced direct traffic from ordinary direct traffic. And you cannot reproduce another person's result, because answers vary by account, location, personalisation and model version.
Is SEO going away with AI?
No. Live retrieval runs on search indexes, so the work that gets you into an index is the work that gets you into an answer. What is going away is the assumption that a ranking converts into a visit.
If you are not in the index, you are not in the answer. LLM SEO is not a replacement for SEO. It is SEO judged by a different scoreboard.
Traffic and visibility have decoupled. A page can be read by more people than ever while receiving fewer sessions, because the reading happens inside someone else's interface. That changes what you report and how you value a page. It does not change crawlability, indexation or the value of being genuinely worth quoting.
For most UK businesses this is one discipline with two scoreboards, delivered inside one retainer. That is how we run it, and it is set out on our generative engine optimisation agency page. For the umbrella view of AI search work alongside conventional search, our AI SEO agency page covers the whole service.
When an LLM gets your brand wrong
Fix the source, not the model. You cannot edit a model's answer, but almost every wrong answer traces back to a page, profile or record that a model can read, and those you can change.
- It describes services you no longer offer.
- Find the page still saying it: usually an old service page, a stale directory listing or a third-party profile. Update or remove it, then request reindexing.
- It gets your location, size or founding date wrong.
- Entity inconsistency. Align your Organization schema, public company record, LinkedIn page and directory listings to one set of facts, then wait for recrawl.
- It attributes a competitor's claim to you.
- Publish a clear, dated, sourced statement of the correct claim on a page you own. Consensus is what models reconcile against, so contribute to it.
- It does not mention you at all in your category.
- An absence problem, not an accuracy problem. Check Bing indexation, then off-site mentions, then content structure, in that order.
- It cites an old article instead of the current one.
- Consolidate. Two pages competing on one topic split the signal. Redirect the weaker one.
Expect weeks, not days. Retrieval refreshes on crawl cycles, and anything recalled from training data will not change until a model is retrained.
How do I improve AI search visibility? Your first 90 days
Start with access and measurement, then structure, then off-site. In that order, because there is no point improving content that cannot be retrieved, or anything you cannot measure.
Days 1-30
Access and baseline
- Audit robots.txt for OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot and Googlebot
- Verify in Bing Webmaster Tools, submit the sitemap, enable IndexNow
- Compare Bing indexed pages against sitemap pages and fix the gap
- Build the AI channel group in GA4 and confirm sessions are classified
- Confirm the enquiry form stores session attribution with each submission
- Run the 20-prompt panel across four assistants. This is your baseline
Days 31-60
Structure and identity
- Rewrite the H2 openers on your ten highest-value pages so each answers its heading first
- Convert prose comparisons into real tables with captions
- Add or correct Organization, Article, FAQPage and Person schema
- Reconcile name, address and description across every profile you control
- Add visible publish and modified dates, and maintain them honestly
- Pick one canonical version of every recurring figure and enforce it site-wide
Days 61-90
Evidence and off-site
- Publish one piece of original data with a stated method
- Pursue mentions on the sources models reconcile against: listicles, reviews, forums, press
- Rerun the identical 20-prompt panel and compare against the baseline
- Review AI channel sessions, landing pages and attributed enquiries
- Consolidate or redirect pages competing on one topic
- Ship llms.txt for the option value, expecting nothing from it
Common LLM SEO mistakes
- 1
Treating AI as one destination. Four indexes, four sets of requirements, four measurements.
- 2
Skipping Bing. The commonest cause of total absence from ChatGPT, and the easiest to fix.
- 3
Writing shorter instead of clearer. Chunking rewards self-contained passages, not thin pages.
- 4
Publishing contradictory figures. Three versions of one number is a reason to be dropped.
- 5
Buying a tracker before setting up analytics. The channel group and prompt panel cost nothing.
- 6
Reporting share of voice to a finance director. Report attributed enquiries instead.
- 7
Blocking every bot in 2023 and never revisiting robots.txt.
- 8
Expecting a fortnight to prove anything. Judge quarterly.
Questions
LLM SEO FAQs
Methodology
Correlation figures are from our analysis of 12,400 Google AI Overview queries, published in full in our AEO statistics study. Cross-engine observations come from our brand tracker covering 8,400 prompts across ChatGPT, Perplexity, Gemini and Claude. Referral behaviour references our dataset of 14.7 million AI-attributed sessions. The attribution walkthrough uses our own GA4 property and enquiry records for July 2026, a sample of two, presented as a repeatable method rather than a benchmark. Crawler names, index relationships and setup steps were verified in August 2026.
Last reviewed: August 2026. Next review: February 2027.