Skip to main content

SBPO Consulting · Search

Generative engine optimization, and the claims we will not make

Answer engines do not maintain a separate index you can be optimised into. They retrieve, and they mostly retrieve from search. So generative engine optimization is largely the work of being findable, unambiguous and quotable — plus a small number of genuinely new decisions about crawler access and measurement. We will tell you which parts of the current GEO market are real and which are a text file with an invoice attached.

Where we come in

The GEO and AI search problems this solves

  • Your board has asked what the company is doing about AI search, and you do not have an answer you actually believe.
  • Ask an assistant who to use in your category and a competitor gets named while you do not.
  • Impressions are holding steady but clicks are drifting down, and nobody can tell you whether AI answers are the reason.
  • You have been quoted for a GEO package whose headline deliverable is a text file, and something about that did not sit right.
  • Legal wants AI crawlers blocked, marketing wants the visibility, and the robots.txt decision keeps getting deferred to the next meeting.
  • When an assistant does describe your business it gets the category, the locations or the product line wrong, and you have no idea where it read that.
  • You cannot see how you would measure whether any of this worked, so the budget conversation goes nowhere.

What generative engine optimization actually is

Generative engine optimization is the practice of making a source more likely to be retrieved, quoted and correctly attributed when an AI system composes an answer. The term comes from a 2023 research paper by Aggarwal and colleagues, later presented at KDD 2024, which built a benchmark of generative-engine answers and tested which changes to a source document improved its visibility in the response. Their headline result — that certain content changes, particularly adding citations, quotations and statistics, improved visibility by up to around forty per cent in their benchmark — is genuinely interesting and routinely misquoted. It was measured on a research harness, it varied substantially by domain, and it is not a claim about what ChatGPT will say about your company next Tuesday.

That gap between an interesting finding and a sales promise is most of what is wrong with the current GEO market. So it is worth being precise about the mechanism.

The retrieval layer is still a search index

Answer engines do not consult a private ranking of the web that vendors can influence. They run retrieval — often several searches — against an index, pull passages from the documents that come back, and generate prose over them. Google is unusually direct about the consequence: there are no additional requirements to appear in AI Overviews or AI Mode beyond being indexed and eligible to be shown with a snippet, and no special files, markup or markdown are needed. ChatGPT’s search results are served by its own search crawler working over pages it has been permitted to fetch. Perplexity’s bot exists to surface and link websites.

If your page cannot be crawled, rendered, indexed and snippeted, it is not a candidate for any of this. That is why the first phase of any engagement here is technical SEO rather than anything AI-specific — a fact that makes for a poor pitch deck and a much better outcome.

GEO, AEO and SEO: what is genuinely new

Three acronyms are competing to describe roughly the same work, and the distinction matters less than the vendors selling it would like. Answer engine optimization generally means being the source of a direct answer; generative engine optimization means being retrieved and cited inside a synthesised response. Both sit on top of search.

Stripped of the marketing, exactly three things are new:

Crawler access is now a business decision. Historically you allowed search crawlers because there was no reason not to. Now training access, search access and user-triggered fetching are separate permissions with different commercial consequences, and the default is whatever your platform shipped.

Extraction has replaced the click as the unit of value. A model lifts a passage. Whether that passage makes sense without the paragraph above it is now a content design question.

Measurement broke. Rank tracking does not describe a non-deterministic answer that varies by user and by run.

Everything else on the typical GEO deliverables list — good content, clear structure, technical health, real authority — is the organic search programme you should already have been running.

Establishing a baseline before touching anything

The most common failure in this work is the absence of a control. These systems return different answers to the same prompt on different days, so “we did GEO and now ChatGPT mentions us” is not evidence of anything unless somebody recorded what it said before.

A baseline means an agreed prompt panel — the questions your buyers actually ask, including the unflattering comparison ones — run repeatedly across the engines that matter to you, with results recorded rather than remembered. What comes out of it is usually more useful than the visibility score attached to it: you find out which competitors are consistently named, which third-party sources the models keep quoting when describing your category, and which specific factual errors about your business are circulating. That last category is often the highest-value finding in the whole engagement, because it is fixable and nobody knew about it.

Entity foundations: being an unambiguous thing

Answer engines are describing entities, not documents. If a model cannot resolve which organisation you are, it will either omit you or blend you with something else — and the second failure is worse, because it produces confident, wrong statements about your business.

Resolving the entity is unglamorous and largely within your control. One consistent legal and trading name. One consistent category description used in the same words on your own site, in your structured data and on the third-party profiles you control. Organization markup with a stable identifier and sameAs links pointing at profiles that genuinely exist and genuinely belong to you. Consistent location and contact details everywhere they appear, which is where this overlaps directly with local search work.

A word on Wikidata, which is regularly sold as a GEO deliverable. Its notability policy admits an item if it has a valid sitelink to a Wikimedia project, or refers to a clearly identifiable entity that can be described using serious and publicly available references, or fills a structural need. That second condition is a real bar, not a formality: an entry created for a business with no independent published coverage is a candidate for deletion, and creating one about yourself is exactly the pattern that attracts scrutiny. We will tell you whether you plausibly meet the criterion. We will not manufacture the sources that would make you appear to.

Content structured so a passage survives being lifted

Content written to be extracted looks different from content written to be read top to bottom, and the difference is mostly about self-containment.

The practical rules are dull and effective. Answer the question in the first two sentences of the section, before the context and the caveats. Make each section stand alone, so a passage pulled out of the middle of the page still says what it means — pronouns pointing at a heading two screens up do not survive retrieval. State definitions plainly and once. Put comparisons in an actual table rather than in prose, because a table is trivially parseable and a paragraph of hedged comparison is not. Attribute claims to a source the reader and the model can follow, which is why the citations block at the foot of this page exists.

There is also a quality floor underneath all of this that no amount of formatting substitutes for. Google’s guidance on generative AI features says the same thing its content guidance has said for years: commodity material restating what is already everywhere has nothing to be retrieved for. The pages that get cited tend to carry something first-hand — a method, a constraint, a number somebody actually measured, an opinion with a reason attached. That is a content problem more than a technical one.

Structured data and machine-readable facts

Worth stating clearly, because it is oversold: there is no schema type that makes a model cite you, and Google has said explicitly that no special structured data is needed for its AI features.

What structured data does is reduce ambiguity. Organisation markup that states your name, description, identifiers and verified profiles gives a consistent machine-readable version of facts that are otherwise scattered across a footer, an about page and a directory listing from four years ago. Breadcrumb markup states where a page sits. Article and service markup state what a page is. None of that is a ranking lever, and all of it makes you cheaper to describe correctly.

The corollary is that inconsistent facts are actively harmful. If your site says one thing, your profiles say another and your markup says a third, you have given a retrieval system three candidate answers and no way to choose. Pick one version of the truth and repeat it everywhere.

Crawler access: the lever most sites get backwards

This is the part of GEO that is genuinely new, genuinely consequential, and genuinely misunderstood — and it is a policy decision, not a technical one.

The agents are not interchangeable. OpenAI documents GPTBot as the crawler that gathers content for model training, and OAI-SearchBot as the one that powers search results inside ChatGPT; its documentation states that sites opted out of OAI-SearchBot will not appear in ChatGPT search answers, and that blocking GPTBot does not have that effect. Google-Extended controls whether crawled content is used to train and ground Gemini, and Google states it does not affect inclusion in Google Search or act as a ranking signal. Anthropic publishes an equivalent split between its training crawler, its search crawler and the fetches Claude makes when a user asks for a page. Perplexity distinguishes its indexing bot from user-initiated retrieval, and notes that user-initiated fetches generally do not follow robots.txt because a person asked for that specific page.

Two conclusions follow. First, the sentence “we blocked the AI bots” is not a description of a policy; it needs a list. We have seen sites remove themselves from assistant search results while intending only to opt out of model training, which is an expensive way to protect an asset nobody was going to pay for anyway. Second, blocking training crawlers does not remove your content from models already trained, nor from pages a user pastes in, so it should be chosen for the reasons it genuinely serves — licensing position, competitive concern, principle — and not on the assumption that it prevents your material being used.

We write this up as a decision record with a date and a rationale, and we verify against server logs that the agents you allowed are the agents actually arriving.

Third-party corroboration, and the tactics we refuse

Models describing a category tend to lean on sources that are not the vendor’s own website: industry publications, comparison sites, forums, standards bodies, established directories. Being present and accurately described in those places matters more here than it does in classic search, because a self-description carries less weight when several independent sources are available to synthesise.

That creates an obvious temptation, and a whole cottage industry servicing it: paid Wikipedia editing, listicles placed on freshly registered “best agencies” domains, reciprocal mention schemes, review-site seeding. These work briefly. They are also trivially recognisable, they violate the platforms’ own policies, and when they unravel the reputational cost lands on your brand and not on the agency that ran them. We do not do them, and we will say so early rather than let it become an awkward conversation later.

The legitimate version is slower: earning coverage by having something worth covering, appearing in the registers and directories your industry genuinely uses, publishing data other people want to cite, and being the source that a journalist or an analyst quotes. It is the same work that has always built authority, which is why it fits inside a wider marketing programme rather than sitting off to one side as an AI project.

Measuring AI visibility honestly

Official data exists and it is thin. Google’s Search Console now includes a generative AI performance report covering AI Overviews and AI Mode, but it reports impressions only — no clicks, no position — and it covers Web search results only, with rollout still incomplete. That is genuinely useful as a trend line and useless as an attribution model.

Everything else is inference. Referral traffic from assistants shows up in analytics where those platforms pass a referrer, and it is typically small in volume and disproportionately high in intent. Server logs show which AI agents are fetching which pages, which tells you about access rather than about citation. Commercial visibility trackers run prompt panels at scale and produce a share-of-citation number, which is directionally useful provided everyone remembers it is a sample of prompts chosen by a vendor, not a measurement of the market.

So we report three things separately and do not blend them into a single score: what the official impression data says, what the repeated prompt panel says, and what happened to branded search, direct traffic and qualified enquiries. That third group is the one that pays for the work. Building it properly is analytics work, and it is worth doing before the GEO budget rather than after it.

What we will not promise

No guarantee of citation, because nobody controls a model’s output. No submission to an AI index, because none exists for organic answers. No claim that a specific file, tag or format causes inclusion, because the platforms have said the opposite in writing. No before-and-after screenshot presented as proof, because a single prompt run proves nothing about a non-deterministic system. And no recommendation dressed up as established practice when it is really a plausible hypothesis — we would rather label it as one and let you decide whether to fund an experiment.

What we will say is that the foundations are unusually stable for a new field: be crawlable, be indexed, be an unambiguous entity, be genuinely useful, be quotable, be corroborated, and decide your crawler policy on purpose. If you want a candid read on whether that work is worth doing for your business right now, start a conversation and we will tell you what we would do in your position, including when the answer is nothing yet.

Scope

What our generative engine optimization services include

Every engagement is scoped in writing before it starts. These are the artefacts that leave our hands and become yours.

  1. AI visibility baseline

    A defined set of prompts covering the questions your buyers actually ask, run across the assistants that matter to you and recorded on a fixed schedule: who is named, what is cited, which pages are pulled from, and what is stated about you that is wrong. Without this, nothing that follows can be judged, because these systems are non-deterministic and memory is a poor control group.

  2. Entity audit and disambiguation plan

    How your organisation is currently represented across your own site, your structured data, third-party profiles and public reference sources — and where that representation is inconsistent, incomplete or confusable with another business. The plan states what to assert, where to assert it, and which corroborating sources are realistically obtainable.

  3. Crawler access decision record

    A per-agent ruling covering the training crawlers, the search crawlers and the user-triggered fetchers, each with what it actually gates, what blocking it costs and what it protects. This is a commercial and legal decision as much as a technical one, so it is written to be signed off rather than quietly implemented.

  4. Extractable content specification

    Template and editorial rules that make an answer liftable: a direct answer near the top, self-contained sections that survive being read out of context, definitions stated plainly, comparisons in real tables, and claims attributed to sources a model can follow. Written as guidance your writers can apply without us.

  5. Machine-readable facts layer

    JSON-LD for organisation identity, breadcrumbs, articles, services and products, mapped to fields your CMS holds, plus canonical fact pages for the details that get asked about most — what you do, where, for whom, and on what terms. Markup is not a shortcut to citation, but inconsistent facts are a reliable route to being described wrongly.

  6. Corroboration plan

    A realistic route to being mentioned in places other than your own website: industry publications, genuine directories, standards bodies, partner and supplier listings, and the reference sources that answer engines lean on. It names what is achievable and explicitly excludes the tactics we will not run.

  7. Measurement framework and reporting

    The Search Console generative AI performance report, referral data from assistants where it exists, server-log evidence of AI agent activity, and the repeated prompt panel — assembled into one view with its limitations stated on the page rather than in a footnote. Reporting says what we can and cannot know.

How it runs

How SBPO Consulting delivers GEO and AI search

  1. Baseline before intervention

    We record how you are currently represented across the assistants before changing anything, because these systems return different answers to the same prompt on different days and any before-and-after claim without a recorded baseline is a story. The prompt panel is agreed with you, so it reflects real buying questions rather than terms chosen because we can win them.

  2. Fix the search foundations first

    A page that is not indexed, or is excluded from snippets, is not a candidate for any of this. Google states plainly that there are no additional requirements to appear in its AI features beyond being eligible for Search with a snippet. So crawlability, rendering, canonical hygiene and snippet controls are checked before anything specific to GEO is attempted.

  3. Resolve the entity

    Make it unambiguous who you are: one consistent name, category, description and set of identifiers asserted on your own site, in your structured data and across the third-party profiles you control, with sameAs links pointing at the profiles that genuinely exist. Ambiguity is the single most common reason an assistant describes a business incorrectly.

  4. Restructure what already answers

    Most sites already hold the answers; they are buried under three paragraphs of preamble. We rewrite the highest-intent pages so the answer arrives first and each section stands alone, then extend coverage to the questions the baseline showed you losing. This is editorial work, not markup work, and it is where most of the movement comes from.

  5. Decide crawler access on purpose

    We present the trade-offs per agent and you decide. The decision is implemented in robots.txt, verified against live requests in server logs, and recorded with a date and a rationale, so the next person to inherit the file knows whether a line was a policy or an accident.

  6. Re-run the panel and report what changed

    The same prompts, the same schedule, the same recording method, reported alongside classic search performance and commercial outcomes. Where a change produced nothing measurable we will say so, because a channel this new needs an honest evidence trail more than it needs an optimistic one.

Tooling

Tools and platforms we use for GEO and AI search

We pick tools for the problem, not for the résumé. Where a platform is a poor fit we will say so before you have paid for it.

Answer engines tracked

  • Google AI Overviews
  • Google AI Mode
  • ChatGPT
  • Perplexity
  • Claude
  • Microsoft Copilot
  • Google Gemini

Search and index

  • Google Search Console
  • Bing Webmaster Tools
  • Screaming Frog SEO Spider
  • Rich Results Test
  • Schema Markup Validator

Entity and identity

  • Schema.org
  • JSON-LD
  • Wikidata
  • Google Business Profile
  • Industry registers and directories

Crawler control

  • robots.txt
  • GPTBot
  • OAI-SearchBot
  • Google-Extended
  • ClaudeBot
  • PerplexityBot
  • Server log analysis

Measurement

  • Search Console generative AI performance report
  • Semrush AI visibility tools
  • GA4
  • Looker Studio

Non-negotiables

The standards SBPO Consulting works to

These are checkable. Ask us to demonstrate any of them on your own project before you sign anything.

  1. No guarantees of citation

    Nobody controls what an answer engine says, and any agency offering a guaranteed mention in ChatGPT or AI Overviews is describing an outcome they cannot deliver. We commit to method and measurement, not to a model output.

  2. No manufactured corroboration

    We do not pay for Wikipedia edits, commission fake comparison listicles on disposable domains, or seed review sites. Those tactics work briefly, are trivially identified, and leave the liability with your brand rather than ours.

  3. Crawler access is your decision, in writing

    We explain what each agent gates and what blocking it costs in visibility, then implement whatever you decide and record the reasoning. We will not opt you out of training crawlers by default and call it a privacy feature, and we will not opt you in quietly either.

  4. We tell you when a tactic is unproven

    Some of this field is evidenced, some of it is plausible, and some of it is folklore that spread quickly because it was easy to sell. Recommendations are labelled accordingly, and anything in the third category is not billed as work.

Questions

generative engine optimization services — questions we get asked

Is generative engine optimization different from SEO?

Mostly it is SEO with a different reporting layer and three genuinely new decisions. Answer engines are retrieval systems: they run searches, pull passages from pages, and compose an answer, so being indexed, being eligible for a snippet and being clearly relevant remain the entry conditions. Google is explicit that there are no additional requirements or special optimisations for appearing in its AI features. What is actually new is crawler access policy for AI agents, content structured so a passage survives being lifted out of context, and measurement, because impressions and citations in generative answers do not behave like blue-link rankings. If someone presents GEO as a wholly separate discipline requiring a new stack, ask them which mechanism they are describing.

How do we find out whether AI tools mention our brand?

By testing systematically rather than anecdotally. Agree a panel of prompts that reflect how buyers actually ask — category questions, comparison questions, "who should I use for" questions, and questions about you by name — then run them on a fixed schedule across the assistants that matter, recording which brands are named and which URLs are cited. Because these systems are non-deterministic, a single run tells you almost nothing; the pattern across repeated runs is the signal. Commercial AI visibility tools automate this and are useful for trend direction, but they sample a prompt set someone else chose, so treat them as a thermometer rather than a census.

Should we block or allow AI crawlers like GPTBot?

It depends what you are trying to achieve, and the important thing is that these are separate controls that many sites conflate. OpenAI documents distinct agents: GPTBot is the training crawler, OAI-SearchBot powers search results inside ChatGPT, and blocking GPTBot does not remove you from those search answers — only blocking OAI-SearchBot does that. Google-Extended governs whether crawled content is used for Gemini training and grounding, and Google states it has no effect on inclusion or ranking in Google Search. Anthropic and Perplexity publish similar splits between training, search and user-triggered fetching, and user-triggered fetches often fall outside robots.txt entirely. The common expensive mistake is blocking the search crawler while intending only to opt out of training. Decide the two questions separately.

Does llms.txt actually do anything yet?

Not for Google, on Google's own account. The proposal — a markdown file at the site root listing your most useful pages for language models — is a reasonable idea with real adoption among documentation platforms, and it costs almost nothing to publish. But Google has said directly that you do not need to create machine-readable files, AI text files, markup or markdown to appear in Search or its AI features, and that such files neither help nor harm. No major answer engine has published a commitment to consume it as a ranking or retrieval input. Our position: publish one if you have a large documentation set and it is cheap to generate, and treat anyone selling llms.txt as the centrepiece of a GEO engagement with appropriate suspicion.

Can you guarantee our brand will appear in AI answers?

No, and neither can anyone else. Answer engines compose responses from retrieved sources using models whose outputs vary between runs, personalise by context, and change without notice when the underlying system is updated. There is no submission process, no paid inclusion for organic citations, and no ranking dashboard. What can be committed to is the work that makes citation more likely — being indexed and snippet-eligible, being an unambiguous entity, publishing content that answers the question directly and can be quoted accurately, and being corroborated by sources other than yourself — plus honest measurement of whether it moved. A guarantee in this space is a claim about someone else's model, made by a party with no access to it.

How does GEO affect traffic if AI answers reduce clicks?

This is the right question and it deserves an uncomfortable answer: for some query types, visibility and traffic have genuinely decoupled. A definitional or how-many-grams question can be answered completely in the response, and the click that used to follow does not happen. That is not fixable by optimisation, and pages whose entire value was capturing that click are worth less than they were. What still converts is the query where someone needs to evaluate, compare, trust or buy — and there the citation is doing work even when the click count falls, because being the named source in the answer shapes the shortlist before anyone visits a website. Practically, it means shifting measurement toward brand-qualified demand, direct and branded search, and assisted conversions, and being realistic that some informational traffic is not coming back.

Do we need GEO if our SEO is already strong?

You are most of the way there, and that is the honest answer rather than a modest one. Strong technical foundations, genuine topical depth and real authority are the same inputs answer engines rely on, so sites that already earn organic visibility tend to be cited disproportionately. What a strong SEO programme usually still lacks is three things: a deliberate decision about AI crawler access rather than whatever the platform default happens to be, content structured so a passage can be lifted without losing its meaning, and any measurement of AI visibility at all. Those are worth doing on their own merits and they are a fraction of the work of building the underlying authority. If your search foundations are weak, GEO is not the place to start.

Adjacent work

Technical SEO

Crawl and indexation audits, rendering and performance work, structured data and migration planning — the layer that decides whether your content is eligible to compete at all.

Content marketing

Strategy, topic clusters, briefs, editorial production and a refresh cycle — built around what buyers are trying to find out, and measured against pipeline rather than published volume.

Digital marketing

Strategy, search, paid media, content and measurement run as one programme, so you can say which channel earned the enquiry rather than which one claimed it.

Part of our SEO practice.

Search

Let's talk about your GEO and AI search work.

Send us the problem, the constraint and the deadline. You will get a considered reply from someone who would actually do the work — not a templated proposal.