Most GEO advice is a list of tips without the pipeline behind them, which makes it hard to tell which ones matter and when. This guide starts with what actually happens inside a generative engine and gives you the steps to move forward.

If you are still deciding whether GEO is a separate discipline, start with what GEO is and GEO vs SEO. This page is the how: the pipeline, the practices that follow from it, a 90-day plan, and the things that do not work.

Table of contents
Five stages, two places to winA generative engine runs five stages: query fan-out, retrieval, chunking and selection, synthesis and attribution. The first three are won on your own site with ordinary information-retrieval work. The last two are won off it, through corroboration on domains you do not own and an entity the model can resolve.Five stages,two places to win01Query fan-out02Retrieval03Chunking andselection04Synthesis05AttributionON YOUR SITERetrieval and selectionSubstance in the HTML, pages fast to fetch, one question persection answered in the first two sentences.OFF YOUR SITECorroboration and entityWhat other domains say about you, andwhether the model can resolve yourbrand confidently.You cannot optimize for the model. You optimize for retrieval, selection and corroboration.

How does generative engine optimization work?

A generative engine answering a real question runs roughly this pipeline. Every practice further down maps to one of these stages.

Query fan-out. The user's question is rewritten into several explicit search queries, usually more specific than what was typed. One conversational question can produce five or six retrievals.

Retrieval. Each query hits an index (the engine's own, a partner's, or a live fetch) and returns candidate documents. Classic ranking still decides most of what enters the candidate pool.

Chunking and selection. Documents are split into passages and scored against the rewritten queries. Only a handful survive into the model's context. A passage that answers on its own beats a better-written one that does not.

Synthesis. The model writes an answer from the surviving passages plus what it already knows. Where sources disagree, it favours the version that appears in more of them.

Attribution. It decides whom to name and link. Named brands are usually the ones whose claim survived synthesis and whose entity the model could resolve confidently.

Two things follow from this. You cannot optimize for "the model", only for retrieval and for selection, which are ordinary information-retrieval problems. And the last two stages are won off your website, because corroboration and entity resolution depend on what other domains say.

GEO best practices, derived from the pipeline

Be retrievable (stages 1–2)

  • Make sure your substance is in the HTML, not painted in after a client-side render.
  • Keep the pages you want cited fast to fetch; live retrieval has a timeout, and a slow page silently drops out of the candidate pool.
  • Decide deliberately which AI crawlers you allow in robots.txt. Blocking them all and then buying GEO is a contradiction. And we see it more often than you would think.
  • Publish an llms.txt pointing at the pages that answer questions, and keep it current.
  • Cover the fan-out, not just the head term: the explicit, situational phrasings the engine rewrites to, which is where long-tail pages earn their place.

Be selectable (stage 3)

  • One question per section, asked in the heading, answered in the first two sentences.
  • Sections that stand alone: no back-references, no orphan pronouns, no units defined three headings earlier.
  • Tables for anything comparative: criteria, specs, prices, timelines. They survive chunking almost intact.
  • Numbers with their source and date attached, in the same sentence.
  • Structured data that repeats the visible text exactly. Our AEO guide goes deeper on this layer.

Survive synthesis (stages 4–5)

  • Say something specific enough to be checked. "Leading provider" cannot be corroborated; "works in eight languages from Andorra since 2019" can.
  • Get the same facts stated on domains you do not own: directories, comparison pages, trade press, partner sites, your clients' own case studies.
  • Keep your entity consistent everywhere: one spelling of the brand, one legal name, one location, one description of what you do.
  • Put real authors on real content, with credentials that exist elsewhere on the web.
  • Fix contradictions between your own pages first. A model that finds two different answers on your site will trust neither.

A 90-day GEO plan

Ninety days to the first honest readingWeeks one and two are the baseline, weeks three and four fix retrieval, weeks five to eight restructure the pages that map to the baseline questions, and weeks nine to twelve work on corroboration before the baseline is re-run.Ninety days tothe first honest readingWEEKS 1-2BaselineThe questions, theanswers, who isnamedWEEKS 3-4RetrievalRendering, speed,indexation, llms.txtWEEKS 5-8SelectionRestructure the pagesbehind the questionsWEEKS 9-12CorroborationProfiles, comparisonpages, original dataA shorter window cannot separate your work from the drift.

Weeks 1–2 — Baseline. Write the 20–50 questions your buyers would actually ask an assistant. Run them across the engines that matter to your market, record the full answers, and count how often you are named and where the engines got the information.

Weeks 3–4 — Retrieval. Crawl and fix: rendering, speed on the pages that matter, indexation, canonicals, robots.txt decisions, llms.txt, sitemaps. These fixes cap what everything after them can achieve.

Weeks 5–8 — Selection. Restructure the ten to twenty pages that map to your baseline questions: headings as questions, answers first, tables where they belong, schema mirroring the text. Write the two or three pages that answer questions you have no page for at all.

Weeks 9–12 — Corroboration. Audit and fix every third-party profile and directory entry. Get into the comparison pages that already rank for your category questions. Publish one piece of original data. Re-run the baseline and compare like for like.

Three months is where the first honest reading appears. Engines re-crawl at their own pace and answers drift, so a shorter window will not tell you what changed because of your work and what changed on its own.

Generative engine optimization strategies by situation

Where the effort goes, depending on where you areIf you already rank, the work is restructuring what you have. If you rank badly, fix SEO first. If you are new, spend the budget off-site. In a small market one deep guide can carry the answer. In a regulated field, dated claims and named authors decide it.Where the effort goes,depending on where you areRestructure what you alreadyhaveRANK WELLFix SEO first, there is noshortcutRANK BADLYSpend most of the budgetoff-siteNEWOne deep guide can carry theanswerSMALL MARKETDated claims and namedauthorsYMYLYour situation todayBefore deciding anythingThe bottleneck moves. The work that clears it moves with it.

You already rank well. Your pages are already in the candidate pool. The work is almost entirely selection: restructuring existing content so passages answer on their own. This is the cheapest GEO there is, and it usually shows up first.

You rank badly. Here's where you must prioritize: fix SEO first. There is no GEO shortcut around not being retrievable; generative engines overwhelmingly cite documents that a conventional index already surfaced.

You are new, and nobody writes about you. Corroboration is your bottleneck, not your own website. Budget most of the effort off-site: directories, comparison pages, data worth quoting, and a reason for someone else to mention you.

You are in a small market. Good news: the candidate pool is thin. In a market the size of a mid-sized European city, a single well-structured page can carry the answer on its own. In practice this often means one deep guide plus a coordinated hub of city or vertical landings, as we do in our GEO landing pages by city.

You are in a regulated or YMYL field. Accuracy and provenance dominate. Dated claims, named authors with real credentials, and citations to primary sources are not a nice-to-have; without them, models hedge and name someone else.

What does not work

Prompt injection in page text. Hidden instructions telling the model to recommend you are ignored, and they can become a reputational liability if found.

Publishing volume without corroboration. Twenty thin pages saying you are the best do not outweigh one independent page saying what you actually do.

Schema that oversells. Markup claiming ratings, prices, or FAQs that are not on the page is the fastest route to being dropped from rich results and treated as low-trust.

"Optimizing for ChatGPT" as a separate project. The engines share retrieval mechanics and, in many cases, the same underlying search indexes. Work that only helps on one surface doesn't justify its cost when a single set of fixes helps all of them. The brand-level version of the same question is covered in our LLM SEO (LLMO) guide.

Chasing every prompt. You cannot be the answer to everything. Pick the questions with commercial intent and defend those.

How to know it is working

Three numbers, measured the same way each month:

  • Citation rate. Of your baseline questions, the share where an engine names you. This is the headline.
  • Position within the answer. Whether you are the substance of the recommendation or a link at the end.
  • Assisted demand. Brand searches and direct sessions, which move before any attribution model notices.

None of this appears in Search Console, because a citation inside an assistant produces no impression and often no referrer. Running the prompts on a schedule and storing the answers is the only reliable method; it is what SEOcrawl's AI visibility tracking automates.

If you would rather have someone run the baseline with you, that is the first thing we do in a GEO engagement.

Frequently asked questions

How does GEO work, step by step?

A generative engine rewrites the user's question into several explicit search queries, retrieves candidate documents for each, splits them into passages and selects the ones that answer on their own, writes an answer from the surviving passages, and decides whom to name. GEO works on all five stages: technical work makes you retrievable, structure makes your passages selectable, and off-site corroboration and consistent entity data decide whether you are named.

What are the best practices for generative engine optimization?

Put your substance in the HTML and keep pages fast to fetch; allow the AI crawlers you want reading you and publish an llms.txt; ask one question per heading and answer it in the first two sentences; keep every section self-contained; use tables for comparisons; attach sources and dates to numbers; add schema that repeats the visible text exactly; state facts specific enough to be checked; and get those facts corroborated on domains you do not own.

How long does GEO take to show results?

About three months for an honest first reading. Retrieval fixes land in weeks, content restructuring takes four to six weeks, and off-site corroboration is slower still. Generative engines also re-crawl at their own pace and their answers drift week to week, so a reading taken sooner than a full quarter after the baseline is noise rather than signal.

Do I need to rank on Google to be cited by AI?

Usually yes. Generative engines overwhelmingly retrieve from conventional indexes, so pages that no index surfaces rarely enter the candidate pool. Ranking is not the goal any more, but it is still the qualifier: if your SEO is weak, fixing it is the first part of any realistic GEO plan.

Can hidden instructions make a model recommend my brand?

No. Hidden text or prompt injection aimed at generative engines is detected and ignored, and it is a reputational liability if it is found. The mechanisms that do work are retrievability, passages that answer on their own, consistent entity data, and claims that other domains corroborate.