Savra
GuideAI SEO

The GEO Content Checklist: Formatting That Wins AI Citations

Five layers of checks that make a page extractable and citable by AI engines, demonstrated live on the page you're reading.

Inge · Content StrategistUpdated Jul 17, 20265 min read
Abstract editorial art: ascending mint rank bars on deep ink (The GEO Content Checklist: Formatting That Wins AI Citations)
On this page

Generative engine optimization (GEO) is formatting content so AI engines can extract and cite it. This checklist covers the five layers that decide it: page structure, passage structure, evidence, technical access, and distribution. Run it before every publish. This page follows every check on it, so you can see each one live.

Work through the layers in order, because they gate each other: perfect passages inside broken HTML cite nobody, and perfect HTML around vague prose cites nobody either. Where this page itself demonstrates a check, I say so. A checklist you can inspect live beats one you take on faith.

Page-level checks: can an engine map this page?

A page passes when an engine can map it in one pass: one topic, one H1, headings phrased as questions, and a direct answer under each. Engines reward legible over clever at this layer. If a stranger could skim your headings and predict each section's answer, you pass. This article runs that exact structure, starting with the question in the heading above.

  1. One H1 that states the topic in plain words. Everything below it is an H2 or H3.
  2. Headings phrased as the questions people actually ask, so a heading alone tells an engine what the section answers.
  3. An answer-first paragraph directly under every heading: conclusion first, support after.
  4. Anchor links on headings, so a table of contents, a reader, and an engine can address each section. The template behind this guide extracts them automatically.
  5. A visible date that changes only when the content materially changes. Fake freshness reads as noise.

Passage-level checks: does each block survive alone?

Passages are the unit engines quote, so each block has to survive being lifted out of the page with its meaning intact. That's also why answer-first ordering matters: the conclusion has to live inside the block where the lift happens, never a scroll away.

Write the paragraph you want quoted. Engines lift passages; the page just carries them.
  1. Key answers are self-contained and 40-80 words long: complete enough to stand alone, short enough to be lifted whole.
  2. Definitions name their subject. "Generative engine optimization is..." survives out of context. "It's..." dies there.
  3. Specifics beat adjectives: names, numbers, and dates where a lazier page writes "powerful" and "fast." Vague superlatives are also the loudest tell of AI slop.
  4. One idea per paragraph. If a block needs its neighbor to parse, merge them or rewrite.
  5. Comparison tables and listicles wherever the content honestly fits one. Engines extract both formats often.

The test for every block: paste it into a chat window with nothing around it. If it still says something true and specific about a named subject, it earned its place on the page.

Evidence checks: does the page show receipts?

Evidence is the best-researched lever on this list. Princeton's GEO paper and the industry studies that followed it found that citing sources, including statistics, and writing quotable passages measurably increase the odds of being cited by generative engines. It also happens to be the layer most content skips, which makes it the cheapest edge here.

  • Every claim that isn't yours points to its source.
  • Every number carries its date. An undated statistic is a rumor with digits.
  • A named human author with a real bio sits on the page. This article carries one; anonymous "team" bylines give an engine nothing to weigh.
  • First-party data is labeled as yours. Original numbers are the most quotable asset you can publish.

Sourcing feels slow for the first week, then becomes the fastest part of drafting. Claims you can support get written once. Adjectives get rewritten forever.

Technical checks: can crawlers get clean HTML?

Technical checks are pass or fail, because the major AI crawlers (GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot, Google-Extended) mostly don't execute JavaScript. If your content only exists after scripts run, it doesn't exist for them. Four things to verify, one to skip.

  1. Server-rendered HTML on every content page. Load the page with JavaScript disabled; if the text survives, you pass. This page is server-rendered, so you can run that test right here.
  2. A robots.txt that allows each AI crawler you want citing you, by name and on purpose. This site leaves them open.
  3. Structured data that names entities: article, author, and organization. This page ships that JSON-LD; view source to check.
  4. Fast, clean loads with no interstitials burying the content.
  5. Skip llms.txt for now. Real crawler adoption is near zero, and the four checks above do the actual gating.

Distribution checks: does your entity story travel?

Distribution covers everything off the page, because engines corroborate before they cite. A page described one way, in one place, on one domain travels worse than a page echoed consistently across feeds, profiles, and third-party posts. This layer is unglamorous, which is why most brands skip it and why it still pays.

  • A full-content RSS feed, so aggregators and engines get whole articles instead of stubs. These guides ship in one.
  • Identical brand and product descriptions across your site, LinkedIn, and directories. Entity consistency is voice discipline, and Savra's free Brand Genome audit shows you in 90 seconds how consistently yours reads.
  • Third-party presence: claimed review profiles, useful community answers, honest comparison posts. Engines cite them often, sometimes ahead of brand domains.
  • Refreshes on pages that earn citations, with dates bumped only for real edits.

That's the whole checklist: five layers, run top to bottom before you hit publish. The first pass on an existing page usually takes under an hour, and the page and passage fixes pay off in regular search too. For the strategy underneath the formatting, start with AEO vs SEO, then run the citation play for ChatGPT and Perplexity.

Frequently asked questions

What is generative engine optimization (GEO)?
Generative engine optimization is the practice of formatting and structuring content so AI engines like ChatGPT, Perplexity, and Google's AI Mode can retrieve it, extract passages from it, and cite it in generated answers. The term comes from Princeton research on generative engine visibility.
How long should an answer-first passage be?
Aim for 40 to 80 words: long enough to state a complete, self-contained answer, short enough for an engine to lift whole into a generated response. Name the subject inside the passage so it still makes sense with the rest of the page stripped away.
Do AI crawlers run JavaScript?
Mostly no. GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot, and Google-Extended largely fetch raw HTML without executing scripts, so content that only renders client-side is invisible to them. Server-side rendering is the fix.
Is llms.txt worth adding?
Low priority. The proposal is tidy, but real crawler adoption is near zero, so it changes little today. Server-rendered HTML, an open robots.txt, entity schema, and fast pages are what decide whether you get cited right now.

Keep reading

Marketing intelligence, weekly.

One email a week: what changed in search and AI marketing, and the play to run. Unsubscribe anytime.

Run the play with Savra.

Audit your brand in 90 seconds and let one agent execute the playbook.