Every AI search guide tells you to “write clear, structured, authoritative content.” True, but useless on its own — it doesn’t explain why one paragraph gets pulled into a ChatGPT answer while an almost identical one, sitting right next to it, gets ignored.
AI search engines pick sources through two layers working together: training data (what the model already “knows” from before its knowledge cutoff) and retrieval (live web searches run at the moment you ask a question). For almost any current, comparative, or recommendation-style question, retrieval does nearly all the work — and within retrieval, there’s a distinction most guides skip entirely: being fetched, being cited, and being mentioned are three different outcomes, and only one of them is what you actually want.
This article breaks down that full mechanism, engine by engine, in plain language — because once you understand how the selection actually happens, most of the “GEO best practices” floating around stop feeling like guesswork.
Two Systems: What the AI Already Knows vs. What It Looks Up
Day 1 of this series introduced retrieval-augmented generation (RAG) briefly. Here’s the fuller picture.
Training data is what got baked into the model during its training process — a compressed, statistical “understanding” built from enormous volumes of web content, books, and other text, all captured before a fixed cutoff date. If your brand or content was well-documented across the web well before that cutoff, the model has some baseline familiarity with you, the same way it has baseline familiarity with any well-established entity. This layer changes slowly. It’s not something you can influence this week.
Retrieval is what happens when the model decides it needs current information and runs a live web search before answering. This is where almost all commercial, comparative, and recommendation queries actually get resolved, and it’s the layer every practical GEO tactic in this article targets.
How does the model decide which path to use? Broadly, it asks itself whether it can answer confidently from what it already knows, or whether the question is time-sensitive, outside its training scope, or explicitly asking for current sources — in which case it triggers a search. Questions like “what’s the best budget espresso machine right now” almost always trigger retrieval, because the honest answer changes over time and the model knows it.
The Step Nobody Explains Well: Query Fan-Out
Here’s where it gets genuinely interesting, and where a lot of confusion about “conversational keywords” comes from.
When a model decides to search, it usually doesn’t search your exact question verbatim. Long, natural questions like “I have two dogs and hardwood floors, what vacuum should I get, budget around $400” rarely return useful results if typed into a search box exactly as written. So instead, the model breaks your question into several shorter, more traditional-looking searches and runs them independently — a process called query fan-out.
That vacuum question might fan out into something like:
- best vacuum for pet hair hardwood floors
- vacuum cleaners under $400 comparison
- cordless vs upright vacuum for pet owners
Each of those sub-queries gets its own search, its own set of retrieved pages, and its own contribution to the final answer. This is genuinely good news for anyone doing SEO already: fan-out queries look remarkably like ordinary long-tail keyword research — topic plus qualifier, comparison phrasing, “best for X” framing. If you’ve already done solid keyword research for your niche, you’ve likely already covered much of the fan-out space without knowing it.
It also explains something that trips people up: content doesn’t need to match a user’s exact conversational phrasing to get retrieved. It needs to match the decomposed version of that question — which is much closer to traditional SEO territory than it first appears.
Which Search Index Each Engine Actually Uses
This is the part that resolves a genuinely contested question — including the skepticism raised back in Day 2 about whether AI platforms are “real” independent search engines or just wrappers around existing search infrastructure.
The honest answer: it’s a mix, and it varies by engine.
| AI Engine | Search Index Used for Retrieval |
|---|---|
| ChatGPT | Bing and Google (hybrid, varies by query and over time) |
| Claude | Brave |
| Gemini | |
| Copilot | Bing |
| Perplexity | Its own in-house crawler/index |
| Google AI Mode | |
| Google AI Overviews |
The practical implication is significant: ranking well in traditional Google search genuinely does help your AI Mode and AI Overviews visibility, because they’re drawing from the same index. For ChatGPT and Copilot, Bing visibility matters more than most site owners realize — and Bing is a channel most WordPress site owners have never bothered to check, let alone optimize for.
This is also why treating “AI search” as one undifferentiated channel is a mistake. Optimizing purely for Google won’t help your Claude visibility much, since Claude draws from Brave’s index instead.
One caveat worth being upfront about: these dependencies aren’t fully disclosed by any of the AI labs and can shift over time as partnerships change. Treat this table as the best current picture, not a permanent architecture diagram — and if you want to verify it yourself, the self-audit exercise near the end of this article shows you how.
Fetched, Cited, and Mentioned: Three Different Outcomes
This is the single most useful distinction in this entire topic, and almost nothing written about GEO makes it explicit.
Fetched means the engine’s crawler pulled your page during a retrieval search. That’s it. It doesn’t mean your content shows up anywhere in the final answer — a huge share of fetched pages never appear in the response at all.
Cited means one specific sentence in the AI’s answer is directly attributed to your page, usually as a clickable footnote tied to a single claim. Citation binds to a precise piece of text, not to your page being generally relevant to the topic. You have to be the best available support for one specific claim the model is making — broad relevance alone doesn’t earn this.
Mentioned means your brand name shows up somewhere in the answer without your page necessarily being the cited source of any specific claim. You can be mentioned purely from training-data familiarity even when your page wasn’t the citation source for anything specific in that answer. You can also be cited (a footnote link) without being prominently mentioned in the visible text — a citation nobody actually reads.
Why this matters practically: if you check your site’s server logs and see AI crawler traffic (fetching), that tells you almost nothing about whether you’re actually showing up in answers. Plenty of fetched pages contribute zero citations. The self-audit later in this article is how you actually check the outcome that matters.
What Actually Earns a Citation Once You’re Fetched
Being fetched is necessary but nowhere near sufficient. Four things consistently separate a fetched page from a cited one.
1. Extractability. A citation has to bind to a clean, self-contained claim. Pages with a direct-answer passage near the top — short, specific, no hedging — convert to citations far more often than pages that build up to the point slowly. Interestingly, a large share of cited passages contain no outbound links at all; a self-contained block of text reads as the authoritative answer itself, while a paragraph full of links reads more like a signpost pointing elsewhere, which seems to make it feel less “quotable” on its own.
2. Specificity. Vague claims don’t bind to anything worth citing. “Our plugin makes your site faster” attaches to nothing useful. “Enabling this plugin’s page cache reduced load time by 1.2 seconds across a 50-page test site” is a specific, attributable claim a model can actually lift and defend.
3. Format. Comparison content — listicles, “X vs Y” structures, tables — accounts for a disproportionate share of citations across the sources examined for this article, more than almost any other single content type. Tables in particular get lifted almost directly, because the structure itself does the organizational work the model would otherwise have to do.
4. Consensus. For recommendation-style queries specifically, the model looks for agreement across the sources it retrieved. If your recommendation shows up across several of the pages it fetched for a query, it gets named with real confidence. If you’re the lone voice among many, the model tends to hedge or skip you entirely. This is a big part of why third-party mentions — a Reddit thread, a comparison roundup on another site — often matter as much as your own page: your own domain is one vote, and the model is effectively counting votes.
Why Reddit Out-Cites YouTube (Even at Similar Fetch Rates)
This one has a direct, actionable takeaway for anyone running a content operation that includes video.
Reddit threads and YouTube pages get fetched by AI crawlers at broadly similar rates for many topics. But Reddit converts into actual citations far more often. The likely mechanical reason: when a crawler fetches a Reddit thread, it gets the real text of the discussion — genuine, extractable, citable content. When it fetches a YouTube page, it typically gets metadata — a title, a description, maybe some tags — not a transcript of what’s actually said in the video.
There’s simply nothing in a bare YouTube page for the model to bind a citation to, even if the video itself is excellent and highly relevant.
The takeaway: if you’re producing video content — including things like the Mokimo character videos or any explainer content — publishing a full transcript alongside the video (on your own WordPress site, not just relying on YouTube’s auto-captions) gives AI crawlers something citable that the video alone doesn’t provide.
Engine-by-Engine Personality
Beyond which index each engine draws from, they behave differently enough in practice that it’s worth a quick behavioral snapshot.
ChatGPT uses a hybrid retrieval approach — part live web search, part other data arrangements that shift over time and aren’t fully disclosed. In practice, a relatively small set of domains, especially Reddit and major comparison hubs, dominate citations for commercial and recommendation queries.
Perplexity runs a live search on nearly every query and shows its full source list by default, which makes it the most transparent engine to study directly — you can literally see what it retrieved and cited for any question you ask it. It leans particularly heavily on community sources like Reddit for recommendation-style questions.
Google AI Overviews and AI Mode draw from Google’s own index — but here’s the counterintuitive part: they often cite pages that never cracked the top ten of ordinary Google search results for that same query. The citation model and the traditional ranking model aren’t optimizing for exactly the same thing, so a page can be a strong AI Overview source without being a top-ranked page.
Claude tends toward a more conservative citation pattern — fewer citations, but higher confidence in each one, favoring clearly structured, well-sourced content over a wide scattershot of sources.
The practical implication: a brand that’s well cited in one engine can be nearly invisible in another for the exact same question. Testing across engines individually, rather than assuming one predicts the others, is the only way to actually know where you stand.
A Self-Audit You Can Run in 20 Minutes
You don’t need special tools to check any of this yourself.
- Pick five buying-intent or recommendation questions in your niche — the kind your ideal reader would actually type into ChatGPT or Perplexity.
- Run each one in a fresh ChatGPT chat with search/browsing on, and separately in Perplexity.
- Expand the sources on each response (Perplexity shows this by default; ChatGPT often does too when it performs a search).
- For each result, note three things: was your site fetched at all, was it cited (a specific footnote tied to a claim), or just generally mentioned — and what format did the winning page use (direct answer, table, listicle, forum thread)?
- Do the same exercise for whichever competitor is winning that query. The structural gap — table vs. wall of text, specific stat vs. vague claim — is usually visible in the first hundred words of whichever page won.
Repeat this every month or two rather than obsessing over any single result — AI responses are somewhat probabilistic, meaning the same question can return different sources on different runs. Tracking your average visibility across many prompts over time tells you far more than any single lucky (or unlucky) result.
One More Layer: Personalization
There’s a wrinkle worth knowing about even though it’s not something you can directly optimize for: AI answers are somewhat personalized. The same question can return different results for different people, influenced by things like earlier messages in the same conversation, memory features that retain a user’s stated preferences across chats, and inferred location or timing. Mention “I have a small budget” earlier in a chat, and a later “best hosting plan” question may get filtered through that constraint automatically.
This is one more reason a single test of a single prompt tells you less than it feels like it should — it’s also shaped by conversational context you don’t control. It’s part of why the self-audit below emphasizes running multiple prompts and tracking patterns over time, rather than treating any one result as definitive.
What This Means Specifically for a WordPress Site
Tying the mechanism back to a typical WordPress setup:
- Check whether your theme or page builder renders key content client-side. Many popular WordPress page builders inject body content via JavaScript after the initial page load — exactly the blind spot the FAQ below describes. If a crawler can’t see it in raw HTML, none of the citation factors above can help you.
- Bing Webmaster Tools is worth setting up even if you’ve only ever used Google Search Console, given ChatGPT and Copilot’s partial reliance on Bing’s index.
- Your comparison and “best X” posts are disproportionately valuable — not just for ordinary SEO, but specifically because comparison/listicle formatting is one of the four factors that correlates most strongly with citation.
Frequently Asked Questions
Does ChatGPT crawl the web itself, or does it use another search engine? It uses a hybrid approach that draws on both Bing and Google for retrieval, alongside other data arrangements that aren’t fully disclosed and can shift over time. This is separate from its training data, which is fixed as of its knowledge cutoff.
Why does Perplexity cite different sources than ChatGPT for the same question? Each engine runs its own retrieval and ranking logic, and in Perplexity’s case, its own in-house index rather than Bing or Google. Track engines separately rather than assuming visibility in one predicts visibility in another.
Can AI search engines read JavaScript-heavy WordPress sites? Most AI crawlers read raw HTML and don’t execute JavaScript the way a browser does. A page that displays perfectly for a human visitor can return as a largely empty shell to these crawlers if key content is rendered client-side. This is a common, invisible blind spot for WordPress sites built with heavy page-builder plugins.
Does schema markup directly improve my citation odds? It helps engines parse your page’s structure and entity identity reliably, which supports everything else in this article, but on its own it doesn’t appear to directly lift citation rates in controlled testing. Treat it as the reliability layer underneath good content, not a citation tactic by itself.
Why would my page get fetched but never cited? Usually because the content couldn’t bind to one specific claim — no self-contained answer passage, no concrete statistic — or because a competing source answered that exact claim more precisely or in a more citable format (like a table). Being fetched confirms your page was relevant enough to retrieve; citation requires winning one specific sentence in the answer, not just being topically related.
Where to Go From Here
This article covered the mechanism. The next steps in this series turn it into action:
- Why Your WordPress Site Is Invisible to AI Search — the specific technical reasons WordPress sites get fetched but never cited
- AI Crawlers Guide: GPTBot, ClaudeBot, and robots.txt — making sure you’re not accidentally blocking the crawlers this whole article depends on
- Answer Capsules: Writing the Self-Contained Paragraph AI Actually Cites — turning the extractability principle above into a repeatable writing habit
Understanding why a citation happens is what makes every other tactic in this series make sense — from here on, the goal is turning that understanding into the actual habits your published content follows.
