Guide

Where do AI engines get their information?

Every AI answer comes from two places: what the model learned in training, and what it finds in a live search. Here is which search each engine uses, according to the companies themselves, which content deals they have announced, and which kinds of sites end up cited most.

Quick answer

Two places. Training data, mostly public web pages learned months earlier, and live search at question time. Gemini and Google's AI answers search Google, Copilot uses Bing, Perplexity runs its own index, ChatGPT uses third-party providers plus partner content, and Claude runs its own web search. Reddit, YouTube and Wikipedia are cited often.

What are the key facts about where AI engines get their information?

EngineWhere live answers come fromSource
ChatGPT"Third-party search providers, as well as content provided directly by our partners." OAI-SearchBot decides what can show in ChatGPT searchOpenAI
Gemini (API)Grounding with Google Search: the model writes Google searches and cites the resultsGoogle
Google AI Overviews and AI ModePages that are "indexed and eligible to be shown in Google Search with a snippet"Google
Microsoft CopilotSends a short generated query to "the Bing search service"Microsoft
PerplexityIts own index, "covering hundreds of billions of webpages" (Sept 2025); PerplexityBot crawls for itPerplexity
ClaudeAnthropic's web search tool, which always returns citations; Claude-SearchBot crawls to improve search resultsAnthropic
Training data (OpenAI)Public internet info, info from third-party partners, and info from users, trainers and researchersOpenAI
Content dealsGoogle (Feb 2024) and OpenAI (May 2024) both announced access to Reddit's Data APIGoogle, Reddit

An AI engine can answer from two places, and they work very differently.

Training data is what the model learned before it was released. OpenAI says its models, including the ones behind ChatGPT, are built from three main sources: "information that is publicly available on the internet," information it partners with third parties to access, and information that users, human trainers and researchers provide. OpenAI says it doesn't intentionally gather data from sites behind paywalls, and it filters out things like spam and adult content. OpenAI Training happens months before you ask a question, so it can't know about your new prices or your new location.

Live search happens at the moment someone asks. The engine runs a web search, reads the results and writes an answer with links. Anthropic's documentation for Claude spells out when this happens: Claude searches when a question needs "Current prices, rates, scores, or statistics" or information about organizations that "might have changed," and answers from memory for stable facts. Anthropic

What this means in plain words: questions about local businesses, prices and "best" options usually trigger a search. So for most business questions, the search index the engine uses matters more than what the model learned in training.

Which search index does each AI engine use?

Here is what each company says in its own documentation. Where a company doesn't name its provider, we don't guess.

ChatGPT

OpenAI says ChatGPT search "leverages third-party search providers, as well as content provided directly by our partners." OpenAI OpenAI also runs its own crawler, OAI-SearchBot, which is "used to surface websites in search results in ChatGPT's search features." Sites that block it "will not be shown in ChatGPT search answers." A separate crawler, GPTBot, collects content that may be used for training. OpenAI For the Bing side of the story, see Bing and AI answers.

Gemini

Google's Gemini developer docs describe Grounding with Google Search. The model decides whether a Google search would help, "automatically generates one or multiple search queries and executes them," then writes an answer with inline citations to the source pages. Google

Google AI Overviews and AI Mode

These are built on Google Search. To be shown as a supporting link, a page must be "indexed and eligible to be shown in Google Search with a snippet." Google says there are no extra requirements or special optimizations. AI Mode and AI Overviews may also use "query fan-out," running several related searches across subtopics. Google

Microsoft Copilot

Microsoft says Copilot "may fetch information from the Bing search service." It turns your prompt into a short search query, sends it to Bing, and uses the results to write the answer. Microsoft

Perplexity

Perplexity runs its own search infrastructure. When it launched its Search API in September 2025, it said the API uses "the same global-scale infrastructure that powers Perplexity's public answer engine," with "an index covering hundreds of billions of webpages." Perplexity Its crawler, PerplexityBot, is "designed to surface and link websites in search results on Perplexity" and is not used to train AI models. Perplexity

Claude

Anthropic's web search tool lets Claude run searches, read the results and answer "with cited sources." Citations are always on, and each result carries the page's URL, title and a page_age date. Anthropic Anthropic's documentation doesn't name an outside search provider. It does run Claude-SearchBot, which crawls the web "to improve search result quality," and ClaudeBot, which collects content that could be used for training. Anthropic

Which content deals have AI companies announced?

Some content reaches AI engines through paid or partner access, not just crawling. These are deals the companies announced themselves.

  • Google and Reddit (Feb 22, 2024). Google said it "now has access to Reddit's Data API," giving it "efficient and structured access to fresher information" that it can "display, train on, and otherwise use." Google
  • OpenAI and Reddit (May 16, 2024). Reddit and OpenAI announced that "OpenAI will access Reddit's Data API" to bring Reddit content into ChatGPT, "especially on recent topics." OpenAI republished the announcement on its own site. Reddit OpenAI
  • OpenAI and news publishers (Oct 31, 2024). When it launched ChatGPT search, OpenAI named its publisher partners: Associated Press, Axel Springer, Condé Nast, Dotdash Meredith, Financial Times, GEDI, Hearst, Le Monde, News Corp, Prisa (El País), Reuters, The Atlantic, Time and Vox Media. It also said "Any website or publisher can choose to appear" in ChatGPT search. OpenAI
  • Wikipedia. Wikimedia Enterprise, the Wikimedia Foundation's paid data service, lists Amazon, Google, Microsoft, Meta, Perplexity and Mistral AI among companies that use it. Wikimedia Enterprise The Foundation calls Wikipedia "one of the highest-quality datasets in the world for training AI." Wikimedia Foundation
What this means in plain words: a deal doesn't guarantee a site gets cited. It means the engine has easy, structured access to that content, which is one more reason those sites show up so often.

Which kinds of sites do AI engines cite most?

No AI company publishes its own citation shares, so these numbers come from SEO software companies that track AI answers. They are their measurements, not ours, and they change month to month.

Share of citations, top domains in Google AI Mode Ahrefs Brand Radar, United States, all topics, September 2026. Mention share among the top sources.
Show the numbers
ItemValue
reddit.com17.9%
youtube.com17.8%
google.com12.5%
facebook.com10.2%
instagram.com5.7%
en.wikipedia.org4.0%
  • Reddit. Ahrefs' September 2026 data puts Reddit first in ChatGPT (16.8% of citations) and Google AI Mode (17.9%), and second in Google AI Overviews (18.5%). Ahrefs ChatGPT Ahrefs AI Mode Ahrefs AI Overviews More in Reddit and AI answers.
  • YouTube. First in Google AI Overviews (22.9%) and second in AI Mode (17.8%), but eighth in ChatGPT (2.5%). More in How to get your videos cited by AI.
  • Wikipedia. Second in ChatGPT (7.0%) and 4.0% in both AI Mode and AI Overviews. More in Wikipedia and AI answers.
  • Review sites. In ChatGPT, Ahrefs lists Consumer Reports third (3.7%), with Tripadvisor (1.5%), Trustpilot (1.5%), Angi (1.4%) and BBB (1.2%) in the top 32. More in Do reviews affect AI recommendations?
  • LinkedIn. Semrush tracked 230,000 prompts from July to October 2025 and found Reddit and LinkedIn among the top five cited domains in ChatGPT search, Google AI Mode and Perplexity; AI Mode cited LinkedIn in nearly 15% of its responses. Semrush More in LinkedIn and AI answers.
Each engine leans differently. YouTube dominates Google's AI answers but not ChatGPT's. Semrush also saw ChatGPT cite Reddit and Wikipedia far less from mid-September 2025, though they were still its two most-cited domains. Don't build your plan around one engine's list.

What should a business do about it?

This is our advice, built on the facts above.

  1. Let the search crawlers in. Check your robots.txt for OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot and Bingbot. Blocking a search crawler can keep you out of that engine's answers. Blocking a training crawler like GPTBot or ClaudeBot is a separate choice. See AI crawlers.
  2. Get indexed in Google and Bing. Google's AI answers need your page indexed and snippet-eligible. Copilot uses Bing. Submit sitemaps to both.
  3. Put your facts in plain text on your own site. Prices, hours, service areas and what you do. Live search reads pages, so make the answer easy to find.
  4. Show up where engines already look. Join relevant Reddit threads honestly, under a labeled business account. Post short, captioned videos on YouTube. Keep your LinkedIn company page and articles public.
  5. Earn reviews on the sites your customers use. Honest reviews on Google, Yelp, Tripadvisor or your industry's review site. Never buy or fake them.
  6. Earn independent coverage before you think about Wikipedia. Our Wikipedia guide explains why most small businesses shouldn't write their own article.
  7. Check each engine separately. They use different indexes, so you can be cited in one and missing from another.

How do I see which sources AI engines use for my questions?

Ask each engine a question your customers ask, like "best [your service] in [your city]," and open the sources it lists. Note which are your pages, which are Reddit threads or review sites, and which are competitors.

Our Prompt Simulator runs one question across 12 engines at once and shows the sources each one cites. For an engine-by-engine playbook, see How to get cited by every AI engine.

Sources

  1. How ChatGPT and our foundation models are developed — OpenAI Help Center. Read Oct 6, 2026.
  2. Introducing ChatGPT search — OpenAI. Read Oct 6, 2026.
  3. Overview of OpenAI Crawlers — OpenAI. Read Oct 6, 2026.
  4. OpenAI and Reddit Partnership — OpenAI. Read Oct 6, 2026.
  5. Reddit and OpenAI Build Partnership — Reddit, Inc.. Read Oct 6, 2026.
  6. An expanded partnership with Reddit — Google. Read Oct 6, 2026.
  7. Grounding with Google Search — Google AI for Developers. Read Oct 6, 2026.
  8. AI features and your website — Google Search Central. Read Oct 6, 2026.
  9. Data, privacy, and security for web search in Microsoft Copilot and Microsoft Copilot Chat — Microsoft Learn. Read Oct 6, 2026.
  10. Introducing the Perplexity Search API — Perplexity. Read Oct 6, 2026.
  11. Perplexity Crawlers — Perplexity. Read Oct 6, 2026.
  12. Web search tool — Anthropic. Read Oct 6, 2026.
  13. Does Anthropic crawl data from the web, and how can site owners block the crawler? — Anthropic. Read Oct 6, 2026.
  14. Wikimedia Enterprise — Wikimedia Foundation. Read Oct 6, 2026.
  15. In the AI era, Wikipedia has never been more valuable — Wikimedia Foundation. Read Oct 6, 2026.
  16. The 50 Most-Cited Websites in ChatGPT (September 2026) — Ahrefs. Read Oct 6, 2026.
  17. The 50 Most-Cited Websites in Google AI Mode (September 2026) — Ahrefs. Read Oct 6, 2026.
  18. The 50 Most-Cited Websites in Google AI Overviews (September 2026) — Ahrefs. Read Oct 6, 2026.
  19. The Most-Cited Domains in AI: A 3-Month Study — Semrush. Read Oct 6, 2026.
FAQ

Common questions

Does ChatGPT use Google or Bing?

OpenAI says ChatGPT search "leverages third-party search providers, as well as content provided directly by our partners," and runs its own OAI-SearchBot crawler. Sites that block OAI-SearchBot won't be shown in ChatGPT search answers. OpenAI OpenAI

Where does Perplexity get its sources?

From its own search index. Perplexity said in September 2025 that its index covers "hundreds of billions of webpages" and powers its public answer engine. PerplexityBot crawls for it. Perplexity Perplexity

Does Gemini use Google Search?

Yes, through Grounding with Google Search: the model writes Google searches, reads the results and answers with citations. Google's AI Overviews and AI Mode use pages indexed in Google Search. Google Google

Does Claude search the web?

Yes. Anthropic's web search tool lets Claude search when a question needs current information, and every answer includes citations. Anthropic runs Claude-SearchBot to improve search results and ClaudeBot for training data. Anthropic Anthropic

Why does AI cite Reddit so much?

Reddit has huge numbers of real conversations, and both Google and OpenAI announced access to Reddit's Data API in 2024. Ahrefs' September 2026 data ranks Reddit first in ChatGPT and Google AI Mode citations. Google Reddit Ahrefs

See which sources AI engines use for your business.

Enter your domain. The free AI visibility check shows where ChatGPT, Gemini, Perplexity and Google's AI answers mention you, and where they don't.