How to see AI bots in your server logs
AI engines read your website before they answer questions about you, but those visits never show up in Google Analytics. Here is where they do show up, which names to look for, how to tell a real bot from a fake one, and what to do with what you find.
Look in your server or CDN logs, not Google Analytics. AI bots name themselves in each request, such as GPTBot, ChatGPT-User, ClaudeBot, Claude-User and PerplexityBot. GA4 only counts visitors who run its JavaScript tag and automatically drops known bots. To confirm a visit is genuine, match its IP address against the company's published IP list.
What are the key facts about AI bots in your logs?
| Fact | Detail | Source |
|---|---|---|
| Bots name themselves | Each request carries a user-agent, e.g. GPTBot/1.4; +https://openai.com/gptbot | OpenAI |
| Three kinds of bot | Training crawlers, search crawlers, and fetchers that act on one user's request | OpenAI, Anthropic, Perplexity |
| User fetchers and robots.txt | OpenAI: "robots.txt rules may not apply" to ChatGPT-User; Perplexity-User "generally ignores robots.txt rules" | OpenAI, Perplexity |
| Some names never appear | Google-Extended "doesn't have a separate HTTP request user agent string"; Applebot-Extended "does not crawl webpages" | Google, Apple |
| GA4 can't see them | GA4 counts visits through a JavaScript tag, and "Traffic from known bots and spiders is automatically excluded" | |
| Verify by IP | OpenAI, Anthropic, Perplexity, Google, Apple and Bing publish IP lists for their bots | Each company |
| Verify by DNS | A real Googlebot IP reverse-resolves to googlebot.com, google.com or googleusercontent.com, and back again | |
| CDN dashboards | Cloudflare's AI Crawl Control shows AI crawler requests per crawler and per page, on all plans | Cloudflare |
| Scale | GPTBot made 569 million requests and Claude 370 million across Vercel's network in one month (Dec 2024) | Vercel (third-party) |
What kinds of AI bots visit my website?
AI companies now run separate bots for separate jobs. Knowing which job a bot does tells you what its visit means.
- Training crawlers read pages that may be used to train future AI models. Examples: OpenAI's
GPTBot, which "crawls content for training foundation models," and Anthropic'sClaudeBot, which collects content "that could potentially contribute to their training." OpenAI Anthropic - Search crawlers build the index an AI engine searches when it answers. Examples:
OAI-SearchBotfor ChatGPT search,Claude-SearchBot, andPerplexityBot, which is "designed to surface and link websites in search results on Perplexity." OpenAI Perplexity - User-triggered fetchers open one page because one person asked a question right now. Examples:
ChatGPT-User,Claude-UserandPerplexity-User. Perplexity puts it simply: "When users ask Perplexity a question, it might visit a web page to help provide an accurate answer." Perplexity
The rules differ by kind. OpenAI says of ChatGPT-User: "Because these actions are initiated by a user, robots.txt rules may not apply." Perplexity says Perplexity-User "generally ignores robots.txt rules." Anthropic says its bots, Claude-User included, honor robots.txt. So a user fetcher can appear in your logs even on pages you told crawlers to skip.
For which of these to allow or block, see Which AI crawlers should you allow. This guide is about seeing them.
Which bot names should I search my logs for?
Search your logs for these names in the user-agent field. We only list bots whose official pages we read on 6 October 2026. Companies rename and add bots, so check the source page before you write a rule.
| Look for | Company | Job | Shows in logs? | Source |
|---|---|---|---|---|
GPTBot | OpenAI | Training | Yes | OpenAI |
OAI-SearchBot | OpenAI | ChatGPT search | Yes | OpenAI |
ChatGPT-User | OpenAI | User-triggered fetch | Yes | OpenAI |
ClaudeBot | Anthropic | Training | Yes | Anthropic |
Claude-SearchBot | Anthropic | Search | Yes | Anthropic |
Claude-User | Anthropic | User-triggered fetch | Yes | Anthropic |
PerplexityBot | Perplexity | Search (not training) | Yes | Perplexity |
Perplexity-User | Perplexity | User-triggered fetch | Yes | Perplexity |
Googlebot | Google Search | Yes | ||
Google-Extended | Gemini training control | No, robots.txt token only | ||
Google-CloudVertexBot | Vertex AI Agents | Yes | ||
bingbot | Microsoft | Bing search | Yes | Bing |
Applebot | Apple | Spotlight, Siri, Safari | Yes | Apple |
Applebot-Extended | Apple | Training control | No, "does not crawl webpages" | Apple |
meta-externalagent | Meta | Training and indexing | Yes | Meta |
meta-webindexer | Meta | Meta AI search | Yes | Meta |
meta-externalfetcher | Meta | User-triggered fetch | Yes | Meta |
A full user-agent is long. OpenAI's search bot, for example, sends a browser-like string that ends compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot. Search for the short name, not the whole string, so version changes don't break your filter.
Why doesn't Google Analytics show AI bots?
Two reasons, and either one alone would hide them.
GA4 only counts visitors who run its code. Google Analytics collects data through the Google tag, a JavaScript snippet you paste into the <head> of each page. Google A visit is only recorded if that script runs. Vercel, a hosting company, studied AI crawler traffic on its network in December 2024 and reported that "none of the major AI crawlers currently render JavaScript," naming GPTBot, ClaudeBot and PerplexityBot. They download the page's text and leave. The exceptions Vercel found were Gemini, which uses Googlebot's infrastructure, and Applebot, which renders JavaScript. Vercel
GA4 throws known bots away on purpose. Google says "Traffic from known bots and spiders is automatically excluded," and "you cannot disable known bot traffic exclusion or see how much known bot traffic was excluded." Google
So GA4 isn't broken. It's built to measure people. The bot visits happen one step earlier, on your server, and that's where they're recorded.
Where do I find my server logs?
Every request to your site passes through a web server or a CDN (a content delivery network, the layer many sites sit behind). That layer writes a line for each request: the time, the page, the status code, the IP address and the user-agent. Where you read it depends on how your site is hosted.
If your site is behind Cloudflare. Cloudflare's AI Crawl Control is "Available on all plans" and needs no setup. Cloudflare Its Overview tab shows total AI crawler requests, status codes, the most requested pages, and crawlers grouped by company. The Crawlers tab lists each crawler with request counts and allow or block controls. Referral data, meaning people arriving from AI services, is on paid plans only. Cloudflare
If your site is on Vercel. The Logs section shows runtime logs, and each request's details include the "Request User Agent." Two limits matter. Runtime logs cover functions and middleware, but for static pages only requests served from cache; Vercel says "to get all static logs check Log Drains." And retention is short: 1 hour on Hobby, 1 day on Pro, up to 30 days with Observability Plus. Log Drains, which export logs elsewhere, are on Pro and Enterprise plans. Vercel Vercel
Other hosts. Most hosts keep raw access logs somewhere in the control panel, or will turn them on if you ask. Ask your host or developer three questions: where are the raw access logs, do they include the user-agent and IP address, and how many days are kept?
How do I know a bot is really who it says it is?
Anyone can type GPTBot into a user-agent. Google's own verification guide exists because site owners worry about "spammers or other troublemakers" claiming to be Google. Google So before you trust a line in your logs, check the IP address it came from.
Method 1: match the IP against the company's published list. These files list the IP ranges each bot uses:
| Bot | IP list |
|---|---|
| GPTBot | https://openai.com/gptbot.json |
| OAI-SearchBot | https://openai.com/searchbot.json |
| ChatGPT-User | https://openai.com/chatgpt-user.json |
| Anthropic bots | https://claude.com/crawling/bots.json |
| PerplexityBot | https://www.perplexity.com/perplexitybot.json |
| Perplexity-User | https://www.perplexity.com/perplexity-user.json |
| Googlebot and other Google crawlers | https://developers.google.com/static/crawling/ipranges/common-crawlers.json (plus special-crawlers, user-triggered-fetchers and user-triggered-agents files) |
| bingbot | https://www.bing.com/toolbox/bingbot.json |
Sources: OpenAI, Anthropic, Perplexity, Google, Bing. Apple publishes an Applebot IP list too, linked from its Applebot page. Apple Google notes the ranges are written in CIDR format, a short way of writing a block of addresses.
Method 2: a two-way DNS check. Google describes this for its crawlers: Google
- Run a reverse DNS lookup on the IP from your logs, for example
host 66.249.66.1. - Check that the name ends in
googlebot.com,google.comorgoogleusercontent.com. - Run a forward DNS lookup on that name.
- Check that it returns the same IP you started with.
Apple says Applebot traffic is "generally identified by using reverse DNS in the *.applebot.apple.com domain." Apple
If a request says ClaudeBot but its IP isn't on Anthropic's list, treat it as an imposter, not as Claude. Anthropic also warns that blocking by IP instead of robots.txt "may not work correctly or persistently guarantee an opt-out," since its IPs can change. Anthropic
What should I do with what I find in my logs?
This is our advice. The point of reading the logs is to find problems AI engines hit on your site that you can't see from your own browser.
- Check that the search crawlers and user fetchers are there at all. If you see no
OAI-SearchBot,Claude-SearchBot,PerplexityBotorbingbotin a week of logs, something may be blocking them. Start with your robots.txt and CDN rules. - Look at the status codes they got. A
200means they got the page. A403usually means a firewall or bot rule turned them away. A404means they asked for a page that no longer exists. A5xxmeans your server failed. Fix the 403s and 5xx errors on important pages first. - See which pages they read. Cloudflare's "Most Popular Paths" table does this for you. Cloudflare If bots spend their visits on old tag pages and never reach your services or pricing pages, improve your internal links and sitemap.
- Treat user-triggered fetches as live demand. A
ChatGPT-UserorPerplexity-Userhit on your pricing page means someone just asked an assistant something your pricing page answers. Make sure that page states its facts plainly in the HTML, not only inside scripts. See Can AI read JavaScript websites? - Separate real bots from fakes before you act. Verify the IP first. Blocking a fake GPTBot is fine. Blocking the real OAI-SearchBot because a fake one misbehaved is not.
- Check again after any change. A new firewall rule, CDN setting or site rebuild can block bots without any visible error to you.
What should be on my AI bot log checklist?
| Check | Done when | Source |
|---|---|---|
| Logs found | You know where your raw access or CDN logs live and how long they're kept | Our advice |
| User-agent and IP included | Each log line shows both | Our advice |
| Bot names filtered | You can filter for the short names in the table above | OpenAI, Anthropic, Perplexity, Google, Apple, Meta |
| No search for Google-Extended | You know Google-Extended never appears as a visitor | |
| Visits verified | Suspicious hits are checked against the published IP lists or by two-way DNS | Each company, Google |
| Search bots present | OAI-SearchBot, Claude-SearchBot, PerplexityBot and bingbot all appear over a week | Our advice |
| Errors fixed | No 403 or 5xx responses to AI bots on pages you want cited | Our advice |
| GA4 kept for humans | You use GA4 for people clicking through from AI, and logs for bots | |
| Rechecked after changes | Logs are reviewed after any firewall, CDN or site change | Our advice |
How do I know if bot visits turn into AI answers?
Logs tell you an AI engine read your page. They don't tell you whether it then cited or recommended you. For that, ask the engines directly.
Our Prompt Simulator runs one question across 12 AI engines at once and shows which sources each one cites. If a page is crawled often but never cited, the problem is the page's content, not access. For the basics of what these bots are, see the AI crawlers glossary entry.
Sources
- Bots — OpenAI. Read Oct 6, 2026.
- Does Anthropic crawl data from the web, and how can site owners block the crawler? — Anthropic. Read Oct 6, 2026.
- Perplexity Crawlers — Perplexity. Read Oct 6, 2026.
- Google's common crawlers — Google Search Central. Read Oct 6, 2026.
- Verifying Googlebot and other Google crawlers — Google Search Central. Read Oct 6, 2026.
- About Applebot — Apple. Read Oct 6, 2026.
- Meta Web Crawlers — Meta. Read Oct 6, 2026.
- bingbot IP ranges (bingbot.json) — Microsoft Bing. Read Oct 6, 2026.
- Set up Analytics for a website and/or app — Google Analytics Help. Read Oct 6, 2026.
- Known bot-traffic exclusion — Google Analytics Help. Read Oct 6, 2026.
- AI Crawl Control — Cloudflare. Read Oct 6, 2026.
- Analyze AI traffic — Cloudflare. Read Oct 6, 2026.
- Runtime Logs — Vercel. Read Oct 6, 2026.
- Logs — Vercel. Read Oct 6, 2026.
- The rise of the AI crawler — Vercel. Read Oct 6, 2026.
Common questions
Can I see ChatGPT visiting my website?
Yes, in your server or CDN logs. Look for GPTBot (training), OAI-SearchBot (ChatGPT search) and ChatGPT-User (a page opened because a user asked). OpenAI publishes IP lists for each, so you can confirm a visit is real. OpenAI
Why don't AI crawlers show up in Google Analytics?
GA4 records a visit only when its JavaScript tag runs, and Vercel reported in December 2024 that the major AI crawlers don't run JavaScript. GA4 also says it automatically excludes known bots and spiders, and you can't turn that off. Google Vercel
Will I see Google-Extended in my logs?
No. Google says Google-Extended doesn't have its own user-agent string. It's a robots.txt token that controls use of content for Gemini training, and the crawling itself is done by Google's existing crawlers. Google
How do I check that a bot calling itself ClaudeBot is real?
Match the request's IP address against Anthropic's published list at claude.com/crawling/bots.json. If the IP isn't on it, the user-agent was faked. Anthropic
Does Perplexity-User follow robots.txt?
Generally not. Perplexity says that since a user requested the fetch, Perplexity-User "generally ignores robots.txt rules." PerplexityBot, its search crawler, does follow them. Perplexity
Related pages.
See whether AI engines can read your site.
Enter your domain. The free AI visibility check shows where ChatGPT, Claude, Gemini, Perplexity and Google's AI answers mention you, and where they don't.