Key takeaways
- AI crawlers from OpenAI, Anthropic and Perplexity index your site to power AI-generated answers.
- A single misconfigured robots.txt line can silently block all of them and kill your GEO visibility.
- Each crawler has a distinct purpose: training data, search results, or real-time user queries.
- Allowing the right bots requires explicit Allow directives, server-side HTML and verified crawlability.
- Google's Google-Extended token controls AI training use; it does not block AI Overviews.
AI crawlers from OpenAI, Anthropic and Perplexity decide whether your site can appear in AI answers. GPTBot, ClaudeBot and PerplexityBot each read your pages before ChatGPT, Claude or Perplexity can ever cite you. One misconfigured line in robots.txt can block all of them at once, and most site owners never notice. Here is what each crawler does, and how to configure robots.txt so AI can see you.
Google's own guidance for AI Overviews and AI Mode starts with the same basics: crawling must be allowed in robots.txt and the page must meet Google Search's technical requirements. Google adds that there are no extra requirements or special files needed to appear in these features (Google Search Central, updated 10 December 2025).
What are AI crawlers, and why should you care?
AI is the new search engine. The brands it cites get bought. The rest get ignored.
Behind every ChatGPT answer, every Perplexity citation, every Claude response that names a vendor, there is a crawler that visited a website first. These are AI crawlers: automated bots that read your pages, extract your content, and feed it into the systems that generate AI answers. If those crawlers cannot access your site, you do not exist in AI search.
This is not theoretical. Researchers Aggarwal et al. (Princeton, 2023) introduced the concept of Generative Engine Optimization (GEO), showing that the way content is structured and indexed by AI systems directly affects how often it surfaces in generative engine responses (arxiv.org/abs/2311.09735). Visibility in AI-generated answers is a measurable, optimizable outcome, and it starts with crawl access.
The good news: controlling which AI crawlers can access your site is straightforward. It lives in one file, robots.txt. The bad news: most sites have it wrong.
Meet the major AI crawlers: GPTBot, ClaudeBot, PerplexityBot and Google
Not all AI crawlers do the same thing. Here is the breakdown.
Official AI crawler user agents at a glance (from each company's documentation, checked on 3 October 2026):
| Company | robots.txt token | What it does | Follows robots.txt | Official IP ranges |
|---|---|---|---|---|
| OpenAI | GPTBot | Collects content for model training | Yes | gptbot.json |
| OpenAI | OAI-SearchBot | Surfaces sites in ChatGPT search | Yes | searchbot.json |
| OpenAI | ChatGPT-User | Fetches pages for user actions in ChatGPT | May not apply | chatgpt-user.json |
| Anthropic | ClaudeBot | Collects content for model training | Yes (+ Crawl-delay) | bots.json |
| Anthropic | Claude-SearchBot | Indexes content for Claude search results | Yes | bots.json |
| Anthropic | Claude-User | Fetches pages when a Claude user asks | Yes | bots.json |
| Perplexity | PerplexityBot | Surfaces and links sites in Perplexity search (not training) | Yes | perplexitybot.json |
| Perplexity | Perplexity-User | Fetches pages when a user asks a question | Generally ignored | perplexity-user.json |
| Google-Extended | Token for Gemini training and grounding (not a separate crawler) | Yes | common-crawlers.json |
OpenAI: GPTBot, OAI-SearchBot and ChatGPT-User
GPTBot is OpenAI's primary web crawler. Its official purpose is to crawl content that may be used to train OpenAI's generative AI foundation models. OpenAI actually operates two distinct crawlers relevant to visibility:
- GPTBot collects content to make OpenAI's generative AI foundation models more useful and safe (training). Disallowing it signals your content should not be used for training.
- OAI-SearchBot surfaces websites in ChatGPT's search features. Block it and your pages will not appear in ChatGPT search answers.
- ChatGPT-User fetches pages for certain user actions in ChatGPT and Custom GPTs. OpenAI notes that because these actions are initiated by a user, robots.txt rules may not apply.
- OAI-AdsBot checks the safety of pages submitted as ads on ChatGPT; it does not follow robots.txt either.
The user-agent string for GPTBot looks like this:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot
GPTBot and OAI-SearchBot are independent: you can allow one and block the other. OpenAI publishes the IP ranges of each bot (GPTBot, OAI-SearchBot, ChatGPT-User). Source: OpenAI bots documentation.
Anthropic: ClaudeBot, Claude-User and Claude-SearchBot
ClaudeBot is Anthropic's crawler for collecting web content that could contribute to training Claude models. Anthropic also runs two additional bots:
- Claude-User is triggered when a Claude user asks a question that requires fetching a live web page. Blocking it reduces your visibility in user-directed Claude queries.
- Claude-SearchBot indexes content to improve Claude's search result quality. Blocking it reduces your visibility in Claude search responses.
Unlike OpenAI's and Perplexity's user-triggered fetchers, all three Anthropic bots honor robots.txt, including Claude-User, and they support the non-standard Crawl-delay directive. Anthropic publishes its IP ranges at claude.com/crawling/bots.json. Source: Anthropic crawler documentation.
Perplexity: PerplexityBot and Perplexity-User
PerplexityBot is designed to surface and link websites in Perplexity search results. Perplexity explicitly states it is not used to crawl content for AI foundation models: it is purely for search indexing. Perplexity also runs Perplexity-User, which fetches pages in real time when a user asks a question. Perplexity-User generally ignores robots.txt because the fetch is user-initiated.
The PerplexityBot user-agent string:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
Perplexity publishes the IP ranges of PerplexityBot and Perplexity-User. Source: Perplexity crawler documentation.
Google-Extended: Google's AI training token
This one confuses a lot of people. Google-Extended is not a separate crawler. It is a robots.txt policy token that tells Google whether content already crawled by Googlebot can be used for AI-related purposes, like training Gemini or grounding AI responses.
Crucially: blocking Google-Extended does not remove you from AI Overviews. Google states that Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal". AI Overviews pull from Google's search index, which Googlebot builds, so Googlebot must be able to crawl and index your pages. Google's crawler IP ranges are published in common-crawlers.json. Source: Google crawler documentation.
How to check if AI crawlers can access your site
Before you configure anything, verify the current state of your robots.txt.
Step 1: Find your robots.txt file
Go to yourdomain.com/robots.txt. Every site has one, or should. If you get a 404, that is already a problem: most crawlers will default to assuming everything is allowed, but some behave differently.
Step 2: Look for these patterns
Scan for any of the following:
- User-agent: * followed by Disallow: / blocks all bots, including every AI crawler.
- User-agent: GPTBot with Disallow: / blocks OpenAI's training crawler.
- User-agent: OAI-SearchBot with Disallow: / removes you from ChatGPT search results.
- User-agent: ClaudeBot with Disallow: / blocks Anthropic's training crawler.
- User-agent: PerplexityBot with Disallow: / removes you from Perplexity search results.
Step 3: Check your server logs
If you have access to server logs, search for the bot user-agent strings listed above. If you are not seeing any hits from GPTBot or PerplexityBot, either you are blocking them or your content is not being prioritized for crawling.
Step 4: Test with Google Search Console
Use the URL Inspection tool in Google Search Console to verify Googlebot can access your pages. For AI Overviews specifically, you need standard Googlebot access, not Google-Extended.
Step 5: Check your WAF
Web Application Firewalls (Cloudflare, AWS WAF and similar) can block bots at the network level, before robots.txt is even read. If your WAF is set to block unknown bots or rate-limit aggressively, AI crawlers may be silently rejected. Perplexity's official docs include explicit WAF configuration guides for Cloudflare and AWS for exactly this reason.
How to allow AI crawlers in robots.txt (with a real example)
Here is a production-ready robots.txt that explicitly allows all major AI crawlers while keeping control over training-data use. The configuration, block by block:
- Allow all standard crawlers: User-agent: * with Allow: / and a Disallow for your private paths such as /admin/ and /private/.
- OpenAI: allow both OAI-SearchBot and GPTBot with Allow: / to keep ChatGPT search and training access open.
- Anthropic: allow ClaudeBot, Claude-User and Claude-SearchBot, each with Allow: /.
- Perplexity: allow PerplexityBot with Allow: / for search indexing.
- Google: allow Googlebot with Allow: /, and optionally set Google-Extended to Disallow: / if you want to keep your content out of Gemini training without affecting AI Overviews.
- Finish with your Sitemap: line pointing at your sitemap.xml.
A few notes on this configuration:
- The Google-Extended disallow is optional. If you are fine with Google using your content for Gemini training, remove that block.
- Each bot directive is independent. Allowing OAI-SearchBot and blocking GPTBot is a valid strategy if you want ChatGPT search visibility without contributing to training data.
- Order matters within a user-agent block, but not between blocks. Each User-agent section is evaluated independently.
Should I block GPTBot in robots.txt?
Usually no. GPTBot is OpenAI's training crawler, while OAI-SearchBot is the one behind ChatGPT search, so blocking GPTBot alone does not take you out of ChatGPT answers. The real risk is a blanket rule that blocks every bot and catches both.
Here is the uncomfortable truth: a lot of sites accidentally block AI crawlers. The most common culprit is a legacy robots.txt that was set up to block scrapers and now catches every AI bot in the process. A blanket Disallow: / under User-agent: * is the single most damaging line you can have in your file right now.
My take, from building Howseen: the sites bleeding the most AI visibility almost never blocked crawlers on purpose. It is a stale robots.txt line, copied from some old scraper-defense setup, that nobody has reopened in two years. Before you touch anything else, go read that one file line by line.
Why does it matter so much?
AI answers are citation-based. When ChatGPT, Claude or Perplexity answer a question, they cite sources. Those sources are pages that were crawled, indexed and deemed relevant. If your page was never crawled, it cannot be cited. If it cannot be cited, your brand does not exist in that answer.
GEO is the new SEO. The research from Aggarwal et al. (arxiv.org/abs/2311.09735) showed that optimizing content for generative engines, including ensuring crawl access, can measurably improve how often a source appears in AI-generated responses. Crawl access is the floor. Without it, nothing else matters. Once the crawlers are in, structure, schema and off-page citations are what move you up the answer, the full method is in our GEO playbook.
The stakes are asymmetric. Allowing AI crawlers costs you nothing. Blocking them costs you every AI-driven referral, every brand mention in a ChatGPT answer, every Perplexity citation that could have driven a click. Get this wrong and AI recommends your competitors instead of you, silently.
The one legitimate reason to block: if you have proprietary content you do not want used for model training. In that case, block GPTBot and ClaudeBot specifically, but keep OAI-SearchBot, Claude-SearchBot and PerplexityBot open so you still appear in AI search results.
Server-side HTML: why it matters for AI crawlers
Fixing robots.txt is step one. Step two is making sure AI crawlers can actually read your content once they get in. Most AI crawlers do not execute JavaScript. They fetch the raw HTML response from your server and parse what is there. If your site renders content client-side, meaning the page arrives as an empty shell and JavaScript fills it in, AI crawlers see a blank page.
This is what server-side rendering (SSR) solves. With server-side HTML, the full page content, your headings, your body copy, your structured data, is present in the initial HTTP response. The crawler does not need to run JavaScript. It reads the HTML directly.
Does this affect you?
If your site is built with a JavaScript framework (React, Vue, Angular, Next.js in client-only mode), check your rendered output. The quickest test: open your page in a browser, right-click and select View Page Source (not Inspect). If the source shows mostly empty div tags and script references, you have a client-side rendering problem. The fix depends on your stack:
- Next.js: use server components in the App Router, or getServerSideProps in the Pages Router.
- Nuxt: enable SSR mode.
- React SPA: consider a pre-rendering service or static site generation for key pages.
Googlebot has invested years in JavaScript rendering. Most AI crawlers have not. GPTBot, ClaudeBot and PerplexityBot are primarily HTML parsers. If your content is not in the initial HTML response, it is invisible to them, regardless of what your robots.txt says. Server-side HTML is non-negotiable for AI visibility.
How to detect GPTBot and other AI crawlers (and spot fake ones)
Anyone can put "GPTBot" in a user-agent string, so detection takes two checks: the user-agent token, then the IP address.
- Find the hits. Search your access logs for the tokens: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot and Perplexity-User.
- Verify the IP. Compare each request IP with the ranges the company publishes: OpenAI's gptbot.json, searchbot.json and chatgpt-user.json, Anthropic's bots.json, Perplexity's perplexitybot.json and perplexity-user.json (links in the table above). A matching token from an IP outside those ranges is not the real bot.
- Watch the pages they hit. Real search crawlers concentrate on pages you link and list in your sitemap. If OAI-SearchBot or PerplexityBot never reach your key pages, check internal links, your sitemap and your WAF rules.
A quick way to count AI crawler hits in a standard access log:
grep -E "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User" access.log | awk '{print $1}' | sort | uniq -c | sort -rn
The first column is the number of requests, the second the IP to check against the published ranges.
How to verify AI crawler access is actually working
Configuring robots.txt and enabling server-side rendering is only half the job. Verification closes the loop.
Check your server access logs
Filter for the bot user-agent strings: GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot and PerplexityBot. If you are seeing regular hits from these bots on your key pages, you are in good shape. If you see nothing, something is blocking them: robots.txt, a WAF rule, or a CDN configuration.
Fetch your page as a bot
Use curl with a spoofed user-agent to simulate what a crawler sees:
curl -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" https://yourdomain.com/your-page
If the response contains your full page content, you are good. If it returns a 403, a CAPTCHA page or an empty body, something is blocking the bot at the server or CDN level.
Monitor AI citation tracking tools
The most direct signal is whether your brand actually appears in AI answers. Tools that track AI visibility, monitoring how often your brand is named and your site cited across ChatGPT, Gemini, Perplexity, Google AI Overviews and AI Mode, will show you whether your crawl configuration is translating into actual GEO visibility. That is exactly what Howseen AI tracks. Prefer to start by hand first? Here is how to check if ChatGPT mentions your brand. Get your AI visibility score.
Keep reading: the GEO playbook, why AI recommends your competitors and what GEO is.
Does blocking GPTBot remove me from ChatGPT answers?
Not on its own. GPTBot is OpenAI's training crawler; OAI-SearchBot is the one that surfaces websites in ChatGPT search. Blocking GPTBot signals your content should not be used for training, while blocking OAI-SearchBot removes you from ChatGPT search answers. To stay visible in ChatGPT, keep OAI-SearchBot allowed.
Will AI crawlers respect my robots.txt?
The crawlers do: GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot and PerplexityBot all follow robots.txt according to their official documentation. User-triggered fetchers differ: OpenAI says robots.txt rules may not apply to ChatGPT-User, Perplexity says Perplexity-User generally ignores them, while Anthropic says Claude-User honors robots.txt.
Where is the official documentation for GPTBot, ClaudeBot and PerplexityBot?
OpenAI documents its bots at developers.openai.com/api/docs/bots, Anthropic at support.claude.com, Perplexity at docs.perplexity.ai/guides/bots and Google at Google Search Central. Each page lists the user-agent tokens and the published IP ranges.
How do I detect GPTBot in my server logs?
Search your access logs for the user-agent token (GPTBot, OAI-SearchBot, ChatGPT-User), then check that the request IP falls inside the ranges OpenAI publishes in gptbot.json, searchbot.json and chatgpt-user.json. A user-agent can be faked; the IP check is what confirms a real OpenAI bot.
Does blocking Google-Extended stop me from appearing in AI Overviews?
No. Google states that Google-Extended does not impact a site's inclusion in Google Search. It only controls whether content can be used for Gemini training and grounding. AI Overviews rely on Google's search index, built by Googlebot.
Do I need server-side rendering for AI crawlers if my site uses React?
Yes, if your content is rendered client-side. Most AI crawlers parse the HTML they receive and do not execute JavaScript, so content that only appears after JavaScript runs can look empty to them. Server-side rendering or static generation puts the content in the initial HTML.
How often do AI crawlers visit my site?
None of the major AI companies publish a fixed schedule. Frequency depends on your site's authority, how often content changes and each crawler's queue. What you control is not blocking them; Anthropic also supports Crawl-delay if you want to slow ClaudeBot down without blocking it.
Start tracking what AI crawlers actually see
Fixing robots.txt and enabling server-side HTML gets you into the game. But you need to know whether it is working. Are you being cited in ChatGPT answers? Does Perplexity link to your pages? Does Claude mention your brand when users ask about your category?
Howseen tracks your visibility across up to six AI surfaces buyers use, depending on plan: ChatGPT, Gemini, Perplexity, Google AI Overviews, AI Mode and Claude. It shows where you appear, which competitors are named instead, which pages the engines cite, and what to fix. Track your AI visibility with Howseen.
