AI crawlers are the automated bots that AI companies use to read web pages. Some collect content to train models; others run a live search when a user asks a question and find sources for the answer. This is the most important distinction when you make your robots.txt decision: blocking a training bot and blocking a search bot have very different consequences.
In this article we cover the main bots, how they are managed with the robots.txt file, and sensible decisions for different kinds of companies.
Training bots, search bots, user bots
Most major providers now separate their bots by purpose. The names below are the ones the providers publish in their own documentation; since new bots can be added over time, we recommend checking the current documentation when you make your decision.
| Provider | For training | For search and sources | On user request |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Perplexity | (no separate training bot announced) | PerplexityBot | Perplexity-User |
| Google-Extended (permission token) | Googlebot | Through Google's own tools | |
| Apple | Applebot-Extended (permission token) | Applebot | — |
| Common Crawl | CCBot (open data archive) | — | — |
Google-Extended is not a separate crawler but a permission token. It determines whether content collected by Googlebot may be used to train Gemini models. According to Google, turning this token off does not affect your site's position in Google Search; however, because the AI overviews in Google Search are fed by regular Googlebot crawling, you cannot opt out of those overviews with Google-Extended.
User-triggered bots arrive when a person tells an assistant "read this page." Some providers note in their documentation that robots.txt rules may be applied differently for this kind of visit.
How to manage them with robots.txt
robots.txt is a plain text file that sits at the root of your site and tells bots which areas they may enter. Each rule consists of a bot name followed by allow or disallow lines. For example, to block only OpenAI's training bot while leaving its search bot open, you add these four lines to the file:
User-agent: GPTBotDisallow: /User-agent: OAI-SearchBotAllow: /
Separate rules can be written the same way for ClaudeBot, CCBot or Google-Extended. You can also block just a specific folder rather than the whole site; for example, a folder of technical documents can stay open while a customer-only download area is closed.
Two important caveats. First, robots.txt is a request, not a security measure. The major providers state that their bots follow these rules, but there are crawlers that don't. Content that truly must stay private should be protected with a password. Second, robots.txt does not erase the past; content that has already been collected is not withdrawn by this file.
What does blocking cost you?
When you block search bots, the assistants in question cannot read your site during live search or cite it as a source. When someone asks about your company, the answer may be built from incomplete or outdated information on other sites rather than from your own. We cover the other conditions for visibility in AI answers in our article on appearing in AI answers.
The short-term effect of blocking training bots on visibility is less clear. In return, you limit the use of your content in model training. For publishers whose business centers on producing original content, this can be a meaningful choice.
There is also an often-overlooked situation: some content delivery and security services may block AI bots by default. Even if your robots.txt is open, bots may be stopped at the firewall. Server logs and security settings need to be checked together.
Sensible decisions by type of company
| Type of company | Sensible starting point |
|---|---|
| Corporate brochure site, service business | Allow search bots; leaving training bots open is usually harmless too |
| Manufacturing, B2B, distributor | Allow search bots; make technical documents visible |
| E-commerce | Allow search bots; block cart, account and filter pages for all bots |
| Publisher earning revenue from original content | Consider keeping search bots open and training bots blocked |
| Customer portal, dealer panel | Login areas should already be password-protected; block them in robots.txt as well |
Practical checklist
- Open your site's robots.txt file: is there a rule for AI bots, and was it set deliberately?
- Have search bots (OAI-SearchBot, Claude-SearchBot, PerplexityBot) been blocked unintentionally?
- Is a bot-blocking setting switched on in your firewall or content delivery service?
- Check your server logs to see which bots are visiting and how much load they create.
- Is content that must stay private actually behind a password?
- Write down your decision and its date; providers change their bot lists.
How we do it at Globya
We make this decision together, based on the company's business model, and make sure robots.txt, the firewall and the server settings all reflect the same decision. In our hosting and maintenance service we also monitor the load bots put on the server; for crawlers that create excessive load, interim measures such as rate limiting can be put in place. We handle the visibility side of the decision through our SEO and GEO work and the data privacy side together with our approach to KVKK, Türkiye's Personal Data Protection Law.
If I block GPTBot, will I disappear from ChatGPT entirely?
Not exactly. GPTBot is a training bot; ChatGPT's search feature runs on OAI-SearchBot. If you block GPTBot and leave OAI-SearchBot open, you can still appear as a source in search results.
If I turn off Google-Extended, will my Google rankings drop?
According to Google, no. Google-Extended only determines the use of content in Gemini models and does not affect search rankings.
How long does a robots.txt change take to have an effect?
Bots re-read the file at regular intervals; a change is usually picked up within a day. However, data collected earlier is not deleted by the change.
The Globya assistant is online 24/7; it answers right away and passes your question to the team if needed.