Guide
Should a UK Business Block GPTBot and Other AI Crawlers?
AI crawlers do different jobs, so treat them separately. Allow the search crawlers, OAI-SearchBot and PerplexityBot, if you want to be shown in those answers. GPTBot and Google-Extended relate to training, and blocking them is a legitimate choice that does not remove you from Google Search. Check that your CDN or firewall is not blocking everything by default, because that is the most common cause of invisibility.
Every few weeks a client forwards me an article telling them to block AI bots to protect their content. The next week another tells them to allow every bot or disappear. Both are too blunt, because AI companies run several crawlers with quite different purposes, and the right answer depends on which one you mean.
I read the crawler documentation from OpenAI, Perplexity and Google directly for this post, because a lot of what circulates online is out of date or wrong. Here is what each company says its bots do, and the policy I would set for a UK local business.
What are the main AI crawlers and what does each do?
Each company separates search from training, and in places from user-triggered fetching. OpenAI says its settings are independent: a site can allow OAI-SearchBot to appear in search results while disallowing GPTBot to signal that its content should not be used to train OpenAI's generative AI foundation models.
| Crawler | Company | What the company says it is for | Respects robots.txt? |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Surfaces websites in ChatGPT's search features | Yes |
| GPTBot | OpenAI | Crawls content that may be used in training generative AI foundation models | Yes |
| ChatGPT-User | OpenAI | Fetches pages when a user asks ChatGPT or a Custom GPT. Not used to decide Search inclusion | OpenAI says robots.txt rules may not apply, because users initiate it |
| PerplexityBot | Perplexity | Surfaces and links websites in Perplexity search results. Not used to crawl for AI foundation models | Yes |
| Perplexity-User | Perplexity | Supports user actions in Perplexity. Not used for crawling or training | Perplexity says it generally ignores robots.txt, since a user requested the fetch |
| Google-Extended | A control token for whether crawled content may be used to train future Gemini models and for grounding in Gemini Apps and Vertex AI | Used as a robots.txt control |
Source: OpenAI, Perplexity and Google crawler documentation, read in October 2026. These pages change, so check them before editing robots.txt, and note that each lists a version number and published IP ranges.
Does blocking GPTBot remove me from ChatGPT?
No. OpenAI says GPTBot is the training crawler, and that disallowing it indicates content should not be used in training, while search visibility is controlled by OAI-SearchBot. The two are independent, which is the single most useful fact on the page.
It also means that the reflex 'block everything AI' does damage to the wrong thing. If you disallow OAI-SearchBot, OpenAI says your site will not be shown in ChatGPT search answers, although it can still appear as a navigational link. For a local business that wants to be named, that is the opposite of the goal.
What does Google-Extended actually control?
Google says Google-Extended manages whether crawled content may be used to train future Gemini models and for grounding in Gemini Apps and Vertex AI, and that it does not affect inclusion in Google Search or act as a ranking signal. Googlebot is what governs Search, including AI Overviews.
Be careful with one detail. The grounding mentioned in that description is the practice of supplying content from the Search index to the model at prompt time. Blocking Google-Extended therefore may reduce how your content is used in Gemini Apps answers, even though Search is untouched. That is a trade-off to make deliberately, not an obviously free protective step.
Google's guide to its generative AI features is explicit that a page must be indexed and eligible for Search with a snippet to appear in AI Overviews, and that crawling must be allowed in robots.txt and by any CDN or hosting infrastructure. If you block Googlebot, you lose all of it. I look at that guide in Google's AI optimisation guide.
What is the commonest reason a UK site is invisible to AI crawlers?
A firewall, CDN or security plugin blocking unfamiliar bots by default, not a decision anyone made. Perplexity's documentation includes setup guidance for Cloudflare and AWS web application firewalls because this is routine: a managed rule treats a new user agent as hostile and refuses it.
- Managed hosting and security plugins that bot-block unknown agents, common on WordPress hosts.
- CDN bot-fight or AI-scraper settings switched on during a spam scare and never reviewed.
- A blanket disallow left in robots.txt from a staging site.
- Rate limiting that drops crawler requests because they arrive in bursts.
Both OpenAI and Perplexity publish IP ranges so you can verify a request genuinely comes from them and allow it by user agent and address together, which is the secure way to do it. Allowing by user agent alone invites spoofing, and blocking by user agent alone catches the genuine bots.
What would a sensible robots.txt look like?
For a local business that wants to be named by AI search, allow the search crawlers and make a separate choice on training. This is an illustration, not a prescription, and it must be tested on your live site.
Remember that robots.txt is a request, not a lock. Reputable crawlers obey it, and the user-triggered fetchers say they may not. If something genuinely must not be read, do not put it on a public page.
Should I block training? What is my view?
It is a legitimate business decision with no single right answer, and it is separate from visibility. My own reasoning is as follows. A local service business gains nothing from being absent from search answers, so allow the search crawlers. Whether you want your pages in training data is a matter of principle and commercial sensitivity, and blocking GPTBot costs a local business very little.
- Allow search crawlers, always, if being named is a goal.
- Consider blocking GPTBot if you publish original research, paid content or proprietary guidance you would rather not feed a model.
- Think twice about Google-Extended, given the grounding detail above.
- Do not block Googlebot, under any circumstances, unless you want to leave Google.
What do I do if I use a managed host or Cloudflare?
Look for bot management and AI crawler settings in the dashboard, and allow the search bots by user agent and verified IP range. Many managed hosts and CDNs now offer one-click toggles that block known AI crawlers. They are easy to switch on during a scare and easy to forget, and they do not distinguish the search crawler you want from the training crawler you may not.
- Cloudflare and similar CDNs: check the bot and AI crawler settings, and add allow rules for OAI-SearchBot and PerplexityBot that combine the user agent with the published IP ranges.
- WordPress security plugins: review firewall and rate-limiting rules, which often block unfamiliar agents by default.
- Managed WordPress hosts: ask support whether AI crawlers are blocked at server level, and for the logs that show it.
- Test it: request a page using the crawler's user agent string and confirm you receive the page, not a block or a challenge.
Will allowing or blocking these crawlers change my Google rankings?
Not for the AI crawlers, and for Googlebot only if you block it. Google says Google-Extended is not a ranking signal and does not affect inclusion in Search, and OpenAI and Perplexity run their own separate crawlers with no connection to Google's index. What Google's own guide requires for its AI features is that a page be indexed and eligible for Search with a snippet, which depends on Googlebot, not the others.
So the risk runs in one direction. Blocking the wrong bot costs you visibility in that product, while allowing a search crawler costs you nothing you were using. If you are unsure, allow the search crawlers, decide on training deliberately and never touch Googlebot.
What about llms.txt?
Google says Google Search does not use llms.txt or similar special machine-readable files. That is covered in the post on Google's AI guide. Other AI products may behave differently, but I would not spend money on it, and it is no substitute for allowing the crawlers to read your actual pages.
How do I check that it is working?
- Fetch your robots.txt in a browser and read it as the crawlers do, group by group.
- Check your server logs or hosting dashboard for the user agents above, and confirm they receive a successful response rather than a block.
- Review CDN and firewall settings for bot management and AI-scraper toggles.
- Verify IP ranges against the published JSON files before allowlisting.
- Wait a day after changes, then re-run your prompts as described in how to check your AI visibility.
This takes an hour and it is the most reliably valuable technical task in AI search, because the failure it fixes is total. A business that blocks its own visibility cannot be rescued by better content.
Straight answers
Questions
Should I block GPTBot?
That is your choice. OpenAI says GPTBot crawls content that may be used in training its generative AI models, and that disallowing it indicates content should not be used for training. It does not control whether you appear in ChatGPT search.
Which crawler controls ChatGPT search visibility?
OAI-SearchBot. OpenAI says sites opted out of it will not be shown in ChatGPT search answers, though they can still appear as navigational links, and recommends allowing it in robots.txt.
Does Google-Extended affect my Google rankings?
No. Google says it does not affect inclusion in Google Search and is not a ranking signal. It controls use of content for training Gemini models and for grounding in Gemini Apps and Vertex AI.
Do ChatGPT-User and Perplexity-User obey robots.txt?
Both companies say they may not, because the fetches are triggered by a user's request. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them.
How long do robots.txt changes take to apply?
OpenAI says it can take about 24 hours from a robots.txt update for its search systems to adjust, and Perplexity gives a similar figure. Check again after a day before judging the result.
