Blocking AI Crawlers vs. Googlebot: How to Protect Content Without Losing Organic Visibility
Veronika VerešováWhat does it mean
The decision to block AI crawlers from accessing content may seem like a simple way to protect company data. However, from September 2026, it may have an unexpected side effect if not set correctly – it could also restrict Googlebot, Bingbot, or Applebot and negatively impact the organic visibility of the website.
The reason is the new Cloudflare rules, which change the way they distinguish between different types of AI crawlers. Some of them are not only used for training artificial intelligence but also ensure content indexing for search engines. If you block them without thorough verification, you may inadvertently restrict traditional SEO as well.
More info
Why Companies Are Starting to Address AI Crawlers' Access to Their Content
Generative AI is fundamentally changing the way web traffic is generated. Along with the growing number of AI tools, there is also an increase in the number of crawlers visiting websites and working with their content.
This naturally raises a question that more and more companies are asking today: Do we want our content to be used for training foreign AI models?
However, the answer is not as simple as it may seem. Not all AI crawlers do the same thing, and their different behaviors are precisely why Cloudflare is changing the way they are managed.
The Difference Between a Classic Crawler, an AI Agent, and a Model Training Bot
Today, it's no longer enough to just talk about "AI bots." Cloudflare distinguishes three basic types of crawlers based on their purpose.
- Search crawler indexes content so it can appear in search results or be used in AI responses.
- AI agent works in real-time for a specific user. A typical example is ChatGPT-User or agents in Gemini or Claude, who visit a page only to fulfill a specific user request.
- Training crawler, on the other hand, collects content for training or fine-tuning language models.
The difference between these three categories is the basis of Cloudflare's new rules.
Content Protection vs. Risk of Losing Organic Traffic
Protecting your content from being used in AI training is a completely legitimate decision. However, it is also important to consider that some crawlers today perform multiple tasks simultaneously.
If you block them without thorough checking, you may unintentionally restrict the indexing of your website in search engines. The result may not only be better content protection but also a gradual decline in organic traffic.
How Cloudflare Divides AI Crawlers: Search, Agent, and Training
Cloudflare no longer treats all AI crawlers the same. The new system divides them based on the purpose for which they visit the web.
Search Crawlers: Indexing Content for Future Responses
Search crawlers are used to index content for both classic search and AI responses. Thanks to them, your content can appear in search results or be used as a trusted source when generating responses.
In most cases, it makes sense to keep their access open.
Agent Crawlers: Bots Acting in Real-Time for the User
Agent crawlers do not collect content into databases or train models. They work for a specific user and visit the web only when they need to fulfill their request.
This includes, for example, ChatGPT-User or agents integrated into Gemini or Claude.
Training Crawlers: Collecting Content for Training or Fine-Tuning Models
Training crawlers download content for the purpose of training or fine-tuning AI models. This is the category that companies most often want to restrict.
The problem is that not all AI operators use separate crawlers for individual activities. Some well-known bots today combine multiple functions at once.
When Blocking AI Crawlers Can Affect Googlebot
The biggest risk arises precisely with crawlers that perform multiple tasks simultaneously.
Why Combined Crawlers Pose a Problem with Strict Rules
Googlebot, Bingbot, and Applebot today are not only used for web indexing. Cloudflare labels them as combined crawlers because they also participate in the AI functions of their platforms.
From September 2026, Cloudflare will evaluate such crawlers according to the strictest rule that applies to them. So if you block the training part of the crawler, you may unintentionally restrict its search functionality as well.
The result may be slower crawling of the web, poorer indexing of new content, and a gradual decline in organic visibility.
The Difference Between robots.txt Recommendation and Network-Level Blocking
When setting rules, it is important to distinguish between the robots.txt file and network-level blocking.
The robots.txt file represents a recommendation that the crawler may respect. However, Cloudflare works directly at the network layer. If you block the crawler there, it simply cannot access the content.
That's why incorrect Cloudflare settings can have a significantly greater impact than just modifying robots.txt.
How Incorrect Settings Can Affect Indexing and SEO Performance
If Googlebot loses access to content, it won't show up overnight. New pages will start indexing more slowly, content updates will reach search engines with delays, and organic visibility of the web may gradually decline.
That's why it's important to verify the impact of each rule before implementing it.
What the SEO and DEV Team Should Check Before Changing Rules
Changing rules for AI crawlers is not just a marketing decision. It requires collaboration between SEO specialists and developers.
Audit of Cloudflare Rules, Verified Bots, and Server Blocks
The first step should be an audit of current settings.
Check the older Block AI Bots rule, the list of Verified Bots, and any server blocks that may affect crawler access to the web.
Checking Logs, Search Console, and Crawlability Tests
Configuration alone is not enough. It is necessary to verify how Googlebot actually accesses the web.
Server logs, Google Search Console, and crawlability tests can help reveal potential problems before they affect organic traffic.
How to Set Content Access Based on Purpose: Search, Agent, or Training
Instead of blanket blocking all AI crawlers, it is more appropriate to decide based on their purpose.
In most cases, it makes sense to keep Search crawlers open, individually assess Agent crawlers, and limit Training crawlers according to the company's strategy.
This approach helps protect content without harming organic visibility.
Recommended Approach for E-Shops and Large Websites
Every website has a different strategy, but there are several recommendations that make sense in most cases.
When to Block Training Bots and When to Keep Search Crawlers
If the goal is to prevent content from being used in AI model training, focus primarily on the Training category.
However, search crawlers should remain available because they ensure the long-term visibility of the web in search engines and AI Search.
How to Combine Content Protection with Long-Term Visibility in AI Search
The best solution is not to completely close the web to AI systems but to set clear rules based on the purpose of individual crawlers.
Companies should decide which content they want to make available for search, which for AI agents, and which they do not want to provide for model training.
Why These Rules Should Not Be Changed by Marketing Alone Without DEV Consultation
Cloudflare settings can affect the functioning of the entire web.
Therefore, they should not be decided by marketing alone. Proper settings require collaboration between SEO specialists, developers, and infrastructure administrators who can verify the actual impact on indexing and organic visibility.
How ui42 Can Help
Managing AI crawlers is becoming part of modern technical SEO. It's not enough to just address indexing and robots.txt – it's increasingly important to know how the web communicates with AI systems and what impact individual settings have on its visibility.
At ui42, we connect technical SEO, development, and AI. Our SEO and DEV specialists, along with AI agents FLUIDUM, continuously monitor indexing, web crawling, and changes in search engines and AI environments. Thanks to this, we can timely detect risky settings, propose the correct Cloudflare configuration, and protect content without harming organic traffic or visibility in search engines.
Latest news
Contact us
Everything for the growth of your business in one place
At ui42, we combine strategy, creativity, technology, marketing, and AI into one entity.
We build brands and visual identities, create websites and e-shops, design UX and CRO, create creative content, and deliver measurable results through performance marketing.
Our solutions are enhanced by FLUIDUM – ui42's AI intelligence, which transforms company data into a competitive advantage.
Thanks to this, you gain a partner who can cover the entire digital ecosystem of your business – from the first contact with the brand to conversion.
Don't miss out on the latest news from the world of UX, programming, analytics, and marketing.