Block AI Training Without Blocking Search: A 2026 Crawler Checklist
A practical guide to training, search, and agent crawlers. Understand Cloudflare's new control, OpenAI's separate bots, and the limits of robots.txt.
“Block AI bots” sounds like one switch. It is not. One crawler may collect material for model training, another may support search, and a user-directed agent may fetch a page during an active task. Blocking the wrong class can protect content and cut off discovery. Allowing everything can leave a publisher with no meaningful training preference.
Cloudflare's September 2026 Disallow AI Training announcement makes this distinction unusually concrete. Here is how to make a deliberate policy without mistaking a robots preference for an enforcement guarantee.
- Write down the use you object to. Training, AI summaries, search indexing and on-demand agent access are different decisions.
- Audit both robots rules and edge blocking. A permissive
robots.txtcannot undo a firewall rule; aDisallowalone cannot physically stop a noncompliant bot. - Test the exact user agents and important URLs after a change. Retain search access if discoverability is part of your business model.
Methodology & sources
Editorial review for factual claims (as of 2026-09-28).
We reviewed Cloudflare's September 15 announcement, OpenAI's crawler documentation and Google's AI Search guidance. This is a policy checklist, not legal advice, proof that every crawler complies or a recommendation to enable a particular Cloudflare setting on an uninspected domain.
Three jobs, three decisions
| Access purpose | Example | Question for the site owner |
|---|---|---|
| Search discovery | Googlebot, Bingbot, OAI-SearchBot | Do we want this service to find and surface our pages? |
| Model training | GPTBot, Google-Extended training control | Do we permit our content to be used for training? |
| User-directed access | ChatGPT-User and other agents | Should an assistant fetch this URL for a user's current task? |
OpenAI explicitly separates OAI-SearchBot (ChatGPT search), GPTBot (content that may be used to train foundation models) and ChatGPT-User (some user-triggered actions). OpenAI says a site can allow its search bot while disallowing its training bot. It also notes that user-initiated fetches may not obey robots.txt in the same way as automatic crawling. None of this means every appearance in an AI answer depends on one bot or that a robots change instantly erases prior content.
What Cloudflare's new setting actually does
For customers using Cloudflare's AI Crawl Control, Disallow AI Training publishes applicable no-training preferences through Bot Preference Sync in robots.txt, keeps designated “Accountable” mixed-use crawlers available for search, and blocks other training crawlers. Cloudflare now says its separate Block and Block on pages with ads settings can block mixed-use crawlers such as Googlebot and Bingbot—therefore affecting search. Cloudflare's “Accountable” label includes both current capabilities and time-bound operator commitments; it is not a certificate that every promised control is live today.
The most important caveat: Cloudflare says its Disallow AI Training setting does not yet automatically convey a no-training robots preference to Bingbot. Microsoft targets support for such a domain-level robots mechanism in early 2027. Cloudflare points publishers to Bing's existing NOARCHIVE control, but Microsoft says that NOARCHIVE excludes the content from Bing Chat answers while preserving conventional search results. Thus “no Bing training” can still cost AI-answer visibility today. Review that trade-off rather than presenting NOARCHIVE as a free substitute for a training-only opt-out.
A safe change sequence
- Inventory current state: save the live
robots.txt, CDN/WAF bot rules, site-access logs, important indexed URLs and Search Console/Bing Webmaster Tools baseline. - Choose the objective: for example, “decline foundation-model training while keeping conventional search and ChatGPT search discoverability.” Document that as a policy, not just a switch name.
- Map each crawler: check official operator documentation for its search, training and user-action agents. Review mixed-use bots separately from training-only bots.
- Apply the narrowest control: if using Cloudflare, inspect the exact Search, Training and Agent settings. If not, use platform-specific robots or publisher controls and understand their compliance limits.
- Verify live behavior: fetch
robots.txtfrom the public domain, inspect WAF decisions and real bot requests, and test representative product/article URLs. Watch index coverage and citations over time; a single successful fetch does not prove continued inclusion.
Do not paste a generic “block all AI” snippet into production. It may contradict your actual discovery goal, especially where search and training use the same crawler. Nor should an AI visibility vendor imply that allowing a crawler guarantees citations. Reachability is only a prerequisite for some routes, not a promise of selection.
When a stricter policy is reasonable
A publisher whose economics depend on visitors may choose tighter controls than a merchant whose catalog must be found. A private area should be protected by authentication, not by robots directives. The defensible decision is the one that names the trade-off, the specific pages, the operator and the verification evidence. Revisit it as crawler policies and platform controls change.
Frequently asked questions
The same Q&A pairs ship as FAQPage structured data so AI engines can quote them verbatim.
- Can I block GPTBot but keep ChatGPT search discovery?
- OpenAI says its crawler settings are independent. GPTBot is associated with content that may be used for foundation-model training, while OAI-SearchBot supports ChatGPT search. You can disallow the former while allowing the latter, subject to your actual robots and firewall configuration. That access choice does not guarantee a citation.
- Does Cloudflare Disallow AI Training stop every training use while preserving search?
- No universal guarantee follows from the setting. Cloudflare says it does not yet automatically convey a no-training preference to Bingbot. Microsoft says its current NOARCHIVE control preserves conventional search results but excludes the content from Bing Chat answers. Review that AI-visibility trade-off and verify operator-specific and edge behavior.
- Is robots.txt enough to protect private content?
- No. Robots rules express crawling preferences to cooperating agents and do not authenticate visitors or enforce privacy. Use proper access control for private pages. For public pages, review robots.txt together with CDN and firewall rules so you do not accidentally block search or leave unwanted access open.
Primary documentation and limitations
- Cloudflare: Disallow AI Training and mixed-use crawlers — setting behavior, migration and Bing exception.
- Microsoft Bing: publisher controls for Bing Chat and training —
NOARCHIVEexcludes Bing Chat answers but preserves conventional search results. - OpenAI: crawler and user-agent roles — independent search, training and user-action controls.
- Google Search Central: generative AI and Search guidance — ordinary crawl and indexing foundation.
Reviewed September 28, 2026. Crawler behavior and contractual commitments can change; verify the live configuration before editing it.
Related articles
Your GEO Score
Establish an AI mention baseline you can defend
GEO Tracker AI runs repeatable checks for supported engines so you can see whether your brand is mentioned, what context shows up, and how that changes week over week — complementary to Search Console, not a replacement for it.