Cloudflare's September 15 AI crawler defaults: what changes and what to check this week
From September 15, 2026, new Cloudflare zones block Training and Agent crawlers by default on pages that show ads, 'Verified' no longer means 'allowed', and robots.txt gains a use= signal. What each change does to your citations, and the three settings to confirm.
On September 15, 2026, Cloudflare’s default bot policy changes for every new domain onboarded to its network: crawlers it classifies as Training or Agent are blocked by default on pages that display ads, Search crawlers stay allowed, and a bot’s “Verified” status stops being an automatic allow. Cloudflare announced the change on July 1, 2026, one year after it introduced block-by-default for AI crawlers and its Pay Per Crawl marketplace.
For a site that wants to be cited, the change is mostly good news with two sharp edges. This article covers the three parts of the announcement, what each does to the crawlers that produce citations, and what to verify before the date.
What are the three categories?
Cloudflare now sorts AI-related bot behavior into three use cases, with different defaults:
| Category | Cloudflare’s definition | Default from Sept 15 |
|---|---|---|
| Search | ”any behavior that collects or indexes your content, so it can answer questions about it later” | Allowed |
| Agent | ”automated behavior that is acting, usually in real time, on a person’s behalf, to get something done right now” | Blocked on pages that display ads |
| Training | ”a crawler taking your content to train or fine-tune a model” | Blocked on pages that display ads |
Cloudflare also tracks Transact, Data Collection, Security Testing, SEO, Ads Verification, Social/Link Preview, Feed Fetching, and Monitoring categories, but the three above are the ones with new defaults.
Map these onto the three kinds of AI crawler and the citation impact is clear. OAI-SearchBot, Claude-SearchBot, and PerplexityBot are Search and stay open. GPTBot, ClaudeBot, and Bytespider are Training and get blocked on ad pages, which costs no citations. The awkward class is Agent: ChatGPT-User, Claude-User, and Perplexity-User fetch a page because a user just asked about it, and blocking them on an ad-monetized page means the answer engine tells the user it could not read your site.
Who does this apply to?
Cloudflare’s wording is specific: “For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default.” Existing zones keep their current settings and can change them under Security. Coverage of the announcement has described the new defaults as also reaching existing free-plan zones; if your site is on a free plan, treat that as unconfirmed and check the setting rather than assuming either way.
“Pages that display ads” is not given a technical definition in the announcement. Cloudflare’s stated rationale is that “an ad is a signal that a website owner meant for a person to land there and see it.” Expect detection to be based on ad-network markup and known ad-serving hosts; expect false positives on pages with third-party embeds.
What does “Verified no longer means allowed” change?
Until now, a bot on Cloudflare’s verified list was allowed by default unless a customer blocked it. From September 15: “we are no longer viewing Verified as ‘default allowed.’ Now, the Verified label makes a bot allowable with its relevant category.” Verification proves identity; the category decides access.
The consequence that matters most is for multi-purpose crawlers. Cloudflare names Googlebot, Applebot, and BingBot as crawlers that combine Search with Training and says they “will be allowed/blocked according to all of their behaviors.” In the announcement’s words, “multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training.”
Read that twice. A zone that opts into blocking Training on ad pages, under the new model, blocks Googlebot on those pages. That removes the page from Google Search, AI Overviews, and AI Mode, since all three depend on Googlebot as the inclusion guide explains. Whether Google changes its crawler declarations in response is the open question of the autumn; until it does, a Training block is a Google block on monetized pages.
What is the new use= signal in robots.txt?
Cloudflare’s Content Signals policy, introduced in September 2025, adds a Content-Signal: line to robots.txt expressing how content may be used. The July 2026 update adds a use= parameter with three values:
| Value | Meaning |
|---|---|
use=immediate | ”interact, but store and reuse nothing” |
use=reference | ”index, excerpt, and link back” (the default in Cloudflare-managed files) |
use=full | ”summarize and reproduce” |
A managed robots.txt will carry a line such as:
Content-Signal: search=yes,ai-train=no,use=reference
Two facts about this line. First, it is a preference, not a control. Google’s John Mueller said of the original Content Signals directive that it has “no effects whatsoever for any crawler or LLM” and “just adds bloat and future maintenance to your robots.txt file”; no operator has committed to honoring use=. Second, Cloudflare says it will track compliance in its bot registry and strip Verified status from bots that abuse a declared preference, which is the only enforcement the signal has. The audit’s robots_ai_bots check ignores Content-Signal lines and reads only the User-agent / Disallow rules that crawlers actually obey.
What should you check before September 15?
Three settings, in this order.
1. Which category your zone blocks. In the Cloudflare dashboard, open Security → Bots (or AI Crawl Control on plans that have it) and read the AI crawler policy. Decide per category: Search should be allowed; Agent should be allowed if you want live-fetch citations; Training is a business decision. If you block Training, confirm what happens to Googlebot on ad pages before saving.
2. Whether ads appear on pages you want cited. The defaults key on ad presence. A content page with an ad slot inherits the Agent and Training blocks; the same page without the slot does not. If a page’s value is in being cited rather than in ad revenue, that is a reason to remove the slot.
3. What your origin returns to an AI user agent. The audit fetches with a browser user agent, so it will pass fetch_direct even when your edge returns a challenge only to ChatGPT-User. Test directly:
curl -sI -A "Mozilla/5.0 (compatible; ChatGPT-User/1.0; +https://openai.com/bot)" https://example.com/ | head -1
curl -sI -A "Mozilla/5.0 (compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot)" https://example.com/ | head -1
A 403 or a 503 with a cf-mitigated: challenge header on either means the edge is blocking the citation path regardless of what robots.txt says. Note that a curl from your laptop will not carry OpenAI’s IP, so a rule that verifies by IP will block your test and allow the real bot; use Cloudflare’s bot analytics to see what the real requests received.
The longer arc
The 2025 change made blocking the default for new sites. The 2026 change makes the default category-aware, which is the first time a CDN’s defaults have distinguished the crawl that trains a model from the crawl that cites your page. That is the right distinction. The cost is that every site owner now has to know which of their traffic is which, and to know that Googlebot sits on both sides of the line.
Further reading
- Cloudflare bot protection and AEO — the silent killer
- Training, index, or live fetch: the three kinds of AI crawler
- Google AI Overviews and AI Mode: what actually decides whether you’re in
Find out whether a direct fetch of your page succeeds today: Run an audit →