Skip to main content
Back to Blog
A kraft paper parcel on a shelf illuminated by a magnifying glass, with an open archway on the left and a solid wall on the right.
Generative Engine Optimization (GEO)Intermediate

The Crawler Split: Why 'Block AI Bots' Is Now Three Decisions

On September 15, 2026, Cloudflare splits the single 'Block AI Bots' switch into Search, Agent, and Training — three decisions that decide whether your products get cited, bought by an agent, or absorbed into a model. Here is what sellers should do before the default changes.

7 min read
AI SEOGEOAI SearchAI VisibilityAI AgentsStructured DataCrawler Hints
TL;DR & Key Takeaways
TL;DR:

On September 15, 2026, Cloudflare splits the single 'Block AI bots' switch into three controls. Search: keep open to stay cited in ChatGPT and Google AI answers. Agent: keep open if you want AI shopping agents to buy from you. Training: block to stop your content feeding a model. OpenAI confirms the same split, and that opting out of its search crawler removes you from ChatGPT answers. Keep Search open, decide Agent, and finish your attributes — crawlable is not the same as understood.

Key Takeaways:
  • 'Block AI bots' is no longer one toggle; it is three independent decisions: Search (be cited), Agent (be bought from), and Training (be trained on).
  • Keep the Search crawler open: OpenAI confirms sites opted out of OAI-SearchBot do not appear in ChatGPT search answers.
  • If your store is on Cloudflare, set your own rules before September 15, 2026 or inherit defaults that block Training and Agent on ad pages.
  • Access only decides whether a bot may read you; your structured attributes decide whether it can use what it reads, so fill brand, GTIN, shipping, returns, and delivery fields.

For thirty years the rule was simple: let search engines crawl your pages, and they sent you visitors in return. That bargain is broken, and on September 15, 2026 it gets a new default for more than 20% of the web. Cloudflare — the network in front of roughly one in five internet domains — is splitting the single “AI bots” switch into three independent controls: Search, Agent, and Training. Two of them will be blocked by default for new sites.

The change matters to marketplace sellers because each control maps to a different kind of AI visibility. Keep the Search control open and your products stay eligible for the answers shoppers read inside ChatGPT, Google AI Mode, and Perplexity. Get it wrong and you disappear from those answers entirely. Here is what changed, why the old one-toggle instinct now backfires, and where sellers should focus.

The 30-year crawler bargain is over

When Cloudflare declared the first “Content Independence Day” in July 2025, it named the problem bluntly: the deal that had held for three decades — we crawl you, and you send back referrals — was no longer true. AI systems were taking content and, in Cloudflare’s words, “sending back nothing.”

The asymmetry is measurable. As Cloudflare’s engineering team wrote a year later, AI crawlers “already request content anywhere from a hundred to tens of thousands of times for every visitor they send back.” A store can be read thousands of times to ground an answer engine and receive not a single click in return.

100-10,000x
more requests AI crawlers make per visitor they send back, by Cloudflare's count
Source: Cloudflare

For a small store, that asymmetry creates a painful choice. Cloudflare describes it as a Faustian bargain: “either show up in search and let AI train on you, or risk losing discoverability.” Block everything and you protect your content but vanish. Allow everything and you hand over your work for nothing. For most of the past year that was a binary switch — “Block AI Bots” or don’t. That binary is now obsolete.

Search, Agent, Training: the new three-way map

In July 2026, Cloudflare replaced the single toggle with a pragmatic taxonomy built around what a bot actually does on your site:

  • Search crawls to index your content so it can answer questions about it later — the behavior that funnels real shoppers back.
  • Agent acts in real time on a person’s behalf, including chat-fetch bots and browser-driving agents.
  • Training absorbs your content to improve a model, permanently.

This is not just Cloudflare’s opinion. OpenAI independently organizes its own traffic the same way, with three separate user agents you can allow or block on your own terms: OAI-SearchBot for search, GPTBot for training, and ChatGPT-User for user-directed actions. Crucially, OpenAI states that each is independent — a site “can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot.” You do not have to choose between being cited and being trained on; they are separate dials.

The defaults are about to move. Starting September 15, 2026, new domains on Cloudflare will have Training and Agent blocked by default on pages that carry ads, while Search stays allowed. Sites that choose to block Training will also block multi-purpose crawlers such as Googlebot, Applebot, and BingBot, which combine search with training. Existing customers can set their own rules before that date. Cloudflare says it now analyzes more than one trillion requests per day across more than 20% of the web, so its defaults effectively redraw the line for a large share of the internet.

20%+
of the web sits behind Cloudflare, which now analyzes over 1 trillion requests per day
Source: Cloudflare

What each decision costs you

Treating the three controls as one is where sellers get hurt, because each carries a different consequence.

Search: blocking it removes you from AI answers. This is the one to keep open. OpenAI is explicit that “sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers,” though they can still appear as plain navigational links. If you blanket-block “AI bots” to protect against training, you may also block the search crawler — and the result is silence inside the answers where shoppers now start their research. That is observation, not speculation: the engines document the trade-off themselves.

Agent: blocking it cuts you off from agents that buy. This is the control that connects to agentic commerce. When Amazon’s Buy for Me or Google’s Agent Payments Protocol completes a purchase on a shopper’s behalf, an agent visits your store and interacts with it. Close the Agent door and you can still be cited, but you cannot be bought through an automated checkout. For sellers who want a shot at agent-mediated orders, that is the more expensive door to lock. (See our agentic commerce breakdown for the purchase layer this sits underneath.)

Training: the content-rights call. This is the only one with no direct visibility downside. Blocking Training protects your original work from being absorbed into a model, and OpenAI’s own separation means you can block GPTBot without losing ChatGPT search placement. For sellers whose listings are mostly commodity attributes, the risk is low; for sellers with original photography, patterns, or written guides, it is worth keeping Training closed while leaving Search open.

Crawlable is not parseable

Even sellers who get the access decision right run into a second problem: being allowed to be crawled is not the same as being understood when you are.

An anonymized FirstShelf audit of 21 listings across 90 days put the average AI-readiness score at 49.7 out of 100.

49.7/100
average AI-readiness score in an anonymized FirstShelf audit of 21 listings over 90 days

The breakdown is more telling than the headline. Listings scored 31.7 on structure quality, 30.8 on entity authority, and 41.9 on semantic density — the dimensions a search crawler or agent reads to decide whether your product matches a query. The most common missing fields were delivery method, license and usage terms, and editability: exactly the attributes a shopper’s agent checks before it will commit to a purchase.

Bar chart: FirstShelf AI-readiness breakdown (n=21, 90 days). Semantic density 41.9/100, Structure quality 31.7/100, Entity authority 30.8/100, Platform compliance 49.5/100 (FirstShelf audit, 21 listings, 90 days)
Anonymized audit of 21 listings; each score is out of 100

The lesson lines up with everything else in this archive: the access switch decides whether a bot may read you, but your structured attributes decide whether it can use what it reads. Sellers who lock down Training and assume the job is done are protecting content that is, in practice, too incomplete for an agent to act on. (Our agent-readiness deep dive covers how these signals are scored.)

What marketplace sellers should do now

  • Do not blanket-block “AI bots.” If you run your own storefront, confirm that your robots.txt allows search crawlers such as OAI-SearchBot and Googlebot even while you disallow training bots such as GPTBot. One toggle is now three; treat them separately.
  • If your store is on Cloudflare, decide before September 15, 2026. The new defaults block Training and Agent on ad-monetized pages for new domains and affect multi-purpose crawlers such as Googlebot. Set your preferences explicitly rather than inheriting the default.
  • Keep the Search door open and the content behind it clean. Visibility in AI answers is worthless if the crawler finds missing attributes. Fill brand, GTIN, size, shipping, returns, delivery method, and license fields — the same fields Google’s Merchant Center and shopping agents compare on.
  • Decide the Agent door deliberately, not by accident. If you want a shot at agent-mediated orders, do not block agent traffic or gate your cart behind flows an agent cannot complete. If you do not want them, closing that door costs you nothing in citations.
  • Marketplace sellers: check whether your platform allows search crawling. On Etsy, Amazon, or Shopify your marketplace controls robots.txt, not you. The lever you do control is listing completeness — and that is what decides whether you are cited or skipped when the crawler does arrive.

How FirstShelf can help

FirstShelf audits your listings against the same machine-readable signals a search crawler or shopping agent reads — structure quality, entity authority, semantic density, and platform compliance — and shows you the exact attributes to fill before a bot ever reaches them. Think of it as the second half of the decision in this article: you set which doors stay open; FirstShelf makes sure there is something coherent behind each one.

See how crawlable your listings really are

FirstShelf scores your listings on structure, entity authority, and semantic density so a search crawler or shopping agent can actually use them.

Audit your store free

Frequently Asked Questions

Does this affect sellers on Etsy, Amazon, or Shopify?

Indirectly. On those marketplaces the platform controls robots.txt and crawler access, not the individual seller, and the large marketplaces currently allow search crawling. What you control is listing completeness, the structured attributes that determine whether you are cited or skipped once the crawler does arrive. The Cloudflare defaults mainly affect sellers running their own storefront on Cloudflare.

If I block GPTBot, will I still appear in ChatGPT answers?

Yes, as long as you keep OAI-SearchBot allowed. OpenAI separates its crawlers: GPTBot handles training, OAI-SearchBot handles search. OpenAI states each is independent, so you can disallow GPTBot to prevent training while still appearing in ChatGPT search answers. Blocking OAI-SearchBot is what removes you from those answers.

Do I need to do anything before September 15, 2026?

Only if your site is on Cloudflare. From that date, new Cloudflare domains will block Training and Agent bots by default on pages that carry ads, while Search stays allowed, and sites that block Training will also block multi-purpose crawlers such as Googlebot. Existing customers can set their own preferences in Security settings before the date to avoid inheriting the default.

Glossary

robots.txt
A plain-text file at the root of a website that tells crawlers which pages they may or may not access. Sellers on their own storefront use it to allow search crawlers like OAI-SearchBot while blocking training crawlers like GPTBot.
OAI-SearchBot
OpenAI's dedicated search crawler. OpenAI states that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links. It is separate from GPTBot, the training crawler.
Pay Per Crawl
Cloudflare's marketplace, launched on Content Independence Day in July 2025, that lets site owners charge AI crawlers for access instead of simply allowing or blocking them. It is one alternative to the binary 'Block AI bots' toggle.

Sources