AI & ML
133 of 10,099 Shopify stores block an AI crawler. Six block the one ChatGPT shops with.
bpmcginley Dev.to (EN Zone)
1 views
Merchants are told two opposite things about AI crawlers: block them, because
they take your content and give nothing back; and admit them, because that is
how a store gets into an AI shopping answer. Both assume a decision is being
made. We wanted to know how many stores have made one, and which way.
Method
Every store in a corpus of 10,099 known Shopify storefronts has its
/robots.txt read as part of a scan. The file is parsed the way the major
crawlers document parsing it: most specific User-agent group wins, * is
the fallback, longest matching path rule wins, Allow beats Disallow on a
tie. Each of twelve crawler names is asked one question: may it fetch
/products/? A store counts as blocking a crawler when the answer is no.
A second reading looks at the product page for a meta name="robots" tag
carrying noai or noimageai.
Readings were taken between 29 August and 2 September 2026. Every store that
blocked at least one crawler, or carried the tag, is one row in the CSV at
the end. The rest of the corpus blocked nothing and is the denominator.
What this cannot see. A robots.txt is a request. A store can also block a
crawler at the edge, with a bot-management rule or a firewall, and that block
is invisible here because the scanner is not the crawler being blocked. Every
count below is a floor.
Result
133 of 10,099 stores block at least one crawler: 1.32%.
Crawler
What it feeds
Fetches at answer time
Stores blocking
CCBot
Common Crawl
no
81
GPTBot
OpenAI training & retrieval
no
77
Bytespider
TikTok / Doubao
no
72
Amazonbot
Alexa+ / Rufus
no
60
Google-Extended
AI Overviews & AI Mode grounding
no
58
ClaudeBot
Claude retrieval & citations
no
53
Applebot-Extended
Apple Intelligence
no
48
meta-externalagent
Meta AI
no
47
ChatGPT-User
Live fetches during a ChatGPT chat
yes
11
PerplexityBot
Perplexity search & shopping
yes
10
OAI-SearchBot
ChatGPT search & shopping results
yes
6
Perplexity-User
Live fetches when a Perplexity user asks
yes
3
"Fetches at answer time" marks the four crawlers that read a page, or index
for a search result, at the moment a person is asking. The other eight crawl
ahead of time, to train or ground a model. The split follows each operator's
published description of the name. The twelve are the names merchants'
block lists actually carry, not every answer-time agent that exists;
Anthropic's and Google's live-fetch agents are not among them.
How they block
130 of 133 block with Disallow: /. The whole site, not the product pages.
3 block /products/ alone.
54 block exactly one crawler.
28 block the same eight, verbatim: Amazonbot, Applebot-Extended, Bytespider,
CCBot, ClaudeBot, GPTBot, Google-Extended, meta-externalagent. 4 more block
those eight and others. Identical lists do not arise independently; this
one reads like a copied snippet, and it contains none of the four
answer-time crawlers.
1 store blocks all twelve.
Training or answering
Of the 133 stores that block anything, 120 block only training and grounding
crawlers. 13 block at least one crawler that fetches at answer time; 1 of
those blocks only answer-time crawlers, 12 block both kinds.
The clearest pair is OpenAI's. GPTBot trains; OAI-SearchBot is what
ChatGPT's search and shopping results are built from. 77 stores block the
first. 6 block the second, and every one of those 6 also blocks the first.
The reverse, keep the shopping crawler out and let the training crawler in,
happens on 0 stores. No store in the corpus has decided to stay out of AI shopping answers.
The stores that are out of them by robots.txt are there because a copied
training opt-out happened to include the name.
The noai tag
8 stores carry a noai or noimageai meta robots tag on their product page.
6 of the 8 block nothing at all in robots.txt. By their names, six of the eight are
musicians' merchandise stores, which again looks like one template rather
than eight decisions. The tag is a separate mechanism and this study makes no
claim about which crawlers honour it.
What it does and does not mean
Blocking AI crawlers is rare on Shopify. 98.7% of stores block none of
the twelve by robots.txt. Where it happens it is mostly a whole-site rule
aimed at training crawlers, and mostly one list.
The shopping crawlers are almost never blocked. 13 stores in 10,099 block
any of the four answer-time crawlers we checked. For the rest, whatever stands between
the store and an AI shopping answer, it is not robots.txt.
It is not a recommendation. A merchant who does not want their catalogue
in a training set has a reason, and a copied list that does the job is a
fine way to act on it. The finding is narrower: the list most of them copied
was written before the answer-time crawlers existed, so a store using it has
made a decision about training and, without noticing, none about shopping.
It is a floor. Edge blocks are real and this cannot count them.
Reproduce it
curl -s https://EXAMPLE.com/robots.txt
Find the User-agent: group that names the crawler; if none does, the *
group applies. Under that group, the longest rule matching /products/
decides, and Disallow: / matches everything.
All 139 rows, one per store, are in
ai-crawler-blocking.csv:
host, grade, date read, one column per crawler, whether the noai tag was
present, and whether the rules were whole-site or product pages only. Every
row names a domain whose robots.txt you can read yourself.
Originally published at
shelfglance.com/research/ai-crawler-blocking,
where the corpus of 10,099 scanned storefronts lives.
Read original: https://dev.to/bpmcginley/133-of-10099-shopify-stores-block-an-ai-crawler-six-block-the-one-chatgpt-shops-with-13ec
← Previous
Scraping 150k+ Instagram followers reliably: batching, resume-on-error, and enrichment
Next →
Shopify's agent-commerce category filter doesn't filter. We checked 190 stores.
Related
M
Machines Can Only Build What Someone Already Imagined
AI & ML
0
DEV Community
W
What breaks when you ship 21 AI tools that never touch a server
AI & ML
0
DEV Community
T
The 404 only we could see: 23.8 hours inside a cache entry we made ourselves
AI & ML
0
DEV Community
W
What Should a Board Ask Before Approving an AI Coding Tool Rollout?
AI & ML
0
Dev.to (EN Zone)
Comments0
No comments yet — be the first