In May, a software directory told us it could not verify our listing. Its crawler kept failing to load our homepage. The homepage loaded fine for me. It loaded fine for everyone on the team, in every browser we tried. We opened a ticket with our host. They checked the firewall settings, found nothing, and suggested the directory was the problem.
It was not the directory. It took us three months, one accidental IP ban and a small piece of software to prove it.
Why this belongs on an MCP blog
Most of what we write here is about the front door: how to build an MCP server so AI agents can work with your business properly, with real tools, sensible auth and data you control. All of that quietly assumes something more basic. That an AI system can reach your website in the first place.
Two kinds of bot matter. Training crawlers like GPTBot, ClaudeBot, Google-Extended and Common Crawl’s CCBot decide what next year’s models know about you. User-triggered fetchers like ChatGPT-User, Claude-User and Perplexity-User run when a person asks an assistant about you right now, and the assistant goes to look. Block the first and the models forget you exist. Block the second and the assistant shrugs at a live buyer.
An MCP server on a site that turns away ClaudeBot is a beautifully designed front door on a street that has been closed off. So before we write more about doors, a story about the street.
The rule nobody had set
Our host was SiteGround. Our first test was the obvious one: send a request that looks like the directory’s crawler.
StackShare/1.0 -> 200
Mozilla/5.0 (compatible; StackShareBot; +https://stackshare.io) -> 200
Mozilla/5.0 (compatible; StackShareBot/1.0; +https://stackshare.io) -> 403
The only difference between the last two lines is /1.0. Any user agent shaped like Mozilla/5.0 (compatible; Name/Version; +URL) was refused. A bot name I made up on the spot, SomeBot/1.0, got the same 403. So this was not a blocklist of known bad actors. It was a pattern.
That pattern is the format every well-behaved crawler is told to use: identify yourself, give a version, link to a page explaining what you do. Scrapers pretending to be Chrome walked straight in. The honest bots were the ones being turned away. The error page itself even carried a noindex tag, so anything that saw it was also being told to forget the page.
I understand why a rule like this exists. Scrapers are real, they cost money, and a few years ago most things announcing themselves as “compatible; Something/1.0” were junk. I don’t think anyone at the host set out to block GPTBot. I also think that in 2026 a regex that cannot tell an AI crawler from a content thief does more harm than good.
Measuring it properly
Knowing one directory bot was blocked told us nothing about the size of the problem. So we wrote a probe: 83 real crawler user agents, taken from the vendors’ own documentation where they publish one, across search engines, AI crawlers, SEO tools, directories, social link previews, web archives and uptime monitors. One request per bot, spaced out, from a clean cloud IP, recording exactly what the server sent back.
Our three sites on that host scored between 67% and 69%. A site we run on Vercel scored 100% on the same probe, the same day.
| Category | On the shared host | On Vercel |
|---|---|---|
| Search engines | 17 of 18 | 18 of 18 |
| AI crawlers | 7 of 23 | 23 of 23 |
| SEO tools | 6 of 9 | 9 of 9 |
| Directories | 5 of 10 | 10 of 10 |
| Social link previews | 10 of 10 | 10 of 10 |
Google and Bing were fine, which is exactly why nobody had noticed. Search Console looked healthy because Googlebot was never the one being blocked.
The AI row is the one that hurt. GPTBot, ClaudeBot, Claude-Web, CCBot, Google-Extended, Meta-ExternalAgent, Applebot-Extended, Amazonbot, Bytespider, Cohere, Mistral, You.com, Diffbot, AI2Bot and two smaller crawlers, all refused. Most were not even given the courtesy of a 403. The connection was simply dropped, which from the outside looks exactly like the site being down.
The second wall, found the hard way
My first version of the probe fired all its requests from my laptop in about twenty seconds. Somewhere around request eighty the host decided I was an attack. Every response turned into an HTTP 202 with a header reading sg-captcha: challenge, a redirect to a captcha page and, for good measure, x-robots-tag: noindex.
My office IP stayed in that state for two days.
That is a second protection layer, separate from the user-agent rule, and in some ways the more worrying one. A crawler that fetches a little too eagerly from one address does not get a polite “slow down”. It gets a page telling it not to index anything. The probe now runs from serverless infrastructure with deliberate spacing between requests, a lesson I would rather have read than learned.
The experiment
Months of tickets had produced a partial fix and a “resolved” status on a problem that was still there. So we stopped asking and tested the obvious hypothesis instead: if the host is the cause, moving the site should fix it, and nothing else should need to change.
We picked our smallest site, a WordPress blog with 28 posts, and moved it to a $6 a month DigitalOcean droplet. Same WordPress, same content, same Cloudflare in front. Only the origin server changed. It took about ninety minutes.
| Before | After | |
|---|---|---|
| Overall | 57 of 83 (69%) | 82 of 83 (99%) |
| AI crawlers | 7 of 23 | 23 of 23 |
| Directories | 5 of 10 | 10 of 10 |
| SEO tools | 6 of 9 | 9 of 9 |
The single remaining miss is Python-urllib, a programming library rather than a crawler, which Cloudflare’s default rules decline. That is a block I am happy to keep. With Cloudflare switched off for a few minutes the score was 83 of 83, which settled the other question: our CDN was never part of the problem.
Then we moved everything else, including our main site and this blog. The main site had been hosted in the UK while most of our customers are in the US and Canada, so the move came with a bonus. Measured from New York, median page load went from 595 ms to 40 ms.
I re-ran the probe this morning, two months on. All three sites: 82 of 83.
What I would tell you to do
Check, do not assume. Your robots.txt is a request. A firewall rule is a wall, and the wall wins. Most site owners have never looked at what their host, their security plugin or their CDN does to a bot that announces itself honestly, because no dashboard shows it.
If you find blocks, ask your host for the specific rule and a written fix, then re-test rather than trusting the ticket status. If they cannot give you either, moving is cheaper than it sounds. Ours took an afternoon and costs less per month than a sandwich.
And if you are building an MCP server, do this first. It is a strange feeling to finish a careful piece of agent infrastructure and then discover the agents were never let in.
The tool
We turned the probe into botsvisibility.com. Paste a URL, wait about a minute, and see which of the 83 crawlers your site lets in, which it refuses and which it challenges. It is free, there is no signup, and the code is open source under MIT at github.com/phwizard/botsvisibility, so you can run it yourself or add the bots we have missed.
Run it on your own site and tell me what you find. I am collecting these, especially the ones where the host swears nothing is blocked.