Can ChatGPT Actually Read Your Website? Three Things That Quietly Block It
A robots.txt line written years ago, a security setting nobody remembers, a menu built in JavaScript. Three ways a site becomes invisible to AI assistants, and how to check each in minutes.
When a business asks why ChatGPT never mentions it, the conversation usually jumps straight to content, reviews and reputation. Those matter. But there is a more basic question underneath them, and it is surprisingly often the answer: can the assistant read your website at all?
A page an AI assistant cannot fetch is a page it cannot quote, summarise or recommend. And the ways a site ends up shut out are rarely deliberate. A line in a file written years ago. A security setting nobody remembers switching on. A menu that only exists once JavaScript has run. Here are the three we check first, and how you can check each one yourself in a few minutes.
1. Your robots.txt file
Every website can publish a small text file at /robots.txt that tells automated visitors what they may and may not read. Well-behaved crawlers, including the ones run by OpenAI, Anthropic, Google and Perplexity, obey it. Open yours now: type your address followed by /robots.txt into a browser.
The complication is that each AI company runs several crawlers with different jobs, and a rule aimed at one does not necessarily touch the others:
- OpenAI documents three:
GPTBotcollects content that may be used to train its models,OAI-SearchBotbuilds the index ChatGPT’s search cites from, andChatGPT-Userfetches a page when someone asks ChatGPT about it. OpenAI’s own help pages on ChatGPT search say a site needs to allow OAI-SearchBot to be eligible to appear. - Anthropic documents three for Claude:
ClaudeBotfor training,Claude-SearchBotfor search, andClaude-Userfor fetching a page a person has asked about. - Perplexity uses
PerplexityBotfor its index andPerplexity-Userfor fetches on request. - Google is different. Its AI Overviews and AI Mode are built from the ordinary Google Search crawl. The
Google-Extendedtoken controls whether your content is used for Gemini models, and Google states it does not affect your inclusion in Google Search.
This split is useful. It means you can make a real choice: some owners are uncomfortable with their content training AI models and block the training crawlers, which is a legitimate decision. What they rarely intend is to block the search crawlers too, because those are the ones that decide whether an assistant can find and cite them when a customer asks.
What to look for
A crawler follows the group of rules that names it. If no group names it, it follows the group under User-agent: *. So look for two things:
Disallow: /under any of the names above. That shuts that crawler out of the whole site.Disallow: /underUser-agent: *, with no specific group allowing the AI crawlers back in. That shuts out everyone who is not named, including every AI crawler, and usually Google as well. It is more common than it should be, often left over from a site that was once in development.
For comparison, this is the relevant part of ours:
User-agent: *
Disallow: /wp-admin/
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /
The explicit Allow lines are belt and braces: with that * group, those crawlers were already allowed. We list them so that anyone reading the file, human or machine, can see the decision was deliberate.
2. Security settings between the crawler and your site
robots.txt is a polite request. Many sites also sit behind a firewall, a hosting security layer or a content delivery network that can turn a crawler away before it ever reads the file. These settings are often switched on by whoever set up the domain, then forgotten.
Cloudflare, which sits in front of a large share of the web, is the most visible example. It has offered site owners controls for AI crawlers for some time, and announced in July 2026 that from 15 September 2026, new domains would block AI training and AI agent crawlers by default on pages that display ads, while search crawlers stay allowed.
Read that carefully before panicking: it applies to new domains, and to pages carrying ads. A typical Cayman business website has no ads and is not affected by that default. But it shows how these decisions now get made at a layer the owner never looks at. If someone has switched on a blanket “block AI bots” option, or an aggressive bot-protection mode, your robots.txt can say “welcome” while the network in front of it says no.
How to check
The simplest test uses the assistants themselves. Open ChatGPT, paste the full address of one of your pages, and ask it to summarise that exact page. Do the same in Claude. If it reports that it cannot access the page, while the page loads fine in your own browser, something between it and your site is refusing the request. Then ask whoever manages your domain or hosting to look at the bot and firewall settings.
This test checks the “fetch on request” crawler, not the search index, so a pass is not a guarantee. A failure, though, is a clear sign.
3. Navigation that only exists in JavaScript
This one is the least obvious and potentially the most damaging, because it hides whole sections of a site rather than the site as a whole.
Browsers run JavaScript; most crawlers do not, or do so only partially. If the links to your pages only appear after JavaScript has run (menus that build themselves when clicked, “load more” buttons, some site-builder templates, single-page apps), a crawler reading the raw HTML may never discover those pages exist.
A test published by Search Engine Land in September 2026 put numbers on it. On a test site, pages linked with ordinary HTML links were crawled by GPTBot and ClaudeBot, while pages linked only through client-side JavaScript were not found by either. Even Google’s own crawlers reached only part of the JavaScript-linked content. When the links were changed back to plain HTML, the AI crawlers found the pages.
For a Cayman business the usual casualties are the menu page, the individual service pages and the rooms or tours pages: exactly the pages with the facts an assistant needs to recommend you.
How to check
Open your home page, right-click, and choose View page source, or press Ctrl+U. That is roughly what a crawler that does not run JavaScript receives. Press Ctrl+F and search for the name of one of your inner pages, or its address. If your menu links appear as ordinary <a href="…"> tags, you are fine. If they are nowhere in the source, your navigation is being built by JavaScript, and it is worth a conversation with whoever built the site.
While you are there, search the source for a sentence from your menu or your price list. If it is only in an image or a PDF, that is a related problem we covered in why ChatGPT cannot read your menu.
Questions we get asked
Should I block AI crawlers from training on my content?
That is a business decision, not a technical one, and reasonable people disagree. What matters is making it on purpose and per crawler. If you block GPTBot and ClaudeBot but allow OAI-SearchBot, Claude-SearchBot and the user-fetch crawlers, your content stays out of training data while remaining findable and citable.
Will an llms.txt file fix this?
No. An llms.txt file is a proposed convention for pointing AI systems at your most important content. It cannot override a robots.txt block, a firewall rule or navigation a crawler never sees. We wrote about what llms.txt does and does not do.
My site is on a website builder. Do I control any of this?
Usually some of it. Several builders let you edit robots.txt or have added a setting for AI crawlers in their SEO or privacy options. It is worth ten minutes in the settings, or a message to their support asking plainly whether AI crawlers are blocked on your site.
If I fix this, will ChatGPT start recommending me?
Not by itself. Access is the entry ticket, not the prize. Once the crawlers can read your site, what they find there, and what other sites say about you, decides the rest. But none of that can work while the door is shut.
Where to start this week
Three checks, ten minutes: open /robots.txt and look for Disallow: /; ask ChatGPT and Claude to summarise your home page; view the source of your home page and search it for the name of an inner page. If all three come back clean, the door is open.
Our free checker at toctoc.ky/seo-checker reads your robots.txt the way these crawlers read it, including the User-agent: * fallback that catches most sites out, and flags any AI crawler you are shutting out. If what you find needs fixing properly, that is what we do: AI search visibility for Cayman Islands businesses, and websites built so crawlers can read them.
Written by
Andre Gutierrez
Toc Toc Marketing builds websites in the Cayman Islands that Google and AI assistants can read, understand and recommend.
Meet the team →