AI crawler checker for ChatGPT, Claude and Perplexity

ChatGPT, Claude, Perplexity and Google's AI features can only quote pages their crawlers are allowed to fetch. This test reads your robots.txt rules for 15 AI and search crawlers, sends requests with real AI crawler user agents to see whether your firewall or CDN turns them away, checks your llms.txt and confirms your content is readable without JavaScript. Blocking AI training crawlers is reported as a policy choice and does not lower your score.

Runs only this test. Want everything? Run all 18 tests from the home page.

What this test checks

AI search and answer crawlers

robots.txt rules for the 8 crawlers that fetch pages to answer users: OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Googlebot and Bingbot.

Home page and inner pages

Rules are checked for your home page and up to 3 inner pages linked from it, to catch sites that open only the home page to crawlers.

AI training crawlers

Rules for GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, meta-externalagent and Amazonbot, reported as a content-licensing choice and not scored.

Firewall and bot protection

Whether requests with OAI-SearchBot, GPTBot, ClaudeBot and PerplexityBot user agents get the same page as a browser, or an error, a challenge page or a cut-down page.

llms.txt and llms-full.txt

Whether /llms.txt exists and follows the llmstxt.org layout (title, summary, sections with links), and whether an optional llms-full.txt is offered.

Content without JavaScript

Whether the main text is in the HTML itself, because most AI crawlers do not run JavaScript.

Quoting allowed

nosnippet or max-snippet:0 directives that tell Google, including AI Overviews, not to quote the page.

Sample result

A real result for devteam.co.kr, tested 1 hour ago.

AI crawler access

8 of 8 AI search crawlers allowed in robots.txt, 4 of 4 served by your firewall, llms.txt present.

100/100A+
AI training crawlers (Info)
Training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot…) are allowed. Blocking them is a content-licensing choice and does not affect AI search answers.
All allowed
AI search crawlers allowed (robots.txt) (Good)
All 8 AI search and answer crawlers may read the site.
8 / 8 allowed
Target: ChatGPT, Claude, Perplexity, Google and Bing search crawlers allowed
AI crawlers not blocked by firewall (Good)
Requests with ChatGPT, Claude and Perplexity crawler user agents received the same page as a browser.
4 / 4 served
Target: Same page for AI crawlers as for browsers
llms.txt (Good)
llms.txt follows the llmstxt.org format: 19 links in 6 sections.
4KB
Target: # Title, > summary, ## sections with links
llms-full.txt (Good)
llms-full.txt offers the full text of your key pages in one file.
Found
Content readable without JavaScript (Good)
The page text is in the HTML itself, so crawlers that do not run JavaScript can read it.
4,756 chars of text
Target: Main text present in the HTML

robots.txt rules for AI and search crawlers

OAI-SearchBotOpenAI · ChatGPT search index
Allowed
ChatGPT-UserOpenAI · Live visits for ChatGPT users
Allowed
GPTBotOpenAI · Model training
Allowed
Claude-SearchBotAnthropic · Claude search index
Allowed
Claude-UserAnthropic · Live visits for Claude users
Allowed
ClaudeBotAnthropic · Model training
Allowed
PerplexityBotPerplexity · Perplexity search index
Allowed
Perplexity-UserPerplexity · Live visits for Perplexity users
Allowed
GooglebotGoogle · Google Search and AI Overviews
Allowed
Google-ExtendedGoogle · Gemini training and grounding
Allowed
BingbotMicrosoft · Bing and Copilot
Allowed
Applebot-ExtendedApple · Apple Intelligence training
Allowed
CCBotCommon Crawl · Open dataset used to train many LLMs
Allowed
meta-externalagentMeta · Meta AI training
Allowed
AmazonbotAmazon · Alexa and Rufus answers
Allowed

Real requests with AI crawler user agents

OAI-SearchBotHTTP 200
Served
GPTBotHTTP 200
Served
ClaudeBotHTTP 200
Served
PerplexityBotHTTP 200
Served

llms.txt

DevTeam Korea (주식회사 뎁팀)
DevTeam Korea(주식회사 뎁팀)는 한국의 소프트웨어 개발사입니다. 스타트업 MVP, 정부지원사업 과제 개발, AI 연동·사내 AI 도입, ERP·그룹웨어·WMS 시스템 연동, Flutter 앱, Electron 데스크톱 앱, 다국어 웹사이트를 기획부터 배포·운영까지 한 팀이 맡고, 운영 중인 사이트의 장애(502·500 에러, 해킹, 인증서 만료)를 긴급 복구합니다. 소스코드와 서버 권한은 모두 고객에게 이관합니다.
Sections: 핵심 사실 (Key facts) · 개발 서비스 (Build) · AI · 시스템 · 운영 · 개선 · 회사 · 자료 · Optional

How it works

  1. Our test server in Virginia, USA, fetches /robots.txt and applies each crawler's rules as RFC 9309 describes; a robots.txt that returns a server error counts as blocking everything.
  2. Your page is then requested at the same moment with a regular Chrome user agent and with the published user agents of OAI-SearchBot, GPTBot, ClaudeBot and PerplexityBot, and the responses are compared.
  3. These requests come from our server, not from the AI companies' IP ranges, so a firewall that verifies crawler IP addresses may treat the real bots differently.
  4. We also request /llms.txt and /llms-full.txt and read your page HTML without running JavaScript.
  5. Every request goes only to public IP addresses on ports 80 and 443, and the check usually takes about 10 seconds.

How it is scored

The score uses our item-weighted formula: a pass earns full credit, a warning half and a fail none. AI search crawler access, the firewall check and readable content without JavaScript count 3×; llms.txt and permission to quote 2×; llms-full.txt 1× when present. Training-crawler rules are reported but not scored, and a missing llms-full.txt is treated as optional.

CheckGoodNeeds workPoor
AI search crawlers allowed in robots.txtAll 86–75 or fewer
AI crawler user agents at your firewallSame page as a browserOther error, no response or a much shorter pageHTTP 401, 403, 406, 429, 451 or 503, or a challenge page
Text in the HTML without JavaScriptReadable text present—Under 250 characters on a script-built page
llms.txtTitle, links and a summary or sectionsMissing or incomplete—
nosnippet or max-snippet:0Not set—Set

Grades: A+ 95–100 · A 90–94 · B 80–89 · C 70–79 · D 60–69 · F below 60.

FAQ

Should I block GPTBot?
It depends on whether you want your content used to train OpenAI's models. GPTBot collects training data; ChatGPT search relies on OAI-SearchBot, and visits made on behalf of ChatGPT users come from ChatGPT-User. You can block GPTBot and still allow OAI-SearchBot, so your pages can appear and be linked in ChatGPT search answers. This test reports training-crawler blocks without lowering your score.
What is llms.txt?
llms.txt is a proposed standard, described at llmstxt.org, for a Markdown file at the root of your site that gives AI assistants a short map of your content: a # title, a one-line > summary, and ## sections that list key pages as links with short notes. Support among AI services is limited and not well documented, so treat it as a low-cost extra, not a replacement for crawlable pages.
How do I get ChatGPT to mention my website?
Nobody can guarantee it. You can make your site easy to find, read and trust: allow OAI-SearchBot and ChatGPT-User in robots.txt and in your firewall, serve your main text in the HTML, state clearly who you are and what you offer, and earn mentions on reputable sites. ChatGPT search draws partly on third-party search providers, so being well indexed in regular search engines also helps.
My robots.txt allows AI bots, so why are they blocked?
robots.txt is only a request; your server, CDN or firewall decides what is actually served. Bot-protection features such as Cloudflare's AI bot blocking, WAF rules and rate limits can answer AI crawlers with 403 errors or challenge pages even when robots.txt allows them. Because our requests do not come from the AI companies' published IP ranges, confirm a block in your firewall logs before changing rules.
Do AI crawlers run JavaScript?
Most do not: the crawlers of OpenAI, Anthropic and Perplexity generally read the HTML your server returns. Googlebot renders JavaScript, but rendering can happen later than crawling. If a framework builds your page in the browser, use server-side rendering or static pre-rendering so the text is in the HTML.