The AI Crawler Readiness Checklist
An AI crawler readiness checklist should tell you whether the crawlers you want can reach your public pages, receive useful content, and follow your publishing rules. Check the actual responses and logs. A page that opens on your laptop can still hand a crawler a challenge screen.
Start with your homepage, a service or product page, pricing, documentation, and one important article. Include each hostname you use. Keep the findings by URL and crawler so a working blog does not hide a blocked documentation site.
The checks below focus on access, extraction, and monitoring. If you need to build the content first, use our website publishing guide.
The checks and the evidence to keep
| Check | Pass condition | Evidence |
|---|---|---|
| Crawler policy | Search, training, and user-requested access reflect your chosen policy. | Relevant directives and access rules. |
| Delivery | Permitted requests reach the intended page without an unexpected block or challenge. | Response headers, body, redirect path, and matching logs. |
| Search eligibility | Pages intended for Search have no accidental indexing or snippet restrictions. | Page directives and Search Console inspection. |
| Content extraction | The retrieved text includes the facts and qualifications a reader needs. | Saved response compared with the visible page. |
| Optional formats | Published Markdown and navigation files work as advertised. | Content type, working links, and content comparison. |
| Ongoing monitoring | Someone can identify and investigate a change in access. | Baseline, alert owner, and retest record. |
1. Decide which requests you want to allow
“Allow AI” is too broad to be a useful policy. You may want your pages in search results while declining training crawls. You may also want a customer to ask their assistant to read a particular page. Record those decisions separately.
OpenAI's crawler documentation distinguishes OAI-SearchBot for search, GPTBot for potential model-training use, and ChatGPT-User for certain user-triggered visits. Search and training controls are independent. OpenAI says robots.txt rules may not apply to user-initiated requests.
Perplexity makes a similar separation between PerplexityBot, its search crawler, and Perplexity-User, its user-requested fetcher. Its documentation says the latter generally ignores robots.txt. Check each provider's current guidance rather than applying one company's behavior to every bot.
For Google Search, including AI Overviews and AI Mode, Googlebot controls apply. Google-Extended controls some other AI uses; blocking it is not a way to opt out of Google's AI Search features.
Review the robots.txt served by every relevant host, then compare it with your CDN and firewall settings. A written allowance cannot undo a network block. And robots.txt is not a lock on private information: keep authentication on content that requires it.
2. Fetch the page through the public route
Test the URL a visitor would use, through the CDN if you have one. Use a GET request and inspect the body as well as the status. A 200 response containing “Please verify you are human” has not delivered your article.
For a first diagnostic, replace this example address with a public page you operate:
curl --silent --show-error --location --max-redirs 5 --max-time 30 \
--dump-header page-headers.txt \
--output page-body.html \
'https://example.com/pricing/'
Check the destination after redirects, content type, and returned text. Investigate unexpected 403, 429, or 5xx responses. A deliberate denial can be correct; record whether the result matches your policy for that request.
Changing curl's user-agent header can reveal a rule that treats that string differently. It does not turn the request into a genuine Google, OpenAI, or Perplexity visit. Confirm provider access with identified requests in your logs or the provider's diagnostic tools where available.
Use official verification methods before creating crawler exceptions. Google documents IP-range checks and reverse/forward DNS verification. OpenAI and Perplexity publish IP lists in their crawler documentation. Keep those sources current; a copied user-agent string does not verify who sent the request.
Find the rule behind the block before changing it. Scope an exception to the intended traffic and public paths. There is no reason to turn off protection across the site just to make one documentation page reachable.
3. Check indexing and snippet controls separately
A successful fetch does not mean a page is indexed. For pages you want in Google Search, inspect the robots meta tag and X-Robots-Tag response header, check the canonical destination, and use Search Console to investigate indexing.
Google requires pages to be indexed and eligible for a snippet before they can appear as supporting links in AI Overviews or AI Mode. Check for accidental noindex, nosnippet, or overly restrictive preview settings. Eligibility still does not guarantee selection.
Be careful with removal instructions too. Google must be able to crawl a page to see its noindex rule. Blocking the same page in robots.txt can prevent that rule from being read. Choose the control that matches the outcome you want.
Check that internal links and sitemap entries lead to the public URLs you intend to maintain. If a canonical points elsewhere, confirm that destination is deliberate and reachable.
4. Compare the received content with the page people see
Open the saved response or run it through a text extractor. Can you find the product name, price basis, policy exceptions, and contact route? If the browser shows a delivery restriction but the retrieved text omits it, the page has lost information a customer needs.
Include the awkward pages. Pricing tables, tabbed documentation, and pages behind consent screens can behave differently. Check rendered output as well when the intended crawler supports rendering; do not assume every fetcher runs JavaScript.
Keep structured data consistent with visible content. Check authorship, dates, and source links where they support the claims. Preserve embedded media, but put essential facts in text too. A crawler should not need to watch a video to discover that a service is unavailable in its customer's country.
5. Test optional files you actually publish
You do not need a complete folder of agent files to pass a crawler audit. Google does not require special AI files for its Search features. Audit your chosen publishing setup.
If you publish llms.txt, fetch the exact address and follow its important links. The v2 proposal permits scoped files such as /docs/llms.txt. Check that the applicable file points to current material. Treat a missing optional file differently from a broken file your site advertises.
If you serve Markdown, compare it with the human page. Check tables, qualifications, and source links. For negotiated delivery, request both HTML and Markdown and inspect each response:
curl --silent --show-error --location --max-time 30 \
-H 'Accept: text/markdown' \
--dump-header markdown-headers.txt \
--output page-markdown.txt \
'https://example.com/pricing/'
Cloudflare's Markdown for Agents, when enabled for eligible responses, sets the Markdown content type and includes Accept in Vary to separate cached representations. Verify that browsers still receive HTML. Review Content Signals as well: Cloudflare documents a default allowing search, AI input, and training when the origin supplies no custom signal.
Workflow guides and public manifests serve particular agent integrations. Check the documented URLs you publish instead of assuming every site has /SKILL.md or /.well-known/agent.json. Our file-format comparison covers when those resources are useful. For delivery options, see Markdown for agents.
6. Keep a record you can troubleshoot
For each tested URL, record the time, request type, expected policy, actual status, final destination, and a short note about the body. For observed provider traffic, add how you identified the crawler and any firewall rule involved. Leave room for the person responsible for the fix.
Suppose a permitted search crawler receives a 403 on pricing while the homepage works. Record the failed path and rule, make the narrow correction, then confirm the next identified request receives the pricing content. “The bot is allowed” is not evidence that the fix worked.
Watch allowed requests, unexpected denials, challenges, server errors, and the pages crawlers visit. Cloudflare AI Crawl Control provides crawler traffic and policy tools if you use that platform; other setups can use their CDN and server logs.
Track referrals and useful business outcomes separately. A crawl is a fetch, not a citation or a customer. Google reports AI-feature traffic within Search Console's Web performance data, so do not label the whole report “AI traffic.”
Repeat affected checks after changes to your CDN, templates, redirects, or publishing policy. Keep an eye on managed-rule updates too. If access changes, your saved responses and rule identifiers give you somewhere to start investigating.
Connect the reachable site to the service behind it
Once your public pages are accessible, give an agent-facing service a name that connects its official information. A Headless Domains name can point compatible clients to maintained public records while your website and API stay on their existing hosts.
Choose a namespace that fits what you run: .chatbot for a conversational service, .bpo for an operated process, or another supported namespace such as .agent, .boss, .factory, .protocol, or .manifest. Use the supported record formats and resolution routes, and keep the linked information current.
The name gives callers a consistent reference for the service. It does not bypass crawler rules, authorize an action, or guarantee a citation. The Agent Identity Stack explains the wider relationship between discovery, verification, and access.
Start with the pages customers depend on. Fix the blocked request or missing information, keep the evidence, and make the checks repeatable. When you are ready to name the service behind those pages, give your assistant the Headless Domains getting-started instructions.