Before rewriting a website for AI search, check whether the systems you want to reach can retrieve its useful pages. A crawler blocked by a firewall cannot benefit from a better introduction. Equally, allowing a training crawler does not buy a place in a search answer. Access, indexing, retrieval, and model training are different processes.
This guide provides a practical audit for public websites. Start with a small sample: your homepage, an important service or category page, a product page, and a detailed guide. The objective is to make deliberate access decisions and verify what your infrastructure actually serves.
Separate search access from model training
OpenAI identifies OAI-SearchBot as its search crawler and GPTBot as a crawler for content that may be used in model training. Their robots.txt settings are independent. ChatGPT-User handles certain user-triggered visits; it is not the automatic search crawler, and robots.txt rules may not apply to those visits. Check the current names and published IP ranges in OpenAI's crawler documentation.
For Google, Google-Extended is a robots.txt control for certain Gemini training and grounding uses. It is not a separate HTTP user agent and does not determine inclusion or ranking in Google Search. Google's crawler reference explains the distinction.
Write your own policy in plain language before changing configuration. For example: public buying guides should be available to search crawlers; customer account information should require authentication; training access should follow the publisher's separate decision. That is more precise than a rule saying “block AI” or “allow everything.”
Build a small access matrix
Create one row per sampled URL and one column for each access path you need to test. Record the following observations, along with the time and environment:
- The normal browser response without an existing login or consent cookie.
- The HTTP status and final URL after redirects.
- The applicable robots.txt group and path rule.
- The robots meta tag and any X-Robots-Tag response header.
- The returned title, main heading, and a distinctive sentence from the main content.
- Any firewall challenge, rate limit, login page, or empty application shell.
A 200 status is not enough. A challenge page can also return 200, as can an error template. Searching the response for a sentence you expect from the article helps distinguish access to the content from access to a generic wrapper.
Check the infrastructure as well as robots.txt
Review your CDN and web application firewall logs for blocked requests to the sampled URLs. Group failures by rule and response code. If one security rule challenges every unfamiliar browser, editing robots.txt alone will not resolve the problem.
A request made with a copied crawler user-agent string is useful for reproducing some behavior, but it does not prove that a real crawler can connect. User agents can be spoofed. When a provider publishes verification methods or IP ranges, use them to investigate genuine traffic rather than broadly trusting any request with a familiar name.
Keep protections for administrative and account routes. The fix for a public article being blocked should normally be scoped to legitimate public access. Disabling site-wide security to make one test turn green creates a much larger problem than the original visibility issue.
Distinguish blocked crawling from excluded indexing
A crawl restriction and an indexing instruction solve different problems. Google's noindex documentation explains that its crawler must be allowed to retrieve a page to see its noindex instruction. Do not block a URL in robots.txt and assume Google has read a noindex tag hidden behind that block.
Also inspect canonical targets. In an illustrative audit, a useful guide might load perfectly but point its canonical to a category page because of a template error. That finding belongs in the indexing part of the audit, not in the firewall report. Keep each diagnosis tied to the evidence that supports it.
Make important information available without interaction
Compare the initial HTML with the rendered page. Does the server return the actual product name and specification, or only a loading indicator? Does a delivery explanation appear only after choosing a region? Document these differences so the development team knows which information depends on browser behavior.
For essential facts, a server-rendered summary is a useful implementation choice. It also helps people on slower connections. This is a resilience recommendation, not a claim that every crawler processes JavaScript in the same way.
Close the audit with evidence
- Save the failing response and the rule responsible for it.
- Apply the smallest change that matches the access policy.
- Repeat the same anonymous request and content check.
- Watch genuine crawler logs for the affected route family.
- Review indexing and discovery separately after recrawling.
Use our guide to content for AI search for the next editorial step. Successful access establishes a technical prerequisite. It does not guarantee indexing, a citation, or a particular answer.