GPTBot never fetched our robots.txt. Another OpenAI bot did, seconds before most of its crawls
Table of contents
In June I wrote that most AI batch crawlers never fetched our robots.txt. That post covered 15 days. I now have 139: from 18 May to 3 October 2026, 32,046 requests to the test bed at next.jsseo.dev.
GPTBot is still at zero. So are Amazonbot and Meta-ExternalAgent.
What the June post missed was happening in the same second. In 135 of GPTBot’s 190 crawl sessions, another OpenAI crawler had fetched /robots.txt less than five seconds earlier.
On this domain, GPTBot does not read robots.txt itself. A separate OpenAI fetcher reads it, and GPTBot starts crawling within seconds.
The setup
One domain, one rulebook. The robots.txt on next.jsseo.dev has not changed since launch: User-agent: *, Allow: /, one Sitemap: line. Nothing here measures whether a bot obeys a rule. It measures who asks for the file, and when.
The tracker reads Cloudflare’s request logs, so it also sees /robots.txt and /sitemap.xml, which never touch the app. Bot claims are checked where the vendor publishes a way: IP ranges for OpenAI, Anthropic and Amazon, reverse DNS for Google, Bing and Apple. Requests that fail the check are left out of every number below unless I say otherwise. Meta publishes neither, so its numbers rest on the user agent plus the network it came from.
139 days in one table
| Bot | Requests | robots.txt fetches | Share | Pattern |
|---|---|---|---|---|
| GPTBot | 805 | 0 | 0% | never, but /sitemap.xml 140 times |
| Amazonbot | 403 | 0 | 0% | never |
| Meta-ExternalAgent | 222 | 0 | 0% | never |
| OAI-SearchBot (verified) | 333 | 333 | 100% | nothing else, ever |
| ClaudeBot | 1,808 | 721 | 40% | every two hours since July |
| Googlebot (verified or pending) | 218 | 100 | 46% | before content |
| Bingbot | 657 | 93 | 14% | from separate machines |
| Applebot | 199 | 27 | 14% | before content, sometimes |
GPTBot’s row is the odd one. A crawler that reads the sitemap once a day for four months, and never the file that declares the sitemap.
What changed over time
robots.txt fetches per month, out of all requests that month, failed verifications excluded. May starts on the 18th; “none” means no verified traffic that month.
| Bot | May | Jun | Jul | Aug | Sep |
|---|---|---|---|---|---|
| OAI-SearchBot | 35 of 35 | 70 of 70 | 74 of 74 | 64 of 64 | 86 of 86 |
| ClaudeBot | 7 of 92 | none | 279 of 701 | 159 of 392 | 260 of 577 |
| Googlebot | 41 of 76 | 20 of 43 | 13 of 31 | 12 of 36 | 14 of 32 |
| Bingbot | 12 of 160 | 28 of 286 | 16 of 53 | 17 of 53 | 18 of 61 |
| Applebot | none | 3 of 6 | 16 of 83 | 6 of 52 | 2 of 2 |
| GPTBot | 0 of 155 | 0 of 132 | 0 of 163 | 0 of 145 | 0 of 155 |
| Amazonbot | 0 of 55 | 0 of 58 | 0 of 29 | 0 of 97 | 0 of 106 |
| Meta-ExternalAgent | 0 of 89 | 0 of 9 | 0 of 3 | 0 of 116 | 0 of 5 |
Two rows moved. ClaudeBot went from 7 robots.txt requests in its first two weeks to the most frequent reader on the site: after six weeks with no verified traffic, it came back on 3 July and settled into robots.txt plus sitemap every two hours, the heartbeat I described in August. Applebot arrived in June.
The zeros did not move. Four and a half months, three crawlers, not one request for the rules.
OpenAI: one bot reads, another crawls
The first OpenAI visit in the logs is at 14:05:30 UTC on 18 May. In that same second OAI-SearchBot requested /robots.txt and GPTBot requested /ssr/clean/product. GPTBot then walked 15 more cells in the next 30 seconds.
That could be a coincidence once. It repeats.
OAI-SearchBot shows up with two user agents. One is OAI-SearchBot/1.0. The other, versions 1.3 and 1.4, has an extra token: compatible; OAI-SearchBot/1.4; robots.txt; +https://openai.com/searchbot. Both pass the IP check against OpenAI’s published list. They never share an address: 23 addresses for the tagged variant, 34 for the plain one.
- The tagged variant made 182 robots.txt requests. 138 of them were followed by GPTBot within 60 seconds.
- The plain variant made 151. None was followed by GPTBot within 60 seconds.
From GPTBot’s side, I split its 753 verified requests into sessions, starting a new session after 30 minutes of silence. That gives 190 sessions. 135 of them began within five seconds of a robots.txt request from the tagged variant. Every one of those 135 followed the tagged variant, none the plain one.
The comparison that matters is chance. OAI-SearchBot asks for robots.txt roughly every ten hours, so any GPTBot request lands near one eventually. I drew 20,000 random moments across the same period and measured the same gap. 0.0% of them fell within five seconds of a robots.txt request. For GPTBot’s session starts the figure is 71%. As a control, Amazonbot’s 293 sessions against the same OpenAI robots.txt requests: 0% within 60 seconds, the same as random.
The sitemap follows the same path. 97 of GPTBot’s 140 /sitemap.xml requests came within five seconds of a tagged robots.txt request. GPTBot could be guessing the standard path, but the timing fits a simpler story: the Sitemap: line sits in a file another OpenAI bot has just read.
The remaining 55 GPTBot sessions had no robots.txt request in the five seconds before them. 40 of those started with /sitemap.xml. My guess is a cached copy of the rules, but the logs cannot show a cache.
What OpenAI says. OpenAI’s crawler documentation says its crawlers share information to avoid duplicate crawling, and that a robots.txt fetch may carry an extra robots.txt marker in the user agent, so logs without full paths can still tell a rules check from a page fetch. The tagged OAI-SearchBot requests fit that description, and the timing above is what a shared robots.txt fetch looks like from the outside.
Bing and Meta split it too, less cleanly
Bingbot fetched content from 169 verified addresses. None of those addresses ever requested robots.txt. All 93 robots.txt requests came from 11 other addresses, all reverse-resolving to msnbot-*.search.msn.com, same user agent. 40 of the 93 were followed by a Bingbot content request from a different address within 60 seconds. Bing appears to read the rules centrally and crawl from elsewhere, with a looser timing link than OpenAI’s: 13% of Bingbot’s content sessions start within a minute of a robots.txt request, against 0.1% for random moments.
Meta’s split runs across two user agents. facebookexternalhit, the link-preview fetcher, requested robots.txt 60 times, all from Meta’s network (AS32934). Not once did it fetch a page of its own within the next minute. Meta-ExternalAgent never requested robots.txt, but on its first visit it arrived eight seconds after a facebookexternalhit robots.txt request, and 8 of those 60 requests were followed by Meta-ExternalAgent within 60 seconds. 13.5% of Meta-ExternalAgent’s sessions start that way, against 0.0% for random moments. Something links them. With no published IP list for Meta, I can only say both sides came from AS32934.
Crawlers that read their own rules
Googlebot, Applebot and ClaudeBot fetch robots.txt themselves, and at least some of the addresses that fetch content read it first.
Googlebot: 8 of the 10 verified addresses that fetched content had requested robots.txt first. The median gap between its robots.txt requests is 15 hours.
Applebot: 18 of 52 content addresses read robots.txt before their first cell.
ClaudeBot: 720 of its 721 robots.txt requests come from Anthropic’s published ranges. 6 of the 13 addresses that fetched content read robots.txt first. 713 of the 721 were followed by more ClaudeBot traffic within a minute, almost always /sitemap.xml, and only 9 by a content page. 79 of its addresses never fetched content at all.
Never
Amazonbot made 403 requests from 254 addresses and never asked for robots.txt. None of those addresses asked under any other user agent either, and Amazon’s other agents, Amzn-SearchBot and Amzn-User, never visited. If Amazon reads the rules, it does so somewhere I cannot see.
The on-demand fetchers also never asked: Google Read Aloud (134 requests), Google-InspectionTool (50), NotebookLM (26), ChatGPT-User (67, of which 9 verified) and Perplexity-User (8). The one exception is Claude-User, whose robots.txt requests in May came from Google Cloud addresses that are not on Anthropic’s list, as covered in the post on Anthropic’s IP list.
What this means for site owners
- Count robots.txt requests per operator, not per user agent. Filtering your logs for
GPTBotandrobots.txtreturns zero rows on this domain, even though OpenAI read the file 182 times on GPTBot’s behalf. - Look for the
robots.txttoken. OpenAI marks its rules fetch in the user agent. A log line withOAI-SearchBot/1.4; robots.txt;is that check. - Think twice before blocking one OpenAI agent by IP. If the tagged OAI-SearchBot fetch is what GPTBot relies on, a firewall rule against OAI-SearchBot’s addresses may also cut GPTBot off from your rules. I have not tested what GPTBot does then.
- Absence in the log still proves nothing. Amazonbot shows no rules check of any kind, and this log cannot tell whether Amazon reads the rules somewhere else.
What this does not prove
- One small domain. A large site with frequent changes may see different fetch patterns, and a different split between agents.
- Timing is not mechanism. 135 sessions within five seconds, and only ever after the tagged variant, is hard to explain otherwise, but I am inferring a shared fetch from the outside.
- Compliance is untested. The rulebook allowed everything for all 139 days. The Disallow test I planned in June is still to come, and it is the only way to see whether GPTBot follows a rule it never fetched.
- Sessions are my construction. A 30-minute silence starts a new session. With a 20-minute cut-off the share is 70%, with 60 minutes 73%. Random moments stay at 0%.
- Meta is unverified. Meta publishes no IP list. The network number is consistent with Meta, nothing more.
Method notes
Window: 2026-05-18 12:37 to 2026-10-03 13:17 UTC. The tracker logs page requests twice, once from Cloudflare’s logs and once from the app’s middleware, and Cloudflare’s feed misses some requests on busy days. I count a request once: every Cloudflare row, plus a middleware row only when Cloudflare has no row with the same address, path and user agent within three seconds. That turns 39,454 rows into 32,046 requests. robots.txt and sitemap requests never reach the middleware, so their counts come from Cloudflare alone. robots.txt means an exact match on /robots.txt; 95 of the 2,096 requests got a 301 redirect, and I count them as requests. Verification: OpenAI and Anthropic by IP range against their published lists, Amazon and Common Crawl by IP range or reverse DNS, Google, Bing and Apple by forward-confirmed reverse DNS. Requests that failed verification are excluded unless stated; that removes, among others, 50 fake OAI-SearchBot requests and 478 fake Googlebot requests. Addresses are compared through a salted hash, never in the clear. The random-moment baseline uses 20,000 uniformly drawn timestamps between the first and last session start of the bot being tested.