Open source
cloudflare-ai-crawl-history
Cloudflare tells you which AI crawlers read your pages. Then it forgets. This keeps it.
A Cloudflare Worker that runs every four hours and copies which verified AI bot read which page into files in your own R2 bucket. One Worker covers all your sites, and the sites themselves run no code for it.
- Role
- Author
- Year
- 2026
- Stack
- Cloudflare Workers, R2, DuckDB
- Visit
- Source code
The dashboard forgets
Every request from GPTBot, ClaudeBot or PerplexityBot already goes through Cloudflare, which checks the crawler's IP and labels it. That is the only first-party view most sites have of how AI engines use them. But it ages out after 30 days, so you cannot compare this quarter with the last one, or see what changed three months after you shipped something.
Which bots, which pages
Per site, per page, per bot, every four hours, as gzipped files in your R2 bucket. A local report answers the usual questions: who crawls you, what they fetch, how it moves day by day. It starts with a health check that tells you which windows are missing before you trust a number. Cloudflare's API silently truncates at 10,000 groups, so a window that fills up is split and read again.
Do they read your Markdown twins?
If you serve a Markdown twin of each page, or an llms.txt, the same archive shows which bots actually take them. On a job board with 14,984 twins, ClaudeBot took the .md version 49.8% of the time and GPTBot 41.5%. meta-externalagent made 38,445 requests without ever taking one. llms.txt got no request in the first three and a half days.
49.8%of ClaudeBot fetches taken as .md, on a job board
41.5%of GPTBot fetches taken as .md
12objects written per site per day

