Your CDN can overrule robots.txt: finding the layer that refuses AI crawlers
This story is from 2026-10-04. It is preserved in the archive; the latest stories are on the live feed.
In the first post I found Japanese tourism sites whose robots.txt welcomes GPTBot while the server answers it with 403. This one is about the next question: which layer is saying no, did anyone mean it, and what do you change. I ran the same kind of check on a different set of sites first, to see w…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-10-04 13:22 · DEV Community — AI
Your CDN can overrule robots.txt: finding the layer that refuses AI crawlers