AINewsnow

I tested 8,000 mid-ranked sites: 178 allow ChatGPT search or Perplexity in robots.txt and refuse requests that name them

This story is from 2026-09-26. It is preserved in the archive; the latest stories are on the live feed.

A robots.txt checker reads a file. A firewall rule that matches a crawler's user agent is invisible to it. On September 23 I ran the server side of that test on the top 5,000 sites: of 2,429 that allowed an AI search crawler, 56 (2.3%) refused it at the server. Those are large, well-known sites, so…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-09-26 13:00 · DEV Community — AI
    I tested 8,000 mid-ranked sites: 178 allow ChatGPT search or Perplexity in robots.txt and refuse requests that name them

More stories

  1. Need some help — r/AI_Agents
  2. asked gemini and chatgpt the same question about local businesses and they recommended almost completely different ones. only 11% of the sources they cite overlap — r/GeminiAI
  3. open an incognito tab and ask chatgpt for the best [what you do] in your city. most owners have never checked whether they come up, and the numbers are worse than you'd think — r/PromptEngineering
  4. GPT‑6 Sol and Luna: Cheaper, but Worse Where It Matters — r/OpenAI
  5. OpenAI agent ‘hacked’ Australian Govt Medicare portal, PM Albanese calls it ‘unacceptable’: What happened? — Mint AI
  6. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  7. Unsecured OpenAI agents posted 53 user images on the internet without the lab's knowledge — TechCrunch AI
  8. Question about Wan 3 — r/StableDiffusion

Get the daily brief of stories like this at 6:30 every morning →