Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
d4rkn0d3z
on March 26, 2025
|
parent
|
context
|
favorite
| on:
Trapping misbehaving bots in an AI Labyrinth
"When we detect unauthorized crawling..."
How did you do that?
mog_dev
on March 26, 2025
|
next
[–]
Simple, you add the trapped paths to robots.txt Well behaved robots will not crawl them.
ccgreg
on March 26, 2025
|
parent
|
next
[–]
Cloudflare's documentation says that Labyrinth is not based on robots.txt.
Epskampie
on March 27, 2025
|
root
|
parent
|
next
[–]
In line 1 of of the linked page: "waste the resources of AI Crawlers and other bots that don’t respect “no crawl” directives".
ccgreg
on March 29, 2025
|
root
|
parent
|
next
[–]
Does that indicate the robots.txt is how "no crawl" is indicated? robots.txt doesn't have "no crawl", it has allow and disallow.
CaffeineLD50
on March 26, 2025
|
parent
|
prev
|
next
[–]
And the misbehaved bots follow the path right into the pit and then...the Void of Infinite AI Abyss.
d4rkn0d3z
on March 26, 2025
|
parent
|
prev
|
next
[–]
Nifty trick.
hombre_fatal
on March 26, 2025
|
prev
[–]
Just consider how you click around HN versus how your crawler would behave if you wanted to crawl every page of HN starting from the homepage.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search:
How did you do that?