Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

"When we detect unauthorized crawling..."

How did you do that?



Simple, you add the trapped paths to robots.txt Well behaved robots will not crawl them.


Cloudflare's documentation says that Labyrinth is not based on robots.txt.


In line 1 of of the linked page: "waste the resources of AI Crawlers and other bots that don’t respect “no crawl” directives".


Does that indicate the robots.txt is how "no crawl" is indicated? robots.txt doesn't have "no crawl", it has allow and disallow.


And the misbehaved bots follow the path right into the pit and then...the Void of Infinite AI Abyss.


Nifty trick.


Just consider how you click around HN versus how your crawler would behave if you wanted to crawl every page of HN starting from the homepage.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: