Hacker News

Favorites Setup
Comment by delichon | original | Creepy Crawlies
[−]delichon · 2026-08-30 Sun 15:30 UTC · link
> Why is git.kernel.org “interesting” to crawlers

Interesting to crawlers is not a narrow scope. We have the same problem on a B2B car wash site.

[−]iririririr · 2026-08-30 Sun 16:12 UTC · link
so true. the article authors wishing crawlers will use git instead is so funny because the crawlers don't care at all. they are scrapping everything with brute force. they don't care about your content or effective alternatives, and one more site driving their real users crazy with Anubis is nothing more than a new blip in their dashboard. the crawler operators will not even look at the url.
[−]wiredfool · 2026-08-30 Sun 16:20 UTC · link
Seeing the exact same thing on (somewhat high profile) open data sites I run.

The crawlers get stuck in a loop requesting the dataset listing page with every. single. combination. of. facets. At essentially as fast as it can be pumped out or blocked.

Had one bot super interested in a single organization to the tune of 1000 r/min for 24+ hours. Realistic rotating user agents, realistic sec-* headers, no ip address seen more than a couple times in 10 minutes. The only commonality was the route.

[−]marginalia_nu · 2026-08-30 Sun 16:32 UTC · link
Yeah my search engine saw traffic of up to 160 queries per second the other day from some bot that was ostensibly searching for information on Jack Parsons. Just variations on the same query in different permutations of filters and site:-terms.