While we're dealing with the same issue at work, I sometimes still wonder exactly who these scrapers are. OpenAI, Google, Anthropic and others are normally fairly well behaved (Minus Anthropic attempting to hide behind a browser-for-hire company). Mostly you can get IP range and user-agents for the large players, while it problems mostly stem from bots pretending to be Chrome.
Our largest offenders seems to be mostly limited to South-East Asia, so probably mostly Chinese AI projects, but that's speculation. I also don't recall ever seeing Grok IP ranges or a specific Grok UA, but that doesn't mean that they're hiding, perhaps they're just not interested.
Well, some websites claimed China is behind it, which could make sense (I would not know either way). At the same time, though, I kind of doubt your carte blanche here for all those companies. Why would you think none of them are responsible for the AI slop spam?
> Our largest offenders seems to be mostly limited to South-East Asia, so probably mostly Chinese AI projects, but that's speculation
Ok. So you also don't know. Well, I don't know either, but I don't make a speculation by claiming x, y, and z companies to be exempt. In my book they are all responsible.
We see constant abuse from the Tencent ASN/associated ACE ASN, and I've memorized the china169 backbone asn as AS4837 because of thr nonstop crawlers splattered across their network ranges. It's not possible to ID the operator running the crawlers running from these networks but there's a clear signal of the origin of some of these entities.
It is time for micropayments integrated in the browser. Pay 5 cents to access each bug report. Not fun, it shouldn't be like this, but better than not having a bugtracker at all.
Not sure it works, isn't it already like that if you have the cursed javascript-only pages? It's basically proof of work to access their content (and if you are scrapping en masse you likely also need some LLM getting involved)
I suppose the cost of hardware, network usage, and electricity is still under 0.05? I mean, we know the unit economics for LLMs don't make a ton of sense, except if enterprise customers really are what keeps the lights on.
With micropayments, the server owner makes 1 million requests times 0.05 cents = ~$500
The status quo with those js PoW pages doesn't really benefit the server owner at all, it's wasted energy.
I mean proof of work is always wasted energy, but I figure it's better to kill two birds with one stone.
CoinHive was one example of this. (I think this is a correct link? https://github.com/cazala/coin-hive). Although I think ideally you would want to have some sort of browser plugin or app that runs on bare metal instead of a proof-of-work in the browser, because RandomX is designed such that it's slow when implemented in JS (https://github.com/tevador/RandomX/blob/master/doc/design.md)
and do you happen to know a zero friction payment system that works internationally like the Internet itself does? because I promise you that unless you already have a captive audience, even a minuscule amount of friction to access your service will cost you an overwhelming percentage of organic human visitors.
Something like Chaum's blind-signature based ecash, or GNU Taler would be ideal in terms of efficiency, but it's still centralized in distribution. Freenet currently uses something like this.
Another option I was thinking of would be a pretty inflationary (or demurrage) cryptocurrency in which you have some sort of RandomX or other CPU-bound PoW. A web server could act as a mining pool and use mining shares interchangibly with micropayments.
You could do this mining-share method with Monero right now, it's just that you have higher transaction size in Monero and no real analogue to Bitcoin's LN-based microtransactions. Also you would want the cryptocurrency to be more inflationary (or demurrage-based) to promote usage.
Monero's FCMP++ lays some groundwork for payment channels, but it still lacks the nessisary timelocks. Also there was DLSAG which could have enabled payment channels I think, but it's no longer relevant. I also insist that you would need to change the tokenomics to favor greater inflation (maybe you could make coinbase scale linearly with hashrate?), otherwise the miner reward would be economicially insufficient.
Lightning - a fast, instant, bitcoin layer 2 network - is perfectly sufficient for micropayments. Volatility is a no-issue in this case, as you can freely trade the 5 cents in realtime into other assets and minimize holding time of BTC. You will loose the spread, ofc.
Friction-free is a big ask. Brave tried something a while back with their Brave Payments, but of course no one trusts them and it’s not going anywhere.
How many pages does the average software developer visit everyday? 1.000? Price at 0.001 per load and it’ll be completely impractical for crawlers but super cheap for the average connected human on earth.
There's a reddit thread from 9 years ago with people complaining the site blocking crawlers and thus not being indexed by search engines. So I'm not sure what they consider "overload." At this point they should dump the db on thepiratebay.
There are definitely patterns you can use against the scrapers. This maintainer just didn't have time for it, which is understandable.
We direct scraper traffic to a bot-specific server using Cloudflare's load balancer, slowly analyzing traffic and adding conditions one at a time. No accidental scraper DDoS in a long time.
Most scrapers are relatively honest in some way shape or form.
What kind of uptime/expectations/etc. are you working with?
Scraper "attacks" don't take down our robot-specific server very often, so it's safe for us to take heavy-handed approaches that sometimes redirect users there. 99.99% (that's a made-up statistic, but it's a very high number) of the time, the misdirected users don't realize anything is amiss.
Start by analyzing your traffic, specifically user agents. Look for "robot" or even "bot" in the user agent and load balance those to a robot-specific server. This can all be done within Cloudflare. The only code is the user agent condition. FWIW I'm very open to input here if anyone reading notices that we're shooting ourselves in the feet. Based on our analysis, the remaining traffic is a good picture of our intended users.
We have loads of other conditions, mostly banning specific IP ranges for entities when we know exactly who they are, but this is a good start.
What are they scraping the gentoo bugzilla for? I'm confused. Unless you're actively using Gentoo why would this be a resource? Very confusing. Also you'd think we'd have LLM BitTorrent by now, where if they want to scrape something we get a DHT hash for the content and share it with one another, rather than melt servers with the millionth request of the day.
going on a bit of a tangent here - some discussion there seems to be imply that AI bots mostly use IPv4s, which makes sense to me given that they are probably bots hosted by some cloud. Whereas IPv6 may be more organic traffic from e.g. mobile users (not in this case probably)... which made me think if at some point IPv6 may at some point win against IPv4 just because its more organic traffic (i.e. not from big tech cloud), leading to pages blocking IPv4? Just speculation on my side
Corrected title: "Gentoo bugzilla closed because the guy running it, who says it is 'unusable anyway', saw a lot of traffic from different IP addresses with no clear pattern and accuses AI"
The year is 2026, somehow peoples basic web apps are not able to keep up with scrapers. Scrapers are not new. My side load is like .01 even with 10x the traffic of last year. It's called "serving static content", "caching" and many other things that are not new concepts.
Then you've got good old cloudflare which is free to use
While we're dealing with the same issue at work, I sometimes still wonder exactly who these scrapers are. OpenAI, Google, Anthropic and others are normally fairly well behaved (Minus Anthropic attempting to hide behind a browser-for-hire company). Mostly you can get IP range and user-agents for the large players, while it problems mostly stem from bots pretending to be Chrome.
Our largest offenders seems to be mostly limited to South-East Asia, so probably mostly Chinese AI projects, but that's speculation. I also don't recall ever seeing Grok IP ranges or a specific Grok UA, but that doesn't mean that they're hiding, perhaps they're just not interested.
Most of the nonsense comes through botnets/residential proxies, so it's very hard to say.
Well, some websites claimed China is behind it, which could make sense (I would not know either way). At the same time, though, I kind of doubt your carte blanche here for all those companies. Why would you think none of them are responsible for the AI slop spam?
> Our largest offenders seems to be mostly limited to South-East Asia, so probably mostly Chinese AI projects, but that's speculation
Ok. So you also don't know. Well, I don't know either, but I don't make a speculation by claiming x, y, and z companies to be exempt. In my book they are all responsible.
We see constant abuse from the Tencent ASN/associated ACE ASN, and I've memorized the china169 backbone asn as AS4837 because of thr nonstop crawlers splattered across their network ranges. It's not possible to ID the operator running the crawlers running from these networks but there's a clear signal of the origin of some of these entities.
It is time for micropayments integrated in the browser. Pay 5 cents to access each bug report. Not fun, it shouldn't be like this, but better than not having a bugtracker at all.
Perhaps the biggest problem is 5 cents or less isn't enough to cover minimum transaction fees.
0.05 cents even. The beautiful thing is, minuscule amounts of micropayments are already enough to fix the incentives
Not sure it works, isn't it already like that if you have the cursed javascript-only pages? It's basically proof of work to access their content (and if you are scrapping en masse you likely also need some LLM getting involved)
I suppose the cost of hardware, network usage, and electricity is still under 0.05? I mean, we know the unit economics for LLMs don't make a ton of sense, except if enterprise customers really are what keeps the lights on.
Problem is I'd rather have the 0.05cents go to however runs the website (the bugtracker in this case) than go to the electricity company.
With micropayments, the server owner makes 1 million requests times 0.05 cents = ~$500
The status quo with those js PoW pages doesn't really benefit the server owner at all, it's wasted energy.
I mean proof of work is always wasted energy, but I figure it's better to kill two birds with one stone.
CoinHive was one example of this. (I think this is a correct link? https://github.com/cazala/coin-hive). Although I think ideally you would want to have some sort of browser plugin or app that runs on bare metal instead of a proof-of-work in the browser, because RandomX is designed such that it's slow when implemented in JS (https://github.com/tevador/RandomX/blob/master/doc/design.md)
and do you happen to know a zero friction payment system that works internationally like the Internet itself does? because I promise you that unless you already have a captive audience, even a minuscule amount of friction to access your service will cost you an overwhelming percentage of organic human visitors.
Something like Chaum's blind-signature based ecash, or GNU Taler would be ideal in terms of efficiency, but it's still centralized in distribution. Freenet currently uses something like this.
Another option I was thinking of would be a pretty inflationary (or demurrage) cryptocurrency in which you have some sort of RandomX or other CPU-bound PoW. A web server could act as a mining pool and use mining shares interchangibly with micropayments.
You could do this mining-share method with Monero right now, it's just that you have higher transaction size in Monero and no real analogue to Bitcoin's LN-based microtransactions. Also you would want the cryptocurrency to be more inflationary (or demurrage-based) to promote usage.
Monero's FCMP++ lays some groundwork for payment channels, but it still lacks the nessisary timelocks. Also there was DLSAG which could have enabled payment channels I think, but it's no longer relevant. I also insist that you would need to change the tokenomics to favor greater inflation (maybe you could make coinbase scale linearly with hashrate?), otherwise the miner reward would be economicially insufficient.
Lightning - a fast, instant, bitcoin layer 2 network - is perfectly sufficient for micropayments. Volatility is a no-issue in this case, as you can freely trade the 5 cents in realtime into other assets and minimize holding time of BTC. You will loose the spread, ofc.
Friction-free is a big ask. Brave tried something a while back with their Brave Payments, but of course no one trusts them and it’s not going anywhere.
I think it's gonna kill off the internet not just for bots but to whole class of lower income countries, especially for younger learners
How many pages does the average software developer visit everyday? 1.000? Price at 0.001 per load and it’ll be completely impractical for crawlers but super cheap for the average connected human on earth.
That would gate the internet to whole lower income countries
Aren't many of those countries basically being de-facto blocked by Cloudflare filtering out spam already anyways?
Genuine question, I'm not up to date on how Cloudflare operates right now
There's a reddit thread from 9 years ago with people complaining the site blocking crawlers and thus not being indexed by search engines. So I'm not sure what they consider "overload." At this point they should dump the db on thepiratebay.
There are definitely patterns you can use against the scrapers. This maintainer just didn't have time for it, which is understandable.
We direct scraper traffic to a bot-specific server using Cloudflare's load balancer, slowly analyzing traffic and adding conditions one at a time. No accidental scraper DDoS in a long time.
Most scrapers are relatively honest in some way shape or form.
One thing that surpised me about gentoo is just how low budget it is as an operation. They are doing everything with a $12k budget. [1]
[1] https://www.gentoo.org/news/2026/01/05/new-year.html
How did you implement this? My site's getting hammered, any tips would be appreciated
What kind of uptime/expectations/etc. are you working with?
Scraper "attacks" don't take down our robot-specific server very often, so it's safe for us to take heavy-handed approaches that sometimes redirect users there. 99.99% (that's a made-up statistic, but it's a very high number) of the time, the misdirected users don't realize anything is amiss.
Start by analyzing your traffic, specifically user agents. Look for "robot" or even "bot" in the user agent and load balance those to a robot-specific server. This can all be done within Cloudflare. The only code is the user agent condition. FWIW I'm very open to input here if anyone reading notices that we're shooting ourselves in the feet. Based on our analysis, the remaining traffic is a good picture of our intended users.
We have loads of other conditions, mostly banning specific IP ranges for entities when we know exactly who they are, but this is a good start.
> Most scrapers are relatively honest in some way shape or form.
Did you miss a "dis" in there?
And people complain about cloudflare/anubis/etc. Unfortunately, it's looking like this is the alternative.
What are they scraping the gentoo bugzilla for? I'm confused. Unless you're actively using Gentoo why would this be a resource? Very confusing. Also you'd think we'd have LLM BitTorrent by now, where if they want to scrape something we get a DHT hash for the content and share it with one another, rather than melt servers with the millionth request of the day.
They're scraping everything. It doesn't matter what. It doesn't matter if it makes sense. They just scrape it all.
Lots of build failure detailed investigations and gcc/kernel expertise in debugging misbehaving or outright ICEs.
That's it, I guess?
going on a bit of a tangent here - some discussion there seems to be imply that AI bots mostly use IPv4s, which makes sense to me given that they are probably bots hosted by some cloud. Whereas IPv6 may be more organic traffic from e.g. mobile users (not in this case probably)... which made me think if at some point IPv6 may at some point win against IPv4 just because its more organic traffic (i.e. not from big tech cloud), leading to pages blocking IPv4? Just speculation on my side
My residential ISP doesn't support IPv6 at all. If a site is ipv6 only, I can't reach it unless I use a VPN.
The hardest-to-mitigate bot traffic tends to come from residential proxies, and most residential connections are still v4-only.
However, it is also easy to get large numbers of v6 addresses cheaply.
This is weird, every VPS I've rented in the last 8 years had IPv6, and some didn't even had IPv4.
If a problematic bot is on any mainstream VPS it will just get shut down.
How would I shut down a bot on Digital Ocean making over a million requests a day?
Send an email to the Digitalocean abuse email with the time ranges and origin IPs.
So where are bugs reported now?
That's the thing, they aren't.
Corrected title: "Gentoo bugzilla closed because the guy running it, who says it is 'unusable anyway', saw a lot of traffic from different IP addresses with no clear pattern and accuses AI"
Every website that wants to, can just charge for dumps. Anubis or even Cloudflare can handle the rest I think.
Scrapers will ignore the dumps.
You can link to them in 10 million HTTP 429 responses, they will still ignore them.
Those scrapers can be blocked. Some scrapers won't ignore the responses, and maybe it'll lead to a meaningful reduction in scraping traffic.
AI skynet is winning. It is stealing time from humans, thus forcing down their activity, as can be seen here.
A headline would be great if those AI companies would close down. I hold them all responsible for this.
Plot twist: the AI wins not through terminators but a kafka nightmare
The year is 2026, somehow peoples basic web apps are not able to keep up with scrapers. Scrapers are not new. My side load is like .01 even with 10x the traffic of last year. It's called "serving static content", "caching" and many other things that are not new concepts.
Then you've got good old cloudflare which is free to use
A bugtracker is inherently dynamic.
Are you volunteering to setup and host that extra complexity for the Gentoo project?