Why not just offer a paid endpoint for the crawlers? It's not like the demand is going to go away anytime soon.
It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs.
I am also curious how the (Russian?) paywall bypass mirror archive.is is doing given that they are probably subject to similar amounts of traffic.
My guess? Because even with a paid endpoint, the type of unscrupulous yahoo that is DDOSing IA today would probably still abuse the free endpoints because they can. The revenue that might come from a paid endpoint could help to scale up, but with how slow IA usually seems, I suspect there is an upper limit to how much traffic they can serve without a LOT more revenue.
This is a major "this is why we can't have nice things" situation in my opinion. IA is one of the most valuable gems of the Internet. The only thing that even comes close to preserving our shared history. The damage being caused (both by the effective DDOSing and by the knock-on impact that abuse has in encouraging publishers to remove their content from the archive) is incredibly serious.
IA might be large enough to earn consideration, but generally scrapers just don't care about being good citizens. I work in the GLAM space and we offer OAI-PMH interfaces for the harvesting of our collections data - which doesn't stop companies from preferring to scrape our website for worse (less complete, less structured, less standardized) data instead.
Do they know it exists? On my site, some types of blocked bots are getting plain-text instructions saying why I'm blocking them and what they can do instead - and it seems to have worked in some cases.
Micropayments would solve so many Internet problems. It's not too late to adopt.
Content creators could charge by page instead of depending on malware/ad/surveillance revenue. Spam is cut if there's a charge per mail. Scraping abuse goes away, along with a bunch of DDOS garbage.
The impact is a few cents per page or mail, negligible for a human. But if you're consuming a trillion pages per day, you'd reconsider.
Micropayments would solve all the problems except for the problem that people absolutely loathe micropayments. Like, vein-popping furiously hate them.
Whenever the topic of micropayments for internet content comes up, a bunch of people start talking about payment processors and their floor on prices, and so on. That's not wrong, but it can be designed around and I think it's a scapegoat to avoid confronting the fact that users despise micropayments and we'd rather blame credit card companies for the lack of adoption.
Micropayments would solve so many problems for the internet. And, cryptocurrencies would solve so many problems for micropayments. But, it's a non-starter because any proposal gets flooded with people popping veins about how crypto can't solve anything.
If it's so easy to solve, go solve it. Set up a test site with micropayments. Maybe scrape CNN and see how many people will pay you micro for a copy of CNN.
It doesn't need to be crypto, or payment processor based.
My ideal experience would be I load $20 into the browser somewhere like a wallet in one block (that could be a payment processor step). If I visit a participating page, it decrements my wallet $.01 or whatever.
The downside is the possibility of abuse and tracking by governments, which would have to be handled at the source, not the symptom.
> Why not just offer a paid endpoint for the crawlers?
Because then you're definitely violating US copyright law. There are four prongs of fair use analysis, and one of them is the "nature of the use." In this case, you'd be turning into a commercial use.
What if you're not charging for the content, but as compensation for the network bandwidth / server resources consumed by serving that content? The idea isn't to profit from content, but just to allow the IA to continue to serve its purpose as an archive of public content without being overwhelmed by bots.
> It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs
The problem with this perspective is that it ignores the victimization which is happening to all sorts of sites right now.
On one hand, you have content owners/suppliers which are trying to place restrictions on how much free bulk use is allowed.
When scrapers go to exotic lengths to evade the blocks, eg by using thousands of ephemeral IP addresses to collect an entire corpus, saying stuff like that makes it sound like it's all a wash.
"Oh, what a silly situation... How did we ever end up like this? It's not good for anyone ..."
No, there is a victim trying to defend themselves from rampant theft of resources, and a corporate asshole which doesn't care about the effects of their actions.
> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.
I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.
In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've ålso already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.
What sites would they be targeting? Generic "just give me anything"? Whenever I check regular sites on IA, the coverage is spotty -- they'll have the homepage and a few important pages, but it quickly fizzles out.
Very understandable, you can't store all 15000 pages of any random website and update them etc etc, but that makes them pretty useless for indirect scraping because you usually don't want a tiny taste, you want everything.
It is. They will most likely eventually need to move to a walled model for Wayback due to scraper aggressiveness (like Reddit deprecating anonymous old.reddit.com), or behind Cloudflare for aggressive bot and scraping protection.
We used to always "scrape" the wayback machine for any sort of news article we actually paid to consume. I was absolutely shocked by major news sites making very important edits to an article without any sort of editorial notice!
Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021 when we were doing this.
Once upon a time some people explored backing up the Internet Archive.
However, that experiment ended. They mention there were some learnings and they then say:
> The Internet Archive continues to explore methods and code to decentralize the collection, to have a mirror running in various ways - these include IPFS, FileCoin, and others. The INTERNETARCHIVE.BAK project also added general mirroring and tracking code to a number of projects that are still in use.
I would really like to know if any sort of thing like that is still ongoing and if it’s accessible to people in general. Would be nice to mirror some data from IA to my local drives, for example via BitTorrent or IPFS, to have it for offline exploration and personal archive.
The Internet Archive's torrents are a sick joke. I've yet to find one that actually manages to complete. They always get stuck at 90-something percent but that final blocks always fail verification and get retried, fail, and the process repeats forever. Because they're web seeds they're hitting IA infrastructure and not offloading to a real swarm. So their broken torrents are just screwing themselves.
The internet archive's decentralization project is paused as far as I can tell. They have too many things to do and too little funding to do it all. Their current strategy seems to be establishing new legal entities outside the us like in Canada and Switzerland, but they don't accept web traffic even though they hold full copies of the internet archive. There used to be a full copy in Egypt at the library of Alexandria and another in the Netherlands. Not sure if they're still in use, but they did accept web traffic. They hold a decentralized web camp every year in the middle of a forest
> I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.
Appalling, yes. But also expected. I'm surprised they haven't been the target of scrapers for years. But sites putting their content behind login walls and other anti-bot mechanisms has certainly exacerbated this. But again, this isn't at all a surprising progression.
> we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.
To be fair, another big motivation was likely users on sites like HN using archive.org (and similar sites) to get around their paywalls. In fact, I'd be surprised if this wasn't a big motivator.
Again, it sucks, but it's not at all surprising to see it progress like this. I wouldn't be surprised to see similar blocks on other archive sites eventually.
I've personally been using the Wayback Machine more often because I increasingly find myself being blocked from websites who are trying to keep out scrapers even though I'm just a regular person with JS disabled (along with a bunch of other stuff)
Sites are getting too overzealous with blocking IMO. I got blocked for several hours by huggingface simply because my download didn't complete and I had to retry. It gave me error 429, suggested I login, and the login page wouldn't load because error 429.
A popular tech news site blocked my phone because of Apple Private Relay. That didn't last long because their traffic fell off a cliff when that happened.
Many sites are throwing more captchas at the problem, without understanding that captchas don't actually help with LLMs, they just hinder normal users and primitive scripts. LLMs solve captchas just fine.
Some big sites have put up improved paywalls. I'm fine with subscribing to a quality site, however, WSJ and all the other big media sites routinely spit out regurgitated garbage that can be had for free elsewhere (and due to political spin, their garbage is less valuable than the free versions of said content).
Some folks are declaring the internet dead. I wouldn't go that far, however, I will say that a reckoning is going to happen, especially when advertisers figure out that most ads served on basically every website are no longer viewed by humans.
Nah, there's actual value in hitting historical versions and with agents the gap between "how long has this product been offered by this company" and "I should go to wayback machine and do a binary search to find the earliest snapshot that contains this product offering " has closed.
I run some small websites, including a tiny forum that’s been a goldmine for scrapers. I had to significantly tweak some firewall rules and configuration after scrapers behind residential proxies suddenly accounted for over 99% off requests.
However, I also relaxed rules for automated traffic that was well-behaved, and I went out of my way to ensure that the Wayback Machine was able to hit everything. I should kick a small donation their way. They provide an incredibly valuable service and I love the benefit that I get from them just for personal side projects.
How do you separate the Wayback Machine from malicious bots that pretend to be the Wayback Machine? Are you whitelisting their IP blocks?
Because I get a ton of scraper requests that forge Googlebot, Bing, and Yandex user-agents that are totally not coming from their IP ranges. In fact, sometimes they all come from the same IP...
Sorry, not following. I thought the user agent is a string that the caller can set to anything. There is no immediate, reliable way to tell whether the string is correct.
Yeah, a real browser would produce certain patterns and never certain others. So in some cases one could clearly say it's not a human using a browser. But a scraper could also make efforts to mimic human browsing. Mostly the frequency of requests can tell with high probability tell it's not human. But such algorithms might occasionally give false positives for real users, exactly like it has obviously happened for the archive.
What google service are you referring to? Not sure whether the archove uses any of Google's tracking. I have pretty strong blocking of trackers and ads. But the archive works for me.
Google's scraper bot at least used to be behind IPs that you could identify via reverse-then-forward DNS. Not sure if that is still up to date, though.
Yep, verifying the IPs is still the way to go. You often see websites that do it wrong when you set your user agent to Google Bot and they give you a different version of the page without validating that.
Nearly every time a link is posted to HN to a site behind some form of wall, a high voted comment on the post will be a link to an archive site bypassing the owners wall. Bot owners are not the only ones routinely circumventing the choices of content owners.
> We’re getting better at telling abusive bots apart from the people who depend on the Wayback Machine every day. If you think you were blocked in error, email info@archive.org with your operating system, browser, and IP address, and we’ll look into it.
I’ve read a few different experiences with hosts having had success with Anubis to cut down on excessive scraping. One that comes to mind is the Dolphin project.[1]
I’ve read a few of those threads, but often it’s people at cross-purposes with the goals of Anubis. Is there a chance you could clarify the “not working” bit?
Basically, the cost of an optimized solution is orders of magnitude lower than the cost of an in-browser solution. Anyone dedicated (AI scraping companies count) can easily afford to solve workloads higher than your users will tolerate.
You might stop casual scrapers, but you're not going to stop someone who cares.
That's not even the fundamental problem. Even if the payload runs optimally in the browser, the cost of CPU is so small that it's basically irrelevant.
If you waste your user's time with something that would take a full minute to run on a datacenter core, you're costing the scraper something like $0.000005: 360 W TDP on a 128-core EPYC 9754 * $0.10/kWh. In reality, it'll be substantially less than that, because CPUs don't use 0W at idle.
The only way this would make any sense is if there were many more scrapers than users and scrapers cared more about latency than real users, but that's the exact opposite of reality. The entire endeavor is so fundamentally misguided that it almost seems like a psyop.
I've not been able to access web.archive.org from my work computer - I always get the 429 error.
But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.
Your workplace is probably redirecting traffic through a datacenter IP range. Especially if they have their own datacenters like google, microsoft, oracle, amazon, etc.
Try making a vpn via digital ocean for example and you'll see similar patterns.
They seem to be aggressively blocking IPv6 source IPs. I ran into this problem over the past month traveling. I got nothing but 429 errors until I switched on my VPN (which is IPv4 only) and magically the Wayback machine worked again. The lack of transparency on the part of IA is very frustrating.
I really appreciate the Archive team's efforts to make the Wayback Machine more responsive, and have donated a few times to support them.
Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages; and from corporate ISPs (for example, on airport WiFi), you often can't access the WM at all. I hope they'll find a way to relax those.
I have had similar experiences just opening archived pages with a couple of embedded images. It's at a level where just using the site normally is painful.
Generally speaking I feel like detecting the bots might be a lost cause. For someone like the Internet Archive I don't know how to deal with it, for smaller sites, cache everything, static pages whenever possible.
Sadly I see rate-limiting usage in general becoming a thing. With residential proxies and more sophisticated bots either pretending to be Chrome or directly piloting Chrome, it's going to become impossible to tell a real user from a bot. Only solution is to pretend that everyone is a bot and design for it.
They wouldn't be paying for the content, just the bandwidth. Like buying a linux OS on a CD ROM was about the cost of media not profiting off of the software.
There was a feature on Amazon Web Services for a while, and I wish it was still there...
Downloader pays.
I make some content and upload it. When you want to download it, you pay Amazon the egress fees. And maybe I get to charge just a bit more, to help me with the Ingress, storage, content creation, etc.
I mean, I know that there's going to be problems with rate limiting, etc. And yes, we have those problems with LLM tokens today. But this just feels like such a useful thing that it baffles me that it doesn't exist already.
unrelated: if a website gets hit with "This URL has been excluded from the Wayback Machine", do existing snapshots get purged or may they still be preserved somewhere?
Mad props to the people at the Archive. You are the heros we need in a formerly open internet that is surrendering to evil big corps and closing down free access.
The Internet Archive is in a really bad spot being attacked from multiple sides at once. But — while service has not been consistent — they have maintained open access. I can still access anonymously from Tor without Cloudflare or some other centralized gatekeeper showing me the middle finger.
If you got some money to spare, consider donating to them. They need it.
I'm expecting a full court legal attack on Wikipedia at some point. It's just too good for information, and companies would rather you use their chatbot to regurgitate that info now that they have their own copies. Google was already built to pretend as if they had some magic system giving you "Answers" that were 95% just the text of the infobox they used to have for wikipedia on the side.
It's going to suddenly be extremely valuable that wikipedia didn't settle for having a small rainy day fund and instead ceaselessly grabbed every fucking donation they could for two decades so they can fight such a legal battle.
Are you disputing that Gaza had more births than reported deaths during this period, or do you have other examples of “genocides” where the population grew during their genocide?
I don't think there's much grounds for a serious legal attack at this point, and also it would look so horrible from a PR standpoint no non-desperate would dare it. Anyway, google has been steering enough traffic away, now that its AI jumps in with an answer and the wikipedia link which 99% of the response is based on is always hidden with a bunch of other overlaid icons. Google has surely managed to kill off a bunch of reddit and stackoverflow traffic with this trick.
Why would any chatbot provider attack Wikipedia *legally*? Captcha is fully solved, and agents are fully capable of acting as editors, pushing any agenda desired by the user.
Strangers can't really edit Wikipedia any more, especially if their edit is suspicious. It's a closed system despite the appearance. An anti-vandal bot or human will quickly revert your edit.
There is a HN article on abusive AI crawlers on the front page almost every week, but we rarely talk about the path forward. Web scraping has been around for as long as the internet, and it was fine because we had established best practices (rate limiting, self-identification, robots.txt, etc.) that the industry agreed upon. Now we have AI labs and their crawlers that don't care about any of this gentlemen's agreement.
So how do we go from here? Is adding more difficult Anubis and Cloudflare bot protection really the solution? How many millions of human hours and billions in infra costs are we willing to spend on this arms race?
Some approaches that I think are promising:
- A robots.txt V2[0] as a standard way for website owners to state how bots and AI crawlers can use their online content and where to go (e.g. distinguish search from AI training use cases, point to a downloadable file instead of crawling everything, etc.).
- Something like Web Bot Auth[1] as a non-centralized standard for self-identifying bots and agents cryptographically. This would allow websites to allow or deny bots very precisely.
Yeah it’s kinda crazy to me that what was once a back alley python script is now accepted as ‘fine, free for all’. The new era of bros really are smth else.
- Find some new way for Cloudflare to acquire paying customers. If their business did not depend on the status quo, they would be exceptionally well positioned to roll out the technical & organizational frameworks that that make massive botnets a thing of the past.
I've been getting this error a lot. Asking users to email them with details of their OS, browser, IP address is just crazy. Their support is supposedly already swamped and they are asking for more!? Changes made by IA shouldn't become my responsibility.
Imagine if Apple or Microsoft introduced a bug and said, ah yes we know about it we did that on purpose and we know it affects a huge number of people, if each of you could email us these details that'd be great. It's just such an insane request.
IA have broken it and have no real idea how to make it better so they are going to whitelist IPs or browsers or entire operating systems? Wild.
No, it's more like you're requesting something from them and they're telling you they may need some technical, non-personally-identifiable info from you to fulfill your request.
What is 'insane' here is the shear level of entitlement displayed here, including lumping a niche, free, volunteer supported service in with billion dollar, for profit corporations and demanding they pander to your inflated expectations.
It doesn't matter who or what the service is, how much they have, or whatever else. They created a problem and now users have to pay for the inconvenience by emailing(!) specific details that could be captured automatically through web logs: OS, browser, IP address. It's ridiculous.
> details that could be captured automatically through web logs
You can't be serious. Are you ok? The entire point is that they're trying to tell bots and humans apart. They're trusting email (and how you write your email) as a good signal that you're human. What are you talking about getting it from the log? The point is to correlate. How do you expect them to know who you are in the log unless you give them that info?
> They created a problem
No, they're dealing with a problem, and compromised that some human users may unfortunately get blocked.
> and now users have to pay for the inconvenience
You don't have to anything. You can just not use them. They don't owe you their service.
Somebody is handing out free apple lollipops, they ran out, compromised on giving grape ones, and now you're complaining you're being forced to eat a grape one and you don't like grape. Don't eat it.
Who do you think the Internet Archive is? They are not Google, they have very low funding, very high expenses and are constantly under legal pressure.
The fact you can even access the Internet Archive for free is a result of tens thousands of human hours striving for one goal. Digital Preservation. If you rely so much on IA, you should consider donating.
Data analysis would be a good start, better blocking heuristics, an off-the-shelf solution used by other organisations that don't have this problem, etc.
OTOH I do appreciate their openness. It beats those times when I try visiting a site, only to get a cryptic 403 error or similar and no suggestion the site would like to hear from me.
Is it established that the scraping scourge of late is primarily driven by AI companies? Anyone aware of any relevant studies?
Beginning to think that the difficulty to browse most websites nowadays due to throttling is yet another negative externality of AI development that society is forced to bear.
It's shame that the AI arms race causes such collateral damage. Free resources were always exploited, but the stakes ($T) and capabilities around AI allow unprecedented abuse. I wish we could go back... :/
I really don't see any solution to this; the scrapers probably wouldn't even mind destroying sources like IA too much, which would leave them as the only "authorative" source of knowledge in the end. Best way is likely regulation incl. hefty (!) fines, but politics are too slow and too fragmented to be effective. So... Enjoy it while it lasts, I guess.
I wonder why doesn’t the Internet Archive require logging-in prior to accessing the Wayback Machine. It would probably help them distinguish humans from bots, at a very little cost to humans.
> If its the difference between the information being available at all
That's the point. The solution should avoid information not being available. Requiring login will incentivize bots to create spam accounts and move the battle to a new frontier, hurting real people in the process.
Yes. It's impossible to tell which because the split is residential proxies, dataset curators, and AI companies all being separate actors. However I fucking guarantee you it's out there and people are too cowardly to be honest about it so they don't get sued out of existence.
The anti virus industry needs to crack down on crawler and proxy malware, plus ISPs FINALLY need to replace CGNATs with iov6 to stop crawlers banning everyone behind a NAT.
This! But only wayback machine, other services I'm not sure.
But maybe the problem is that they can't serve the data from the other websites like this, if they use it commercially. Right now they have non-commercial use, from what I understand.
> Shouldn't the solution be to gate bulk access for automated services for a price?
The problem is that many of the people who are scraping this data doesn't want to pay. These are organisations who would rather not clone your git repo, and instead scrape every single page on your Forgejo installation. These are NOT nice people.
177 comments
[ 0.18 ms ] story [ 14.4 ms ] threadIt serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs.
I am also curious how the (Russian?) paywall bypass mirror archive.is is doing given that they are probably subject to similar amounts of traffic.
This is a major "this is why we can't have nice things" situation in my opinion. IA is one of the most valuable gems of the Internet. The only thing that even comes close to preserving our shared history. The damage being caused (both by the effective DDOSing and by the knock-on impact that abuse has in encouraging publishers to remove their content from the archive) is incredibly serious.
Content creators could charge by page instead of depending on malware/ad/surveillance revenue. Spam is cut if there's a charge per mail. Scraping abuse goes away, along with a bunch of DDOS garbage.
The impact is a few cents per page or mail, negligible for a human. But if you're consuming a trillion pages per day, you'd reconsider.
Whenever the topic of micropayments for internet content comes up, a bunch of people start talking about payment processors and their floor on prices, and so on. That's not wrong, but it can be designed around and I think it's a scapegoat to avoid confronting the fact that users despise micropayments and we'd rather blame credit card companies for the lack of adoption.
My ideal experience would be I load $20 into the browser somewhere like a wallet in one block (that could be a payment processor step). If I visit a participating page, it decrements my wallet $.01 or whatever.
The downside is the possibility of abuse and tracking by governments, which would have to be handled at the source, not the symptom.
You visit website A,A,A,B,C,D,A,A
At the end of the month, you send your entire 20$ randomly to one of the websites you visited.
This will level out everyone's contribution and reward websites with lots of traffic. It eliminates the need for micropayments.
You can look into international call termination fee fraud in the public telephone network, for more on this.
Because then you're definitely violating US copyright law. There are four prongs of fair use analysis, and one of them is the "nature of the use." In this case, you'd be turning into a commercial use.
The problem with this perspective is that it ignores the victimization which is happening to all sorts of sites right now.
On one hand, you have content owners/suppliers which are trying to place restrictions on how much free bulk use is allowed.
When scrapers go to exotic lengths to evade the blocks, eg by using thousands of ephemeral IP addresses to collect an entire corpus, saying stuff like that makes it sound like it's all a wash.
"Oh, what a silly situation... How did we ever end up like this? It's not good for anyone ..."
No, there is a victim trying to defend themselves from rampant theft of resources, and a corporate asshole which doesn't care about the effects of their actions.
I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.
In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've ålso already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.
Very understandable, you can't store all 15000 pages of any random website and update them etc etc, but that makes them pretty useless for indirect scraping because you usually don't want a tiny taste, you want everything.
https://en.wikipedia.org/wiki/Tragedy_of_the_commons
On the other hand, the Internet Archive is a non-profit offering a free public resource.
> Hacker News and the Rails forum are blocking the text fetcher, so I'm using the browser workflow to inspect the pages directly
Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021 when we were doing this.
However, that experiment ended. They mention there were some learnings and they then say:
> The Internet Archive continues to explore methods and code to decentralize the collection, to have a mirror running in various ways - these include IPFS, FileCoin, and others. The INTERNETARCHIVE.BAK project also added general mirroring and tracking code to a number of projects that are still in use.
https://wiki.archiveteam.org/index.php/INTERNETARCHIVE.BAK
I would really like to know if any sort of thing like that is still ongoing and if it’s accessible to people in general. Would be nice to mirror some data from IA to my local drives, for example via BitTorrent or IPFS, to have it for offline exploration and personal archive.
Appalling, yes. But also expected. I'm surprised they haven't been the target of scrapers for years. But sites putting their content behind login walls and other anti-bot mechanisms has certainly exacerbated this. But again, this isn't at all a surprising progression.
> we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.
To be fair, another big motivation was likely users on sites like HN using archive.org (and similar sites) to get around their paywalls. In fact, I'd be surprised if this wasn't a big motivator.
Again, it sucks, but it's not at all surprising to see it progress like this. I wouldn't be surprised to see similar blocks on other archive sites eventually.
A popular tech news site blocked my phone because of Apple Private Relay. That didn't last long because their traffic fell off a cliff when that happened.
Many sites are throwing more captchas at the problem, without understanding that captchas don't actually help with LLMs, they just hinder normal users and primitive scripts. LLMs solve captchas just fine.
Some big sites have put up improved paywalls. I'm fine with subscribing to a quality site, however, WSJ and all the other big media sites routinely spit out regurgitated garbage that can be had for free elsewhere (and due to political spin, their garbage is less valuable than the free versions of said content).
Some folks are declaring the internet dead. I wouldn't go that far, however, I will say that a reckoning is going to happen, especially when advertisers figure out that most ads served on basically every website are no longer viewed by humans.
Open access doesn't seem sustainable.
But I might just grumpy about spending another hour this week adjusting rules to prevent bots.
However, I also relaxed rules for automated traffic that was well-behaved, and I went out of my way to ensure that the Wayback Machine was able to hit everything. I should kick a small donation their way. They provide an incredibly valuable service and I love the benefit that I get from them just for personal side projects.
Because I get a ton of scraper requests that forge Googlebot, Bing, and Yandex user-agents that are totally not coming from their IP ranges. In fact, sometimes they all come from the same IP...
Yeah, a real browser would produce certain patterns and never certain others. So in some cases one could clearly say it's not a human using a browser. But a scraper could also make efforts to mimic human browsing. Mostly the frequency of requests can tell with high probability tell it's not human. But such algorithms might occasionally give false positives for real users, exactly like it has obviously happened for the archive.
What google service are you referring to? Not sure whether the archove uses any of Google's tracking. I have pretty strong blocking of trackers and ads. But the archive works for me.
https://developers.google.com/search/blog/2006/09/how-to-ver...
https://www.peeringdb.com/asn/7941
I often cannot get past captchas, and archive.org is one of the fallbacks I try.
However, archive.is, etc are more reliable.
I wish the internet archive acted more like a library system, where multiple organizations could mirror the content.
They are a big single point of failure, and I’m shocked Trump/SCOTUS haven’t intentionally burnt the archives down yet.
* BAD: AI companies scraping the web
* BAD: Actors scraping Wayback Machine
* GOOD: xcancel scraping twitter
I’ve read a few of those threads, but often it’s people at cross-purposes with the goals of Anubis. Is there a chance you could clarify the “not working” bit?
[1] https://dolphin-emu.org/blog/2025/06/04/dolphin-progress-rep...
Basically, the cost of an optimized solution is orders of magnitude lower than the cost of an in-browser solution. Anyone dedicated (AI scraping companies count) can easily afford to solve workloads higher than your users will tolerate.
You might stop casual scrapers, but you're not going to stop someone who cares.
If you waste your user's time with something that would take a full minute to run on a datacenter core, you're costing the scraper something like $0.000005: 360 W TDP on a 128-core EPYC 9754 * $0.10/kWh. In reality, it'll be substantially less than that, because CPUs don't use 0W at idle.
The only way this would make any sense is if there were many more scrapers than users and scrapers cared more about latency than real users, but that's the exact opposite of reality. The entire endeavor is so fundamentally misguided that it almost seems like a psyop.
I see anubis, 90% of the time I close the tab before it finishes.
I've not been able to access web.archive.org from my work computer - I always get the 429 error.
But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.
Try making a vpn via digital ocean for example and you'll see similar patterns.
Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages; and from corporate ISPs (for example, on airport WiFi), you often can't access the WM at all. I hope they'll find a way to relax those.
Sadly I see rate-limiting usage in general becoming a thing. With residential proxies and more sophisticated bots either pretending to be Chrome or directly piloting Chrome, it's going to become impossible to tell a real user from a bot. Only solution is to pretend that everyone is a bot and design for it.
Are the abusive bots in the room with us?
Downloader pays.
I make some content and upload it. When you want to download it, you pay Amazon the egress fees. And maybe I get to charge just a bit more, to help me with the Ingress, storage, content creation, etc.
I mean, I know that there's going to be problems with rate limiting, etc. And yes, we have those problems with LLM tokens today. But this just feels like such a useful thing that it baffles me that it doesn't exist already.
Note that Archive Team is separate from the Internet Archive.
The Internet Archive is in a really bad spot being attacked from multiple sides at once. But — while service has not been consistent — they have maintained open access. I can still access anonymously from Tor without Cloudflare or some other centralized gatekeeper showing me the middle finger.
If you got some money to spare, consider donating to them. They need it.
The future is bleak :\
It's going to suddenly be extremely valuable that wikipedia didn't settle for having a small rainy day fund and instead ceaselessly grabbed every fucking donation they could for two decades so they can fight such a legal battle.
It's since been renamed to "Gaza genocide", as it should have been all along - but it took forever to get there.
This is because of their policy of only mirroring what mainstream media outlets are saying.
Some approaches that I think are promising:
- A robots.txt V2[0] as a standard way for website owners to state how bots and AI crawlers can use their online content and where to go (e.g. distinguish search from AI training use cases, point to a downloadable file instead of crawling everything, etc.).
- Something like Web Bot Auth[1] as a non-centralized standard for self-identifying bots and agents cryptographically. This would allow websites to allow or deny bots very precisely.
- what else?
[0] https://datatracker.ietf.org/doc/draft-vaughan-machine-reada...
[1] https://datatracker.ietf.org/doc/html/draft-meunier-http-mes...
It's surely to serve as data to help tell humans apart from bots.
> Changes made by IA shouldn't become my responsibility.
They're a free service. It's not their responsibility to service you either.
IA have broken it and have no real idea how to make it better so they are going to whitelist IPs or browsers or entire operating systems? Wild.
Wild.
You can't be serious. Are you ok? The entire point is that they're trying to tell bots and humans apart. They're trusting email (and how you write your email) as a good signal that you're human. What are you talking about getting it from the log? The point is to correlate. How do you expect them to know who you are in the log unless you give them that info?
> They created a problem
No, they're dealing with a problem, and compromised that some human users may unfortunately get blocked.
> and now users have to pay for the inconvenience
You don't have to anything. You can just not use them. They don't owe you their service.
Somebody is handing out free apple lollipops, they ran out, compromised on giving grape ones, and now you're complaining you're being forced to eat a grape one and you don't like grape. Don't eat it.
The fact you can even access the Internet Archive for free is a result of tens thousands of human hours striving for one goal. Digital Preservation. If you rely so much on IA, you should consider donating.
They are a nonprofit with a mission and continue to solicitate donations based on that mission.
Beginning to think that the difficulty to browse most websites nowadays due to throttling is yet another negative externality of AI development that society is forced to bear.
I really don't see any solution to this; the scrapers probably wouldn't even mind destroying sources like IA too much, which would leave them as the only "authorative" source of knowledge in the end. Best way is likely regulation incl. hefty (!) fines, but politics are too slow and too fragmented to be effective. So... Enjoy it while it lasts, I guess.
That's the point. The solution should avoid information not being available. Requiring login will incentivize bots to create spam accounts and move the battle to a new frontier, hurting real people in the process.
If anything, it will make it harder to block due to the vastness of the IPv6 address space.
Cross verify hashes to prevent cheating.
Ez.
My blogs are getting slammed and there are issues with cloudflare or captchas.
Fine in theory but determined scrapers will use residential proxies in bulk.
But maybe the problem is that they can't serve the data from the other websites like this, if they use it commercially. Right now they have non-commercial use, from what I understand.
The problem is that many of the people who are scraping this data doesn't want to pay. These are organisations who would rather not clone your git repo, and instead scrape every single page on your Forgejo installation. These are NOT nice people.