683 points by ChrisArchitect 5 days ago | 359 comments | View on ycombinator
simonw 5 days ago |
basilikum 5 days ago |
The Internet Archive is in a really bad spot being attacked from multiple sides at once. But — while service has not been consistent — they have maintained open access. I can still access anonymously from Tor without Cloudflare or some other centralized gatekeeper showing me the middle finger.
If you got some money to spare, consider donating to them. They need it.
robotmay 4 days ago |
Still can't remember what my Tripod site address was, but that might be lost to time.
Thank you, Archive.org.
BeetleB 5 days ago |
I've not been able to access web.archive.org from my work computer - I always get the 429 error.
But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.
timpera 5 days ago |
Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages for a while; and from corporate ISPs (for example, on airport WiFi), you often can't access the WM at all. I hope they'll find a way to relax those.
userbinator 4 days ago |
Thank you for not immediately blaming it on "AI bots". I suspect there's some entity manufacturing consent for strong identity/age verification/sanctioned-browser-OS "walled garden" Internet, and these random DDoSes are part of that.
I knew something was up when a few alternative YouTube front-ends I use suddenly put up the 'nubis and complained about the high volumes of traffic they were getting flooded with; of course someone actually going after that data would be aiming their "AI bots" at YouTube directly instead of trying to suck it through a tiny little-known proxy-site, so it really strained the credibility of the argument.
hubraumhugo 5 days ago |
Some approaches that I think are promising:
- A robots.txt V2[0] as a standard way for website owners to state how bots and AI crawlers can use their online content and where to go (e.g. distinguish search from AI training use cases, point to a downloadable file instead of crawling everything, etc.).
- Something like Web Bot Auth[1] as a non-centralized standard for self-identifying bots and agents cryptographically. This would allow websites to allow or deny bots very precisely.
- what else?
[0] https://datatracker.ietf.org/doc/draft-vaughan-machine-reada...
[1] https://datatracker.ietf.org/doc/html/draft-meunier-http-mes...
CqtGLRGcukpy 5 days ago |
emaro 5 days ago |
I really don't see any solution to this; the scrapers probably wouldn't even mind destroying sources like IA too much, which would leave them as the only "authorative" source of knowledge in the end. Best way is likely regulation incl. hefty (!) fines, but politics are too slow and too fragmented to be effective. So... Enjoy it while it lasts, I guess.
delis-thumbs-7e 4 days ago |
I really so through some money their way, they do wonderful work.
thimabi 5 days ago |
xacky 5 days ago |
pelican0 5 days ago |
Beginning to think that the difficulty to browse most websites nowadays due to throttling, is yet another negative externality of AI development that society is forced to bear.
roughly 4 days ago |
sicktriple 4 days ago |
1vuio0pswjnm7 4 days ago |
Thank you
https://news.ycombinator.com/item?id=49571448
I had a feeling it was due to "AI" companies and developers using "agents"
Not surprised
ilamont 5 days ago |
My blogs are getting slammed and there are issues with cloudflare or captchas.
GaryBluto 4 days ago |
Mr_Minderbinder 3 days ago |
petterroea 4 days ago |
tech234a 5 days ago |
cranberryjoe 4 days ago |
vlyan 5 days ago |
msephton 4 days ago |
All that information is available at the point of failure, the user should not need to email it in.
int32_64 5 days ago |
Roark66 4 days ago |
The amount of data in archive.org is about 100PB. We're talking 10 racks of disks.
I think archive.org should sell "archive as a service" for let's say $15mln. Half of that would be hardware cost and the deliverable could be 12 DC racks containing entire archive.org.
brador 5 days ago |
Cross verify hashes to prevent cheating.
Ez.
tgtweak 4 days ago |
potato-peeler 4 days ago |
ignoramous 5 days ago |
MattCruikshank 5 days ago |
Downloader pays.
I make some content and upload it. When you want to download it, you pay Amazon the egress fees. And maybe I get to charge just a bit more, to help me with the Ingress, storage, content creation, etc.
I mean, I know that there's going to be problems with rate limiting, etc. And yes, we have those problems with LLM tokens today. But this just feels like such a useful thing that it baffles me that it doesn't exist already.
xbar 4 days ago |
UltraSane 5 days ago |
undefined 4 days ago |
Onavo 5 days ago |
It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs.
I am also curious how the (Russian?) paywall bypass mirror archive.is is doing given that they are probably subject to similar amounts of traffic.
lousken 5 days ago |
mrhcon 4 days ago |
halfblood_walks 4 days ago |
josefritzishere 5 days ago |
sehw 4 days ago |
unkeen 5 days ago |
xyst 5 days ago |
swingandamiss 5 days ago |
msephton 5 days ago |
I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.
In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.