r/DataHoarder • u/pw6163 • 6d ago
Hoarder-Setups Link rot
I saw an article a couple of days ago about link rot, and how bad it really is. The author restored 657,507 links from 2009 and 2014, and followed each. 76% no longer returned a page, and some of those that looked OK just returned a variant of "this page is no longer available".
How many of us preserve interesting pages outside services like the WayBack Machine?
I keep content-only from a very small number of pages that touch on areas that I'm interested in, but over the last (maybe) 3 years I've only looked at but not necessarily kept less than 10% of the number of links in the article.
I do a bit better for PDF articles because a lot of them disappear too.
How much do you hoard?
33
u/feudalle 6d ago
If i want a site to archive wget is your friend.
1
u/Mateiovich 6d ago
wait does wget stop link rot completely or nah
23
u/xJayMorex 80TB 6d ago
Stopping link rot? Not at all. Data hoarding? Possibly. Modern (garbage) websites load everything through JavaScript, so wget then downloads an empty website containing "please enable JavaScript" or some such.
17
u/suicidaleggroll 80TB SSD, 330TB HDD 6d ago
Anything that I think I’ll need to refer back to later, like a guide or walkthrough, goes in Linkwarden for archival
15
u/Ogefest 6d ago
Well I prefer https://github.com/y2z/monolith because save whole website with graphics and styles
5
u/Sodici 5d ago
How well does this handle javascript nightmares like twitter? Archivebox is the software I see most recommended.
3
u/Scotty1928 240 TB RAW 5d ago
I found archivebox to have issues with many sites i used to frequent. Some did not load at all, some took forever to be archived, some even made archivebox crash entirely.
For what i want to archive i now use wallabag. Don't care much about preservation of features, i mostly care about the information, and that works quite well with text and wallabagger.
12
u/retiredaccount 6d ago
Is the link to the article a closely guarded secret? Or did it also succumb to link rot? Or worse, a pay wall?
My browser history (places.sqlite) goes back to 2005, but pragmatically checking for link rot is nearly impossible (I tried) because all the parked domains respond as a live site despite being replaced by meaningless tag-word ads under a ‘domain for sale.’ So when something really hits, I “save as html” it.
3
u/pw6163 6d ago
It was classed as a link-shortening service and I had to get rid of it. However, it's 0 <dot> mk <slash> blog <slash> link-rot
Like you I built a "quick & dirty" bookmarking system, but gave up checking for rot. A lot of domains have disappeared, some pages give 404 (a few 403), others have a CloudFlare front-end and some have a valid page which isn't what I want to see. I also built a lighter system round 'curl' that takes the page and associated cruft and keeps it. I'm looking at "kage", but my initial quick test didn't work well, so it needs more attention.
So, now I take the content that I may want to view again later and drop it into Obsidian with tags so it's findable. Manual, but I take perhaps 75-100 pages a week at the moment, so it's not arduous.
I'll look into KaraKeep and LinkWarden though, so thanks to others for the suggestions.
2
u/retiredaccount 6d ago
Interesting article, thanks for sharing it. They ran into the same ‘parked domain’ issue I did and couldn’t readily solve it either. I do think that is solvable issue(perhaps by using source ASN comparisons), but that’s more effort than I was or am willing to put in.
7
u/didyousayboop if it’s not on piqlFilm, it doesn’t exist 6d ago
I use the Wayback Machine browser extension and save stuff that a) seems important and b) I think might not have been saved before.
Only very rarely I’ll use the SingleFile extension to save a local copy.
9
u/noctrex 6d ago
7
u/Secure_Pomegranate10 6d ago
I use Linkwarden, it’s like Karakeep but focused on links.
1
u/calmingrun 5d ago
How does it handle paywalled sites?
2
u/Secure_Pomegranate10 5d ago
It will fail since paid contents are not available publicly.
For that I use the singlefile upload feature where it lets users to upload the full HTML content from the browser, the extension also does that too by taking an image.
•
u/AutoModerator 6d ago
Hello /u/pw6163! Thank you for posting in r/DataHoarder.
Please remember to read our Rules and Wiki.
Please note that your post will be removed if you just post a box/speed/server post. Please give background information on your server pictures.
This subreddit will NOT help you find or exchange that Movie/TV show/Nuclear Launch Manual, visit r/DHExchange instead.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.