I found an old database backup of 0.mk on a disk I had kept.
0.mk started in 2009 as a passion project built by three of us. We worked on it for a few hours each week around our regular jobs. We eventually closed it in 2014 because the revenue (hint: no revenue) could not cover hosting, development, and the constant work of fighting spam and reviewing abuse.
The recovered historical corpus contains 657,607 links. For this analysis, we followed every one of them.
Of the 655,178 links with safe, crawlable targets, 76.7% no longer returned a loading page. After removing repeated destinations, 78.7% of the 492,620 distinct crawlable URLs still did not load. So duplicate links are not creating the result.
I use “did not load” rather than “gone” deliberately. Some URLs returned 403 or 429 and may have blocked the crawler. Pages that returned 2xx or 3xx count as loading even when they now lead to parked domains, login walls, or removed-content notices.
There is one large distortion in the yearly data. A single account created 83,398 URLs pointing to one hostname in 2011. At URL level, 92.5% of that year did not load. Count each hostname once and the result becomes 61.7%, almost identical to 2010 and 2012.
A few things I did not expect:
- 835 restored links point at Facebook’s old photo CDN. None loaded.
- The first link ever shortened was a CSS stylesheet on a WordPress blog.
- Someone shortened localhost on the second day.
- The longest stored URL is 38,753 characters and repeatedly says TRYING_THE_MAXIMUM_URL.
Most users came from one regional online community, so this is not a census of the whole web. It is a record of what that community shared between 2009 and 2014.
I brought 0.mk back to test whether AI can now handle enough development, spam filtering, abuse review, monitoring, and support to make the service sustainable where the original economics failed.
Happy to answer questions about the crawl, the old data, or the rebuild.
Intriguingly, the current Republic of North Macedonia (and its neighbors/constituent parts) has undergone a protracted and contentious naming dispute of its own.
Webpages dying is probably one of the biggest design flaws of the original web.
I am not saying old content needs to be preserved forever, but so much content has factually been lost over time. Old logs from text-based MUDs for instance, even for MUDs that still exist today.
> Webpages dying is probably one of the biggest design flaws of the original web.
I'd say it's one of the biggest design flaws of the current web, what with more and more content hidden behind paywalls, increasingly restricted WAFs, and rendered client-side via convoluted JavaScript.
Archiving and mirroring of old-style websites, delivered as static HTML, is simple and straightforward. 20 years from now, most web content from ~1996 to ~2015 will still be accessible, but much of today's web content probably won't.
> Webpages dying is probably one of the biggest design flaws of the original web.
Let me introduce you to the alternatives: print media, film, stone engravings. That stuff tends to get burned and shattered and it takes FOREVER to make copies.
I'm being cheeky but I don't know what design change you could possible make to the web to make webpages not die.
Am I getting old? 09-14 is not even close to the old web for me. The old web, to me, was back when people still published physical 'phone' books for websites.
The old web doesn't necessarily mean the oldest web. 12-17 years ago was very much an older fairly different era of the web that's worth analyzing even if it's on the younger side of the old web. I can definitively sympathize with your reaction though, it doesn't feel like that era was that long ago yet.
I vaguely recall those, but they were more for normies trying to get online. For me the old web is what I saw when I logged into my university gopher server and saw the advertisement for something called the World Wide Web which I could browse via lynx. Soon enough I got a PPP connection and then Mosaic/Netscape 1.0. However everything after javascript shipped (let alone CSS) is new new new. I'd almost go as far as saying if it doesn't have a tilde in the URL it's not old web... almost...
I remember the naive implication/expectation that a URL would be permanent. That once a file is identified by its URL, you could bookmark it and it would always be there. And absolute worst case, if someone really, really, really had to change a file's URL, they would politely return 301 Moved Permanently, and feel very bad about it.
Now people don't give a shit about URLs. Webmasters casually move files around all the time because they feel like it, and if links get broken, who cares, that's the referring site's problem! Their beautiful file hierarchy is more important than the web staying connected!
It seemed reasonable, because keeping things up on the internet was becoming cheaper. We failed to see that capitalism would even demand payment for such trivial things.
The blog mentions "0.mk's revenue did not cover hosting" but then goes on to implement expensive AI integration. Not counting cost for tokens to do the development.
Also:
> Reply to any 0.mk email and the message lands in a feedback queue the AI reads, triages, and acts on
Is this dangerous? What about jailbreaking AIs and having it delete everyone's account?
Old web was kind of dumb anyway. You can put on the rose tinted glasses and feel elite about browsing some shitty site 20 years ago or enjoy the fruits of modern design.
I'd define it as the web until the time that Facebook truly took off and conquered the hearts and minds of so many (that was the first huge shot in the losing war of the old web). For example, a key part of the old web was what used to be called the "blogosphere", and the blog's height was those final years before FB (in the ascendancy) and the first years of the FB era (descendancy).
Here's a contrarian timeline: maybe the old web will return? My reasoning: when the "old web" was great, most people thought the internet was for nerds. Sure even casual users would forward funny emails to their friends, but when it came time for news most people would read the paper or watch the nightly broadcast. When they paid for their Big Mac they'd lay down cash. There were plenty of people using the internet, but again, they were nerds or nerd-adjacent. So, maybe the LLM-ification of everything could end up being like a huge filter, where all the non-nerds no longer see the point of anything and move on to gardens with even higher walls. And then what you have left will be a small subset of people who know how to reach their desired corners of the internet, just like before.
48 comments
[ 1.4 ms ] story [ 22.6 ms ] thread0.mk started in 2009 as a passion project built by three of us. We worked on it for a few hours each week around our regular jobs. We eventually closed it in 2014 because the revenue (hint: no revenue) could not cover hosting, development, and the constant work of fighting spam and reviewing abuse.
The recovered historical corpus contains 657,607 links. For this analysis, we followed every one of them.
Of the 655,178 links with safe, crawlable targets, 76.7% no longer returned a loading page. After removing repeated destinations, 78.7% of the 492,620 distinct crawlable URLs still did not load. So duplicate links are not creating the result.
I use “did not load” rather than “gone” deliberately. Some URLs returned 403 or 429 and may have blocked the crawler. Pages that returned 2xx or 3xx count as loading even when they now lead to parked domains, login walls, or removed-content notices.
There is one large distortion in the yearly data. A single account created 83,398 URLs pointing to one hostname in 2011. At URL level, 92.5% of that year did not load. Count each hostname once and the result becomes 61.7%, almost identical to 2010 and 2012.
A few things I did not expect:
- 835 restored links point at Facebook’s old photo CDN. None loaded. - The first link ever shortened was a CSS stylesheet on a WordPress blog. - Someone shortened localhost on the second day. - The longest stored URL is 38,753 characters and repeatedly says TRYING_THE_MAXIMUM_URL.
Most users came from one regional online community, so this is not a census of the whole web. It is a record of what that community shared between 2009 and 2014.
I brought 0.mk back to test whether AI can now handle enough development, spam filtering, abuse review, monitoring, and support to make the service sustainable where the original economics failed.
Happy to answer questions about the crawl, the old data, or the rebuild.
Intriguingly, the current Republic of North Macedonia (and its neighbors/constituent parts) has undergone a protracted and contentious naming dispute of its own.
https://en.wikipedia.org/wiki/Macedonia_naming_dispute
I am not saying old content needs to be preserved forever, but so much content has factually been lost over time. Old logs from text-based MUDs for instance, even for MUDs that still exist today.
I'd say it's one of the biggest design flaws of the current web, what with more and more content hidden behind paywalls, increasingly restricted WAFs, and rendered client-side via convoluted JavaScript.
Archiving and mirroring of old-style websites, delivered as static HTML, is simple and straightforward. 20 years from now, most web content from ~1996 to ~2015 will still be accessible, but much of today's web content probably won't.
Let me introduce you to the alternatives: print media, film, stone engravings. That stuff tends to get burned and shattered and it takes FOREVER to make copies.
I'm being cheeky but I don't know what design change you could possible make to the web to make webpages not die.
Now people don't give a shit about URLs. Webmasters casually move files around all the time because they feel like it, and if links get broken, who cares, that's the referring site's problem! Their beautiful file hierarchy is more important than the web staying connected!
Also:
> Reply to any 0.mk email and the message lands in a feedback queue the AI reads, triages, and acts on
Is this dangerous? What about jailbreaking AIs and having it delete everyone's account?
0.mk, you had one job…
But maybe that's more a measure of my own age and perceptions rather than an accurate representation of the various eras of the internet/web...
Have an LLM “guess” random URLs seeded with words from a dictionary, iterating over each word and guessing a URL.
It guesses a lot of correct URLs. This is one method of “URL hunting” that doesn’t involve a 3rd party list or index.
Then just scan those pages for other URLs, visit them, and add a tally every time you come across a URL (for page rank).
Then search anything, see what the results are. You have invented a dark web search engine.
[1] https://wiki.archiveteam.org/index.php/URLTeam#cite_note-1 [2] https://lwn.net/Articles/683880/