Where did the old web go? We followed 657,607 links to find out
179 points - yesterday at 5:49 PM
SourceComments
For me, this doesn't indicate that those young folks don't know what they are talking about. Rather, that missing the "old web" is more a cultural thing in itself. We got introduced to the web and discovered these communities we became part of, but something changed and our community died.
There had to be a cutoff that changed something! That killed our community. In reality the change, destruction of communities happened all the time. Some survived, some changed that returning to them wouldn't feel the same, some moved. If you don't feel part of that community anymore, it is destroyed, even if someone else thinks it is still alive and just moved to new-forum/teanspeak3/dig/reddit/discord/tumblr/facebook/tiktok/next thing.
The small Ultima online private server I used to be part of died, because a few people moved on with life, even though it was still there for a long time. Not because of WoW. My WoW guild suffered the same fate, despite Wow still being there. Now the few people I became close friends with are still left, wearing the guild name as a server name on discord, but that's a very different community.
Even if the exact people suddenly came back, it wouldn't be my community anymore.
My old web can't come back, because it is not a matter of technology.
I am sure the people that join the web today will lament the death of the old web of today in a decade.
But maybe that's more a measure of my own age and perceptions rather than an accurate representation of the various eras of the internet/web...
0.mk, you had one jobâŠ
I am not saying old content needs to be preserved forever, but so much content has factually been lost over time. Old logs from text-based MUDs for instance, even for MUDs that still exist today.
Also:
> Reply to any 0.mk email and the message lands in a feedback queue the AI reads, triages, and acts on
Is this dangerous? What about jailbreaking AIs and having it delete everyone's account?
0.mk started in 2009 as a passion project built by three of us. We worked on it for a few hours each week around our regular jobs. We eventually closed it in 2014 because the revenue (hint: no revenue) could not cover hosting, development, and the constant work of fighting spam and reviewing abuse.
The recovered historical corpus contains 657,607 links. For this analysis, we followed every one of them.
Of the 655,178 links with safe, crawlable targets, 76.7% no longer returned a loading page. After removing repeated destinations, 78.7% of the 492,620 distinct crawlable URLs still did not load. So duplicate links are not creating the result.
I use âdid not loadâ rather than âgoneâ deliberately. Some URLs returned 403 or 429 and may have blocked the crawler. Pages that returned 2xx or 3xx count as loading even when they now lead to parked domains, login walls, or removed-content notices.
There is one large distortion in the yearly data. A single account created 83,398 URLs pointing to one hostname in 2011. At URL level, 92.5% of that year did not load. Count each hostname once and the result becomes 61.7%, almost identical to 2010 and 2012.
A few things I did not expect:
- 835 restored links point at Facebookâs old photo CDN. None loaded. - The first link ever shortened was a CSS stylesheet on a WordPress blog. - Someone shortened localhost on the second day. - The longest stored URL is 38,753 characters and repeatedly says TRYING_THE_MAXIMUM_URL.
Most users came from one regional online community, so this is not a census of the whole web. It is a record of what that community shared between 2009 and 2014.
I brought 0.mk back to test whether AI can now handle enough development, spam filtering, abuse review, monitoring, and support to make the service sustainable where the original economics failed.
Happy to answer questions about the crawl, the old data, or the rebuild.
Have an LLM âguessâ random URLs seeded with words from a dictionary, iterating over each word and guessing a URL.
It guesses a lot of correct URLs. This is one method of âURL huntingâ that doesnât involve a 3rd party list or index.
Then just scan those pages for other URLs, visit them, and add a tally every time you come across a URL (for page rank).
Then search anything, see what the results are. You have invented a dark web search engine.
[1] https://wiki.archiveteam.org/index.php/URLTeam#cite_note-1 [2] https://lwn.net/Articles/683880/