Where did the old web go? We followed 657,607 links to find out

179 points - yesterday at 5:49 PM

Source

Comments

ketzu today at 8:28 AM
What I find really interesting is that people can heavily disagree what and when the old internet was. It may have ended anywhere from 1989 to 2023, brought by the introduction of the www, mass accessibility, social media, or Ai.

For me, this doesn't indicate that those young folks don't know what they are talking about. Rather, that missing the "old web" is more a cultural thing in itself. We got introduced to the web and discovered these communities we became part of, but something changed and our community died.

There had to be a cutoff that changed something! That killed our community. In reality the change, destruction of communities happened all the time. Some survived, some changed that returning to them wouldn't feel the same, some moved. If you don't feel part of that community anymore, it is destroyed, even if someone else thinks it is still alive and just moved to new-forum/teanspeak3/dig/reddit/discord/tumblr/facebook/tiktok/next thing.

The small Ultima online private server I used to be part of died, because a few people moved on with life, even though it was still there for a long time. Not because of WoW. My WoW guild suffered the same fate, despite Wow still being there. Now the few people I became close friends with are still left, wearing the guild name as a server name on discord, but that's a very different community.

Even if the exact people suddenly came back, it wouldn't be my community anymore.

My old web can't come back, because it is not a matter of technology.

I am sure the people that join the web today will lament the death of the old web of today in a decade.

MetaWhirledPeas yesterday at 10:47 PM
Here's a contrarian timeline: maybe the old web will return? My reasoning: when the "old web" was great, most people thought the internet was for nerds. Sure even casual users would forward funny emails to their friends, but when it came time for news most people would read the paper or watch the nightly broadcast. When they paid for their Big Mac they'd lay down cash. There were plenty of people using the internet, but again, they were nerds or nerd-adjacent. So, maybe the LLM-ification of everything could end up being like a huge filter, where all the non-nerds no longer see the point of anything and move on to gardens with even higher walls. And then what you have left will be a small subset of people who know how to reach their desired corners of the internet, just like before.
morganf yesterday at 9:35 PM
I'd define it as the web until the time that Facebook truly took off and conquered the hearts and minds of so many (that was the first huge shot in the losing war of the old web). For example, a key part of the old web was what used to be called the "blogosphere", and the blog's height was those final years before FB (in the ascendancy) and the first years of the FB era (descendancy).
bradley13 yesterday at 8:02 PM
2009-2014? That's not the old web. Or I'm old. Take your pick.
jperras yesterday at 9:00 PM
I would posit that "old web" could be defined as the period before Google Search became public (<~1997).

But maybe that's more a measure of my own age and perceptions rather than an accurate representation of the various eras of the internet/web...

z_rho_one yesterday at 7:15 PM
Remember the good ol' days when we all thought that everything on the web would exist for eternity and over.
Anon4Now today at 12:15 AM
When I think of "old web", I always think of "Britney Spears' Guide to Semiconductor Physics". Thankfully, it still lives:

https://britneyspears.ac/lasers.htm

mryall yesterday at 8:41 PM
Quite ironic that a link shortener which went offline for a decade or so is now posting about other sites not staying online.

0.mk, you had one job


ErigmolCt today at 5:05 AM
The interesting number here isn't 76.7% dead, it's that 23.3% alive is an upper bound. A server answering HTTP doesn't mean the thing somebody shared in 2011 still exists
6c696e7578 yesterday at 8:29 PM
It seems putting anything on the web that allows submit is screaming to get spammed these days. Is there any sort of spam filter that's worth using?
TrinityFate today at 8:07 AM
The cruelest version of this is the page that survives while everything it links to dies. Looks alive, leads nowhere.
tokai yesterday at 6:40 PM
Am I getting old? 09-14 is not even close to the old web for me. The old web, to me, was back when people still published physical 'phone' books for websites.
WA today at 7:01 AM
Websites always have been shopping windows. They change with the seasons, they come and go. The idea that websites are forever and archived like newspapers might be laudable, but it really applies only to
 online newspapers.
levocardia yesterday at 7:18 PM
100% AI generated text. The irony.
hmhrex yesterday at 7:20 PM
Purevolume mention made me sad. I miss that community.
shevy-java yesterday at 6:29 PM
Webpages dying is probably one of the biggest design flaws of the original web.

I am not saying old content needs to be preserved forever, but so much content has factually been lost over time. Old logs from text-based MUDs for instance, even for MUDs that still exist today.

deleted yesterday at 11:56 PM
deleted yesterday at 6:14 PM
hmartin yesterday at 6:49 PM
Site got hugged? Is there a torrent?
Lord_Zero yesterday at 7:49 PM
The blog mentions "0.mk's revenue did not cover hosting" but then goes on to implement expensive AI integration. Not counting cost for tokens to do the development.

Also:

> Reply to any 0.mk email and the message lands in a feedback queue the AI reads, triages, and acts on

Is this dangerous? What about jailbreaking AIs and having it delete everyone's account?

exitnode yesterday at 6:19 PM
Wow, that is a great domain!
SwellJoe yesterday at 7:33 PM
"The old web" is 1993 to 2007. It's all been downhill ever since.
nipunaeka89 yesterday at 8:10 PM
I was wondering what happened to the old web too
kindawinda today at 12:09 AM
Not enough links
tdx yesterday at 5:50 PM
I found an old database backup of 0.mk on a disk I had kept.

0.mk started in 2009 as a passion project built by three of us. We worked on it for a few hours each week around our regular jobs. We eventually closed it in 2014 because the revenue (hint: no revenue) could not cover hosting, development, and the constant work of fighting spam and reviewing abuse.

The recovered historical corpus contains 657,607 links. For this analysis, we followed every one of them.

Of the 655,178 links with safe, crawlable targets, 76.7% no longer returned a loading page. After removing repeated destinations, 78.7% of the 492,620 distinct crawlable URLs still did not load. So duplicate links are not creating the result.

I use “did not load” rather than “gone” deliberately. Some URLs returned 403 or 429 and may have blocked the crawler. Pages that returned 2xx or 3xx count as loading even when they now lead to parked domains, login walls, or removed-content notices.

There is one large distortion in the yearly data. A single account created 83,398 URLs pointing to one hostname in 2011. At URL level, 92.5% of that year did not load. Count each hostname once and the result becomes 61.7%, almost identical to 2010 and 2012.

A few things I did not expect:

- 835 restored links point at Facebook’s old photo CDN. None loaded. - The first link ever shortened was a CSS stylesheet on a WordPress blog. - Someone shortened localhost on the second day. - The longest stored URL is 38,753 characters and repeatedly says TRYING_THE_MAXIMUM_URL.

Most users came from one regional online community, so this is not a census of the whole web. It is a record of what that community shared between 2009 and 2014.

I brought 0.mk back to test whether AI can now handle enough development, spam filtering, abuse review, monitoring, and support to make the service sustainable where the original economics failed.

Happy to answer questions about the crawl, the old data, or the rebuild.

saadyousfi yesterday at 9:35 PM
[dead]
twotwigs yesterday at 10:38 PM
Fun experiment:

Have an LLM “guess” random URLs seeded with words from a dictionary, iterating over each word and guessing a URL.

It guesses a lot of correct URLs. This is one method of “URL hunting” that doesn’t involve a 3rd party list or index.

Then just scan those pages for other URLs, visit them, and add a tally every time you come across a URL (for page rank).

Then search anything, see what the results are. You have invented a dark web search engine.

bakoserge yesterday at 10:08 PM
[flagged]
colenikol2 yesterday at 8:09 PM
[flagged]
juleiie yesterday at 9:21 PM
Old web was kind of dumb anyway. You can put on the rose tinted glasses and feel elite about browsing some shitty site 20 years ago or enjoy the fruits of modern design.
HDBaseT yesterday at 11:09 PM