I appreciate that decisions have to be made, but it's odd that several of the things I think of as HN classics don't appear to be in this list. Things like "The Story of Mel" ... how can that not be there?
What submissions do you think of as classics, and yet don't appear? How can the selection criteria be improved?
================
Edit (After OrangeMonkey's comment):
Sorry, OrangeMonkey's reply to this comment made me realise that "The Story of Mel" isn't necessarily recognised by everyone, so I did the HN search[0] and here are the top few hits. I haven't made all these into links, because you can do the search for yourself. If you do so you'll see that the exact same URL is submitted many times, but doesn't always get 30 or more votes each time. Possibly different URLs are found and submitting because the previous ones are rejected by the dupe-detector.
Finding the true HN classics is not a simple job, so consider this a useful test case. If your criteria don't find this one, maybe a more complex selection criterion is needed.
The Story of Mel (1983) (https://www.cs.utah.edu/~elb/folklore/mel.html)
445 points|thunderbong|17 days ago|167 comments
The Story Of Mel (http://www.catb.org/jargon/html/story-of-mel.html)
169 points|jacquesm|8 years ago|77 comments
The story of Mel (1983) (http://www.pbm.com//~lindahl/mel.html)
80 points|vaksel|13 years ago|22 comments
The Story of Mel Explained (http://jamesseibel.com/blog/?p=109)
68 points|seibelj|7 years ago|25 comments
The story of Mel, a Real Programmer (http://www.pbm.com/~lindahl/mel.html)
43 points|donw|14 years ago|9 comments
The Story of Mel (1983) (http://www.cs.utah.edu/~elb/folklore/mel.html)
22 points|thefreeman|8 years ago|8 comments
The Story of Mel ... (http://www.jargon.net/jargonfile/t/TheStoryofMel.html)
19 points|urlwolf|12 years ago|17 comments
These are just the top few in the search engine's rankings. If you search for the URLs you'll see that some of them have been submitted many, many times. For example:
This is the tip of a much bigger set of related, often frustrating problems. It's a problem with annotations in a loose sense, for example, which involves anything that relies on URLs as identifiers. The reality is that URLs _are_ identifiers, but it's necessary to recognize that they identify different "printings" (e.g. a work carried by many different "publishers"—The Story of Mel carried by Lindahl is not the same printing as the one carried by Brunvand, for example). There's a lot we're missing out on by abandoning the lessons (and conventions) of legacy print media in the move to the Web. The addition of "accessed on $DATE" in bibliographic citations is an indicator of how we've truly messed things up with this abandonment and by our desire to treat the Web as a sui generis medium unlike anything that came before. Both this and the industry at large has normalized/legitimized (professionalized, even) unhygienic practices in publishing.
What's needed is probably something like the way OpenLibrary functions as a registry for works and particular editions, or an open database like graph.global where you can can mint relations linking individual pieces of content—this thing (identified by URL $X) and this other thing (identified by URL $Y) are the "same" thing.
However, I think we'd most benefit as a first step from a technological and social re-calibration of the way we interact with URLs entirely:
(It's telling that the "past" links at the top of all HN submissions point to a search by title, rather than a search by URL. I've personally hit minor hurdles in my own Algolia searches that turn up 0 results, only to eventually realize that it's because the URL I've entered differs from the one that was submitted in their respective "https"/"http" schemes.)
The convention I outline in that comment is probably the way to go. If I could refer to
<https://apress.com/P. Seibel. Coders at Work: Reflections on...>
then that would help a lot. It wouldn't fix the deduplication/canonicalization problem, exactly, but it would reinforce some good habits and the way we conceptualize the things we're working with, which could lead to this sort of thing becoming more tractable.
I'm not sure how the ordering is done but I feel like the dataset is incomplete. Many "classic" links that get re-posted every few months don't show up and I barely recognize any in the first ~50.
I also think the threshold is too low. I'd go with something more strict like 50 upvotes and reposted at least 10 times.
Meta: this seems to be a classic case of an insufficiently dated web page, making it age not so well. There's no way this was just created, right? It seems so weird to say "until January" in August the year before ...
Most links seem to be at least ~5-7 years old, or more. That also doesn't feel accurate, I mean a post by jwz (that you can't even visit from here (beware nsfw)) at number 1?
I also feel like that SpaceX hyperloop link could use this context:
> As I’ve written in my book, Musk admitted to his biographer Ashlee Vance that Hyperloop was all about trying to get legislators to cancel plans for high-speed rail in California—even though he had no plans to build it.
>
> Several years ago, Musk said that public transit was “a pain in the ass” where you were surrounded by strangers, including possible serial killers, to justify his opposition. But the futures sold to us by Musk and many others in Silicon Valley didn’t just suit their personal preferences. They were designed to meet business needs, and were the cause of just as many problems as they claimed to solve—if not more.
I see that many links were posted to like 5 times, 8 times, 11 times. What are the rules around reposts?
On one hand I see these classics being reposted again and again. But on the other hand we also see some posts marked [dupe]. What are the rules? What is allowed to stay and what is marked as [dupe]?
30 comments
[ 1.9 ms ] story [ 77.5 ms ] threadWhat submissions do you think of as classics, and yet don't appear? How can the selection criteria be improved?
================
Edit (After OrangeMonkey's comment):
Sorry, OrangeMonkey's reply to this comment made me realise that "The Story of Mel" isn't necessarily recognised by everyone, so I did the HN search[0] and here are the top few hits. I haven't made all these into links, because you can do the search for yourself. If you do so you'll see that the exact same URL is submitted many times, but doesn't always get 30 or more votes each time. Possibly different URLs are found and submitting because the previous ones are rejected by the dupe-detector.
Finding the true HN classics is not a simple job, so consider this a useful test case. If your criteria don't find this one, maybe a more complex selection criterion is needed.
[0] https://hn.algolia.com/?q=story+mel* The Story of Mel (1983) - 17 days ago: https://news.ycombinator.com/item?id=32395589
* The Story of Mel Explained - July 20 2015 - https://news.ycombinator.com/item?id=9913835
* 12 times: http://www.catb.org/jargon/html/story-of-mel.html
* 6 times : https://www.cs.utah.edu/~elb/folklore/mel.html
* 4 times : http://www.pbm.com//~lindahl/mel.html
* 3 times : http://www.jargon.net/jargonfile/t/TheStoryofMel.html
It's just that subsequent submissions have not always scored as highly, exactly because people recognise them.
Perhaps "classic" should be:
* Has scored more than 50 points on some submission[0];
* The URL has been submitted more than 10 times[1];
========
[0] For some value of "50"
[1] For some value of "10"
What's needed is probably something like the way OpenLibrary functions as a registry for works and particular editions, or an open database like graph.global where you can can mint relations linking individual pieces of content—this thing (identified by URL $X) and this other thing (identified by URL $Y) are the "same" thing.
However, I think we'd most benefit as a first step from a technological and social re-calibration of the way we interact with URLs entirely:
<https://news.ycombinator.com/item?id=29803419>
(It's telling that the "past" links at the top of all HN submissions point to a search by title, rather than a search by URL. I've personally hit minor hurdles in my own Algolia searches that turn up 0 results, only to eventually realize that it's because the URL I've entered differs from the one that was submitted in their respective "https"/"http" schemes.)
The convention I outline in that comment is probably the way to go. If I could refer to <https://apress.com/P. Seibel. Coders at Work: Reflections on...> then that would help a lot. It wouldn't fix the deduplication/canonicalization problem, exactly, but it would reinforce some good habits and the way we conceptualize the things we're working with, which could lead to this sort of thing becoming more tractable.
I also think the threshold is too low. I'd go with something more strict like 50 upvotes and reposted at least 10 times.
Example: Currently at position 0 in the list: "Watch a VC use my name to sell a con (jwz.org) Posted 2 times, created in 2011". When I check the same link in https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... I get 4 results not 2.
Edit: Maybe can share the exact metrics you used? Did you take an average of the votes?
Large of number repeated submissions would result in a low average score.
Edit 2:
Here you can find some big recurring submissions via "by: dang previous news.ycombinator":
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
More recently "by: dang macroexpanded news.ycombinator":
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
- https://posobin.com/hn_classics (this post) which uses inflation-adjusted scores
- https://jsomers.net/hn with only posts from before 2019
- https://hn.lindylearn.io (built by me) with live-updating data ranked by the number of reposts per URL
Most links seem to be at least ~5-7 years old, or more. That also doesn't feel accurate, I mean a post by jwz (that you can't even visit from here (beware nsfw)) at number 1?
> As I’ve written in my book, Musk admitted to his biographer Ashlee Vance that Hyperloop was all about trying to get legislators to cancel plans for high-speed rail in California—even though he had no plans to build it.
>
> Several years ago, Musk said that public transit was “a pain in the ass” where you were surrounded by strangers, including possible serial killers, to justify his opposition. But the futures sold to us by Musk and many others in Silicon Valley didn’t just suit their personal preferences. They were designed to meet business needs, and were the cause of just as many problems as they claimed to solve—if not more.
https://time.com/6203815/elon-musk-flaws-billionaire-visions...
It seems like the actual link posted is no longer valid; http://www.spacex.com/hyperloop/ redirects me to spacex.com’s homepage.
[As I recall, you’ll likely want to copy and paste that url into your address bar to avoid some referrer detection on his site.]
I learned a new word there: https://www.urbandictionary.com/define.php?term=Ground%20Sco...
On one hand I see these classics being reposted again and again. But on the other hand we also see some posts marked [dupe]. What are the rules? What is allowed to stay and what is marked as [dupe]?
>Are reposts ok?
>
>If a story has not had significant attention in the last year or so, a small number of reposts is ok. Otherwise we bury reposts as duplicates.
>
>Please don't delete and repost the same story. Deletion is for things that shouldn't have been submitted in the first place.
https://www.cbc.ca/news/canada/british-columbia/passengers-v...
Not sure how this one ended up here?