Tell HN: After six days of DupDetector, time to take stock.
I come to HN for high-quality links, and even more high-quality discussion, and anything that dilutes that discussion is, to me, a bad thing.
I originally intended to run the DupDetector for a week, but after 5 1/2 days there's enough information to tell me what I want. The thing that has caught me by surprise is the way it has had such widely differeing reactions. Some people have accused me of "a novelty account", something I'd never heard of, but which appears to be associated with Reddit. I thought of it more as a robot assistant.
In particular, I thought more people would be interested in the technology and the hacking. More than anything it's the lack of a response on that level that's made me pause.
And the final factor is today, when DupDetector's karma has fallen from 27 to 12. I don't care about the karma, but it's an indication of people's feeling about the exercise.
I'll run it for just a few hours longer as I tidy things up, but basically I'm stopping, explaining, and I'll see what the response is. I thought it would be, and it would be thought to be, cool, interesting and useful.
Maybe I've misjudged my audience. I am reviewing the situation.
43 comments
[ 2.8 ms ] story [ 152 ms ] threadI think it would be more useful if integrated with HN at the submission phase. Seeing something is a dupe of something else after its already on the frontpage seems too late.
It almost felt petty, that you were scolding the submitter for submitting content that had already been submitted. The comments section of Hacker News is one of the last places on the web that has not been hammered with noise.
I would be somewhat interested in how you solved this problem automatically though.
Disagreeing with something is no reason to down vote a comment. I'm only at ~100 karma, but I'm not waiting to get to 500 to down vote people I disagree with. Disagreement spurs discussion. It'd be pretty boring if everyone agreed on everything.
Since you have to be "qualified" to use the down vote, it should be reserved for instances where the user is not being a respectful fellow hacker.
All that being said, maybe you can include a line saying "Just a bot doing research" before it lists the duplicates.
I've been seeing an increasing number of duplicates myself, including entries using the same URL with an octothorpe added at the end -- a pretty obvious ploy, in my mind.
When HN started up, it was about the conversation. If someone else had already posted a link, great: People just joined in the conversation there if they had something to say. From my perspective, it saved the time of having to post it. And, in that I was often learning as much or more from the conversation on HN than from the links themselves, I was happy to find that conversation focused in one thread.
I'm not sure how, but I'd like to see the site steered back in that direction, if possible. (I have a few half-baked "ideas", but PG and crew have already demonstrated themselves to be more insightful than me -- in my own mind.)
I do think identifying the HN member behind DupDetector might be a benefit, to demonstrate their investment in the community and therefore, in my mind at least, credibility. Yes, I see the email address now, and I half remember off the top of my head whose domain that is. But I might have been a bit more supportive if I knew who was behind it and that they had an established, positive history with HN.
Anyway, just my 2¢; spend them before you need a wheelbarrow full.
The functionality is imho highly desirable, but a more dynamic solution (regularly updating the number of comments, for instance) would be a better implementation.
I hate them. Why? Because the entire point of social bookmarking is to find things you find interesting, not things that you find unique. That little arrow to the left of the title means "I found this link interesting. I think other people will find it interesting as well."
It absolutely does not mean "This link is unique. Nobody has seen it before." If that were the point, we could just pipe RSS feeds into the URL submitter, couldn't we?
The very fact that links are appearing on the front page means that a lot of people haven't seen them yet. It means that they got some utility out of reading them, and it means that they thought others would too.
Sometimes, dupes are good. I don't remember who said it, but that Louis CK interview that gets posted every once in a while, where he is talking about how we're surrounded by wonderful technology and yet nobody cares, they said something to the effect of "I wouldn't care if this was stickied to the top of the page and everybody had to watch it every single day before they post. He is making an excellent point."
Now, I think this is a bit excessive, but the point stands. It isn't about being unique, it is about being good.
Assuming the theory works. People may upvote for a number of reasons, not necessarily because they received value from the link. There remains a lot of work to do in the unexplored space of social, user-powered spaces like Reddit and HN. In theory, the upvote/downvote method maximizes utility of the site users, because those who received value from the article upvoted. Unfortunately, what maximizes utility for one set of users isn't the same for others, which is why we can have such a diversity of sites like HN, Reddit, and Digg.
HN has established itself as a high quality, technically and entrepreneurially oriented site, which is why we don't see the same content as say, reddit.com/r/funny. Defending this position by submitting and upvoting high quality links is critical to maintaining the caliber of HN. This is part of why I don't like seeing resubmitted or duplicate content. Duplicate content (even if very high quality) still takes up a valuable front page location. If enough of the community received new value from the piece (ie. haven't seen it before), then they can upvote it and the article can stay. On the other hand, if you upvote something you've already seen for the benefit of others, you aren't necessarily helping them out. Instead, you are tampering with the ranking algorithm.
If enough new users see an old article and upvote it, then the utility for the aggregate user is maximized, and I am happy with duplicate content when this happens, even if I'm not receiving any value.
Still, there is the problem of users new and old not seeing or knowing about these great quality articles from years ago. Inevitably someone will see an old repost and say "I didn't see this the first time around. I'm glad it was reposted." However, I think using a site like HN to get these old articles new life is using the wrong tool to solve the problem.
"Synonym control is not as wonderful as is often supposed, because synonyms often aren’t. Even closely related terms like movies, films, flicks, and cinema cannot be trivally collapsed into a single word without loss of meaning, and of social context. (You’d rather have a Drain-O® colonic than spend an evening with people who care about cinema.) So the question of controlled vocabularies has a lot to do with the value gained vs. lost in such a collapse. I am predicting that, as with the earlier arc of knowledge management, the question of meaningful markup is going to move away from canonical and a priori to contextual and a posteriori value."
http://many.corante.com/archives/2004/08/25/folksonomy.php
I think this can apply to anything where similarity can seem like a problem.
Just because items seem simliar doesn't take into account the context in which said items are used.
I wonder what people would make of you using your 'real' account with a note that it's a robot posting. On some sites this would be considered karma whoring, but maybe here it would get a better response?
Then, what is it?
I think you have not been clear enough about what you achieved.
I also think you should work on the bot’s politeness. It’s easy to perceive the waltzing in of a bot that says nothing but “This submission has ended up with the points and comments” and provides a link as rude. Technically, sure, that sentence is not or at least not overtly negative but it is easy to perceive it as such.
I would formulate the bot’s phrases consciously positive, showing that your intent is not to be smug about submitters of dupes but that you just want to help, you just want to provide a service.
I used it as a way to find more comments on a topic, and liked it pretty well for that. Of course, I liked it when RiderOfGiraffes was posting them as well; although I wondered how he was able to actually read any of the stories because he seemed to be everywhere with dupe reports.
I had hoped that the DupDetector would give me more time to read stories, but I found two things:
1. There were many, many more dupes being found than I expected. Checking the robot's output before confirming the posting took about the same amount of time as it used to take finding fewer, but doing it by hand. Fully automating it would only take a little more work, and maybe I'd then get the time back.
2. I was reading a bit more, but I found that the additional material wasn't interesting. I was probably already reading everything I found useful, interesting, instructive, or engaging. I'm working on a script to help me with that as well.
The problem I'm finding is the sheer volume, and most of it is repeats, politics, repeats, TSA, repeats, wikileaks, repeats, Assange, repeats, etc. The proportion of material with deep technical content is much less than I remember. Consider - as I write this anything older than 40 minutes old has already fallen off the "newest" page. At that pace it can't all be worth reading.
It's suggested that newcomers read the "news" page, and perhaps the "over" page so they get enculturated with what this site is about. Similarly, it's suggested that older hands inhabit the "newest" page so we can vote up those things that deserve it, and flag the inappropriate.
But I can't keep up with "newest" any more, and much of what I would find interesting is vanishing before I can find it. Searching deeper will find it sometimes, but the 'bots were intended to help.
So I'm working on trying to help, working on trying to find the good stuff (by some definition), and working on adding value.
So, for what it's worth, that's what I think. I hope it sparks off some interesting or useful thoughts.
If you post a link, and it exists, it simply flags: here's all the posts that already have this link. Do you still want to proceed?
Gives you the choice of jumping into the existing discussions or re-posting (eg, if the old subs are years out of date)
But too many URLs are different, sometimes subtly, sometimes not so. I've seen submissions with just an extra hash on the end. There are the submissions with all the feedburner crap cluttering it up, and so on. This was an attempt to be more thorough about detecting duplicates, doing "more properly" a job that's already done, and therefore presumably desirable.
The function is definitely desirable, but it seems better suited as a browser plugin or such. Although it's not literally, botting it and having it contribute almost makes it mandatory for all users, versus being an optional component.
It's the root of the use case here - the functionality is great, but not all users find it desirable.
You can submit links that nobody's looked at recently (modulo the consistent crashiness of Arc). As a hack I think it's clever and terrible all at once.
My claim was that HN should bake in a "dup" button that let users flag duplicate posts, not with the goal of removing dups, but with the goal of cross-referencing them (and possibly sharing karma or some other idea). I won't recap the whole thing here, but I find DupDetector an interesting complement to those ideas.
Dups that are distant in time are more complex. One aspect is that people might not have seen the previous stroy - so it would be good to show it again. Another aspect is that the previous discussion is lost, leading to the same points being repeated, instead of (possibly) being built upon - so it would be better to resurrect the previous submission and therefore discussion.
One solution is to detect and combine dups, but enable them to launch the story fresh, if sufficient time has passed.
There is in fact already a discrete implementation of this: stories over a year old (I think that's the period) can be resubmitted as a new story. So I'm suggesting a continuous version of this idea, where the "newness" of a story gradually increases, til it becomes completely new after a year. "Newness" could be implemented with a factor on the story-score. This would enable old stories (and their discussions) to return to the front-page.
Change the focus from strict dupe-finding to "add additional context to the article people are already looking at." Copy some data about the comments, note the age of the previous discussions, even reproduce the highest-rated comment. These things will give it a more positive/helpful image.
In the sense of an interesting hack, it was fun and pragmatic. Thanks for doing it. But it seems that the community doesn't see the need for the service.
my 0.02 c