Am I mistaken in thinking that they outsource to other people? (I'm having trouble loading their website due to the load.)
Also, doesn't HN rank upvotes based on karma and account age? You can't downvote until reaching 500(?) and having your account for a certain duration of time.
> Also, doesn't HN rank upvotes based on karma and account age? You can't downvote until reaching 500(?) and having your account for a certain duration of time.
That sounds about right.
There seem to be some additional rules, too. For instance, there are certain comments I don't see a downvote button for, and I have no clue why.
If there is less demand then there are most likely fewer "quality" accounts that have been set up for HN which in turn, makes the account more valuable.
I can create 100 accounts for Reddit and they'll be used 100x/month, or I can create 100 accounts for HN and they'll be used twice/month. Both take the same amount of time to create. Which should be more expensive?
Wouldn't it rather depend on the value to the buyer? I mean as a general rule of operating businesses you try to pay production costs for things you need and price for customer value for things you sell?
Well, both matter - it's a supply/demand curve. The seller here seems to be assuming that the demand is relatively inelastic, and the opportunity cost of the seller's time (which is apparently finite) needs to be factored in.
So in essence what you have is a seller determining which type of the two accounts to make. He will not be able to re-use the HN accounts as much, and therefore he makes less per unit of time from them, and must charge more to make them worth his time.
If you're ever curious how the folks at reddit's r/the_donald maintain control of the sub and why they so often disable moderation, services like this are both how and why. Intermediates who sell upvotes are happy to run jobs for both sides of a given debate.
This site is really expensive compared to some options I've seen sold.
I think my favorite part of all this is the billing screen. "Doesn't have to be your real name."
There are unique groups there and it's hard to tell their origin. Many more seem like they're not exclusive to the_donald. Much of the pro-trump digital infrastructure is quite mercenary in origin.
I don't use Reddit much, and I've never moderated a sub, so can you explain more? Why would the mods need vote bots to control their own sub? Can't they just moderate it?
The bots (if there are any, I don't know) were designed to upvote the posts enough that they'd flood the "all" page, which is an aggregation of many subreddits.
No, reddit made a special rule that their "stickied" posts wouldn't show up on /r/all because the mods would sticky things to "slingshot" them to the front page one at a time. And the users of the sub knew to upvote anything stickied.
They still will show up on /r/all of the post was never stickied.
So are there any reasonable methods to combat this?
Is it a legal possibility (regardless of feasibility) to bring charges against people for manipulation of this kind? (Does HN even have a User Agreement? I can't remember if I had to do anything when I signed up)
Everyone should be vigilant about flagging content that feels forced trying to convert HN users into newsletter subscriptions, SaaS subscriptions, job applicants, routed through affiliate links etc. If you click the domain name beside submissions it shows you domain history, there's a similar page for users' submission history.
Proof of work systems were originally designed to combat these types of 'spammy' behavior.
Give a challenge string and ask the user submitting an upvote to compute a nonce that when added to the string would produce an hash with 10 leading zeros. User submits the calculated nonce, the server computes the hash for challenge string + nonce and the upvote is registered. This makes it computationaly expensive for an upvote provider to submit hundreds of upvotes
I don't think this idea works in practice. Average user has much less computing power available than such services. CPUs are really cheap to rent nowadays, so it would hardly influence the price of an upvote while possibly making quality of experience much worse for mobile users.
Yeah but if you're using a botnet comprised of other people's cpus to submit your upvotes, you already have a distributed computing environment for which you pay zero dollars for computing time. So your cost is still zero.
Upvote manipulation doesn't work on Hacker News because both the voting ring detector is strong, and flagging is significantly more powerful and serves as counteracting force to upvotes and serves as a check to submissions which legit aren't good for the site. (both of these are in contrast with Product Hunt, which turns a blind eye to voting manipulation and asking-friends-for-upvotes-but-not-explicitly is often recommended as a promotion strategy on that site: https://medium.com/startup-grind/my-startup-launch-on-produc...)
There is also a Hacker News "growth hacking" tactic of linking to /newest on Facebook/Twitter instead of linking to the submission directly to avoid the voting ring detector, which also doesn't work and is blatantly obvious: http://i.imgur.com/08pAFOw.jpg
I'm pretty sure that's unequivocally false. It's difficult enough to make it work that it probably wouldn't be worth it, but I'm close to 100% sure I could force something to the top if I wanted to. I don't particularly want to spend the time automating accounts to make them "look natural" and I don't really want to shell out the money to spend on IPs.
But I know for a fact the same thing is happening on Reddit with great success every day. I could spell out how, but I don't think the HN mods would be too keen on that. Any hacker should be able to figure it out, though.
Furthermore, if you produce something solid, all you'd really need is the first <10 upvotes to get on the front page and let it catch on "organically." I can promise you that is happening _all the time_ on HN.
I'm glad you keep your Twitter search running and keep calling people out that blatantly ask for upvotes; the easiest way to kill spam is just to make it difficult, but to say that manipulation doesn't work is just not true.
> I just don't have the time or the money to shell out and buy IPs.
You need a Hacker News account to upvote, and obviously having a large number of new accounts upvote something would trip a voting ring detector.
> Furthermore, if you produce something solid, all you'd really need is the first <10 upvotes to get on the front page and let it catch on "organically."
Much easier said than done. The median score of a HN submission is 1-2 points last time I checked. (the discoverability of HN submissions on /newest is an issue, but that's another topic)
And if it's actually solid, it's not the market for vote brigading.
I agree manipulation does happen, but it doesn't mean it works and instead silently penalizes the submission to death.
> You need a Hacker News account to upvote, and obviously having a large number of new accounts upvote something would trip a voting ring detector.
Correct, but what if I gave myself a couple or few months to prepare accounts? Do a search on blackhatworld (or another spam forum) for "aged accounts" and you'll see that you can generally buy them in tiers - the older they are the more expensive they are.
I'm not saying it's worth it or I think it's something people should do, but saying it's not possible or doesn't work is absurd.
I think it'd be fairly obvious that something suspicious was going on if a large number of dormant "aged accounts" just happened to login for the first time in months to upvote a specific thread.
I'm pretty sure upvotes from different users are unequally weighted when it comes to getting a story onto the front page, and that upvotes from fresh users aren't counted with nearly the same weight as someone who's been here a while and has a track record of predicting which stories will perform well on the front page. There've been times that I've given a story languishing on /newest its 4th upvote and it pops onto the front page, and there've also been times that I've seen stories get 10+ upvotes in short succession and get stuck on page 3. It also wouldn't surprise me if receiving a lot upvotes from newly-created accounts, or accounts that have all upvoted the same things, ends up tripping the voting-ring detector.
If you look at how PG originally created the algorithm in Arc this behavior makes a lot of sense, especially when alongside it is the explanation that HN will cancel out votes that look shady. (I could be wrong, but I believe this is the case).
Basically the algorithm looked like:
Score = (p-1) / (t+2)^g
p = points (the -1 is to negate the point of the original submitter)
t = time
g = gravity (or some gravitational constant)
So the p value matters a lot; especially when there's a smaller t (a new story). This gives newer stories a chance to bubble up with fewer points against the ones that have 1,000 points but have been on the front page for days. It's a pretty ingenious algorithm, actually.
It makes sense, given that formula, that even if all p values are weighted equally, moving p from 3 to 4 would make a big difference, and would be able to push something new to the front page. You'll see them now and then with 2-3 upvotes on the front page, because they're very new and the front page is relatively stale.
Anyway, I'm ~100% sure spammers can get around the vote ring detection, but I don't think that's happening on HN in a major way.
These were cases where the submission was both older and the point value was lower than other submissions that were not on the front page, though.
It's possible that flags, voting ring detector, or spamkiller were negatively impacting the other story, but it definitely isn't a straight votes & time formula anymore.
> I'm pretty sure that's unequivocally false. [...] I could spell out how
What makes you sure? If you know something we don't, we need to hear it—the community relies on us to get this right. It's also personal: I've poured countless hours into combating vote manipulation, and if you know my code isn't working, you bet I want you to spell out how.
The upside of a public discussion outweighs the downside because (1) if you tell us things we don't know, we can start investigating for them—it's what we don't know that's worrisome; and (2) community members might respond with valuable knowledge of their own. That's also why we haven't buried the OP (which, in case users are wondering, we do know about and have lots of data on).
For what it's worth, I don't think HN is in any immediate danger, nor do I know of any active, large-scale vote manipulation happening currently. But to say something like "vote manipulation on HN does not work" is a very bold claim.
I understand that HN's ranking algorithm isn't 100% disclosed because people post startups here and this forum needs to be a fair environment. I say this to make clear that I do get the security ramifications of revealing too much.
But it sounds like you have a bit of understanding when it comes to this topic, which I'm very interested in understanding more about generically speaking. I've always been curious how vote manipulation detection _actually works_, and the kinds of things that can be done to combat vote hacking of various kinds.
I say this as someone vaguely considering the idea of making a small HN-like forum at some point (still figuring out the why and how).
>Upvote manipulation doesn't work on Hacker News because both the voting ring detector is strong, and flagging is significantly more powerful and serves as counteracting force to upvotes and serves as a check to submissions which legit aren't good for the site.
You say this, and it may or may not be true, but of course, Hacker News isn't releasing the code to prove it.
This isn't crypto, and there's no concrete definition of manipulation that would support proof (unlike the kind of threats crypto deals with); it's a fuzzy problem domain.
If that's true, it implies that vote manipulation detection is already easy to get around. It's just relying on security through obscurity and probably direct human intervention to cover for unpublished flaws.
Aren't you holding us to an impossible standard here? The problems we're talking about are not cryptographic; they're much closer to, say, law enforcement. Has anybody figured out how to model such things with mathematical guarantees?
It's just depressing that a site that calls itself Hacker News, and that centers around discussions about software, web applications and security, by so many intelligent and capable people, is less open and less hacker friendly than Wordpress.
And I understand why - the problem is that votes mean something to startups that post here, they have a potential real money value. But it's unfortunate that any insight into how the HN staff deals with the problem of being gamed is denied to the community at large, out of fear of being gamed even more. Plenty of people here would be willing to learn Arc, audit and contribute to the codebase if you let them, and that scrutiny could be beneficial, not just harmful.
The idealism here is really noteworthy, and (as just another commentator/user here) I must say I like it very much.
As you yourself note, however, I can't see any possible way to make what you're describing practically work out: there is absolutely NO way to tell the good guys from the "growth hackers". :(
The other problem is that the bunch of people willing to do this work - which (I say this non-critically, please note!!) is similar to IRL work for a community organization or similar, in that it's not attractive or "widely scoped shiny impact that looks awesome on paper" - is much, much smaller than the bunch of people with evil grins just waiting for more tidbits to use against everyone else.
Sure. Ask in the spirit of good conversation, not cross-examination, e.g. "how did you come to that conclusion" or "why do you think that"—the way people put things in a friendly in-person conversation. Make it clear that you're just curious, and find some way to indicate that you assume good faith on the part of the other.
In face-to-face conversation, most good-faith signals are sent with tone of voice and physical expression. We don't have those channels in text so it's necessary to encode them some other way.
FYI I really appreciate the way you moderate dang, thanks for all the hard work. You're clear, transparent, informative and paitent. It really makes all the difference.
I feel like the example you provided vs. the person's detached comment is splitting hairs over a stylistic issue. In the detached comment, ErikVandeWater is precisely asking for what he wants. In a community heavily weighted with engineers, I think this behavior should be expected, and not penalized.
I understand, but we have a lot of experience with "stylistic issues" triggering destructive flamewars on HN, and there were signs of that in this case.
They run a botnet of users, utilizing browser mods, that upvote a whitelist of users. Sometimes you can see it in normal /r/news submissions, the comments from certain users are voted above 100 immediately as they say crazy racist things and the person they are arguing with has a normal number of upvotes. You can find old lists of the users to upvote and downvote.
I don't know how this website actually works, but if I wanted to make a website like this, here's how I would do it.
The first thing on my mind would be avoiding detection. You can't spin up 5 DO droplets and do your work from there-- every IP you use needs to be a residential IP, and you need thousands of IP addresses. You could go to the trouble of building your own botnet, but that's illegal and very difficult.
Luckily, there's an application called Hola that has convinced 20 million people to willingly join their botnet, and you can buy yourself access to it right here: https://luminati.io .
You'd think that would be the hardest part, but that service means it's the easy part. Now I'd need to go learn PhantomJS and start creating accounts-- being sure to keep each account appearing at a particular IP address. The hardest part is generating some credibility for these accounts. They should always look like they're active-- so I'd be sure to have each account load far more posts than they vote on, and I would upvote plenty of things that I wasn't paid to upvote. Just to generate some randomness.
But that only gets me part of the way there-- these accounts would also need to contribute. My first thought is that you could just scan for reposts in the default subreddits and then repost popular top-level comments from the previous post. I've actually seen this pointed out a few times on reddit. With the amount of comments reddit gets, they probably don't have the resources to detect duplicates.
Hackernews would be significantly harder-- I have no idea how you could generate relevant comments without doing it by hand. That's my guess as to what they're doing.
Luminati is brilliant and evil idea at the same time, also is one of the reasons why I believe that at some point the voting manipulation will go a bit too far and the norm would become to connect all of the accounts to a single real person because otherwise it would be impossible to collect meaningful data about the user behaviour and the quality will nosedive.
If bots become very intelligent and do create quality content, the money would be in manipulating social media not in running a social media service because nobody will be happy showing ads to bots - no matter how insightful comments they post. Reddit/Facebook/Twitter etc. will want their algorithm to cure the content even if the content itself is create by machines.
As fraud becomes ubiquitous, the harder the tech giants will fight.
74 comments
[ 3.6 ms ] story [ 180 ms ] thread"Why so expensive compared to Reddit": http://upvotes.club/buying-hacker-news-growthhackers-upvotes...
HN's also more expensive because there is less demand.
Am I mistaken in thinking that they outsource to other people? (I'm having trouble loading their website due to the load.)
Also, doesn't HN rank upvotes based on karma and account age? You can't downvote until reaching 500(?) and having your account for a certain duration of time.
That sounds about right.
There seem to be some additional rules, too. For instance, there are certain comments I don't see a downvote button for, and I have no clue why.
So in essence what you have is a seller determining which type of the two accounts to make. He will not be able to re-use the HN accounts as much, and therefore he makes less per unit of time from them, and must charge more to make them worth his time.
This site is really expensive compared to some options I've seen sold.
I think my favorite part of all this is the billing screen. "Doesn't have to be your real name."
They still will show up on /r/all of the post was never stickied.
Is it a legal possibility (regardless of feasibility) to bring charges against people for manipulation of this kind? (Does HN even have a User Agreement? I can't remember if I had to do anything when I signed up)
Give a challenge string and ask the user submitting an upvote to compute a nonce that when added to the string would produce an hash with 10 leading zeros. User submits the calculated nonce, the server computes the hash for challenge string + nonce and the upvote is registered. This makes it computationaly expensive for an upvote provider to submit hundreds of upvotes
But that gets user-hostile, so it's not the best idea.
There is also a Hacker News "growth hacking" tactic of linking to /newest on Facebook/Twitter instead of linking to the submission directly to avoid the voting ring detector, which also doesn't work and is blatantly obvious: http://i.imgur.com/08pAFOw.jpg
I'm pretty sure that's unequivocally false. It's difficult enough to make it work that it probably wouldn't be worth it, but I'm close to 100% sure I could force something to the top if I wanted to. I don't particularly want to spend the time automating accounts to make them "look natural" and I don't really want to shell out the money to spend on IPs.
But I know for a fact the same thing is happening on Reddit with great success every day. I could spell out how, but I don't think the HN mods would be too keen on that. Any hacker should be able to figure it out, though.
Furthermore, if you produce something solid, all you'd really need is the first <10 upvotes to get on the front page and let it catch on "organically." I can promise you that is happening _all the time_ on HN.
I'm glad you keep your Twitter search running and keep calling people out that blatantly ask for upvotes; the easiest way to kill spam is just to make it difficult, but to say that manipulation doesn't work is just not true.
You need a Hacker News account to upvote, and obviously having a large number of new accounts upvote something would trip a voting ring detector.
> Furthermore, if you produce something solid, all you'd really need is the first <10 upvotes to get on the front page and let it catch on "organically."
Much easier said than done. The median score of a HN submission is 1-2 points last time I checked. (the discoverability of HN submissions on /newest is an issue, but that's another topic)
And if it's actually solid, it's not the market for vote brigading.
I agree manipulation does happen, but it doesn't mean it works and instead silently penalizes the submission to death.
Correct, but what if I gave myself a couple or few months to prepare accounts? Do a search on blackhatworld (or another spam forum) for "aged accounts" and you'll see that you can generally buy them in tiers - the older they are the more expensive they are.
I'm not saying it's worth it or I think it's something people should do, but saying it's not possible or doesn't work is absurd.
It directly affects the cost of the promotional service
Basically the algorithm looked like:
Score = (p-1) / (t+2)^g
p = points (the -1 is to negate the point of the original submitter)
t = time
g = gravity (or some gravitational constant)
So the p value matters a lot; especially when there's a smaller t (a new story). This gives newer stories a chance to bubble up with fewer points against the ones that have 1,000 points but have been on the front page for days. It's a pretty ingenious algorithm, actually.
It makes sense, given that formula, that even if all p values are weighted equally, moving p from 3 to 4 would make a big difference, and would be able to push something new to the front page. You'll see them now and then with 2-3 upvotes on the front page, because they're very new and the front page is relatively stale.
Anyway, I'm ~100% sure spammers can get around the vote ring detection, but I don't think that's happening on HN in a major way.
It's possible that flags, voting ring detector, or spamkiller were negatively impacting the other story, but it definitely isn't a straight votes & time formula anymore.
What makes you sure? If you know something we don't, we need to hear it—the community relies on us to get this right. It's also personal: I've poured countless hours into combating vote manipulation, and if you know my code isn't working, you bet I want you to spell out how.
The upside of a public discussion outweighs the downside because (1) if you tell us things we don't know, we can start investigating for them—it's what we don't know that's worrisome; and (2) community members might respond with valuable knowledge of their own. That's also why we haven't buried the OP (which, in case users are wondering, we do know about and have lots of data on).
For what it's worth, I don't think HN is in any immediate danger, nor do I know of any active, large-scale vote manipulation happening currently. But to say something like "vote manipulation on HN does not work" is a very bold claim.
But it sounds like you have a bit of understanding when it comes to this topic, which I'm very interested in understanding more about generically speaking. I've always been curious how vote manipulation detection _actually works_, and the kinds of things that can be done to combat vote hacking of various kinds.
I say this as someone vaguely considering the idea of making a small HN-like forum at some point (still figuring out the why and how).
You say this, and it may or may not be true, but of course, Hacker News isn't releasing the code to prove it.
And I understand why - the problem is that votes mean something to startups that post here, they have a potential real money value. But it's unfortunate that any insight into how the HN staff deals with the problem of being gamed is denied to the community at large, out of fear of being gamed even more. Plenty of people here would be willing to learn Arc, audit and contribute to the codebase if you let them, and that scrutiny could be beneficial, not just harmful.
As you yourself note, however, I can't see any possible way to make what you're describing practically work out: there is absolutely NO way to tell the good guys from the "growth hackers". :(
The other problem is that the bunch of people willing to do this work - which (I say this non-critically, please note!!) is similar to IRL work for a community organization or similar, in that it's not attractive or "widely scoped shiny impact that looks awesome on paper" - is much, much smaller than the bunch of people with evil grins just waiting for more tidbits to use against everyone else.
We detached this subthread from https://news.ycombinator.com/item?id=13676504 and marked it off-topic.
I'm legitimately curious if KirinDave has evidence that r/t_d buys upvotes.
In face-to-face conversation, most good-faith signals are sent with tone of voice and physical expression. We don't have those channels in text so it's necessary to encode them some other way.
The first thing on my mind would be avoiding detection. You can't spin up 5 DO droplets and do your work from there-- every IP you use needs to be a residential IP, and you need thousands of IP addresses. You could go to the trouble of building your own botnet, but that's illegal and very difficult.
Luckily, there's an application called Hola that has convinced 20 million people to willingly join their botnet, and you can buy yourself access to it right here: https://luminati.io .
You'd think that would be the hardest part, but that service means it's the easy part. Now I'd need to go learn PhantomJS and start creating accounts-- being sure to keep each account appearing at a particular IP address. The hardest part is generating some credibility for these accounts. They should always look like they're active-- so I'd be sure to have each account load far more posts than they vote on, and I would upvote plenty of things that I wasn't paid to upvote. Just to generate some randomness.
But that only gets me part of the way there-- these accounts would also need to contribute. My first thought is that you could just scan for reposts in the default subreddits and then repost popular top-level comments from the previous post. I've actually seen this pointed out a few times on reddit. With the amount of comments reddit gets, they probably don't have the resources to detect duplicates.
Hackernews would be significantly harder-- I have no idea how you could generate relevant comments without doing it by hand. That's my guess as to what they're doing.
If bots become very intelligent and do create quality content, the money would be in manipulating social media not in running a social media service because nobody will be happy showing ads to bots - no matter how insightful comments they post. Reddit/Facebook/Twitter etc. will want their algorithm to cure the content even if the content itself is create by machines.
As fraud becomes ubiquitous, the harder the tech giants will fight.
That's brilliant.