54 comments

[ 6.5 ms ] story [ 142 ms ] thread
Replications often don't happen, because it's not considered "new," and who wants to pay for something that's not new? Who can get accolades and titles and journal articles and tenure and HIGHER IMPACT FACTORS out of the drudgery of replication?

There are a lot of science popularizers out there who focus a lot on the power of science, and its ability to self-correct. And given their audience, it's probably a good message to get out there. It's a lot better than most other approaches we've had throughout history. But I think they can give the wrong impression of science (often spoken of as SCIENCE!!!! with exclamation marks) as this process delivered to humanity on stone tablets from on high, and with a little too much trust in the idea that "the arc of science is long but it bends toward correctness."

Science is a frustratingly, painfully, agonizingly human process, made up of people in various organizations, with their own lives and goals and objectives. It's a process where there's been enough of a correlation between what its participants want, and what accuracy/truth/improvement demands, that it's worked pretty well so far, and it gets good results eventually. But there are also perverse incentives in organizations and personal failings in its participants. Things like funders' quest for short-term results, the desire to only back proven winners, the drive to publish or perish, etc etc.

Goodhart's law: "When a measure becomes a target, it ceases to be a good measure."

I guess it's a little like the slow, slow process of natural selection. Given enough time, you can get some pretty impressive, complex stuff out of it, but you can also get stuck in a local maximum for millions of years.

I don't know what the solution is. Perhaps it would be nice for graduation requirements for master's/PhD programs to require attempted replication of some existing works -- but then there's the risk of selecting things that are easy to replicate. Or maybe some mandatory percentage of research grants going toward a common replication pool. Setting up a good incentives structure is hard, because any institution that constrains itself by following new rules "for the good of science" will find itself falling behind in comparison with institutions who only work on the new exciting stuff and get more of the media attention.

I feel fortunate that I'm in computer science, where a lot of our work is inherently easily replicable due to the nature of code (of course, once it gets into user studies there's so much wiggle room that the same problems arise), and I always appreciate when a research group releases testable, compilable, documented code alongside their papers. I'd like to see that sort of practice be more mandatory in CS, but when official prestige comes from journals only, the incentive structures don't reward CS researchers for making maintainable, understandable code -- just code that gets the results and pretty pictures. Academic code is infamous for being poorly-architected.

----

Along some of these lines, it's always a good read to go through Feynman's "Cargo Cult Science": http://neurotheory.columbia.edu/~ken/cargo_cult.html

There is a profound problem with the scientific process that I have been worrying about that rarely gets mentioned:

The appeal and power of science is that you propose a model and then test it with evidence. With strong evidence you can validate the predictive efficacy of your model.

However, the model is not "true", it is only predictive. Furthermore, it's only predictive to the extent that the evidence backs it up.

The problem with this process is that it hangs entirely on the evidence. The models themselves do NOT emerge from the scientific process; they emerge from whatever - prejudice, confabulation, logic, supposition, etc. In other words, the original conjecture of a scientific model is merely a product of the pre-existing beliefs of the investigator.

Coupling poor evidence to this sort of conjecture produces "science" that is simply supposition or bias with a veneer of credibility. Given the widespread abuse of methodology (p-hacking, poor understanding of statistics, small or unrepresentative sample sizes, not reporting negative experiments, bad experimental design) this means a lot of "science" is merely varnished prejudice.

I don't see this as a problem. Perhaps I define some concepts differently.

As I see it, the scientific method is a good way to test a model against reality. It is less suited to be a "generative" process for making models. (There are other ways to test models too: internal consistency, interpretability, and so on.)

Iteration, intuition, creativity, variation, intelligence, luck, guesswork, and more are what create theories.

I'm not saying one could not using more "technical" or "rigorous" ways to construct models -- there are, of course, such methods. I'm just saying that "science" (in the Chalmers sense) is about mostly about testability of a model against evidence.

>Iteration, intuition, creativity, variation, intelligence, luck, guesswork, and more are what create theories.

These are all very positive. How about: racism, misogyny, bias, jealousy, and stubbornness. These, also, can feed into the creation of theories, and do.

And, again, the point is that this process is only as good as the evidence; when the evidence is lacking, we're left with only the investigator's expertise to justify the model. That's not science.

Excellent point. Yes, the sum total of human experience filters into model-making. People often construct theories that correspond with their preconceived biases.

I would like to think that (a) "opening the lid" on how these models work and (b) testing them does a decent job of revealing and removing inaccurate or not-useful aspects. What do you think are some good ways to minimize the negative influences on models?

This feels particularly acute in the 'soft sciences' (e.g. Educational research) where the bar for truth-y results is even that much more diminished. I look forward to the day that educators use a common platform for planning/performance/assessment allowing for large-scale data analysis... But that's not even on the radar for in-person educational environments.
I saw a comic a while back, which I can't now locate, that was "science through the eyes of laymen" vs "science through the eyes of scientists".

The "laymen" side was filled with experiments that worked perfectly; the "scientists" side was filled with things like "this data doesn't make any sense" and "I've run this experiment 5 times, and gotten 5 completely different results".

The punchline is the laymen flipping out about something fairly well established (probably vaccines and autism), and the scientists eventually making something that works.

That's where the difficulty lies -- science is a very messy, very human process that's not nearly so clean or accurate as most people expect. But it also occasionally generates useful results, which are then studied in more detail and lead to things like "planes fly" or "we can detect this type of cancer earlier, and operate before it metastasizes". The results of one study in science are of little value, but the results of the long process of science are often tremendously valuable.

For the first 100,000 to 250,000 years (depending on where you draw the boundary) of the existence of modern humans, we could do no better than astrology, acupuncture, bloodletting, and throwing virgins into volcanoes to make it rain. Everything we know and have today is the result of a methodology for gaining reliable knowledge that is about 300 years old. It works. Some day, it may even embrace psychology.
I dabbled in the research of computer vision, a branch of CS, before. Most of the papers published do not have any codes attached, and given the "black magicness" of many of the training approaches, a lot of the papers are hard to replicate. In fact, I find it dubious that many papers claim significant improvement when the same amount of improvement can be brought about by simply better parameter tuning. Without codes I can never be sure that their methods actually worked.
It's psychology. How is this surprising or even news at all?
Just because something isn't surprising does not mean it isn't newsworthy.
One of the things that bother me about the news, say, npr, is that they often cite studies to bolster their positions, but they never mention the confidence in the studies nor alternate theories which may explain some of the undesired state. No, it's always in terms of things which will whet their middle brow supporters and pander to their pet agendas. Be it water rationing, gentrification, organics, immigration issues, taxation, etc.

I mention npr because I listen to them but I'm assured by others it's the same case with news on the right and abortion, religion, guns and the like.

Is there a way to escape the echo chambers?

Do you have infinite time and omniscence? If not there isn't. If Morning Edition is an echo chamber, all media is.
My feeling I that it's a middle brow feelgood machine. It says the things and supports the things which make middle brow listeners like myself happy. Oh, good, they will take care of my guilt for me simultaneously supporting their/our pet causes, so now I don't have to.

Npr is not self critical beyond the surface. It's a given their causes are the right causes. Let's say the housing crisis in major metros. There is never a comprehensive look. What's good, what's bad. It's gentrification is driving out the poor wheelchair bound immigrant who teaches math at the local school, of the market, therefore we know it's bad by definition. I want to know What is done in china, India, Britain, Germany, Nigeria when they face "gentrification" which they do too. Nope, it's a given their take is the right take. Make me happy, make me feel good, feel guilty for me.

Do you ever listen to On The Media[0]? It's also an NPR show, so perhaps it runs into similar issues, but I think they tend to be fairly skeptical and ask questions about these kind of issues (media bias, etc.).

Warning: it can be extremely depressing to listen to specific examples of how poor a job the media does of actually covering "The News".

[0]: http://www.onthemedia.org/

Thanks for the link, I listened to the Ramos bit. To me, it was noticeable the narrator already took a point of view. It also advocated taking points of view and saw taking a point of view on certain issues necessary (in this case advocating for undocumented immigrants).

Immigrations is a tough issue. There is the human factor. When confronted with the individual cases, it's hard to make a case against open borders. But if we took it to the extreme, we know neither the US nor Europe can accommodate the rest of the world without becoming like the world those people are trying to avoid. (Most of the people are economic immigrants).

In addition, suppose we could sustain a European or North American lifestyle. What would resource demand look like if either sustained 100 million or 500 million immigrants?

Also, what is the responsibility of the bad governments for their poor stewardship of their own people? And what about the questions of a North American or European working freely in anyone of those countries? I think it's a two way street. If we open our borders, so should they. And it should be reciprocal. We open it up to Cuba, Cuba opens it up to us. We open it up to the Netherlands, they open it up to us, etc.

There certainly is: don't pay attention to the "news". Easier said than done I admit.
If you find a way then let me know...

I find the best I can do is listen to/read multiple sources (even ones that I strongly disagree with most of the time) and try to separate bias from facts. This is by no means a guaranteed way to get the real story as I'm sure even while trying my personal biases push me to believe some things as "facts" that others would not. It reminds me of a character in the book "The Foundation" who says he reads various historians accounts of things and from that tries to decide what really happened and when pushed on this and asked why he doesn't visit the places himself to see which is right he is taken aback and responds with something along the lines of "how crude" while he clings to the notion that what he does is the "scientific method".

Great things can be done by building on those who come before you (I do it daily when I write code while I couldn't begin to write ASM nor do I have a good mental model of the lower level stuff up to where I work) but there is a danger in that as well. I worry about the "echo chamber" quite a bit (as I used to hold wildly different views when I was growing up and living in the echo chamber that was my household/parents) and I worry that the same thing could be happening again.

Sometimes I wonder if Aaron Swartz had it right by ignoring the news [0] altogether... Voting is where I disagree with him as while a "voting guide" would be nice, finding a completely unbiased one is probably impossible. The closest I think you can get is a quiz/survey on stances you personally take and then matching you up with politicians that vote the same way (consistently, which is a whole other issue) like ISideWith [1].

[0] http://www.aaronsw.com/weblog/hatethenews

[1] https://www.isidewith.com/

The multiple sources theory seems flawed to me. If you add a contaminant to a mixture, its still contaminated. Averaging out historical accounts does not really tell an accurate description story. Nor does cherry picking facts based on your own personal bias or judgement.
> I find the best I can do is listen to/read multiple sources (even ones that I strongly disagree with most of the time) and try to separate bias from facts.

Journalism is an area where past performance is relevant. Verify some of the reported facts yourself, and you'll have a nicely selected set of sources that you can trust to some degree.

This is why we need effective media watchdogs that record past performance in a consumable fashion. I don't know (m?)any websites that do it well, though.
Why would you listen to something you seem to despise as middlebrow?

I subscribed to and tried to keep up with The Weekly Standard for a year to try to see if there was some "secret sauce" I was missing because of my personal political leanings. It was a maddening experience. As it turns out, spending time hanging out in the other side's echo chamber doesn't make you more well-informed or empathetic, just angrier. Do you listen to NPR for similar reasons?

Oh and I guess by way of offering an alternative for getting out of the echo chamber... check out The Economist. They certainly have their institutional preferences too (e.g. U.K. Conservatives vs Labour, central bank monetary policy) but even when dealing with their pet subjects I find them more levelheaded than most mainstream U.S. outlets.
I picked one up in a coffee shop. Instantly noticed it has zero bylines. Not on any article; not even an editors' list in the front. Is that a Brit thing? Just publish stuff, don't say who said it, and expect folks to believe it?

Since I have a policy of never reading anything without attribution, I could read only the letters to the editor (ok they had bylines on those) so I only know about the article quality indirectly.

It's an Economist thing.

All of their letters to the editor used to be addressed simply to "SIR" as well, a practice they've only done away with just this year. Aside from a couple of pseudonymous regular features in each issue (Bagehot, Schumpeter) there are no bylines.

In 2013 they talked a bit about why they choose to operate this way:

http://www.economist.com/blogs/economist-explains/2013/09/ec...

That's an Economist thing. They feel that "what is written is more important than who writes it"[0].

[0]: http://www.economist.com/blogs/economist-explains/2013/09/ec...

How can they pretend to write with authority on a subject? They don't even publish their editorial board. Its little better than a random guy on the internet quoting made-up facts without SOME idea where they came from. And don't say "The Economist" because they don't say who that is!
Maybe the editorial staff assumes the responsibility for all the correctness issues? That would actually makes sense, since they have the final say in what gets published and should do some additional verification. Just a theory. I don't read The Economist very often and never offline.
Except I don't know who that is - they don't publish them either.
"don't say "The Economist" because they don't say who that is!"

It is The Economist. That's the whole point. There is a (good) reputation attached to it. If they fail to live up to their own standards, then that reputation will crumble and people will stop reading and quoting it. It's a brand.

We used to have brands for products, like RCA, Westinghouse, Philips, etc. But these have become just nametags that are bought and sold, and no longer mean much, or anything. But there are a few brands left that still mean something.

I listen to them because they seem to be the least bad option, but far from good. So bad, but not as bad as others. On the other hand they like to think themselves as source of truth, to some extent. In other words that their politics are good politics.

Another thing about npr is that they handicap issues. So, there might be a small issue but it affects one of their agenda, it gets handicapped and given more exposure than something like that would warrant, under even treatment.

I reached the conclusion several years ago that journalism in all forms is nothing more than a specialized branch of entertainment and should be given about the same amount of weight. I think the damage that the news media does is the result of the general population not understanding this. There is an implication that journalists are somehow experts and intellectually a step above the rest of the population similar to doctors and lawyers. This is a mistake. If there were no distinction made between Fox News, NPR, and the latest Reese Witherspoon film, we'd all be better off.
Sadly, you're right about treating them all as different types of entertainment rather than one of them as a rigorous source of news.
>Is there a way to escape the echo chambers?

So far I haven't found a reliable way out. Social media and personal blogs definitely aren't it. In fact, they're feeding the badness.

Interestingly, I feel that paper editions of newspapers and magazines have less broken articles than their websites.

For me, the scary thing isn't that there are biases. It's that pretty much all major media outlets occasionally participate in Culture Wars™. That is, they take an extreme position on some polarizing issue -- whichever appeals to most of their core audience -- and then use every method available to reinforce their "side" and either ignore or try to discredit the "opposition".

But even that isn't the worst of it. The core problem is that both "sides" of such issues are often objectively wrong. Is it even sensible to try extract The Truth™ from two equally full buckets of bullshit?

I feel like this has become much worse recently. Maybe it's just me changing. Then again, there were no Twitter outrage campaigns before Twitter. There were normal outrage campaigns, but they were mostly wars by proxy fought mostly in the media. People observed them and maybe agreed. Right now, almost everyone is an active participant. Which generates news in its own right and skews coverage even more.

"Is it even sensible to try extract The Truth™ from two equally full buckets of bullshit"

That brought to mind the upcoming election cycle.

After spending time sysadmining in biotech and as a side-effect reading lots of papers in the field, my confidence in the validity of many scientific papers has shot way down. Coming from outside the scientific/academic arena, my view is that the system has become a resume/job padding tool for far too many people, who have simply learned how to A) hit all the main logic points the publishers/journals look for so they get published, and B) find ingenious ways to get their names attached to papers even if they were barely involved in at all.

So tired of hearing X has been cited in over Y Z-journal papers! as a way of touting their scientific merit.

Right. How often do you hear people say "_ studies support my claim, _ studies controvert my claim, and _ studies are agnostic on my claim?" That would be more intellectually honest.
I suppose it's like any other metric, like software developers preening their GitHub activity logs so they always look active and constantly working). With so many PhDs and not enough positions, those in charge of hiring need ways to quantify the relative worth/ROI of applicants. This leads to a focus on things like impact factor, which correlate with but don't equate to what we actually care about. As with everything, you get what you measure.

I'd like to learn more about the history of the Scientific Revolution, the Royal Society and all that. Probably the politics and human-nature is constant (Newton/Leibniz conflict is well-known), but I wonder if there was an impact on the quality of science by the fact that a lot of the research was done by old rich aristocrats whose livelihoods didn't depend on their scientific results -- although I guess their reputations did depend on it. Is that similar to or different from modern academics with tenure? Does the fact that there's a tenure selection process end up selecting good academics, or those who can work the system?

I wish more scientists studied meta-analysis: https://en.wikipedia.org/wiki/Meta-analysis; an area of study specifically tasked with the challenge of synthesizing various study results. Roughly speaking, each study is treated as a (weighted) observation and then combined using statistics.

To put it another way, interpreting empirical results requires (or should require) statistics from bottom to top.

To put it yet another way, overemphasis on statistics is a huge part of the problem.

It is extremely difficult to get statistics right. Am I applying a relevant test? Even if it is relevant, is it too sensitive? Is my dataset unbiased or have I controlled for all of the significant biases?

Multiple test corrections are also some of the most useless statistical tests you can perform. They're basically only useful for reducing confidence to reasonable levels. Hundreds of bad tests cannot yield a useful scientific inference through sheer volume.

Meta-analysis is interesting, but it's ultimately a shoddy attempt to patch over the much larger problem of widespread methodological failure. We can't do bad science for decades and then expect meta-analysis to yield new insights; at best it can serve to expose the poor quality of that work.

> Meta-analysis is interesting, but it's ultimately a shoddy attempt to patch over the much larger problem of widespread methodological failure.

I think I know what you are getting at. Let me try to unpack what I mean, and I would like to see if you would agree...

I agree that "widespread methodological failure" is a problem with science in the real world.

However, by your writing, it would seem that you attack meta-analysis as a technique.

1. Meta-analysis (MA) is a solid technique and not shoddy. I'm not saying it works equally well in all situations, but I am saying that it works better than not using it. :)

2. MA cannot "correct" one underlying study that is flawed. (This is obvious; doing statistics with a sample size of one is silly.)

3. Nor can MA correct for a systemic problem in underlying studies.

4. However, in the same way that statistics can be resilient over random errors in measurement, MA can handle studies of varying quality, provided that they do not have significant errors pointing in the same directions. (This is the often claimed, but less-often tested, assumption that errors are randomly distributed.)

The Reddit thread on this article is pretty good: https://www.reddit.com/r/science/comments/3imphg/scientists_...

There's a comment from one of the paper's co-authors that will probably need repeating here (https://www.reddit.com/r/science/comments/3imphg/scientists_...): "To those wanting to dismiss psychological science as "cult science" based on these findings, note how ironic your response is. You're discrediting the very people whose data you are using to back up your claim. This massive, groundbreaking project was conducted on psychological science by psychological scientists. In my view, psychological scientists are among the most dedicated and rigorous scientists there are. No other field has had the courage to instantiate a project like this. And I am sure that many of you would be shocked to find out how low the reproducibility rates are in other fields. Problems of non-reproducibility, publication bias, data faking, lack of transparency, and the like plague every scientific field. The people you are labeling as "cult" scientists are leading the movement to improve science of all types in a much needed way."

Also, the current top comment quotes the following from the paper: "Any temptation to interpret these results as a defeat for psychology, or science more generally, must contend with the fact that this project demonstrates science behaving as it should. Hypotheses abound that the present culture in science may be negatively affecting the reproducibility of findings. An ideological response would discount the arguments, discredit the sources, and proceed merrily along. The scientific process is not ideological. Science does not always provide comfort for what we wish to be; it confronts us with what is"

And another co-author chimed in (https://www.reddit.com/r/science/comments/3imphg/scientists_...): "The real interesting part is seeing how other disciplines hold up in terms of reproducibility. A new project has been started: Reproducibility Project: Cancer Biology, they will try to replicate 50 studies. I am very curious how this will turn out, I highly encourage other disciplines to also start a reproducibility project to test how consistent their findings actually will be. I don't see these results as discouraging, instead, I see it as a big step in developing scientific methods. Now we know which methods and standards might be wrong, we can try to fix it (for example by developing guidlines)."

Worth repeating:

> Problems of non-reproducibility, publication bias, data faking, lack of transparency, and the like plague every scientific field.

Yes. There's a good article on this from a few years back that began to popularize a term for this, called the "Decline Effect": http://www.newyorker.com/magazine/2010/12/13/the-truth-wears... (oblig.: author is disgraced science journalist Jonah Lehrer, but this is one of his pieces that has stood up to all scrutiny).

The most straightforward explanation, mentioned in the article and in some of the /r/science thread comments, is that null results are much harder to publish than positive results. This is not limited to the field of psychology.

All that said, I'm really excited by this paper and by the gradually increasing interest in this problem.

It's a particular problem with the "soft" sciences, though. For example, psychology is tough because you can't read people's minds, you have to go with how it affects their behavior in some measurable way or how they feel about it, and there's a tremendous number of confounding variables. Economics is another tough one, because experiments are not isolated from exogenous conditions and thus cannot really have a control. The whole thing is also highly interrelated with psychology as well. I'd throw poli-sci in that category too.

At least in physics somebody else can set up the same experiment and run it again. Sure, the same tendencies to publication bias, etc apply, but usually within a few years to a decade such things are discovered, corrected, and the problem identified. It's much harder to identify what resulted in the change of outcome for a soft-science experiment.

This is addressed a bit in the Reddit thread, where the comments from practicing scientists seem to mostly agree with you. Mathematicians are claiming they aren't affected at all, physicists are claiming that disproving a major theory in physics would get you a Nobel instead of scorn, and the bioscientists saying it's a bigger problem in their field.

That said, I think it's dangerous to give any branch of science an automatic pass on this. The fundamental causes of the Decline Effect aren't limited to the "soft" sciences, they just currently seem to have a more susceptible culture.

(comment deleted)
>An ideological response would discount the arguments, discredit the sources, and proceed merrily along.

But that's exactly what happens in the media where most of the people get their information about these studies.

What's the significance of the picture at the top of the article?
It is a stock photo that has the title "Personality Disorder"[1] and has the tag "psychology." So the editors or whomever searched shutterstock for "psychology" and scrolled through, saw that they liked that one, and chose it for the article. I doubt if there is any greater significance than that.

1: http://www.shutterstock.com/pic-174194297/stock-photo-person...

Hello everybody, I was an intern at the Center For Science, who are organizers of this reproducibility study. It's kind of strange to me that the COS's first mention on HN is this study when they have been working on a large open source project for a few years now, but it's obviously a welcome surprise. Anyway check it out, https://github.com/CenterForOpenScience/osf.io