52 comments

[ 3.1 ms ] story [ 134 ms ] thread
>And the 62% figure “certainly is consistent with there being a problem” in the field, he says. “It seems funny that there’s been a drift in standards to the point where 62% seems very respectable.”

I'd sure say so. Statistics in science is just so broken. Lots of inappropriate techniques chosen on tradition, ignorance (or concealment) of multiple comparisons, the base-rate fallacy, etc, etc. It's not just social sciences, but medicine, biology, and just about every field that relies greatly on lower-power studies.

Bayesian techniques are really called for in those situations.

I'm not a scientist, nor am I familiar with statistics. I could imagine this being troubling for a scientist or someone familiar with the field. But for a non-scientist it's fairly devastating if you consider the implications.

What it means is that if I trust a study - any study (from my perspective) - I'm essentially flipping a coin. If non-scientist citizens can't rely on what is coming out of the field, then it seems like a massive problem that needs solving before any other.

> that needs solving before any other.

Sadly, I think there is huge incentive to maintain status quo. It is essentially a loophole to manipulate the scientific consensus without the actual scientists being dishonest. Some wise once said "When a measure becomes a target, it ceases to be a good measure.". I think that what we call as "science" today, has thus that fallen prey to this.

You should be considering not just one study but how independent studies fit together into a grander picture.
That's not a useful thing to say to the general public--they're not capable of that. They count on scientists to police themselves.
I get his point, though in the end I agree with you. The interface between the scientific community and the public at large should be reliable. At this point in time it doesn't seem to be.
Scientific papers aren’t written for the lay audience, it’s a conversation amongst scientists.
You shouldn't trust any single study. Researchers are not perfect, reviewers are not perfect, the systems they observe are complex. Even in the "hard" sciences even if the experimental data itself is solid there can be confounders that were overlooked and will only be uncovered by follow-up studies, depending on importance of the result it might take years.

That's why have meta-analysis and systematic reviews and even those aren't exactly bullet-proof.

(comment deleted)
Meta analysis is just the weighted average of expert opinion. Which makes it worthless when the experts don’t actually know anything. https://jamanetwork.com/journals/jama/fullarticle/2698337

You should be skeptical of any so-called science that cannot actually rigorously apply the scientific method, be it for practical, ethical, or any other reasons.

> Even in the "hard" sciences even if the experimental data itself is solid there can be confounders that were overlooked and will only be uncovered by follow-up studies, depending on importance of the result it might take years.

I work in fluid dynamics. In one subfield I work in there is at least one confounder in most studies (Weber number and Reynolds number are confounded, in particular) and an important variable is often omitted (turbulence intensity). Many people don't seem aware that turbulence intensity matters, either, and most who do think it matters don't want to measure or estimate it, so it tends to be ignored. Just because it's hard to measure does not mean it's not important! This is supposed to be "hard" science, but sometimes I feel it's not much better than social science.

How long has this been a problem? I'd say over 50 years, easily!

At a recent conference, one of my papers addressed these issues briefly (I tried to avoid them in a data analysis), but I think it'll take more than that to correct these problems. So I'm planning a new article specifically addressing these concerns. I still expect progress to be slow, but any movement in the right direction would be appreciated...

That's really interesting. Seems like good science is always going to be an uphill battle against human nature/complacency.
Actual meta-analyses and systematic replication are so rare as to basically not matter. In addition, many actual scientists are completely unaware of these fundamental issues in science. Not to mention scientist's statistical misconceptions. The vast, vast majority (over 90% from surveys I've seen) commit the inverse-probability fallacy, for example.
The first step is to not call science what isn’t science. Social sciences is a heavily politicised field which is more akin to journalism than anything else. It doesn’t mean it is not interesting but it does not carry the authority of hard sciences.

Medecine is more problematic. Many studies are at the border of statistical significance but it is impossible to conduct large scale experiments with humans, so it’s not like if we have a better alternative.

> Social sciences is a heavily politicised field which is more akin to journalism than anything else.

You're being downvoted, but my understanding is that lots of social science researchers openly refer to themselves and their discipline as activism. Am I mistaken? Is this observation controversial?

EDIT: Now that people are downvoting me, perhaps someone could take a moment to explain where I'm mistaken?

I haven't downvoted you, but "Science" and "Activism" are not mutually exclusive.
Fair point, but isn’t it cause for skepticism? It seems that science-minded folks are skeptical of the research done by various politically-oriented think-tanks; isn’t a politically oriented field or department similar?
The question is what direction that's in. Activism that is, in essence, urging the adoption of ones findings, if obtained through valid experiments, isn't really biased. Politically-oriented think-tanks are looking to build research that supports a particular political outcome.

That's a pretty strong contrast to say "We have found that X is harmful, and will advocate for people having reduced exposure to X."

Consider, for example, the number of infectious disease epidemiologists, vaccine developers and doctors who have essentially been forced to engage in "activism" thanks to anti-vaxx movements.

I absolutely agree, but it seems valid to question the “direction” of these fields, especially when there appears to be a growing body of research that suggests a strong political bias.
Neither are "Justice" and "Revenge". But there's a good reason we treat them differently.
The prediction-market aspect is fascinating; they asked a group of experts to predict which would and wouldn't replicate, and while we don't see the experiment-level results, the overall average was apparently very close.

Is this common in modern replication studies? Do the results of the prediction-market ever get announced/published prior to the replication results themselves?

> experts also participated in an online “prediction market,” trading shares that corresponded to studies, which paid out only if the given study was replicated.

> Both approaches did well at predicting the outcome for individual studies, and they predicted an overall replication rate very close to the actual figure of 62%.

I think the real headline here is that we should supplement peer review with replication prediction markets.

The real headline is that this is a more formalized version of what science already does.

Science, as a process, naturally replicates studies that are worth replicating, and drops those that are not. The idea of "prediction markets" is a new-tech way of doing what scientists already do: decide where to invest their (limited) budgets to maximize long-term career success.

Is there any information on these prediction markets? Is there a reason only "experts" can participate? This would be an interesting application for ML side projects that are monetarily beneficial while also benefiting society as a whole.
The experts aren't necessarily credentialed experts as we would typically think of them. You can participate in forecasting tournaments like Good Judgement Project and if you score high enough you can participate in invite-only events that are in partnership with entities like DARPA. But the people come from all backgrounds- I got all this from the book Superforecasting by Philip Tetlock.

Edit: got the title wrong

It is too bad that society does not support a robust system of prediction markets on everything that would be interested in knowing the truth about. I think the excuse is that it is too much like gambling. I would say there would definitely be some gambling going on in such markets, but we already allow them for much the more mundane reason of predicting the future time-discounted profits of a business. The whole country participates and has a good deal of its wealth tied up in them (ie. stock markets).
> If an initial replication attempt failed, the researchers added even more participants.

This is an obvious way to tamper with the results; it's just more of the same kind of p-hacking that bad researchers are so often doing. They are using "we re-do the study with a larger population" as a way to re-roll the dice if the first die roll doesn't come up the way they want. (Note that if the die roll did come up the way they want, they don't re-do the study with a larger population in order to see if the replication fails).

Nobody should be taking this seriously.

There are known methods to account for this, and if you adjust the statistics properly to account for the repeat, it isn't p-hacking.
Did they? I can't tell from the article!
I think the point is actually more the opposite. If your study failed to reproduce even with the generous "keep adding more participants until it works" method, then your first study is beyond a doubt BS.
Except no, because rhetorically they are using the 62% figure as a good sign, rather than a bound on how bad things are.
I'd say it's a good article with a bad title. my take away is that even with their generous approach only 62% of studies replicate. That means 1) probably over 1/3 of results are completely bogus and 2) of the ones that were real, about half the effects were substantially weaker than claimed.
This article is Science telling us that "look, over half of the high impact science we publish is actually real! Our peer review works!".
It's not 100% unproblematic but I wouldn't call it p-hacking.
Repeating a study if, and only if, the study failed to provide the desired result is the most basic p-hacking technique there is. If this isn't p-hacking, I don't know what is!

To give a simple model, suppose you decide an effect is significant if there's only a 5% chance that you'd see this data if the effect didn't exist. If you run the experiment again with the same threshold for significance when don't get the desired result, then the probability of seeing an effect that doesn't exist rises to 9.75% (= 0.05 + 0.95 * 0.05).

The effect isn't merely "not 100% unproblematic" it's a serious problem! You've gone from what looks like a p-value of 5% to a p-value of 9.75%.

The fact that the second study is done with a larger sample is pretty much irrelevant unless it also comes with a higher p-value threshold for you to accept the result.

> If experts [in prediction markets] can instinctively spot an irreproducible finding, “that kind of begs the question of why that doesn’t seem to be happening in peer review,” says Fiona Fidler, a philosopher of science at The University of Melbourne in Australia. But if future studies can identify and weigh the best predictors of replicability, reviewers might be given a rubric to help them weed out problematic work before it’s published.

That's a troubling suggestion. Results are valuable to the extent that they're both accurate and surprising. To systematically suppress surprising results as a negative predictor of accuracy sounds like a formula for suppressing surprisingly valuable papers.

> to systematically suppress surprising results

"Weeding out" doesn't mean systematically suppressing. It simply means exhibiting caution before putting the paper in a reputable journal. The Internet provides the entire system with an escape valve in scientists' capacities to publish papers on their own websites. If the finding is surprising enough, it shouldn't be problematic attracting some attention, particularly in the social sciences.

> "Weeding out" doesn't mean systematically suppressing.

It does when I'm weeding my garden. A result that is weeded out of a journal is suppressed from taking root in the minds of its readers. I agree that extraordinary claims require extraordinary evidence to accept. But as long as they meet the ordinary standard of rigor it's just those claims that inquiring readers want to entertain.

Think of Thikonov regularization. Or the LASSO. The basic idea is to start with a prior that the world is uninteresting and let the evidence try to prove otherwise.
(comment deleted)
I believe the thought is extraordinary results require extraordinary proof.
The danger with surprising results is that they often use the surprise part as an excuse for small sample sizes, etc.. Another way to think about this situation is to flip it and say that surprising (and spurious) results are more likely from exploring a small dataset.

Andrew Gelman talks a lot about this issue on his blog:

https://andrewgelman.com/2014/08/01/scientific-surprise-two-...

(comment deleted)
I assume Science knows what it's talking about, so I'm confused on a couple of points. Perhaps some practicing scientists could clear them up for everyone:

For any arbitrary paper, what is your assumption about its accuracy? How much do you rely on it? Can you put a number to it? The null hypothesis that research papers, especially in very difficult fields like social science, are near-infallible seems to be the error, AFAICT. It seems like something people outside science would assume, but I see scientists who are surprised by a 62% replication rate. I'm not surprised or concerned, but maybe I should be.

Why do they call increasing the number of participants, "generous"? Doesn't that increase accuracy, and isn't accuracy the whole point of the replication study? Generosity implies some sort of favor, something above and beyond, while in this case it seems necessary based on the fact that the results changed - it would be failure of the replication study to not increase the number of participants.

When you repeat things over and over, even with an increasingly large sample, you artificially increase the chances of replicating something purely by chance. When papers are published they tend to publish ranges for the expected strength of their findings. What replication studies have found is that these ranges are not representative of reality in most cases. It's not about 'these papers are not infallible' but 'these papers seem to have misrepresented the strength of their findings.' And like the article mentions, even with this rather generous method of trying to 'replicate' studies, they found huge chunks of papers could not be replicated to anything like the strength originally suggested.

The unstated implication here is that it seems a very large chunk of social science researchers are faking their numbers, likely through p-hacking, to create attention (and journal) grabbing headlines that are, for lack of a better word, simply fake.

>For any arbitrary paper, what is your assumption about its accuracy?

From what I've seen in academia, it varies depending on a number of variables, like who wrote the paper and whether their results help or hurt my research.

Scientists are human, after all. And in my discipline, it was rare that the paper would have enough concrete details to replicate, so this way of thinking was the norm.

"For any arbitrary paper, what is your assumption about its accuracy? How much do you rely on it? Can you put a number to it? The null hypothesis that research papers, especially in very difficult fields like social science, are near-infallible seems to be the error, AFAICT. It seems like something people outside science would assume, but I see scientists who are surprised by a 62% replication rate. I'm not surprised or concerned, but maybe I should be."

There is no such thing as "an arbitrary paper". It will depend very much on the study in question - including soft factors like who wrote it, but also the study design, sample size, if I think they approached the statistics appropriately, etc.

"Why do they call increasing the number of participants, "generous"? Doesn't that increase accuracy, and isn't accuracy the whole point of the replication study? Generosity implies some sort of favor, something above and beyond, while in this case it seems necessary based on the fact that the results changed - it would be failure of the replication study to not increase the number of participants."

The reason it might be thought of as generous is the criteria they determine for saying something is replicated is a significant effect in the same direction as the original study (for the record, I hate this criterion). A larger study is more likely to find a significant result if indeed there is one there, so they're giving studies that report an effect a very strong chance of seeing that replicated if indeed it was real.

The facts are spread throughout the article, so this might help:

* Prior replication studies replicated 39% of papers in psychology journals [Ed note: That doesn't mean the other 61% were complete failures; most just didn't produce results as statistically strong as the originals IIRC] and 61% in economics journals.

* This replication study greatly increased the number of participants in some experiments, and for two papers, that changed the replication results from failure to success. Overall 62% of the studies were replicated successfully.

* One researcher "points out that the project repeated only one experiment from each paper, and in his case, it wasn’t the strongest or the most important."

You're missing arguably the single most important sentence in this article:

* If an initial replication attempt failed, the researchers added even more participants.

Which is not unlike doubling-up at blackjack, which will get you thrown out of a casino.

Start with 25 participants, then if you don't get the result you want, add 50 and try again. Carry on doubling until you get the result you want.

Not a statistician here, but as each addition is essentially a new trial, shouldn't they have applied Bonferroni correction to the results?

I would appreciate a reader-friendly list of the studies that were successfully replicated, being it the 39% of the 2015 paper or the 62% of this paper.

Knowing that I have no strong reason to trust most of the conclusions is useful. But know which papers I can trust is more useful.

Anyone knows of such a list?