13 comments

[ 2.4 ms ] story [ 50.5 ms ] thread
>In the era of Big Data, we’ve come to believe that, with enough information, human behavior is predictable. But number crunching can lead us perilously wrong.

The article isn't actually negative about quantitative approaches overall, but the subheading has to conform to the New Yorker's ideology.

"The dangers of making individual predictions from our collective characteristics" is a major area of research, especially in adtech, and the "Great Hack" documentary has showed us how well some have gotten at this for a slice of the population. A critical question for the future of data science is: "can a person really change?" And "how can we determine if they have?"
I hate to be one of those "I been complaining about this all along", but I have since my early college years in 2000!

I didn't have quite the sophistication in the arguments today, though they were essentially the same (1/20 is arbitrary, effect size is very important but ignored, false positives bound to happen, need to consider how rare the event is, etc).

I also lamented something that I feel is the core of the problem to begin with, people treated p-value = science and science = p-value. Once you do that, its easy to put blind faith into a single study that can have huge negative impacts on our freedoms and way of life. statistical significance is a tool that can be used in science, but saying that at the time produced huge defensiveness, especially from statisticians. I am glad this has finally begun to change

The article is interesting, but only touches on the very tip of a very large iceberg, and your comments are spot on.

This issue has been a central one in psychology for decades, especially in clinical psychology. There are many variants of the problem, with lots of corollary problems.

Once central challenge is that even when you say you are interested in inferences about an individual, you usually actually evaluate your inferential strategy across individuals. This becomes problematic because it's easy to identify a serial killer post hoc; it's harder to avoid the inevitable avalanche of false positives if you apply this to thousands or tends of thousands of individuals regularly. In this way, you're not actually interested in one single individual; it becomes critical to be really clear about what your inferential population is, and what you're actually trying to generalize to.

Like you're pointing out, a lot of this too reduces to the significant challenge of knowing which predictive model to use, when you're effectively faced with thousands, if not an infinite range of models (one for each person/situation) to choose from. You might improve your prediction by using a more tailored model for an individual, but then you increase the risk of model selection error. Even if you have a lot of data on a person, they will probably change, circumstances will change, and so forth. The challenge is in knowing how to decide which model to use and when, when to decide that a particular predictive case is an exception.

This is sometimes known as the "broken leg" problem in clinical psych, so called because of a thought experiment in which you're tasked with predicting commuting behavior. You might have volumes of data on people, even individual people, but if you know that someone suddenly breaks their leg, or there is some other anomalous event, it compels you to logically alter your prediction. The trouble arises in knowing when to shift your prediction, when a scenario is different enough from that under which your default predictive model was developed. In the age of big data, you might be able to capture an increasing variety of scenarios, but there will always be cases that are not sufficiently captured, or where logic impels one to switch predictive strategies. But how do you do that?

I personally think what's undervalued is a focus on predictive uncertainty rather than the point value of the prediction. That is, realistically capturing and recognizing the true uncertainty in prediction for individuals, rather than improving the accuracy in prediction itself, which might actually have some fundamental limit.

I like how it has an audio option, much easier to consume that way
Sigh, I was hoping to hear Hannah Fry. To this yankee, she has a wonderful accent (that sounds to me, very much like Pearl Mackie's Bill Potts)

The Curious Cases of Rutherford & Fry Series 12 The Horrible Hangover https://www.bbc.co.uk/programmes/m0001r7k

As someone working in the field I struggle with this type of writing because it is both helpful but also misrepresents the fundamental problem with how statistics are messaged. Few statisticians would argue that big data driven statistics demonstrates determinism in action. Simply that certain markers are more likely than others to predict likely outcomes. None of this explains why something happened simply that it was likely it would.

In the first example of the doctor purposely killing his patients a statistician would say this doctor fit the statistically predicted outcome. But what isn't discussed is that this doctor must have existed as an anomaly within his career group. He had to have been ABNORMALLY excellent as a doctor, it is the only way he could hide his victims within his normal morbidity rates.

This is one of the reasons anomaly detection is such an interesting field because it requires dimensional understanding of a subject.

"Then a biostatistician at Cambridge, Spiegelhalter found that Shipman’s excess mortality—the number of his older patients who had died in the course of his career over the number that would be expected of an average doctor’s—was a hundred and seventy-four women and forty-nine men at the time of his arrest."
Just to add to your comment: Excess mortality is the number of deaths, or mortality, caused by a specific disease, condition, or exposure to harmful circumstances such as radiation, environmental chemicals, or natural disaster. It's a measure of the deaths which occurred over and above the regular death rate that would be predicted (in the absence of that negative defined circumstance) for a given population.

Had the doctor not murdered any patients, he would have had the expected number of patient deaths, but the Cambridge statistician found that the number of patients that died under his care almost perfectly matched the expected number plus the number he murdered, meaning he was a statistically average doctor when he wasn’t murdering his patients.

statistics always involves selection and/or generalization of data. both of those acts always involves some kind of bias (even random selection) in what, how, and why, criteria used in process
The article omits the problem number one here: it is the most surprising results that get the most attention and among the most surprising many would be false. Extraordinary claims get extraordinary attention and this is not cancelled out by the requirement of extraordinary evidence.