57 comments

[ 3.6 ms ] story [ 137 ms ] thread
Full title (not possible to put as title because of HN limits):

An international comparison of the second derivative of COVID deaths after implementation of social distancing measures

It's been a long time since I took calc. What does the second derivative show us?
(comment deleted)
(comment deleted)
Acceleration.

The rate of change of the rate of change of deaths.

The word exponential only occurs once in the paper, and the rate of change of an exponential is an exponential, so take as many derivatives as you want and it's still going up. Why they don't take the log of deaths over time isn't explained.

(comment deleted)
I've seen very little coverage that really uses the term exponential in a way that suggests those reporting know the implications of that. I've seen one plot with a logarithmic y-axis, and that was in the Guardian some time ago. From it, it appeared a crude heuristic could be drawn that from the point that lockdown is implemented, the cumulative deaths appear to grow between two and three orders of magnitude.

I've been grabbing the data put out by Johns Hopkins [0] for the last few days to update some plots in a very hacky Jupyter notebook.

Crucially, the plots are normalized to (estimated) populations and the y-scale is logarithmic. I've put them on github in case anyone else is interested [1].

[0] https://github.com/CSSEGISandData/COVID-19

[1] https://github.com/dpwm/covid19-analysis/blob/master/Coronav...

edit: Clarified that though exponential growth is widely used, the implications of that are not followed up. It's past my bedtime!

Even fewer people understand logarithms than exponential growth. If you publish graphs with a log scale you will certainly confuse some people into thinking that the illness is tapering off.
yep. My first thought was why the hell they didn't they use the word acceleration for better readability.
It's not an exponential. It's a shifted sigmoid.
The logistic function specifically, which is a blend of an exponential growth and an exponential asymptote. So initially it behaves indistinguishably from an exponential function, until you get close to the inflection point.
No it's not the logistic function. Sigmoid is a general term of the shape, logistic specifically refers to an equation that's almost certainly not correct. It's probably close to stretched exponential, which is a reasonable approximation of the curve of an autocatalytic process that is controlled by stochastic collisions that can self-exhaust, but even that's not exactly correct.
Rate of change of rate of change.

They explain this in the paper: a constant (flat line across the graph) would correspond to exponential growth. The actual graphs drop (mostly), indicating sub-exponential growth, which is a good thing.

A constant 2nd-derivative means quadratic growth, not exponential.

I'm bit confused by the graphs, because they show the 2nd-derivative going to 0. But to stop this thing, we actually need to decelerate, eg going negative on the 2nd-derivative.

I think they're talking about the 2nd-derivative of the cumulative deaths. As they can (presumably) only increase, the first derivative is >= 0. The second derivative tends to zero as the cumulative deaths flattens off.

Edit: Now that I think about this some more, you're right. There has to be a decrease in the daily cases, which implies a negative 2nd-derivative. They mention "relative 2nd-derivative." It seems to be defined in the paper in a way that leaves me more confused:

> Daily fatality rates from the included countries were then used to calculate estimates of the relative second derivative of total deaths, N, 1/N d²N/dt², for a period of at least ten days.

Does this mean they are taking the second derivative of the reciprocal of the cumulative deaths?

They are taking the 2nd derivative of total deaths and dividing that by the number of total deaths and calling it the relative second derivative (rate of increase of rate of increase of total deaths, relative to total deaths).
The only way the velocity (number of cases) would go negative is if they discovered a disproportionate number of false positives or (number of deaths) if deaths were discovered to be due to something else [or people came back to life]. As time goes on, the velocity will reach a constant of 0, presumably. The acceleration at a constant velocity is 0.
Eventually acceleration would go to 0, but it needs to go negative first. Right now acceleration is positive and velocity is positive, hopefully soon acceleration will become negative, then eventually velocity will approach 0 from above and acceleration will approach 0 from below.
I believe that you are correct in your calculation, but incorrect in your assumption.

You are treating the curve of daily deaths as the thing to be differentiated. They are treating the curve of total deaths as the thing to be differentiated.

While the curve of daily deaths certainly would have a negative acceleration, the curve of total deaths is monotonically non-decreasing and therefore the second derivative is always positive but trending towards zero.

Like I mentioned above, the only way for total deaths to decrease (making the function not monotonically increasing and therefore capable of having a negative acceleration) would be through errors of accounting.

The goal was never to "stop" Coronavirus. Just slow it down enough that it doesn't crush the medical system.
> A constant 2nd-derivative means quadratic growth, not exponential.

A constant relative second derivative of total deaths is what they are talking about, meaning that they are flattening the 2nd derivative of the exponential by dividing by the exponential, making it constant.

Of course, this is all numerical methods, not analytical methods, so it doesn’t necessarily make pure analytical sense that they are even talking about an exponential.

1st: Velocity

2nd: Acceleration

3rd: Jerk

4th: Jounce

As others have mentioned, acceleration, or the rate at which the rate is changing. Now... I admit to being an armchair epidemiologist. I've been graphing the data from the JHU website myself, but I've just been using semi-log plots. This is an option at some of the "dashboard" sites.

In a semi-log plot (vertical axis is logarithmic), a constant rate of exponential growth is seen as a straight line, the slope of the line is proportional to the doubling rate, and changes in that slope show that the exponential growth rate is changing. This is a way to "eyeball" the graphs without trying to read anything too profound into them. But just comparing the graphs of the US, Italy, and South Korea is interesting.

What I'm not doing is publishing conclusions from this armchair analysis.

This looks like one of those P versus NP papers.
Very little details on the methodology
How so? They said they computed the second derivative. What more is there to say?
You can’t compute the second derivative of data from May 2020 when it is currently March 2020. Where did the data come from?
China data time shifted to match the country of interest.
That seems to make a pretty big assumption western countries can achieve that? My understanding even a small amount of non compliance has big consequences...
FTA: The cumulative deaths for each country, N(t), were estimated by deriving a multiplier Nfinal/N(tconv) at the time of convergence, tconv, to the Chinese trajectory, Nfinal being the total number of deaths in China and applying this multiplier to Chinese data for times beyond convergence.
Can't expect China-like second derivative when countries are not doing these counter measures:

Isolation of suspected cases and close contacts

Universal masking

Restrictions on travel inside country

Sending doctors from the rest of the country to epidemic centers.

This presumes all of these measures are highly effective, which is far from certain or obvious.
No, it only requires that one of the methods is slightly effective.
Not taking any of these measures is highly effective at failing to control the virus. We're currently seeing it play out across Europe and the US.
Or that they have significant additional effect over social distancing alone, which we won’t know for some time yet.
>Can't expect China-like second derivative when countries are not doing these counter measures

* not reporting cases accurately.

I have a hard time believing Spain will have a higher peak than USA.
You can't look at Spain and USA without considering the total population.

As someone in this thread pointed out: http://91-divoc.com/pages/covid-visualization/

Take a look at the charts of cases/1M people. It is obvious that US will have more cases simply because of the population.

it is notable that this work is done by an electrical engineer and a cardiologist, not epidemiologists. More than anything, the arm chair epidemiology is the our current second biggest danger. Epidemiology is hard. incredibly hard. It isn't viral marketing. It isn't electrical engineering. Data isn't pure or assumed to be correct. There data source was basically websites.
Their
thanks. there are a couple other typos I noticed this morning as well. shouldn't HN with Bourbon.
Based on what we've seen so far, electrical engineering and cardiology are significantly more rigorous than epidemiology, which appears to be more like economics or social psychology in terms of the robustness of its methods and quality of its work.
Define rigor.

In context I suspect you mean something like "treat data as objective." I can go to France and pickup and hold the literal kilogram. U can't do that, yet, with people's brains to the level needed for psych measurement.

Personally, I prefer social sciences and epi methods (get economics out of here...) Because they are more transparent about the role of the researcher and the limitations of their data. They don't bluntly trust it...the engineers I work with largely do. They assume data represents truth and is largely without meaningful error.

I've elaborated here on what I mean here: https://news.ycombinator.com/item?id=22737948

It's not just how much data is trusted (though note: Professor Ferguson at Imperial appears to trust the data coming out of Italy almost completely). It's the whole set of problems.

Trusting data and data being objective

>Right now I absolutely want to see papers written by physicists studying COVID-19. Why - because physics is a significantly more rigorous field than epidemiology. I trust the average physicist to have at least slightly higher standards, for instance I trust them to at least pretend to care about statistical uncertainty, and I suspect many of them will upload their source code. I don't expect any of them to email their paper straight to known-friendly newspapers. (from your other post)

Is a really...bad idea. Let's not presume knowledge transfers from one discipline to another, that standards look the same, or that Alan Sokal was anything other than an egomaniac.

Are you an epidemiologist? If so, can you be more specific about the problems and propose alternatives? Thanks.
I am not...however i am married to one. The ranting recently is kind of fun to listen too. My partner is working on covid. I learn a lot, I'm no expert in epidemiology...but...I also do/teach statistics in a biology field. My degrees are is in mechanical engineering and education statistics. (And that's probably enough info to out me to any real life friends in HN ::waves::)

Basically the problem is the authors don't know enough about the data they are getting to run analysis on it. They are not clear, and absolutely need to be clear, on all things data for this paper to work.

Specific examples:

Are the websites they gather from reporting presumptive positives or confirmed positives?

Are they Getting information on when specific deaths occurred? Or are they using the latest update time, and the total number of deaths to date and then assuming a connection?

Are countries accurately reporting? Are they even capable of accurately reporting?

Are they testing enough (all) of the dead or are they relying on presumptive positives?

It's a data thing. Epidemiology is really really good at high quality data and analysis...but that takes time, more time than we have had to do good science during a crisis. It's why you don't do brain surgery in an ER. They are also really good about not over interpreting data...because data can be hinky.

One of the biggest things people in my undergraduate stats course cover with is specificity and sensitivity...false positives and negatives exist. That gets more complex when you consider that when talking about true positives versus false positives you have to have a gold standard test to reference against. So largely you are looking at one test versus another, even if one of those tests is really accurate, or pathology based, then things take time. That test doesn't typically produce a truly binary result but some chemical threshold that we treat as a binary. Decisions at each of those changes in data representation affect your outcomes. Then we talk about how it's impacted by prevalence and I try and teach them ROC curves and someone tells at me about how if you test positive for a disease you have that disease. (Seriously this happens about once a semester)

Where the students are really struggling isn't math, it's philosophical. The idea that a test result isn't objective truth is kinda bonkers the first time you encounter it. Especially for my engineers. I jokingly talk in class about how it's not a math class it's an estimation and bs detection class. I've seen this so much with covid...test results are not perfectly accurate. You can, and often should, bias them towards certain clinical goals. The first tests from cdc were problematic because they were overly sensitive (or contaminated...don't know yet, probably just too sensitive).

The reality more than anything is that there is so much miscommunication about covid, unavoidable communication, that I'm not sure anything is fully trust worthy yet. Weeks ago one of the first rapid studies published in JAMA ( or was it NEJM?) By a German group was submitted, reviewed, published, and withdrawn in about a week. People use words like 'positive test result' and can mean fundamentally different things when taking to each other and don't realize it.

And apologies I'm not following my normal anal hn comment style and citing this a bunch. Responding from my phone.

Thanks for the detailed reply. I agree with your comments regarding data quality, I assumed you had an issue with the methodology. I just took this paper to be a heuristic that could help inform decisions with the limited data we have so far, and not as a definitive answer.
The methodology don't really explain how they aligned the second derivative curve to China or why that's reasonable.

I wouldn't put much faith in their predictions.

I tried fitting a third order polynomial for JHU Covid data, just as a way to kill time. For all countries, the 95% confidence interval of the second derivatice overlapped zero (i.e. it can't be estimated very well). We only have about three weeks of data from the 10th death, for most countries, you can't fit a very good curve with that.