60 comments

[ 8.9 ms ] story [ 175 ms ] thread
I am not so sure how universal this is if it varies between 7ms and 470ms, the article says "there is variation but it's minuscule", to me a two-order of magnitude variation is all but small.

And growing up in Europe and moving to North America I can definitely feel that the "polite" amount to pause in a conversation here feels way longer, it might be small in terms of milliseconds, but if you are used to, say, 50ms, having to pause for 200ms is very, very noticeable

> I am not so sure how universal this is if it varies between 7ms and 470ms, the article says "there is variation but it's minuscule", to me a two-order of magnitude variation is all but small.

Agreed completely. Given the scale, and the mention that the average conversational segment lasts 2 seconds, 7-470ms seems like a major difference.

It also doesn't seem like "the minimum human response time to anything", as the article quotes, if it has that much variation. Humans across all cultures can react far faster than 200ms.

> And growing up in Europe and moving to North America I can definitely feel that the "polite" amount to pause in a conversation here feels way longer, it might be small in terms of milliseconds, but if you are used to, say, 50ms, having to pause for 200ms is very, very noticeable

I've had the same experience in reverse: if you've come to expect longer conversational pauses, short ones give the conversation a very different tone.

And on the flip side, the necessary pause can grow noticeably shorter between people who know each other very well.

It could be that the minimum is 7, the maximum is 470, but almost all values cluster around 200. Then it would be fair to say that variation is minuscule.
That's what I thought they meant, too, but it is not. From the research article (http://www.pnas.org/content/106/26/10587.full):

"The means display somewhat more variation, as shown in Fig. 2. Danish has the slowest response time on average (+469 ms) and Japanese has the fastest (+7 ms). The mean response offset for the full dataset is +208 ms, and the language-specific means fall within ≈250 ms either side of this cross-language mean, approximately the length of time it takes to produce a single English syllable (37)."

The point, I think, is that these means are close compared to the tails: http://www.pnas.org/content/106/26/10587/F1.large.jpg

I think the Atlantic article conveyed the wrong conclusion. The full text of the article is here: http://www.pnas.org/content/106/26/10587.full

The abstract has their conclusions:

"Informal verbal interaction is the core matrix for human social life. A mechanism for coordinating this basic mode of interaction is a system of turn-taking that regulates who is to speak and when. Yet relatively little is known about how this system varies across cultures. The anthropological literature reports significant cultural differences in the timing of turn-taking in ordinary conversation. We test these claims and show that in fact there are striking universals in the underlying pattern of response latency in conversation. Using a worldwide sample of 10 languages drawn from traditional indigenous communities to major world languages, we show that all of the languages tested provide clear evidence for a general avoidance of overlapping talk and a minimization of silence between conversational turns. In addition, all of the languages show the same factors explaining within-language variation in speed of response. We do, however, find differences across the languages in the average gap between turns, within a range of 250 ms from the cross-language mean. We believe that a natural sensitivity to these tempo differences leads to a subjective perception of dramatic or even fundamental differences as offered in ethnographic reports of conversational style. Our empirical evidence suggests robust human universals in this domain, where local variations are quantitative only, pointing to a single shared infrastructure for language use with likely ethological foundations."

The research article claims that all languages have this gap, and that within a culture, the gap is quite consistent. Across cultures and languages, however, the gap has a decent spread, as you observed. In the discussion, the research article authors say,

"Amid a strong universal pattern, we do see measurable cultural differences. However, the range that we show, mean offset of next turn in each language departing no more than a quarter-second from the overall mean, is not of the kind that would imply fundamentally different types of turn-taking systems in the different languages, as the cultural variability hypothesis would suggest."

Honestly, I'm still trying to reason through how their data and conclusions line up, but their conclusions are more subtle than what the Atlantic article explains. (Which is not surprising to me.)

edit: Looking at the distribution graphs, it is true that these means are close compared to the tails of the distributions: http://www.pnas.org/content/106/26/10587/F1.large.jpg

Thanks for this -- the "single shared infrastructure for language use with likely ethological foundations" seems to be the interesting bit; looks like the Atlantic missed the point altogether.
>I am not so sure how universal this is if it varies between 7ms and 470ms, the article says "there is variation but it's minuscule", to me a two-order of magnitude variation is all but small.

Only if we are comparing it with itself.

If we compare it with the times involved in the general context of speaking (the time to say even a small phrase) the variation is insignificant.

In my original family this gap is negative: we start speaking before the other person finishes the sentence. But even as a child I understood that this is not a norm and learned to adjust to other settings.
I understood that some aboriginal American speakers had 'conversational timers' that were in the minutes. That is, when somebody (an elder?) stopped talking, you waited for 5 minutes of silence before you could be sure they were done talking. Navaho perhaps?
Wouldn't be surprised. My family is Finnish but this is basically how we talk in a lot of situations. It's fun watching chatterboxes adapt to the long periods of quiet between phrases.
I like to give a little thought to the words that I speak. The conversational gap timing may be the reason I never seem to get a turn to speak in some conversations.

Sometimes, when speaking with the spouse, I sense a "your turn to speak" gap, but by the time I actually start speaking, I apparently missed my turn, and I get dinged for interrupting the continuation. Then, later, I get double-dinged for not being communicative enough.

I never quite got the hang of "seizing the conch" (as in _Lord of the Flies_) and holding onto it with a continuous stream of priority-holding conversational nulls while my brain works out what I actually want to say next. I'll sometimes pause to take a breath and before I am able to resume, someone else jumps in.

So I'd love it if I had a cultural conversation timeout even just 15 seconds long. I might actually get to talk and finish a line of thought without someone else trampling over me.

Edit (sibling post): So maybe I should try to get some Finnish friends, then? But it's been a while since I've been to Minnesota or da Yoop, doncha know.

I've had exactly the same experiences regarding mis-timed gaps among different groups; it can create an incredibly jarring experience, as well as a perception of rudeness by both sides.

In a more formal setting, you may have no better recourse than just learning to adapt, but with someone you're close to, it seems worth talking through explicitly.

I've definitely learned how to maintain priority with conversational no-ops when it's a situation I want/need to have my say. I dislike doing it, but it's a necessity given the bias our culture has towards fast conversation.

Hopefully your spouse can learn that's your conversational style and adapt to it. Once my partner was made aware of it she started waiting longer before answering. She even apologized once for "cutting me off" during a long pause between thoughts!

There's some interesting things in this article, other than the "universal" gap. Tidbits like this:

> "The brevity of these silences is doubly astonishing when you consider that it takes at least 600 milliseconds for us to retrieve a single word from memory and get ready to actually say it. For a short clause, that processing time rises to 1500 milliseconds. This means that we have to start planning our responses in the middle of a partner’s turn, using everything from grammatical cues to changes in pitch. We continuously predict what the rest of a sentence will contain, while similarly building our hypothetical rejoinder, all using largely overlapping neural circuits."

Makes sense. I (and I'm assuming everyone) will have stories/anecdotes/arguments at the ready during a conversation, and they'll vary and change based on what the other parties are saying. I just didn't realize that this was a necessary thing.

Beyond this, it's doubly unsurprising when you remember that the brain isn't anything like a context-switching CPU, but much more like a bunch of discrete DSP circuits. Much like computers can offload video decoding to a dedicated DSP "pipeline" that will continue performing each step without being coordinated by the CPU, language processing and construction—word-lookup, grammar, etc.—are steps in an offloaded "pipeline" of brain areas.

Though, highlighting a specific sentence of your quote, the analogous operation becomes quite clear:

> We continuously predict what the rest of a sentence will contain[...]

Speculative execution via branch prediction!

(The important thing to highlight with that analogy, though, is that the brain has multiple "branch predictors." In fact, one could think of speculative execution as an inherent low-level feature of each node in a recurrent neural network. This isn't our executive function "predicting the future" generally; this is each part of the brain going on with its work independently by making assumptions that its speculative work isn't going to go to waste.)

I think we typically already know how we are going to respond to the earlier section(s) of the other person's dialog, and this is why our response is ready so quickly. If your counterpart speaks 3 sentences, I suspect that in many cases, your response really only takes into account their first and maybe the second sentence. Their third sentence likely often adds extraneous details that don't require you to adjust your response. Or, in the case of debating (ie: your response is a rebuttal in an argument), we often largely ignore what the other person has said and ramble off a response that doesn't even address the new information given to us by the other party. I find a lot of conversations to be more like two separate one-sided discussions where each person is selfishly just pushing their own talking points, rather than wrapping the other party entirely within the discussion.

Unfortunately, something about my wiring doesn't seem to match these timing cues up with my peers. In a group discussion involving three or more people, I'm desynchronized from the others. Far too often when I begin to speak, I get pre-emptively cut off by someone else who "beats me" to the cue. I think I take a little more time to plan what to say before I open my mouth, whereas my peers are ready to ramble unpolished responses the moment someone else stops speaking. Too many people just like to hear themselves talk and be the centre of attention the millisecond someone else's moment in the spotlight is finished.

>> a sign that we’re spending most of our “listening” time actually prepping what we are going to say (As Chuck Pahlaniuk once wrote, “The only reason why we ask other people how their weekend was is so we can tell them about our own weekend.”)

The article covers my beliefs with this quote, but then goes on to say this study proves otherwise. I guess I'm a pessimist.

Edit: As a side anecdote, I am 30 and have a much easier time hanging out with and speaking with 45+ year olds. Conversation flows naturally and my timing cues are much more in line, as opposed to the rush-to-be-first situation I run into with people my own age.

Your anecdote is inconsistent with the "universal" short gap, and I think you are more correct. The "universality" comes from averaging that smears away individual differences
I suppose if you're going to draw a bell curve for the gaps people exhibit, someone has to fall on the outer edges. I'm one of the outliers on the right side of the graph in terms of milliseconds. :)
Interesting. My algorithm is a bit optimistic as I tend to interrupt people (getting better at it!).
Realizing you did it goes a long way to stopping--kudos!
Having spent much of my life around people speaking languages I often had trouble understanding, the quoted phenomenon becomes even more obvious. I got better and better at maintaining the rhythm of a conversation, and occasionally even fooling someone into thinking I understood them, simply based on verbal cues, snippets of grammatical cues and pitch (and perhaps other signifiers).

And perhaps this is an even better (and much cuter) example of the phenomenon:

https://www.youtube.com/watch?v=_JmA2ClUvUY (talking babies)

> "It’s the minimum human response time to anything...It’s the time that runners take to respond to a starting pistol—and that's just a simple signal"

i.e. the gap is typically exactly as long as it takes for us to react to the "simple signal" that the other person has stopped speaking and now it is our turn.

It does prove that we must be planning our response before the other person has finished speaking, but I wonder if that's such a surprising thing to prove: words and sentences take a lot longer to say out loud than they do to comprehend.

Well, in Brazil there is no gap :D

The conversation is just an endless stream of word from several sources.

There is no discernible gap on my wife's side of the family... they are Colombian. Further, the volume tends to increase (rapidly) as they speak over one another. I am surprised to find they are actually communicating with understanding.
same with my puerto rican wife's family. could be a hispanic thing.

slightly tangential, but on the loudness part of it, funny theory I have is that, and I noticed this when I've visited her family in their hometown, since they live outside of any big city, nature itself is so loud, they have to talk louder and louder to hear each other! ;-) especially at night, there are so many noises coming from nature itself. here in mainland US, in so many places I've been to, from Maine down to the south, from El Paso to San Fran, night time is so much more quiet! anyway, that's my theory until proven otherwise haha :)

> San Francisco

Try living downtown. :)

I've lived all over the world and I suspect it's a class thing. Downtown Chicago people tend to talk over you nonstop with no regard for courtesy. Where I've been in France and the rest of Europe, it's more what you'd expect (a pause). In my rural midwestern working class background there tends to be a pause, we didn't talk over one another. My wife's family is from Guadalajara, were privately schooled and in the Mexican government's "political class", they don't interrupt.

It's ingrained in me enough that if someone talks over me, I'll just stop talking to them and subconsciously consider them to have a low standard of manners.

My brother does this. Interrupts you, and if you don't stop speaking, his volume increases. I personally lose my train of thought and so I just can't keep up with him. But it's a fun game to play sometimes.
That’s why any site should load in under 200ms.
(comment deleted)
Interestingly I've read that people conversing in sign language don't leave gaps in the same way, since "talking" doesn't get in the way of "listening", so both participants tend to "speak" simultaneously. This dovetails with the article: the limiting factor is not how long it takes us to think of our replies, but how much we can convey through speech without it becoming unintelligible.
Along similar lines, though on a longer timescale, people on IRC or similar text-based chat mediums commonly take part in many real-time conversations simultaneously, and don't have any trouble mixing them up.

In theory, if you didn't mind breaking cultural norms, the same thing can work with speech.

The barrier to that is the shared medium - when two people speak at the same time, neither can hear the other clearly.
Uh, I don't know: IM chatting is anything but "instant". It's pretty slow, actually. I don't really feel it while in conversation, but if I look at times each message was posted afterwards — it's easy to notice that conversation of the same volume would be done much quicker verbally. Waiting for everyone to respond on the top of it would be so slow it would kill the conversation.

I don't speak sign language, but I observed others doing that and it seems to be almost as fluent as talking verbally. While both parts are involved in conversation most of the time — it's something we do as we speak verbally as well: making facial expressions, nodding, saying "uh-huh" and inserting little jokes.

I never actually noticed that deaf people really tell stories at the same time: one of them usually seems like listening. So, I'd actually want some reference that would prove that deaf people really do "speak simultaneously".

> “It’s the minimum human response time to anything,“ says Stephen Levinson from the Max Planck Institute for Psycholinguistics. It’s the time that runners take to respond to a starting pistol—and that's just a simple signal.

This seems to match the front end world where CSS and JS animations are typically set to a number between 100-200ms.

And this is one of the main reasons I hate talking on cell phones. Those have a latency up to a quarter of a second, which totally confounds this turn taking.
Or having cross-continental calls, which introduces similar latency.
The article is about culture, but I'm more interested in brain chemistry. Do people on the autistic spectrum have the same gap? People with ADHD? People with Alzheimer's?
I've had it pointed out to me that there's something weird about the show Gilmore Girls in the actors' delivery of lines. The actors don't seem to be speaking particularly fast, but something about the structure of the dialogue still seems to be perceived as very "dense" by most people. I'm guessing the effect is from the actors shortening the conversational gaps between their line deliveries. It seems like the each reply comes near-instantly after the actor "heard" the previous line, which may be a super-stimulus for making the portrayed characters seem extremely "witty."

I also have a feeling that actually increasing verbal velocity would hinder this effect; you wouldn't get the impression of a witty character, but rather a manic one.

I always found it amazing that Lauren Graham was able to remember that much dialog for each episode. I would like to see some stat on how many total lines she had during the course of the show.
They talk about this in one of the special features on the DVDs. The average script length for a typical "hour" show (42 minutes, running time) is one page per minute. Gilmore Girls episodes were closer to 80 pages. They even talk about bringing in a speak coach for the actors so they could learn to speak faster while still being clear.
Do they explain why they did this?
Shows with Aaron Sorkin dialog also have this feel. (A "verbal Normandy" in the words of Stephen Cobert.)
That and the walk-and-talk approach were what made The West Wing so much more enjoyable to me. It audibly and visibly made everything seem more hectic and exciting in an unusual way.
> It seems like the each reply comes near-instantly after the actor "heard" the previous line, which may be a super-stimulus for making the portrayed characters seem extremely "witty."

Not just witty, but in tune with each other. Shows that pull this off give the impression of characters that actually know each other, work together, spent years together, etc.

Not as a universal trait as to be present in my family. One needs to talk louder and interrupt one another to be allowed for a turn.
If you intentionally prolong the gap, you can force the other side to feel discomfort. Also known as awkward silence. Most people will instinctively fill the gap by starting to talk again, even if they don't have anything more to say.

You can use this in negotiations. After someone makes an offer, they'll stop and expect you to respond. If instead you act as though you're thinking and hold the silence, they may get the natural urge to fill the gap and start speaking again, sometimes offering better terms.

One time I used this technique at a negotiation but the other party was an MBA grad so they might've known about this. This resulted in very long pauses, maybe several of 20-30 seconds each, where we both waited for the other person to fill the gap. If someone were watching us they'd think we're both crazy.

A simple way out of that trap is to self-followup with "Take as long as you want to think it over", or "Do you have any questions about my proposal?", or the aggressive "Any objections, or do we have an agreement?" or of course the 21st Century [take out your phone while you wait].
This is one of the most-used and very effective 'tricks' in Ricky Gervais' shows (like The Office and Extra's). I truly believe that the camera staying 'on' after a moment of embarrassment is the single thing that elevates them to the level of excruciating (to some of us). I've been meaning to cut together some scenes where these 'gaps' are left out to see if the character's behavior would still be as excruciating.
(comment deleted)
I fail at this, miserably. Some of us are on the other end of the spectrum. Tending to wait "too long" to respond thus end up talking over someone.

I think this is because some of us prefer to take time to actually think about a response, rather than just react.

I think about all my responses before broadcasting, and I end up making excessively long pauses.
I had a colleague who used to have very long pauses, but only after saying a few words very quickly first in a way that ensured he held the conversational baton. So he would say something like "I don't agree because..." followed by a 15 second pause before continuing the sentence.

Pretty sure it was not intentional, it was just what he had converged on as his optimal conversational strategy, but it was incredibly frustrating for others. Partly just the endless pause, but also I think (after reading this article) that everyone is busily optimizing a response in real time while he talks, but when he stops talking mid sentence, suddenly the algorithm has no more input to process yet still knows there will be more coming at some point soon, so cannot complete and also cannot swap out, so instead goes into an uncomfortable spin wait state that gets quite exhausting after a while.

His issue isn't non-blocking IO. Processing is preventing his single threaded event loop from doing anything else.
This is why we have verbal cues for conversation that indicate that the listener-now-speaker has understood what was said and is formulating a reply: "hmmm, I see. But have you considered, perhaps, that..."

Conversation, as opposed to written text, has a lot of "filler" designed to give the conversant time to formulate their words.

Your implication is that people who respond "on time" have thought less about their response than those who do not. More likely, they have just a better intuition for conversation that allows it appear that way.

That's a fair point but it's hard to take seriously the suggestion that someone who puts more thought into a response isn't likely to have a more insightful answer.

I cut out most filler words out of my vocabulary years ago as I find them annoying but occasionally do use the "mhm" hum. That said, I will fully admit I'm a traditional nerd at heart and do lack intuition for nonsensical, "how's the weather" conversation.

But if you want to speak about something meaningful? A conversation that would flip the roles and leave the majority of the public at a loss for words, you would probably prefer to speak with me.

Most of my social interaction is at places where naturally curious people gather (like user groups / meetups).

This is why in most situations I let most people do the talking, and I focus on questioning them. I can judge what level they're on intellectually and as a bonus, people like to talk about themselves.

I wish they had included India in their study - after watching a panel debate on NDTV (a major news channel there) I'm not so sure this turn taking short gap would be considered 'universal'. I hope no one finds this offensive but it is perplexing how anyone can understand what is being said when everyone is talking at the same time. I wish it was universal, and I am guessing it is true in most places.