Everything falls into step

Clapping crowds, two clocks on a wall, a riverbank of fireflies, a bridge in London and a rack of routers all do the same strange thing without being told. Nature wants it. Computers really don’t.

· 12 min read

You know the end of a good concert. Everyone is clapping, and it sounds like rain on a tin roof: thousands of hands, no pattern at all. Then, a few seconds later, without anyone deciding anything, the whole hall is clapping together. Clap. Clap. Clap.

I always assumed somebody starts it. One loud person in the front row sets a beat and the rest follow.

Nobody starts it.

In 2000 a group of physicists put microphones in theatres in Romania and Hungary and recorded what audiences actually doZ. Néda, E. Ravasz, Y. Brechet, T. Vicsek and A.-L. Barabási, “The sound of many hands clapping”, Nature 403 (2000). They found synchronised clapping runs at roughly double the period of wild clapping, and that the loss of loudness is what pushes audiences out of sync again.. The hall drifts into rhythm on its own. And then something odd happens: synchronised clapping is slower than wild clapping, almost exactly half the pace, so the room gets quieter. People want to be loud, so they speed up. Speeding up breaks the rhythm. The rhythm comes back a few seconds later. A crowd will go round this loop several times in one encore.

I went looking for why this happens, and found the same behaviour in a seventeenth-century bedroom, in mangrove swamps, in your chest, on a bridge over the Thames and inside the internet. In some of those places it’s the whole point. In the others, engineers spend real money stopping it.

Two clocks on a beam

February 1665. Christiaan Huygens, the man who invented the pendulum clock, is sick in bed, staring at two of his clocks hanging from the same wooden beam. He notices their pendulums are swinging in perfect opposition: one goes left exactly when the other goes right.

He disturbs one. Within half an hour they’re back in opposition. He moves the clocks apart, and the effect goes away. He writes to his father about this “odd kind of sympathy” and guesses, correctly, that the beam was doing itHuygens described it in a letter to his father in February 1665. The clocks settled into anti-phase, swinging opposite each other, which is the arrangement where their shoves on the beam cancel out..

Each swing gives the beam a tiny shove. The other clock feels that shove and is nudged a little earlier or a little later. Neither clock is in charge. Each one just keeps getting nudged by the other, a tiny bit, every swing, and after a thousand or so swings the nudges have pulled them into a shared rhythm.

That’s the whole recipe, and it’s going to keep showing up:

  1. Things with their own rhythm: a pendulum, a hand, a heart cell, a router.
  2. A weak way for each one to feel the others: a beam, a sound, a wire.

Nothing else is needed. No leader, no plan, no signal saying “now”.

In the mangroves of Thailand and Malaysia there are trees where thousands of fireflies flash at the same moment, over and over, all night. The whole tree lights up, goes dark, lights up.

For a long time Western scientists didn’t believe the travellers who described this. Insects don’t coordinate. One letter to Science in 1917 called it “certainly contrary to all natural laws” and decided the observer must be blinkingPhilip Laurent, in Science (1917). Steven Strogatz tells this story, along with most of the biology here, in his book Sync (2003), which is the best long read on the subject..

It’s real. Each firefly has an internal clock that winds up and fires a flash when it runs out. When a firefly sees a neighbour flash, its own clock jumps forward a little. That’s the beam again, made of light this time.

In 1990 two mathematicians, Renato Mirollo and Steven Strogatz, proved something stronger than “this can happen”. For clocks that work like this, where seeing a flash pushes you closer to your own, a population where everyone can see everyone ends up flashing in unison from almost any starting pointR. Mirollo and S. Strogatz, “Synchronization of pulse-coupled biological oscillators”, SIAM Journal on Applied Mathematics (1990). The proof assumes every oscillator sees every other one and they’re all identical; real fireflies are messier, but the tendency survives.. They built on a model Charles Peskin had made in 1975 for something closer to home: the ten thousand or so pacemaker cells in your heart that have to fire together for it to beatC. Peskin, Mathematical Aspects of Heart Physiology (1975). He conjectured that identical pulse-coupled oscillators always synchronise; Mirollo and Strogatz proved it fifteen years later..

So the fireflies aren’t being clever. They can’t avoid it.

The equation for all of them

A Japanese physicist, Yoshiki Kuramoto, found the cleanest way to write this down in 1975. Give every oscillator a phase θi, the position of its hand around a clock face, and a natural speed ωi. Then let each one be pulled towards the others:

dθidt=ωi+KN∑j=1Nsin⁡(θj−θi)

The first term is “run at your own pace”. The second is the beam. If a neighbour is slightly ahead of you, sin⁡(θj−θi) is positive and you speed up; if it’s behind, you slow down. K is how strongly everyone feels everyone else.

What surprised me is what happens as you turn K up. You’d expect the crowd to drift into sync gradually: a little coupling, a little order. It doesn’t. Below a certain coupling nothing happens at all, everyone runs at their own speed and the population looks like noise. Past a critical value Kc a cluster suddenly forms, grabs its neighbours and grows. It’s closer to water freezing than to a dimmer switch.

And Kc depends on exactly one other thing: how different the oscillators are to begin with. If their natural speeds are spread out more, you need stronger coupling to pull them togetherFor natural frequencies drawn from a bell-shaped distribution g(ω), Kuramoto showed the critical coupling is Kc=2/(πg(0)): a wider, flatter distribution has a smaller g(0) and needs a bigger K..

Here’s a field of them. Each one only sees its nearest neighbours, the way a firefly sees the bushes around it. Turn the coupling up slowly and watch for the moment it tips. Then drag across them to scare a patch out of step, and see how long the rest take to pull it back.

This figure is interactive.

A hundred and twenty-eight fireflies, each with a slightly different natural rhythm, each watching its eighteen nearest neighbours. Drag the slider to raise the coupling; drag across the field to scramble the ones under your pointer.

Two knobs decide everything: how strongly things feel each other, and how different they are. Keep both in mind, because every story from here on is somebody turning one of them.

The bridge that made people walk in step

On 10 June 2000 London opened the Millennium Bridge, a thin steel footbridge across the Thames. Thousands of people walked over it on the first day. It started swaying sideways, enough that people grabbed the railings. Two days later it was closed.

The engineers at Arup worked out what had happened, and it’s Huygens’ clocks again, with people as the pendulums and the bridge as the beam.

When you walk, your weight shifts a little left, a little right, every step. On solid ground that does nothing. On a bridge that can sway, a tiny sideways movement makes you widen your stance and time your steps to the sway, without thinking about it, to stay balanced. Now your steps push the bridge in time with its own motion, which makes it sway more, which makes more people adjust.

Arup tested this with crowds walking across the closed bridge, adding people a few at a time. The wobble didn’t grow gradually. Below a certain crowd size the bridge was fine; past it, the sway switched on. That’s Kc, measured in pedestrians. Strogatz and colleagues later modelled the bridge with exactly the equation above and got the same thresholdS. Strogatz, D. Abrams, A. McRobie, B. Eckhardt and E. Ott, “Crowd synchrony on the Millennium Bridge”, Nature 438 (2005)..

The fix was dozens of dampers under the deck, about £5 million worth, and the bridge reopened in 2002. Look at what the dampers do in Kuramoto’s terms: they don’t change the walkers at all. They make the bridge move less, so each walker feels everyone else less. They turned down K.

The internet falls into step too

Here’s where it gets uncomfortable for people like me. Computers are full of things with their own rhythm, and they’re all connected to each other. That’s the recipe.

Routers

In the early 1990s, Sally Floyd and Van Jacobson were looking at a boring test: one ping a second between Berkeley and MIT. The pings kept getting lost in bursts, and the bursts arrived on a regular beatS. Floyd and V. Jacobson, “The Synchronization of Periodic Routing Messages”, SIGCOMM 1993, and the longer version in IEEE/ACM Transactions on Networking (1994)..

The culprit was routing updates. Routers running RIP send each other their routing tables on a timer, every 30 seconds. Each router set its timer independently, so you’d expect the updates to be spread out. But when a router received an update from a neighbour, it spent time processing it before resetting its own timer. A router that was slightly ahead got delayed by the busy one next to it. That’s a nudge. Floyd and Jacobson showed that a network of routers, each nudging the others a little, falls into step, until every router sends its update in the same instant and the links choke on all of them at once.

The best part of their paper is how it happens. The change is abrupt: in their model, adding a single router could flip a network from completely unsynchronised traffic to completely synchronised traffic. That’s Kc again, measured in routers this time.

Their fix was to add randomness to the timer. The RIP specification now says it plainly: when many routers share a network, “there is a tendency for them to synchronize with each other”, so every time a router sets its 30-second timer it adds a random offset of up to 5 seconds either wayRFC 2453, the RIP version 2 specification, section 3.8, which also notes the synchronised updates “can lead to unnecessary collisions on broadcast networks”.. In Kuramoto’s terms, they couldn’t remove the coupling (routers have to talk), so they widened the spread: they made the routers deliberately different from each other.

TCP

The same two people found it again in how TCP shares a congested link. When a router’s queue fills up, the old behaviour was to drop whatever arrives next. That hits every connection passing through at the same moment. Every connection then halves its speed at the same moment, the link goes quiet, everyone speeds up together, the queue fills, and round it goes. It’s called global synchronisation, and it wastes a lot of the link.

Their answer, Random Early Detection, starts dropping packets before the queue is full, and picks which ones at randomS. Floyd and V. Jacobson, “Random Early Detection Gateways for Congestion Avoidance”, IEEE/ACM Transactions on Networking (1993).. Different connections get the bad news at different times, so they back off at different times. More dice.

Two computers on one wire

Old Ethernet put every computer on a shared cable. If two of them talked at once, both messages were garbled, so both had to wait and try again.

Wait how long? If both machines used the same rule, say “wait one millisecond”, they’d retry at the same instant, collide again, wait the same again, and collide forever. Two identical machines following an identical rule are perfectly synchronised, which is the worst thing they can be.

So the rule has a dice roll in it. After the nth collision in a row, a machine waits a random number of time slots between 0 and 2n−1 (capped at 210−1)This is truncated binary exponential backoff from the IEEE 802.3 standard. A machine gives up after sixteen collisions in a row.. The doubling makes room as collisions pile up. The randomness is what actually separates the two machines.

The herd at the door

This is the version you’re most likely to cause yourself.

A service goes down for a second. Every client that was talking to it gets an error at the same moment. Every client has been written, sensibly, to retry after a short wait, and double the wait each time it fails. Every client uses the same numbers.

You can see where this goes. They all failed together, so they all retry together. The service comes back up and immediately gets hit by every client in the same millisecond, far more than it can answer. Most get refused. They all wait twice as long, together, and hit it again, together.

This figure is interactive.

Two hundred and forty clients fail at once when a server goes down for a second, then retry with exponential backoff. Green is served, red is refused, and the dashed line is how much the server can take. Drag the jitter up to randomise each wait.

With no jitter the server spends most of its time idle and the rest of it drowning, and some clients are still waiting when the chart runs out. Notice how little randomness it takes to fix it. In this simulation, randomising a quarter of each wait cuts the refusals after the outage from 660 to 83, and randomising half of it cuts them to 5. Marc Brooker at AWS wrote the well-known explanation of this, and put it in two sentences: “The solution isn’t to remove backoff. It’s to add jitter.”Marc Brooker, “Exponential Backoff And Jitter”, AWS Architecture Blog (2015). His simulations compare several jitter schemes; the one he calls “full jitter” waits a uniformly random time between zero and the backoff ceiling.

The same thing happens with time itself. Ask a room of engineers to schedule an hourly job and most of them will pick minute zero, so plenty of servers see a spike at the top of every hour from jobs that have nothing to do with each other. That’s why systemd timers have a setting called RandomizedDelaySec: its whole purpose is to make your job run at a slightly different time from everyone else’s.

Two knobs

I keep coming back to how small the toolbox is. Every story here is one of two moves:

turn down K (feel each other less) widen the spread (be more different)
Millennium Bridge dampers
Routing updates random timer offsets
TCP queues random early drops
Ethernet random backoff slots
Retry storms jitter
Hourly jobs randomised start times

Engineers almost never get to turn down the coupling, because the coupling is the job. Routers must talk to routers, clients must reach servers. So we reach for the other knob, and the cheapest way to make machines different from each other is to let them roll dice.

Nature mostly goes the other way. Fireflies evolved to watch each other. Heart cells are wired together so tightly that they fire as one. When synchrony is what you want, you turn K up until nothing can resist it.

What I like about this is that both sides are working with the same equation. A firefly on a riverbank in Malaysia and a retry loop in a payment service are solving the same problem from opposite ends: one is trying as hard as it can to fall into step, and the other is rolling a die every time it waits, so that it never does.

Worth passing on?