Everything falls into step
Clapping crowds, two clocks on a wall, a riverbank of fireflies, a bridge in London and a rack of routers all do the same strange thing without being told. Nature wants it. Computers really don’t.
You know the end of a good concert. Everyone is clapping, and it sounds like rain on a tin roof: thousands of hands, no pattern at all. Then, a few seconds later, without anyone deciding anything, the whole hall is clapping together. Clap. Clap. Clap.
I always assumed somebody starts it. One loud person in the front row sets a beat and the rest follow.
Nobody starts it.
In 2000 a group of physicists put microphones in theatres in Romania and
Hungary and recorded what audiences actually do
I went looking for why this happens, and found the same behaviour in a seventeenth-century bedroom, in mangrove swamps, in your chest, on a bridge over the Thames and inside the internet. In some of those places it’s the whole point. In the others, engineers spend real money stopping it.
Two clocks on a beam
February 1665. Christiaan Huygens, the man who invented the pendulum clock, is sick in bed, staring at two of his clocks hanging from the same wooden beam. He notices their pendulums are swinging in perfect opposition: one goes left exactly when the other goes right.
He disturbs one. Within half an hour they’re back in opposition. He
moves the clocks apart, and the effect goes away. He writes to his father
about this “odd kind of sympathy” and guesses, correctly, that the beam was
doing it
Each swing gives the beam a tiny shove. The other clock feels that shove and is nudged a little earlier or a little later. Neither clock is in charge. Each one just keeps getting nudged by the other, a tiny bit, every swing, and after a thousand or so swings the nudges have pulled them into a shared rhythm.
That’s the whole recipe, and it’s going to keep showing up:
- Things with their own rhythm: a pendulum, a hand, a heart cell, a router.
- A weak way for each one to feel the others: a beam, a sound, a wire.
Nothing else is needed. No leader, no plan, no signal saying “now”.
Trees that blink
In the mangroves of Thailand and Malaysia there are trees where thousands of fireflies flash at the same moment, over and over, all night. The whole tree lights up, goes dark, lights up.
For a long time Western scientists didn’t believe the travellers who
described this. Insects don’t coordinate. One letter to Science in 1917
called it “certainly contrary to all natural laws” and decided the observer
must be blinking
It’s real. Each firefly has an internal clock that winds up and fires a flash when it runs out. When a firefly sees a neighbour flash, its own clock jumps forward a little. That’s the beam again, made of light this time.
In 1990 two mathematicians, Renato Mirollo and Steven Strogatz, proved
something stronger than “this can happen”. For clocks that work like this,
where seeing a flash pushes you closer to your own, a population where
everyone can see everyone ends up flashing in unison from almost any
starting point
So the fireflies aren’t being clever. They can’t avoid it.
The equation for all of them
A Japanese physicist, Yoshiki Kuramoto, found the cleanest way to write this down in 1975. Give every oscillator a phase , the position of its hand around a clock face, and a natural speed . Then let each one be pulled towards the others:
The first term is “run at your own pace”. The second is the beam. If a neighbour is slightly ahead of you, is positive and you speed up; if it’s behind, you slow down. is how strongly everyone feels everyone else.
What surprised me is what happens as you turn up. You’d expect the crowd to drift into sync gradually: a little coupling, a little order. It doesn’t. Below a certain coupling nothing happens at all, everyone runs at their own speed and the population looks like noise. Past a critical value a cluster suddenly forms, grabs its neighbours and grows. It’s closer to water freezing than to a dimmer switch.
And depends on exactly one other thing: how different the oscillators
are to begin with. If their natural speeds are spread out more, you need
stronger coupling to pull them together
Here’s a field of them. Each one only sees its nearest neighbours, the way a firefly sees the bushes around it. Turn the coupling up slowly and watch for the moment it tips. Then drag across them to scare a patch out of step, and see how long the rest take to pull it back.
This figure is interactive.
Two knobs decide everything: how strongly things feel each other, and how different they are. Keep both in mind, because every story from here on is somebody turning one of them.
The bridge that made people walk in step
On 10 June 2000 London opened the Millennium Bridge, a thin steel footbridge across the Thames. Thousands of people walked over it on the first day. It started swaying sideways, enough that people grabbed the railings. Two days later it was closed.
The engineers at Arup worked out what had happened, and it’s Huygens’ clocks again, with people as the pendulums and the bridge as the beam.
When you walk, your weight shifts a little left, a little right, every step. On solid ground that does nothing. On a bridge that can sway, a tiny sideways movement makes you widen your stance and time your steps to the sway, without thinking about it, to stay balanced. Now your steps push the bridge in time with its own motion, which makes it sway more, which makes more people adjust.
Arup tested this with crowds walking across the closed bridge, adding people
a few at a time. The wobble didn’t grow gradually. Below a certain crowd size
the bridge was fine; past it, the sway switched on. That’s , measured in
pedestrians. Strogatz and colleagues later modelled the bridge with exactly
the equation above and got the same threshold
The fix was dozens of dampers under the deck, about £5 million worth, and the bridge reopened in 2002. Look at what the dampers do in Kuramoto’s terms: they don’t change the walkers at all. They make the bridge move less, so each walker feels everyone else less. They turned down .
The internet falls into step too
Here’s where it gets uncomfortable for people like me. Computers are full of things with their own rhythm, and they’re all connected to each other. That’s the recipe.
Routers
In the early 1990s, Sally Floyd and Van Jacobson were looking at a boring
test: one ping a second between Berkeley and MIT. The pings kept getting lost
in bursts, and the bursts arrived on a regular beat
The culprit was routing updates. Routers running RIP send each other their routing tables on a timer, every 30 seconds. Each router set its timer independently, so you’d expect the updates to be spread out. But when a router received an update from a neighbour, it spent time processing it before resetting its own timer. A router that was slightly ahead got delayed by the busy one next to it. That’s a nudge. Floyd and Jacobson showed that a network of routers, each nudging the others a little, falls into step, until every router sends its update in the same instant and the links choke on all of them at once.
The best part of their paper is how it happens. The change is abrupt: in their model, adding a single router could flip a network from completely unsynchronised traffic to completely synchronised traffic. That’s again, measured in routers this time.
Their fix was to add randomness to the timer. The RIP specification now
says it plainly: when many routers share a network, “there is a tendency for
them to synchronize with each other”, so every time a router sets its
30-second timer it adds a random offset of up to 5 seconds either
way
TCP
The same two people found it again in how TCP shares a congested link. When a router’s queue fills up, the old behaviour was to drop whatever arrives next. That hits every connection passing through at the same moment. Every connection then halves its speed at the same moment, the link goes quiet, everyone speeds up together, the queue fills, and round it goes. It’s called global synchronisation, and it wastes a lot of the link.
Their answer, Random Early Detection, starts dropping packets before the
queue is full, and picks which ones at random
Two computers on one wire
Old Ethernet put every computer on a shared cable. If two of them talked at once, both messages were garbled, so both had to wait and try again.
Wait how long? If both machines used the same rule, say “wait one millisecond”, they’d retry at the same instant, collide again, wait the same again, and collide forever. Two identical machines following an identical rule are perfectly synchronised, which is the worst thing they can be.
So the rule has a dice roll in it. After the th collision in a row, a
machine waits a random number of time slots between and
(capped at )
The herd at the door
This is the version you’re most likely to cause yourself.
A service goes down for a second. Every client that was talking to it gets an error at the same moment. Every client has been written, sensibly, to retry after a short wait, and double the wait each time it fails. Every client uses the same numbers.
You can see where this goes. They all failed together, so they all retry together. The service comes back up and immediately gets hit by every client in the same millisecond, far more than it can answer. Most get refused. They all wait twice as long, together, and hit it again, together.
This figure is interactive.
With no jitter the server spends most of its time idle and the rest of it
drowning, and some clients are still waiting when the chart runs out. Notice
how little randomness it takes to fix it. In this simulation, randomising a
quarter of each wait cuts the refusals after the outage from 660 to 83, and
randomising half of it cuts them to 5. Marc Brooker at AWS wrote the well-known
explanation of this, and put it in two sentences: “The solution isn’t to
remove backoff. It’s to add jitter.”
The same thing happens with time itself. Ask a room of engineers to schedule
an hourly job and most of them will pick minute zero, so plenty of servers
see a spike at the top of every hour from jobs that have nothing to do with
each other. That’s why systemd timers have a setting called
RandomizedDelaySec: its whole purpose is to make your job run at a slightly
different time from everyone else’s.
Two knobs
I keep coming back to how small the toolbox is. Every story here is one of two moves:
| turn down (feel each other less) | widen the spread (be more different) | |
|---|---|---|
| Millennium Bridge | dampers | |
| Routing updates | random timer offsets | |
| TCP queues | random early drops | |
| Ethernet | random backoff slots | |
| Retry storms | jitter | |
| Hourly jobs | randomised start times |
Engineers almost never get to turn down the coupling, because the coupling is the job. Routers must talk to routers, clients must reach servers. So we reach for the other knob, and the cheapest way to make machines different from each other is to let them roll dice.
Nature mostly goes the other way. Fireflies evolved to watch each other. Heart cells are wired together so tightly that they fire as one. When synchrony is what you want, you turn up until nothing can resist it.
What I like about this is that both sides are working with the same equation. A firefly on a riverbank in Malaysia and a retry loop in a payment service are solving the same problem from opposite ends: one is trying as hard as it can to fall into step, and the other is rolling a die every time it waits, so that it never does.