Complexity, and how to measure it

Good statistics is about making good decisions. Since we all make decisions, it’s important we can all get good intuitions about what the numbers mean. This article is designed to help you understand the big picture of complexity, so there won’t be any equations or jargon. The formal paper is published here, but it’s still recommended you read this first to get a sense of the big picture.

What is science?

Science is about identifying cause and effect, then using it to predict what happens next, and systems differ enormously in how easy that is. At one end of the spectrum are systems like a game of snooker. Strike the red cleanly with the cue ball and, as every snooker player learns, the two balls split off at a right angle. If you know how fast the cue ball was going and where it hit, Newtonian mechanics tells you exactly where both balls will go, so you can trace every effect back to its cause. We call this simple deterministic causality.

Now imagine hundreds of thousands of balls on a giant table. You can’t follow every collision, and it turns out you don’t need to. Knowing just how many balls there are and how much energy they have between them, statistics can predict how many will be moving at each speed, and the chart beside the table settles onto that prediction within moments. This is how we handle the air in a room: nobody tracks the molecules, but a thermometer and a pressure gauge tell you everything you need. We call this disconnected causality, and say the balls have linear interactions, because the effects of individual collisions simply add up and average out across the whole table.

How many balls are moving at each speed, slowest on the left The spread predicted from just the number of balls and their total energy

Then there are complex systems. By 1926 hunters had killed off the last wolves in Yellowstone National Park. With no predators to hassle them, the park’s elk and deer spent more time by the rivers, eating the young willows and aspens along the banks. Without those trees holding the soil together the banks eroded, and in places the rivers themselves changed course. Ecologists still argue about how far the chain reached, but the direction is clear: removing one predator reshaped the landscape.

Wolves removed Deer graze by the river Saplings eaten, bank erodes The river changes course

Unlike the giant snooker table, you couldn’t have predicted this from a few high-level measures like the number of animals or the average rainfall. The details matter, because a small change in one part of the system cascades into a large change somewhere else instead of averaging out. We say these systems have intertwined causality or non-linear interactions.

No system is simply complex or simple, though, since every system sits somewhere on a spectrum between the two. Knowing where matters, because most of our statistics quietly assumes a system sits at the snooker end, where a few averages describe the whole. If we could measure how complex a system is, we would know when those averages can be trusted.

How do we define complexity?

The clearest way into this is a thought experiment from the physicist Sean Carroll, which minutephysics turned into a 2-minute video. Pour cream onto black coffee and at first the two sit in separate layers, which you can describe in five words: cream on top, coffee below. Start stirring and the cream pulls out into swirls and streaks, and describing exactly where it all is would take pages. Keep stirring and everything blends into one even light brown liquid, which is easy to describe again.

Ordered, easy to describe Swirly, hard to describe Mixed, easy to describe again

The video uses this to give a working definition: complexity is roughly how hard something is to describe, which is another way of asking how much the details matter. The layered cup and the mixed cup are both simple, because a single sentence tells you everything about them. Only the swirly cup in the middle needs its details, in the same way the rivers of Yellowstone did.

What changes as you stir is the coffee’s entropy, which the next section explains properly. The layered cup has low entropy and the mixed cup has high entropy, and stirring only ever goes one way: a few turns of a spoon will mix the cream in, but no amount of stirring will separate it out again. Complexity appears on the way from one to the other, then fades away.

There have been many attempts to define complexity, but none that could actually be measured. So in 2012 the complexity scientists Carlos Gershenson and Nelson Fernández turned the coffee cup into a measure. Their score is zero for a system with the lowest possible entropy, zero again for one with the highest, and peaks exactly halfway between, just like the swirly cup. It was an important step, but as the rest of this article shows, it wasn’t enough on its own.

What is entropy and information?

To see why, we need to be completely clear what entropy is and how it links to information.

Imagine someone hands you a closed box. You know it has four spaces in a two by two square and holds four different coloured balls, but you can’t see inside. What you know from the outside is called the macro state. Each particular way the balls could be sitting inside is a micro state, and for this box there are 24 of them.

All 24 ways to arrange the four balls in the small box

Now picture a bigger box with 16 spaces in a four by four square, but the same four balls inside. Because the balls have so many more places they could be, the number of possible arrangements, or micro states, jumps from 24 to 43,680.

A few of the 43,680 ways to arrange them in the bigger box

Entropy measures how many micro states fit the macro state, which is the same as asking how uncertain you are about how it’s arranged inside. Guessing the setup of the small box is relatively easy (1 in 24). But in the big box it’s much harder (1 in 43,680). The big box has the higher entropy, not because anything inside it is messier, but because knowing the big picture tells you less about the details.

The same idea works for anything you can make guesses about, including text. If you’ve just read “AAAAAAAAAA” you’d bet the next letter is another A. Guessing what comes after “ABJSTOMXGW” is much harder, so that string has a higher entropy. This is also why entropy measures information. Seeing yet another A tells you nothing you hadn’t already guessed, while every letter of the random string is a surprise, and surprise is what carries information. Claude Shannon worked this out in 1948 while studying communication at Bell Labs, which is why this version is called information entropy.

The problem

We previously said that complexity was simply in the middle between low entropy and high entropy. Which means if we suddenly made a snooker table bigger, its entropy would definitely rise.

But does the system change? Does it actually momentarily become more complex?

No. The balls still interact in exactly the same way they did before.

As we’ve seen, complexity is about causality, where a small change can lead to a big one. So how do we tease apart what is happening in the coffee cups from snooker balls?

The solution

The trick is that entropy is all about the micro states, and we need to look at the collection of them (known as an ensemble). No matter how the balls are arranged, each time you expand the table, the way they spread will look (at a macro level) basically identical.

Now do the same with the coffee. Before stirring, every cup has the same neat layer of cream on top, so if you overlay them all, the average cup looks just like each other one. Once they’re fully mixed the same is true, since every cup is the same light brown liquid and so is their average.

Four cups, each stirred its own way, before stirring (top) and once mixed (bottom) All four overlaid

The middle is different. When we overlay them, the swirls blur into each other. The average looks much more like the fully mixed cup than like any of the cups that made it.

Four cups part way through stirring, each with its own swirls All four overlaid

A more mixed cup has a higher entropy. This is the key insight: complex averages have higher entropies. So we can subtract the entropy of the individual cups from the entropy of the averaged cup and measure exactly how much a system is complex. The paper covers the exact details, but this is the broad idea.

A key characteristic of complexity is whether the average of a system can accurately represent a typical instance of that system.

Practical example

Say you’re crossing the country by train to get to a wedding. There are two trains, each taking an hour, with a 10-minute change in the middle, and you need a taxi waiting at the other end, so you have to tell the driver when you’ll arrive.

On paper the journey takes 2 hours 10 minutes. The first train is usually a few minutes late, though, and if it’s more than 10 minutes late you miss the connection and wait an hour for the next one. On this line that happens about two times in five.

That makes this a complex system, even though nothing about it is complicated. A first train that is 9 minutes late and one that is 11 minutes late are almost the same train, but one gets you there at 2 hours 10 and the other at 3 hours 10, so two minutes at the start become an hour at the end.

The 10-minute window to make the connection Made the connection Missed it, and waited an hour for the next train The average journey time

Run the journey enough times and the arrival times pile up into two separate peaks, a tall one at 2 hours 10 minutes and a shorter one at 3 hours 10. The average comes to 2 hours 34 minutes, which lands in the empty gap between them.

The average journey time is one that almost never actually happens.

You know instinctively it would be stupid to book a taxi to pick you up based on the average journey time. You’d either be needlessly waiting around on the platform for 24 minutes, or the driver would be waiting around for 36 minutes for you. The optimal strategy is to be hopeful and book for 2 hours 10, the time you’ll arrive if all goes well, then rebook if you miss the connection.

This is a simple example, and so one most of us will get right through experience. However, in more complex situations it’s surprising how often we rely on statistics which don’t apply.

Real example

In the late 1940s the US Air Force had a serious problem: its pilots kept losing control of their planes. This was the start of the jet age, so the aircraft were faster and harder to fly than anything before, and the accidents piled up. At the worst point, 17 pilots crashed in a single day.

The first suspect was the pilots themselves, and many crashes were put down to pilot error. But the pilots were experienced, the training checked out, and when engineers went over the controls and electronics they found nothing wrong. So attention turned to the one thing nobody had questioned, the cockpit. Its layout dated from 1926, when engineers measured hundreds of pilots and built the seat, pedals and controls around their average. Perhaps, the thinking went, pilots had simply grown since then.

In 1950 researchers at Wright Air Force Base in Ohio set out to find the new average. They took 140 measurements of 4,063 pilots, from thumb length and crotch height to the distance between eye and ear. On the team was Lt. Gilbert Daniels, who had studied physical anthropology at Harvard. For his undergraduate thesis he had measured the hands of 250 students from very similar backgrounds and found that their hands were not similar at all, so he had his doubts about the average pilot.

Daniels took the ten measurements that matter most for a cockpit, such as height, chest size and sleeve length, and asked how many pilots were average on all ten. He was generous about what counted as average, allowing anyone in the middle 30% of the range for each measurement, and most of his colleagues expected the bulk of pilots to qualify. The answer was zero. Not one of the 4,063 pilots was average on all ten, and even on just three measurements fewer than 3.5% were.

The cockpit built for the average pilot had fitted no pilot at all. “The tendency to think in terms of the ‘average man’ is a pitfall into which many persons blunder,” Daniels wrote in his 1952 report. The human body is complex: your arm length is related to your leg length, but only loosely, so being average in one way tells you very little about the rest of you. Just like the taxi booked for the average journey, a design built for the average ends up suiting almost nobody.

The Air Force’s answer was to stop designing for the average. It told manufacturers that cockpits had to fit every pilot from the 5th to the 95th percentile on every measurement. The manufacturers said this would be too expensive, but the fixes turned out to be cheap: adjustable seats, adjustable foot pedals, and adjustable helmet straps and flight suits. Pilot performance improved, and the same idea is why the seat in your car slides back and forth today. The best cockpit, it turned out, was an adjustable one.

Business example

Say you run a silicon chip factory and you’re choosing between two expensive upgrades. You can only afford to build one, so before you commit you commission a digital twin of the factory, a simulation detailed enough to tell you how many minutes each chip will take to make under each option. Today it takes about 100.

Current setup Option A Option B
Each line is one run of the digital twin All five runs averaged together The average time, in minutes

The tempting thing is to run the twin once for each option. If you had, Option B might well have come back at 83 minutes, a 17% improvement, and you would have signed off on it that afternoon. You would only have found out after building it what the result usually is.

So you do the careful thing and run each option five times, then compare them the way any analyst would, by the average and the spread. Option B averages 94 minutes against Option A’s 96, and its spread is tighter too. Every standard number points to B.

Now look at the individual runs. Option A’s five runs sit almost on top of each other, so whichever one turns out to be the real factory, you get about 96 minutes. Option B’s runs disagree with each other. One is brilliant, and the other four land between 93 and 99.5, which is no better than A and at worst barely better than what you have today. Its average of 94 falls in a gap that none of the runs actually produced, just like the average journey that no train passenger ever takes.

This matters because you only get to build one factory, which means you only get one run. Choosing B is a one-in-five bet on a big win, and otherwise you’ve paid for an upgrade that delivers what A would have, or less. When customers are counting on the chips you’ve promised them, knowing what you’ll get is worth far more than a couple of minutes on paper, so Option A is the better choice.

Traditional statistical tests, such as ANOVA or chi-squared, wouldn’t have caught this. They are built to decide whether two distributions differ, not how much, and they are poor at spotting systems where an occasional run behaves completely differently from the rest. The same problem helps explain why the replication crisis has hit complex fields like psychology so hard: a single study is a single run, and in an incoherent system the next run can tell a different story.

Taking the average and variance is necessary but not sufficient.

Naming it Incoherence

We call this Incoherence. First, because complexity is such an overloaded term. But also because to be coherent means being consistent and understandable. This is echoed in physics, where coherent waves align with each other.

The link with frequentist statistics

Most of the statistics we learn at school is frequentist. It says the probability of something is how often it happens when you repeat it many times, so a coin is fair if, over enough flips, half of them land heads. To get those numbers you repeat the experiment, put all the results together and read off the average, exactly as we did with the train journeys and the factory’s digital twin.

This relies on the assumption we flagged right at the start, that the average describes a typical case. With an ordinary coin it does, because every coin behaves like every other, so throwing all their flips into one pile loses nothing. With the trains it didn’t, because a journey that makes the connection and one that misses it are so different that their average matches neither.

Imagine a sticky coin. Its first flip is a fair 50:50, but whichever way it lands it stays stuck, so every flip after that comes up the same. Flip it 100 times and you get either 100 heads or 100 tails, never anything in between. Now hand out a thousand of these coins, have everyone flip theirs 100 times, and put all the flips together. You get an almost perfect 50:50 split, so a standard test would call the coin fair.

In one sense it is fair, because before the first flip heads and tails really are equally likely. But once you’re holding a particular coin, 50:50 tells you nothing useful. Like the average train journey, it’s an outcome that never actually happens. A system like this, where following one coin over time gives a different answer from averaging across many coins, is called non-ergodic, and it’s exactly where frequentist numbers mislead.

Incoherence catches this with the same subtraction we used on the coffee cups. Each sticky coin on its own has no uncertainty at all, because once you’ve seen its first flip you know every flip after it, so its entropy is zero. All the coins put together are as uncertain as a coin can be, so the gap between them is as large as it can get. Ordinary coins are the opposite: each one looks like the whole pile, and their Incoherence is zero.

Most real systems sit somewhere in between, like the trains, and that is where Incoherence goes further than ergodicity. Ergodicity is a yes-or-no property: a system either is ergodic or it isn’t, and almost nothing in the real world passes the test perfectly. Incoherence is a measure, so instead of just telling you the average is off, it tells you how far off it is, and whether it’s close enough to rely on.

Incoherence tells you directly how much you can rely on frequentist statistics.

Thanks

I want to give particular thanks to Carlos Gershenson, who first proposed the 2012 measure of complexity that this work builds on. Carlos also generously edited the special issue of Entropy in which the Incoherence paper was published.