sciandu
Data & AI

Data & AI

Distributions

Two cities, the same average, completely different weather. Why a single number never tells the whole story.

The average temperature is 15 degrees. Do you pack swimming trunks or a winter jacket? The honest answer: you do not know. Maybe the weather swings between 14 and 16 degrees, maybe between 0 and 30. A single number only tells you the middle, not how the values are spread around that middle. That is what distributions are about, and once you can read them you see far more in the very same numbers.

Nobody is the average

A statistic reports 1.4 cars per household and 1.5 children per family. Of course no household anywhere owns 1.4 cars, and half a child does not exist at all. The average is a calculated quantity, not a real case. In reality there are households with zero, one, two or three cars, and the distribution tells you how many of each kind there are. If you only know the average, you do not know a single real household.

Spread: how far apart the values sit

Back to the two cities. In city A you measure 14, 15 and 16 degrees on three days, in city B 5, 15 and 25 degrees. Both end up with the same mean, which is just another word for the average, because the mean is nothing but the sum of all values divided by how many there are: 45 divided by 3 gives 15 degrees. Yet life there feels completely different. The simplest measure of this difference is the range: largest value minus smallest value. City A has a range of 2 degrees, city B a range of 20 degrees. It gets more precise when every value gets a say: for the average deviation you take each value's distance from the mean, always without its sign, and divide the sum of those distances by how many values there are. For the four daily values 12, 14, 16 and 18 degrees with a mean of 15, the distances are 3, 1, 1 and 3, together 8, divided by 4 that is 2 degrees. The bigger the spread, the less the mean tells you about any single day.

Data bars
BarAnna
Mean 11
AnnaBenCemDanaEli
Anna: 12mean: 11
Try it: build two data sets with the same mean, one tightly packed and one widely spread, and compare the picture.

The bell curve

Many quantities in nature follow a famous pattern: most values crowd around the middle, and the further you move away from the middle, the rarer they get. Drawn out, this gives a bell shape. The heights of adult women are the classic example: very many are roughly average height, and very few are extremely short or extremely tall. Once you know a quantity is bell shaped, you instantly know that extreme values are possible but rare.

Why this makes you smarter

Whenever someone quotes only an average, you can now ask the crucial follow up question: how much do the values spread? A medicine that works well on average may work strongly for some people and not at all for others. A school with a good average grade may consist of uniformly decent classes or of very strong and very weak ones. Same mean, completely different story. The distribution is the story.

Exercises

0 of 6 solved

Time to try it yourself. You can't break anything, every attempt counts.

Two cities have the same average temperature. Can their weather still differ a lot?

Five daily values: 12, 15, 18, 21 and 24 degrees. What is the range in degrees?

A machine delivers exactly 5 grams four times: 5, 5, 5, 5. What is the range in grams?

In a bell curve, most values crowd around the .

In a bell curve, where do most values sit?

All three data sets have a mean of 15. Order them by their range, from smallest to largest.

  1. 110, 15, 20 (range 10)
  2. 214, 15, 16 (range 2)
  3. 35, 15, 25 (range 20)