sciandu
Data & AI

Data & AI

Samples & bias

One spoonful reveals the whole soup, if you stir properly. Why samples work and when they lead you astray.

What you need first

To know whether a soup is salty enough, you do not have to eat the whole pot. One spoonful is enough, provided you stirred well beforehand. Election , quality control and drug trials work on exactly this principle: a small, well chosen part reveals a surprising amount about the whole. The catch lies in the stirring.

Why not just ask everyone?

Asking every person in a country is expensive, slow and often simply impossible. So you draw a sample: a small selection from the whole group you are interested in. Surprisingly, a good sample does not need to be huge. Around a thousand randomly chosen people already give usable estimates for an entire country, whether it has 8 million inhabitants or 80. What matters is not how big the sample is, but how it comes about.

The power of randomness

The golden rule is the random sample: every member of the whole group must have the same chance of being selected. Randomness is not the enemy here but the best ally, because it favours nobody. Picture a bag full of marbles in two colours. If you draw a few marbles blindly from a well mixed bag, your handful roughly mirrors the mix in the whole bag. Small deviations remain, but they shrink as the sample grows and can even be calculated: with a thousand people asked at random, the estimate usually stays within about 3 percentage points of the true value.

Marble bag
Blue3
Yellow5
P(Blue) = 3/8 = 37.5%
Draw marbles from the bag and watch how well even a small sample matches the true mix.

Systematic bias

Things get dangerous when the selection is not random but systematically favours or excludes certain groups. An online survey only reaches people who are online, and volunteers often hold particularly strong opinions. A survey about bicycle friendliness conducted only on a bike path wildly overstates the share of cyclists. Such distortions do not vanish with more data: a huge, skewed sample is worse than a small, clean one. A famous US election forecast failed for exactly this reason in 1936: a magazine counted more than 2 million returned ballots and still named the wrong winner. Its addresses came from telephone books and car registers, which in the middle of the Great Depression led mainly to wealthy households.

The planes that came back

The most famous example of an invisible bias comes from the Second World War. Returning bombers were examined for bullet holes, and most hits were found on the fuselage and wings. The obvious plan: more armour exactly there. The statistician Abraham Wald disagreed: only the planes that had made it back were being examined. Aircraft hit in the engine or cockpit barely appeared in the data, because they had crashed. So it is precisely the spots with few holes that need reinforcing. This fallacy is called survivorship bias: we only see the cases that made it into the data.

Exercises

0 of 6 solved

Time to try it yourself. You can't break anything, every attempt counts.

Why do we work with samples instead of asking everyone?

A survey about online shopping runs exclusively as a pop-up on a shopping website. What is the main problem?

From a well mixed bag 30 marbles are drawn, 12 of them are red. What percentage of the sample is red?

In a random sample, every member of the whole group must have the chance of being selected.

At a school with 800 students, 50 are asked at random, 20 of them cycle. How many cyclists do you estimate for the whole school?

Match each case to the fitting bias or consequence.

Only returning bombers examined
Survey held only on the bike path
Survey held only online