Data & AI
Collecting data
Surveys, measurements, logs: how data comes into the world, and why the method shapes the result.
What you need first
Imagine two newspapers reporting on the same school. One writes that 80 percent of students are satisfied, the other says 45 percent. Who is lying? Maybe nobody. Maybe the two simply asked different questions, at different times, or reached different groups. Before believing any data, it is always worth asking: how was it actually collected?
Three ways to get data
Most data comes about in one of three ways. First, by asking: surveys and interviews record what people say. Second, by measuring: thermometers, scales and sensors record what instruments show. Third, by logging: logs automatically record what happens, such as every visit to a website or every purchase at the till. Each way has strengths and weaknesses. People can be mistaken or embellish, instruments can be badly , and logs only capture what the system sees at all.
Asking good questions
In surveys, the wording helps decide the result. The question 'Don't you agree that the break is far too short?' pushes the answer in one direction, which is called a leading question. More neutral would be: 'How do you rate the length of the break?' The answer options matter too: if 'don't know' is missing, undecided people are forced to pick a side they do not actually hold. Good questions are worded neutrally, easy to understand and offer all the answers that realistically occur. And then there is the question of how many of the people contacted answer at all. That share is called the response rate. If it is small, the result mainly shows the opinion of the few who bothered to reply.
No measurement is perfect
Measurements never deliver pure truth either. Random errors make values come out a little too high or a little too low, for instance because a hand trembles while reading a scale. Repeated measuring helps against this: the average of several measurements is more reliable than a single value. Systematic errors are trickier: a scale that always shows 2 kilograms too much still delivers the wrong average after a hundred measurements. The only cure for systematic errors is to check the instrument and subtract the known error.
The method shapes the result
Who is asked or measured, when, where and how, is baked into every result. A satisfaction survey right after a maths test turns out differently from one on the last day of school. A traffic count on Sunday morning says little about rush hour. This does not mean data is worthless. It just means: an honest result always states the method that produced it. Anyone who presents you with a number but will not reveal how it came about is demanding trust without foundation.
Exercises
0 of 6 solvedTime to try it yourself. You can't break anything, every attempt counts.
Which question is worded neutrally?
A scale shows exactly 2 kg too much on every measurement. What helps against this error?
In a survey 50 people are asked, 30 answer yes. What percentage is that?
Match each collection method to what it records.
A rod is measured three times: 12.1 cm, 11.9 cm and 12.0 cm. What is the average in centimetres?
A scale that shows 2 kg too much on every measurement makes a … error.
Where this leads