Tutorial

Critical Thinking: Statistics

How to avoid being fooled by numbers — inspired by Daniel Levitin's A Field Guide to Lies and Statistics.

Before you check anyone's sources, sample sizes, or spreadsheets, there is a much cheaper test: how likely is this to be real at all? Plausibility is the smell test you run in your head before doing any actual maths. Most misleading statistics don't need debunking with data — they fall over the moment you ask what the world would have to look like for them to be true.

Plausible vs not plausibleIf someone tells you a glass fell onto a rug and didn't break, that's plausible, that would appear pretty plausible. If they tell you a glass fell off the top of a building and didn't smash, you don't need to run an experiment to doubt the claim. It's clearly bullocks.

Statistics work the same way. A number can be delivered with total confidence, printed in a serious-looking report, and still describe a world that cannot exist. The first step to thinking critically should always be is it even remotely plausible.

You spot this in a newspaper. What would you make of this claim?

The Daily Figure · Health, page 4

In the last 35 years, alcohol-related deaths in Ireland have doubled every year

A few pages later, the same paper runs a poll. Anything bothering you about this chart?

The Daily Figure · Reader poll, page 11

Which Newspaper do people trust most?

70%63%60%
The Daily Fail70%
The Smellygraph63%
The Boredian60%

One more from the same paper — but careful, the test cuts both ways.

The Daily Figure · Markets, page 2

ACME has lost 95% of its value

The lessonIn A Field Guide to Lies and Statistics, Daniel Levitin's first step for evaluating any number is exactly this: check it for plausibility before doing anything else. If a claim implies billions of dead, infinite growth, a pie that adds up to 193%, or a glass that survives a ten-storey drop, no amount of official-sounding sourcing should rescue it.

"The average" sounds like a single, settled fact. It isn't — it's a choice between three different numbers, and the choice changes the story. When a report says average without saying which one, it almost always means the mean.

MeanAdd every value up and divide by the count. It uses all the data — which also means a single extreme value drags it up or down.
MedianLine everyone up in order and take the person in the middle. Half sit below, half above.
ModeThe value that appears most often.
Worked exampleFive salaries: €24k, €31k, €31k, €38k and €96k. The mean is €44k — higher than what four of the five people actually earn. The median is €31k (the middle person), and the mode is also €31k (it appears twice). Same room, three different "averages".

Income is the classic case. Because a handful of very high earners can pull the mean far above what a typical person makes (and unusually low values can drag it the other way), the mean is easily skewed — while the median simply shows the person in the middle. That's why honest reporting about salaries, house prices, or wealth usually quotes the median. Try it yourself:

You may have heard a statistic that is casually thrown around: that if the UK were a US state, it would rank 51st in terms of GDP per capita. This would rank the UK below a state like Mississippi, which is the poorest in the US. What is actually happening here is that the statisticians responsible for this data used the mean average rather than the median. There are a handful of very wealthy people in Mississippi, while around 18% of residents live below the poverty line.

One room, two averagesNine people in a room. Drag the highest earner's salary up and watch which "average" follows.
Highest earner€60,000
Mean ("the average")€39,556
Median (middle person)€38,000

Nine ordinary salaries — mean and median tell the same story.

There's one more way an average can mislead: when the data has two "typicals". Lunch spending in the City of London is a classic bimodal distribution — a tall spike of supermarket meal deals around £3–£5, and a second, broader bulge around £35–£40 where businesspeople are taking clients out or dining at higher-end restaurants. Any single "average" has to land somewhere between the two humps.

Lunch in the City of London1,000 lunches: a spike of supermarket meal deals, a bulge of client lunches. Drag to change the mix and watch where the "average" lands.
MEAL DEALSCLIENT LUNCHESMEAN £19.61MEDIAN £8.39£0£10£20£30£40£50£60spend per lunch · bar height = number of people
Expense accounts45%
Mean ("the average")£19.61
Median£8.39

The "average lunch" is now £19.61. Point at the chart where those people are — there aren't any.

The lessonWhen a headline says "average", ask which average. On skewed data like income, the mean chases the outliers while the median stays with the middle person — so the same room can be described as ordinary or wealthy depending on which number was picked. Neither is a lie; quoting the flattering one without saying which is. And when a distribution has two humps, like City lunches, no single average describes anyone at all — look at the shape before trusting the number.

A chart can look precise while withholding the information needed to read it. Before drawing conclusions from its shape, check what each axis actually represents.

What is wrong with this chart?

A report uses this chart to celebrate an improvement. Does the second result really tower over the first?

27%29%31%33%35%28%Before35%After

Average house prices rose from £125,000 to £250,000 over 20 years. What impression does this 100-year chart create?

Average house price · 1925–2025£0k£62.5k£125k£187.5k£250k192519451965198520052025£125k£250k

Probability is how we describe uncertainty without pretending it has disappeared. Levitin separates several meanings that are often bundled together: the known symmetry of a fair coin, the frequency with which something happens across many observations, and a judgement about how likely a future event seems. Trouble begins when we switch between those meanings or reverse the condition in a probability statement.

ClassicalAll outcomes are known and equally likely, such as either side of an ideal fair coin.
FrequencyEstimate likelihood from how often an outcome appears across many comparable observations.
SubjectiveA reasoned estimate based on the evidence currently available, revised when new evidence arrives.

A fair coin has landed heads five times in a row. What is most likely on the sixth toss?

HHHHH?
Conditional probability does not run backwards

Suppose a report says that many road accidents happen during rush hour. That describes the probability of rush hour given that an accident happened. It does not tell us the probability of an accident given that it is rush hour.

Given an accidentWas it during rush hour?
Given rush hourWill this journey end in an accident?

Rush hour contains far more journeys than quieter periods, so it can contain many accidents while any individual journey remains very unlikely to end in one. Always ask which fact comes after the word given.

A headline says: "93% of women with breast cancer are in the high-risk group." What is the chance that a woman in the high-risk group has breast cancer?

The lessonA probability is incomplete until you know what is being conditioned on. Do not reverse "A given B" into "B given A", and when percentages become slippery, turn them into counts out of 100 or 1,000. A simple table often exposes what an alarming headline leaves out.

Data does not have an agenda, but the person choosing which data you see may have one. A political campaign may want your vote. A company may want your money. A pressure group may want attention, donations, or action. The easiest way to move people is often to make them feel something first — fear, anger, pride, relief, or urgency — and let the number provide a coat of authority.

SourceWho selected and presented the number?
FeelingWhat emotion does the framing encourage?
ActionWhat does the source want you to do next?
OmissionWhich comparison would weaken the story?

Two campaigns describe the same fictional labour-market report. Which headline can be true?

Illustrative dataset
People in work30.0m → 30.2m
Unemployment rate4.0% → 4.3%
Campaign in government"200,000 more people are now in work"Reassurance · confidence
Opposition campaign"Unemployment rises to 4.3%"Concern · anger

An investment advert gives you one impressive number. Is it enough to judge the product?

Illustrative investment advertOur fund returned 12% last yearDon't miss the next rise
The lessonAsking who benefits is a prompt to investigate, not permission to dismiss inconvenient evidence. A motivated source can publish accurate data, and a neutral-looking source can still make mistakes. Check the complete dataset, relevant comparisons, definitions, and methods — then decide whether the emotional story survives the context.
Critical Thinking: Statistics — The Plausibility Test | Testing Throughout