Maths●●●●●Difficulty 3 of 5

How can a 95%-accurate test be wrong 98% of the times it says yes?

A breathalyzer that never misses a drunk driver can still be wrong 98 times out of 100 when it says someone is drunk, and the same kind of reasoning error helped convict a British mother whose conviction was later quashed.

▶ Start the story

A breathalyzer that catches every single truly drunk driver and only falsely accuses 5% of sober ones sounds nearly perfect. But if only 1 in 1,000 drivers on the road is actually drunk, a positive result from that test is correct only about 2% of the time, not 95%. Out of 1,000 random drivers, the 1 truly drunk one produces 1 true positive, but the 999 sober drivers produce roughly 50 false positives between them. Many people guess 95% instead, because they judge the test's accuracy and forget how rare drunk driving is among the drivers tested. That mistake is called the base rate fallacy: ignoring how common or rare something is in favor of focusing only on the specific case in front of you.

~2%

chance a flagged driver is really drunk, with a 95%-accurate test and 1-in-1,000 base rate

The same error has played out in court. In 1998, Sally Clark, a British woman, was accused of murdering her two infant children after each died unexpectedly weeks after birth. An expert witness testified that the odds of two children in one family dying of sudden infant death syndrome (SIDS) by chance were about 1 in 73 million, and argued it was therefore far more likely she had killed them. But that figure assumed each death was an independent, unrelated accident, when there's good reason to think a family already affected by one SIDS death is more vulnerable to a second, for genetic reasons, which breaks that assumption entirely. And even taken at face value, the number meant nothing on its own: it had to be weighed against the probability of the alternative, double homicide. Clark was convicted in 1999, prompting the Royal Statistical Society to point out the mistakes in a press release; her conviction was quashed in 2003, after a court found that a forensic pathologist had withheld evidence in her favor.

The same pattern shows up anywhere a rare event is being tested for in a huge population: a facial recognition system that's 99% accurate but scans 10,000 people a day will flag far more innocent people than actual wanted criminals, simply because innocent people vastly outnumber criminals to begin with. A test's accuracy alone never tells you whether a positive result is trustworthy. You also have to know how rare the thing you're testing for actually is.

Quiz me

0/3

  1. 1.Why can a breathalyzer that never misses a truly drunk driver still be wrong most of the time it flags someone as drunk?
  2. 2.What was the key statistical error in the case against Sally Clark?
  3. 3.What is the 'false positive paradox' illustrated by the facial recognition camera example?

Recap

A test's accuracy alone can't tell you whether a positive result is trustworthy; you also need to know how rare the thing being tested for actually is.

Surprising fact · A test that's wrong only 5% of the time can still be wrong about 98% of its positive results if the condition it's testing for is rare enough.

Sources (1)

No source, no claim. Every fact in this lesson (14 claims) cites at least one of these.

  1. [1]Base rate fallacy · Wikipedia
More lessons in ➗ Maths (3) See all maths lessons →

One more light on your map.

Get one lesson like this every day, about the things you love. Free, in two or five minutes.

Get the share card for this lesson ↗