Skip to content
Expedify
12 min

Base rates — what a piece of evidence is actually worth

A face-recognition system that is 99.9% accurate and almost entirely wrong, and a number that sent an innocent woman to prison. Both are the same arithmetic, and it is the arithmetic that decides what any piece of evidence is worth.

This course opens with questions rather than definitions, and the questions have right answers.

Answer each one before reading past it. The answer is directly underneath, so the only way to get anything from this is to commit first.

A face-recognition system at an airport

The system is 99.9% accurate in both directions. It recognises 99.9% of the people on its watchlist, and clears 99.9% of everybody else.

It watches 50,000,000 passengers a year. About one in ten million of them is somebody it is looking for.

The system raises an alarm. How likely is it to be right?

50,000,000 passengers, of whom 5 are on the watchlist.

Actually wanted

System raises an alarm
5
It does not
0
All
5

An ordinary passenger

System raises an alarm
50,000
It does not
49,949,995
All
49,999,995

All

System raises an alarm
50,005
It does not
49,949,995
All
50,000,000

It finds 5 of the 5 it wanted, and stops 50,000 people who had done nothing. That is 1 in 10,001.

Roughly 137 innocent people pulled aside every day, to find five in a year.

Nothing is wrong with the system. The 0.1% it gets wrong is 0.1% of fifty million people. The 99.9% it gets right is 99.9% of five.

This is why deployed face recognition keeps collapsing, and the reason is arithmetic rather than engineering. A model can be excellent and its deployment worthless, and no amount of improving the model fixes it.

The number that convicted Sally Clark

In 1999 Sally Clark was convicted of murdering her two infant sons. Both had died suddenly, months apart, in what she said were cot deaths.

An expert witness told the court that the chance of two cot deaths in one family was 1 in 73 million.

How much should a jury be convinced by that number?

The first error: the figure came from squaring. Take one cot death at about 1 in 8,500 and multiply it by itself, and you get 1 in 72,250,000. That is where 73 million came from.

Squaring assumes the second death is no more likely after the first. Whatever raised the risk for one baby — genes, a household, anything shared — is still there for the other.

The same two deaths, if the first makes the second more likely.

no more likely (what the court was told)

Chance of two
1 in 72,250,000

5 times more likely

Chance of two
1 in 14,450,000

10 times more likely

Chance of two
1 in 7,225,000

A modest link between the two deaths moves the number by an order of magnitude. The court was given the version that assumed no link at all.

The second error is worse, and it is the one this lesson is about. A rare event only counts as evidence if the alternative is rarer still. Nobody asked how rare it is for a mother to murder both of her babies. That is also extraordinarily rare.

The question is never how unlikely the evidence is. It is how much more likely it is under one explanation than the other. Only one of those two numbers was ever put to the jury.

The Royal Statistical Society publicly criticised the use of the figure. Sally Clark's conviction was quashed in 2003 after three years in prison. She died in 2007.

It has a name — the prosecutor's fallacy — and it is the airport question in a courtroom. A number that sounds overwhelming, describing a test rather than a person.

A bird, and a thousand mornings

One more, and this one has no numbers at all.

A bird has been fed by the same farmer at nine o'clock every morning for a thousand days. Nothing has ever gone wrong.

How confident should the bird be about tomorrow morning?

Taleb tells this in The Black Swan, after a chicken of Bertrand Russell's from 1912. The record is not evidence about the farmer. It is evidence about the days that have already happened.

Every model in this course is the bird. It has seen the past. It has never seen the thing that has not happened yet, and it cannot tell you which of the two it is confident about.

Three questions, one shape

  • An excellent system, almost every alarm wrong. Because the people it hunts are one in ten million.
  • An overwhelming number that proved nothing. Because it was never weighed against the alternative.
  • A thousand quiet mornings that said nothing about the farmer. Because they were evidence about mornings, not about farmers.

Evidence is worth nothing on its own. It is worth something only against how common the thing is that it claims to have found. One sentence, and it decides all three answers.

You will meet this again as a number a model has to beat. It is called the baseline, and it is the first thing to ask for when somebody shows you an accuracy figure.

One last thing, and it is just an ordinary task

An inbound lead arrives at your company. Here is everything you know about it.

  • Head of Operations at a 400-person logistics company.
  • Downloaded the pricing page twice this week.
  • Arrived from a comparison article, not from search.
  • Company uses a competitor's product today.
  • No budget mentioned anywhere.

That is everything you know about them.

How good is this lead? Score it out of ten.

Keep your answer. This one comes back later.

No answer to that one, because there is not one. Keep the number in your head. It comes back at the end of this module, and by then it will mean something.

Related lessons