ENFR

The counsellors who were 100% sure

Stat Fails collects documented cases where a statistical error had real consequences, with the reasoning worked through. References at the end.

In 1998, psychologist Gerd Gigerenzer and colleagues ran an experiment that deserves a poster in every clinic. A researcher posing as a client - male, heterosexual, no risk factors - visited 20 public AIDS counselling centres in Germany and asked, before taking the test, one simple question:

If my result comes back positive - could it be a false alarm?

Most counsellors said no. Not unlikely - no. These were professionals whose entire job revolved around this one test, and the typical answer sounded like this:

A false positive is absolutely impossible.
several counsellors put the certainty at 99.9% - or flatly at 100%

Here is the arithmetic they were paid to know:

1 in 10,000 prevalence among low-risk men at the time
99.9% the test's sensitivity; specificity around 99.99%
≈50% the real chance a positive client was infected. A coin flip.

Because in 10,000 clients like him there is one real infection, which the test catches - and about one false alarm. That coin flip was routinely delivered as a verdict, and in the 1990s it was often received as one.

The same mechanism has reached delivery rooms: case reports describe pregnant women given positive HIV results during routine prenatal screening whose confirmatory tests later came back negative. Some had started making irreversible decisions in the days between.
The mechanism fits in one sentence: when almost everyone tested is healthy, even a tiny false-positive rate produces more false alarms than true detections.

That is exactly why confirmatory tests exist - and why a screening result is never a verdict.

To stop making exactly this mistake, we built an interactive playground: 1,000 dots, three sliders, and a preset with the counsellors' numbers already dialed in. Train the intuition on your own eyes.

Open the PPV/NPV playground

What to remember

A positive screening result in a low-risk group is a reason to order the confirmatory test, not to make an announcement. And if the person whose job is the test says 100% - ask them what the prevalence is. If they can't answer, neither can the test.

The full breakdown, in plain English

The quick version above is enough to change what you do. This is the same thing, slowly, with every number shown - for when you want to actually understand it, not just believe it.

The two dials - and why 99% is a lie you tell yourself

Every test has two dials, and neither is the number a patient actually wants. Sensitivity measures the test on the sick: of everyone who truly has the disease, the fraction it flags. Specificity measures it on the healthy: of everyone truly without it, the fraction it correctly clears. That is all - two conditional rates, both measured on people whose status is already known.

The textbook habit of writing both as 99% is where the trouble starts, because real tests are nowhere near that tidy. A screening mammogram catches roughly 85-90% of cancers and clears about 88-90% of healthy women. A modern HIV immunoassay is far sharper, around 99.7% and 99.8%. A rapid antigen COVID test might find only 70% of infections while almost never crying wolf. Accuracy is not one number, and it is never just 99%.

The third number the test can't see: prevalence

Now the number that decides your patient's fate, and it is not on the box. Prevalence is how common the disease is among the people actually being tested - not in the world, in your clinic's queue. It is not a property of the test at all, which is exactly why the counsellors could recite the test's specificity and still get the answer wrong: they were quoting a dial, and the question turned on prevalence.

The math, once, slowly

Here is the whole calculation with the case's own numbers: sensitivity 99.9%, false-positive rate about 1 in 10,000 (specificity 99.99%), prevalence 1 in 10,000.

Imagine 10,000 low-risk men like him.
• 1 is truly infected. The test almost certainly catches him: 1 true positive.
• 9,999 are healthy. At 1 false alarm per 10,000, about 1 also tests positive: 1 false positive.
• So 2 men test positive, and only 1 is actually infected.
PPV = 1 / 2 = about 50%.

The positive result did not move him from unlikely to certain. It moved him from 1-in-10,000 to 1-in-2. Enormous news, and still a coin flip.

So how do you build a test that isn't a trap?

It depends on two things: how rare the disease is, and how much a false alarm costs.

When the disease is rare and a false positive is expensive - frightening the patient, triggering a biopsy - one test is never enough. You screen with a sensitive test to miss nobody, then confirm the positives with a second, highly specific one. That is not bureaucracy; it is the only way to get a positive that means something when almost everyone is healthy. It is literally why HIV diagnosis is a two-step protocol.

When the condition is common and a false alarm is cheap - you will retest tomorrow anyway, or a wrong positive just costs one careful day - you can afford a worse test that scales. A rapid antigen test that misses a third of cases but costs a euro and runs at home is a good public-health tool during a wave, and a poor one for clearing a single worried traveller.

The same logic names two jobs. A very specific test is for ruling in: a positive is hard to fake, so it confirms. A very sensitive test is for ruling out: a negative is hard to fake, so it clears. Clinicians compress this into SpPIn and SnNOut - Specific, Positive, rules IN; Sensitive, Negative, rules OUT.

Keep just this

Sensitivity is measured on the sick; specificity on the healthy.
What your patient wants - the chance a positive is real - is PPV, and it needs prevalence.
Rare disease: even superb specificity yields a poor PPV. Confirm before you conclude.
Rule in with specificity, rule out with sensitivity.

What actually happened in 1998

With that in hand, the case is no longer surprising. The client was chosen to be low-risk on purpose: a heterosexual man with no risk factors. That profile is not a detail - it is the prevalence. Risk factors do not nudge prevalence, they multiply it. The same test on a man from a higher-prevalence group tells a completely different story:

50%low-risk profile, prevalence 1 in 10,000
99.3%risk factors present, prevalence 1.5%

Same test, same positive result. The only thing that changed was the queue the man walked in from - and it swung the answer from a coin flip to near-certainty. The counsellors were not lying about the test. They answered a question about the patient using a fact about the test, and never asked what the prevalence was.

Does this still happen?

The 1998 study was not an isolated result. The underlying error - base-rate neglect, in which the predictive value of a positive is conflated with the sensitivity or specificity of the test - has been documented repeatedly since. In a 2013 study, asked to estimate the probability of disease given a positive result at 1% prevalence, 95% sensitivity and 95% specificity, most surveyed physicians answered around 95%; the correct value is about 16%. Reviews of risk communication (see Gigerenzer et al., 2007, below) report the same pattern across clinicians, patient leaflets and media coverage. The documented consequences include avoidable anxiety, confirmatory procedures that carry non-trivial risk, and clinical decisions taken on a positive whose predictive value was well below certainty.

Where else the same trap lives

Cancer screening: the overdiagnosis debate around mammography and PSA is this arithmetic at population scale.
Forensic DNA: a one-in-a-million match in a database of millions expects false hits.
Security screening: airport and fraud alarms hunting rare events drown in false positives.
Spam and anomaly detection: the same false-alarm math, lower stakes.
Direct-to-consumer genetic tests: rare variants, huge tested populations, shaky positives.

The same formula, six real settings

Test / settingSeSpPrevalencePPV
HIV immunoassay, low-risk screening99.7%99.8%0.010%5%
HIV immunoassay, high-risk group99.7%99.8%1.5%88%
Mammography, women 50-6987.0%89.0%0.600%5%
Fecal immunochemical test (colorectal)79.0%94.0%0.400%5%
COVID antigen, symptomatic, high circulation73.0%99.5%15.0%96%
COVID antigen, asymptomatic mass screen58.0%99.5%0.300%26%

Typical reported values (they vary by test brand and population); PPV is what the formula gives. Sit with the last column.

HIV, low-risk5%HIV, high-risk88%Mammography5%FIT, colorectal5%COVID Ag, symptomatic96%COVID Ag, mass screen26%

The PPV of the scenarios above. The same positive result means wildly different things.

Sp = 99.9% Sp = 99% Sp = 95% 0.001%0.01%0.1%1%10%025%50%75%100% Prevalence (log scale) PPV - chance a positive is real

One test, fixed sensitivity, three levels of specificity. As the disease gets rarer (left), even a 99.9%-specific test loses its positive predictive value. This is the whole story in one picture.

Sources

  • GIGERENZER, G., HOFFRAGE, U., & EBERT, A. (1998). AIDS counselling for low-risk clients. AIDS Care, 10(2), 197–211. https://doi.org/10.1080/09540129850124451
  • Gigerenzer, G., Gaissmaier, W., Kurz-Milcke, E., Schwartz, L. M., & Woloshin, S. (2007). Helping Doctors and Patients Make Sense of Health Statistics. Psychological science in the public interest : a journal of the American Psychological Society, 8(2), 53–96. https://doi.org/10.1111/j.1539-6053.2008.00033.x

Stat Exam Pro drills exactly this reasoning: the Paradoxes & Intuition Traps section is a whole set of cases like this one, as exam questions with worked answers.

Get it on iPhone