Known studies & datasets
The sources for the numbers used across our calculators, with a citation and a worked calculation for each.
AIDS counsellors quoted 100% certainty on a coin flip
Base-rate neglect does not spare specialists: the counsellors ran this test daily and still reasoned from its accuracy instead of the prevalence.
A researcher posing as a low-risk client asked 20 German counselling centres whether a positive screening result could be a false alarm. Most said no; the true predictive value was about 50%.
GIGERENZER, G., HOFFRAGE, U., & EBERT, A. (1998). AIDS counselling for low-risk clients. AIDS Care, 10(2), 197–211. https://doi.org/10.1080/09540129850124451
A positive screening mammogram is right about 1 time in 10
A screening test is tuned to miss as few cancers as possible, which necessarily produces many false positives; a positive is an indication to biopsy, not a diagnosis.
At a screening prevalence near 1%, a mammogram that catches ~90% of cancers and gives ~9% false positives still yields a positive predictive value around 9%. It is the reason a positive screen leads to a biopsy, not a diagnosis.
Figures of the order used in Gigerenzer, G., et al. (2007). Helping Doctors and Patients Make Sense of Health Statistics. Psychological science in the public interest, 8(2), 53–96. https://doi.org/10.1111/j.1539-6053.2008.00033.x
A 25% risk reduction that can mean treating about 100 for one
The relative and absolute reductions describe the same trial; only the absolute one shows how much an individual patient stands to gain.
This Cochrane review reports a risk ratio of 0.75 for major cardiovascular events — a 25% relative risk reduction. What that means for one person depends on baseline risk: at a 4% five-year risk it is a fall to 3%, an ARR of about 1 point and an NNT near 100. The relative figure is the same for everyone; the absolute benefit shrinks as baseline risk falls.
Taylor, F., et al. (2013). Statins for the primary prevention of cardiovascular disease. Cochrane Database Syst Rev, (1), CD004816. Reported: RR 0.75 for major CVD (25% RRR); the ARR and NNT here use a 4% baseline for illustration. https://doi.org/10.1002/14651858.CD004816.pub5
A sepsis AI graded 0.63, sold as 0.83, in hundreds of hospitals
On another hospital's patients the advertised 0.83 became 0.63; external validation exists precisely to catch this.
The Epic Sepsis Model, live in hundreds of US hospitals, was sold with an AUC of 0.76 to 0.83. An independent external validation on about 38,000 hospitalizations at Michigan measured 0.63: at the alert threshold it caught a third of sepsis cases, missed two thirds, and fired on 18% of all patients. The vendor's higher figure came largely from not accounting for when the predictions were made.
Wong, A., Otles, E., Donnelly, J. P., et al. (2021). External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Internal Medicine, 181(8), 1065–1070. https://doi.org/10.1001/jamainternmed.2021.2626
Habib, A. R., Lin, A. L., & Grant, R. W. (2021). The Epic Sepsis Model Falls Short—The Importance of External Validation. JAMA Internal Medicine, 181(8), 1040–1041. https://doi.org/10.1001/jamainternmed.2021.3333
This list grows as we add calculators. Have a study we should include? support@statexampro.com
Stat Exam Pro turns these into exam-style questions with worked answers, in English and French.
Get it on iPhone