Cover of The Art of Statistics

The Art of Statistics

David Spiegelhalter

6 ideas

  1. PPDAC cycle for learning from data

    Statistical inquiry runs as a cycle of Problem, Plan, Data, Analysis, Conclusion, and each stage can go wrong in its own way. Defining the question and designing how data will be collected usually matter more than the analysis technique. Conclusions should feed back into new problems, not end the inquiry.

  2. Expected frequencies make probabilities intuitive

    Restating a probability as a count of people, such as '4 out of 100 people like you', makes risks easier to understand and harder to misread than percentages or conditional probabilities. The approach works especially well for Bayesian problems like screening tests. Picturing a crowd of 1,000 and tracing how many test positive and how many truly have the condition shows directly that most positives can be false when the condition is rare.

  3. Relative versus absolute risk framing

    A headline that bacon raises bowel cancer risk by 18% states a relative risk. In absolute terms, that increase moves lifetime risk from about 6 in 100 to about 7 in 100. Relative figures inflate perceived danger or benefit, so honest communication gives the absolute baseline and the absolute change, ideally as expected frequencies.

  4. Harold Shipman caught too late by data

    The GP Harold Shipman murdered more than 200 patients over decades. His excess death rate, especially among elderly women dying at home in the afternoon, was visible in routine mortality data. A sequential monitoring method like a CUSUM chart would have flagged him years earlier, showing how systematic surveillance of outcomes can detect harm that individual cases conceal.

  5. Probability is epistemic uncertainty, not objective fact

    Most probabilities do not exist as physical properties of the world. They express someone's uncertainty given their knowledge and assumptions, even for coin flips once the coin has already landed unseen. Treating probability as personal and conditional on judgment clarifies why different analysts can legitimately give different numbers, and why those numbers must be justified.

  6. P-values do not measure hypothesis truth

    A p-value gives the probability of data this extreme if the null hypothesis were true. It is not the probability that the hypothesis is true, and it is not the size or importance of an effect. Running many tests, fishing through analyses, and publishing only significant results together generate many false discoveries, which drives the reproducibility crisis.

Save and mark ideas in the app