p-value

On this page
  1. Definition
  2. How to compute
  3. Pitfalls
  4. Related

Definition

A p-value is the probability of observing a test statistic at least as extreme as the one obtained, assuming the null hypothesis is true. It is not the probability that the effect is random, nor P(H0∣data)P(H_0 \mid \text{data}). A small p only shows the data are hard to reconcile with the null; it says nothing about the size of the effect.

How to compute

Compute it from the test statistic and its null distribution: p=P(T≥tobs∣H0)p = P(T \ge t_{\text{obs}} \mid H_0). Analytically for t- and z-tests, or via a permutation test and the bootstrap when the distribution is unknown. The two-sided version doubles the tail.

Pitfalls

A p-value does not replace a confidence interval or an effect size: a significant effect can be trivial. Multiple testing inflates the false-discovery share, so apply a correction (BH, Holm). Peeking and stopping on significance inflate the Type I error. A p-value does not prove a hypothesis, it merely fails to reject it.

Related terms

Mentioned in

Connections map

Terms, articles and projects connected to this concept. Hover a node to see its name; click to open.

← All terms