Other meanings of p-value
STATISTICS
A p-value is a probability measure used in statistical hypothesis testing. It quantifies how incompatible observed data, or results more extreme, are with a specified null hypothesis and statistical model; it does not give the probability that the null hypothesis is true.
A p-value measures the compatibility of data with a null hypothesis under an assumed statistical model. Formally, it is the probability, assuming the null hypothesis and model are correct, of obtaining a test statistic at least as extreme as the one observed.1 A small value indicates that the observed result would be relatively unusual under the null; it is evidence against that null hypothesis, not proof that an alternative explanation is true.
The direction of “at least as extreme” depends on the test. A one-sided test considers departures in one prespecified direction, whereas a two-sided test considers departures in either direction. The p-value therefore depends on the hypothesis, test statistic, sampling design, and model assumptions, rather than being an intrinsic property of the data alone.
A p-value is commonly compared with a prespecified significance level, written as α. If p is less than or equal to α, researchers may reject the null hypothesis; otherwise, they generally report that the evidence is insufficient to reject it. The result is not a declaration that the null hypothesis has been proved, and a large p-value does not establish that two effects are identical.1
Interpretation should be paired with an effect estimate and uncertainty interval. A tiny effect can produce a small p-value in a large sample, while a practically important effect can fail to reach a chosen threshold in a small or noisy study. Confidence intervals, study design, measurement quality, and subject-matter consequences provide information that a p-value alone cannot supply.2
P-values become misleading when researchers treat a threshold as a sharp boundary between discovery and no discovery. The American Statistical Association advises against describing p-values as the probability that a hypothesis is true, the probability that results occurred by chance, or a measure of effect size or importance.1 Statistical significance and practical significance are different judgments.
Repeated testing can also make small p-values appear by chance. Searching many outcomes, subgroups, models, or analysis choices raises the chance of at least one nominally significant result; optional stopping can have a similar effect. Multiple-comparison procedures, preregistered analysis plans, transparent reporting, replication, and appropriate modeling help address these problems.34
P-values have two partly different intellectual roots. Ronald Fisher treated them as measures of evidence against a null model, while the Neyman–Pearson framework emphasized long-run error rates, decision rules, and prespecified alternatives; modern practice often combines ideas that were not originally identical.5 This history helps explain why a p-value is not itself a decision rule.
Exact tests and randomization tests can calculate p-values without relying on the same large-sample approximations used by common tests, although their null distributions may be discrete. In discrete settings, attainable p-values may be coarse and conservative. Bayesian posterior probabilities and Bayes factors answer different questions, because they incorporate prior distributions rather than calculating data extremeness under a null alone.6 “P-hacking” refers to analysis practices that exploit flexibility to obtain appealing significance results, not to a special kind of p-value.
A p-value should be interpreted in the context of the prespecified hypothesis, sampling design, model assumptions, multiplicity, effect size, and uncertainty; no universal cutoff can replace that context.
Help improve the encyclopedia. Reports go straight to the site manager.