Topics
Few statistical quantities are used as widely, or interpreted as carelessly, as the p-value. The problem is not that the definition is especially complicated. A p-value is a tail probability computed from a reference distribution under a null hypothesis. The difficulty is that this probability is conditional on a statistical model, while the scientific questions people usually want to answer concern the plausibility of hypotheses, the magnitude of effects, the probability of replication, or the consequences of a decision. Those are different questions. When the distinction is ignored, the p-value is asked to provide information that it does not contain.
A rigorous treatment therefore has to begin before the p-value itself, with the model that generates the reference distribution. Suppose observed data are represented by a random variable $X$ with distribution indexed by a parameter $\theta$, and suppose the null hypothesis is
A test statistic $T(X)$ is chosen so that values in some direction are increasingly incompatible with $H_0$. If the null hypothesis is simple and large values of $T$ are more extreme, then after observing $T(X)=t_{\mathrm{obs}}$, the one-sided p-value is
The probability is taken over hypothetical repetitions generated under the null model. The observed data are fixed at the stage when the p-value is reported; what varies conceptually is the data that could have been observed had the experiment been repeated under the assumptions defining $H_0$.
The conditional logic of the p-value
The most common interpretation error is a reversal of conditional probability. A p-value concerns
whereas researchers often want
These are not the same quantity. Bayes' theorem relates them only after specifying prior probabilities and a full probability model for the alternatives. A p-value of 0.03 therefore does not mean that the null hypothesis has a 3% probability of being true, nor does it mean that there is a 97% probability that the alternative is correct.
The phrase “the probability that the result occurred by chance” is equally misleading. The p-value does not partition the observed result into a random component and a real component. Randomness is already part of the sampling model. The p-value asks how unusual the test statistic would be if the null model, together with the assumptions used to construct the reference distribution, were valid. A small value indicates incompatibility between the observed statistic and that model. It does not, by itself, identify which assumption failed.
This qualification matters because statistical tests rarely depend on the null hypothesis alone. A classical two-sample t-test, for example, depends on assumptions about independence, the sampling mechanism, the form of the statistic, and either exact distributional assumptions or an asymptotic approximation. A very small p-value may reflect a genuine departure from the parameter restriction in $H_0$, but it can also be influenced by model misspecification, dependence, measurement error, selection, or an analysis chosen after inspecting the data.
From test statistics to reference distributions
The p-value is meaningful only relative to the test statistic and its reference distribution. Consider a one-sample Gaussian problem in which
and suppose $\sigma$ is unknown. To test
the usual statistic is
Under the null and the Gaussian model,
The two-sided p-value is then obtained by measuring how far $|T|$ lies into both tails of the Student distribution,
The phrase “at least as extreme” is therefore not universal. It is defined by the ordering induced by the test statistic. For a symmetric continuous statistic, a two-sided tail calculation is often straightforward. For discrete distributions, composite null hypotheses, nuisance parameters, permutation tests, or non-standard statistics, several legitimate definitions may exist. A p-value is not a primitive property of a dataset; it is the output of a particular testing procedure.
Under a simple continuous null hypothesis and a correctly calibrated test, the p-value has a Uniform$(0,1)$ distribution,
This property explains the repeated-sampling interpretation of a significance level. If a valid level-$\alpha$ procedure is repeated indefinitely under the null, the probability of rejection is at most $\alpha$. The statement is about the long-run behavior of the procedure, not about the probability that a particular rejected hypothesis is false.
For discrete tests, the null distribution of p-values is generally not exactly uniform; valid exact tests are often conservative, so
This small technical detail matters because slogans such as “p-values are uniformly distributed under the null” are true only under conditions that are often omitted.
Statistical significance is a property of a decision rule
A significance threshold such as
belongs to a decision procedure. The conventional rule rejects $H_0$ when $p\le\alpha$. In the Neyman-Pearson framework, $\alpha$ controls the Type I error rate of a prespecified procedure. It is not a universal boundary separating real effects from nonexistent ones, and there is no mathematical discontinuity between p-values of 0.049 and 0.051.
The distinction between evidential interpretation and decision theory is historically important. Fisher emphasized p-values as continuous measures of discrepancy between data and a null hypothesis, whereas Neyman and Pearson developed tests as repeated-sampling decision procedures characterized by Type I error, Type II error, and power. Modern practice often combines these traditions without acknowledging that they answer different questions. Reporting an exact p-value suggests a graded evidential interpretation; applying a fixed threshold treats the same number as part of a binary decision rule.
The threshold should therefore be chosen, when a threshold is genuinely required, in relation to the consequences of false positives and false negatives, the scientific setting, prior evidence, and the design of the study. Treating 0.05 as a law of nature encourages mechanical decisions and obscures the inferential assumptions that matter more.
Effect size, precision, and sample size
A p-value does not measure the magnitude or practical importance of an effect. This follows directly from the structure of most test statistics, which compare an estimated effect with its standard error. For a parameter estimate $\widehat\theta$ tested against $\theta_0$, a statistic often has the approximate form
The same effect estimate can produce a small or large p-value depending on its precision. With a sufficiently large sample, a scientifically negligible effect can be estimated so precisely that the null value is strongly rejected. Conversely, a large effect estimate can fail to reach conventional significance when the sample is small or the outcome is noisy.
This is why effect estimates and uncertainty intervals should normally accompany, and often take interpretive priority over, p-values. A confidence interval shows both the estimated magnitude and the range of parameter values compatible with the data under the procedure. For many standard two-sided tests, rejecting
at level $\alpha$ is equivalent to the corresponding $(1-\alpha)$ confidence interval excluding $\theta_0$. The interval contains substantially more information because it makes the scale of the effect visible.
Confidence intervals are themselves frequently misinterpreted. A 95% frequentist confidence interval does not mean that there is a 95% probability that the fixed parameter lies inside the realized interval. The 95% refers to the long-run coverage of the interval-construction procedure under repeated sampling. If a probability statement about the parameter is required, a Bayesian posterior interval is a different inferential object and depends on a prior distribution.
Large p-values, power, and equivalence
A large p-value should not be interpreted as evidence that the null hypothesis is true. The observed statistic may be close to the null expectation because the true effect is small, but the same result can arise because the study is underpowered or the estimate is imprecise. The distinction becomes especially important when researchers write that two groups are “the same” simply because a conventional significance test failed to reject equality.
Power formalizes the ability of a test to detect a specified alternative. For a rejection region $R$ and parameter value $\theta$,
is the power function. A study with low power over scientifically important alternatives can easily produce large p-values even when meaningful effects are present. Consequently, non-significance alone says little about equivalence.
If the scientific objective is to demonstrate that an effect is sufficiently small, an equivalence or non-inferiority framework is more appropriate. Suppose differences within $[-\Delta,\Delta]$ are considered practically negligible. An equivalence analysis tests whether the parameter lies inside that interval rather than testing whether it is exactly zero. This reverses the inferential logic: the burden of evidence is placed on showing that the effect is small enough to be considered equivalent.
Multiplicity and selective analysis
The validity of a p-value depends on the entire procedure that generated it. If twenty independent null hypotheses are each tested at the 0.05 level, the expected number of false rejections is one, and the probability of at least one false positive can be much larger than 0.05. Searching across outcomes, subgroups, transformations, lag structures, model specifications, or stopping rules produces the same general problem even when the multiplicity is less visible.
Family-wise error rate procedures, such as Bonferroni or Holm adjustments, aim to control the probability of making at least one false rejection within a family of hypotheses. False discovery rate procedures, such as Benjamini-Hochberg, control a different quantity: the expected proportion of false discoveries among rejected hypotheses under their assumptions. Neither method is automatically correct for every analysis because the definition of the hypothesis family and the scientific purpose of the investigation matter.
Selective reporting creates an even deeper problem. If a researcher explores many analyses and reports only the one producing the smallest p-value, the published number no longer has the reference distribution assumed by the nominal test. The same issue appears with optional stopping. Repeatedly collecting data, testing after each batch, and stopping the first time $p<0.05$ inflates the Type I error rate unless the sequential design is accounted for. Sequential tests, alpha-spending methods, always-valid p-values, and related procedures exist precisely because ordinary fixed-sample p-values are not automatically valid under continuous monitoring.
Preregistration does not solve every inferential problem, but it can make the distinction between confirmatory and exploratory analyses more transparent. Exploratory analyses are scientifically valuable; the problem arises when a data-dependent search is reported as though it had been specified in advance.
Exact, asymptotic, permutation, and bootstrap p-values
A p-value inherits the assumptions of the reference distribution used to calculate it. Exact tests derive that distribution from finite-sample probability theory. Asymptotic tests replace the finite-sample distribution by a limiting approximation, often Gaussian or chi-squared. Permutation tests generate a reference distribution by rearranging labels or observations under an exchangeability assumption. Bootstrap tests approximate sampling behavior through resampling from an estimated distribution.
These procedures are not interchangeable. A permutation test can be exact under a randomization design but invalid if the required exchangeability is absent. A Wald test can have poor finite-sample calibration when the likelihood is asymmetric or the parameter lies near a boundary. A bootstrap can improve approximation in some problems but does not rescue an unidentified parameter or a fundamentally biased design.
The phrase “the p-value is 0.02” is therefore incomplete without knowing how the value was generated. A serious statistical report should make the test statistic, null hypothesis, tail convention, reference distribution, and any multiplicity adjustment explicit.
Bayesian evidence is a different inferential object
Bayesian inference addresses a different conditional probability. Given prior distribution $\pi(\theta)$ and likelihood $p(y\mid\theta)$, the posterior is
A posterior probability such as
is a probability statement about the parameter conditional on the observed data and the prior-model specification. It cannot be recovered from a frequentist p-value alone.
Bayes factors also differ fundamentally from p-values. A Bayes factor compares the marginal probability of the observed data under two models,
Its interpretation depends on the prior distribution under the alternative. Because p-values and Bayes factors condition in different directions and encode different assumptions, they can disagree substantially, especially in large samples or when the alternative prior is diffuse. Such disagreement is not paradoxical; the methods answer different questions.
A small simulation
The repeated-sampling calibration of p-values can be illustrated directly. The following simulation generates data from a standard normal distribution, tests the true null hypothesis that the mean is zero, and records the p-values from repeated one-sample t-tests.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
from __future__ import annotations
import numpy as np
from numpy.typing import NDArray
from scipy import stats
def simulate_null_pvalues(
*,
repetitions: int = 10_000,
sample_size: int = 30,
seed: int = 42,
) -> NDArray[np.float64]:
"""Simulate p-values from a valid two-sided t-test under H0."""
if repetitions <= 0:
raise ValueError("repetitions must be positive")
if sample_size < 2:
raise ValueError("sample_size must be at least two")
rng = np.random.default_rng(seed)
pvalues = np.empty(repetitions, dtype=float)
for i in range(repetitions):
sample = rng.normal(loc=0.0, scale=1.0, size=sample_size)
result = stats.ttest_1samp(sample, popmean=0.0)
pvalues[i] = float(result.pvalue)
return pvalues
If the test is correctly specified, approximately 5% of these p-values should fall below 0.05. That result does not mean that 5% of the null hypotheses are false; in the simulation every null hypothesis is true. It means that the testing procedure is calibrated so that approximately 5% of repeated samples cross the rejection threshold when $H_0$ holds.
Conclusion
A p-value is best understood as one component of a statistical procedure rather than as a self-contained measure of truth. It quantifies the tail behavior of a chosen test statistic under a specified null model. Its validity depends on the sampling design, the test statistic, the reference distribution, the stopping rule, the multiplicity structure, and the distinction between analyses specified in advance and those selected after seeing the data. A small p-value can indicate substantial incompatibility with the null model, but it does not quantify the probability that the null is false, the probability that a result will replicate, or the scientific importance of an effect.
Good inference therefore requires more than asking whether $p<0.05$. The effect estimate should be reported on a scientifically meaningful scale, uncertainty should be quantified, power and design should be considered, multiplicity should be handled explicitly, and the inferential framework should match the question being asked. In many scientific problems, the most informative result is not that a null value has been rejected, but that the data support a particular range of plausible effect sizes with known uncertainty and under clearly stated assumptions.
The p-value remains useful when interpreted within those limits. Its weakness is not that it is mathematically incoherent, but that it is routinely asked to answer questions it was never designed to answer.
References
- Cox, D. R. (2006). Principles of Statistical Inference. Cambridge University Press.
- Fisher, R. A. (1925). Statistical Methods for Research Workers. Oliver and Boyd.
- Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical Tests, P Values, Confidence Intervals, and Power: A Guide to Misinterpretations. European Journal of Epidemiology, 31, 337-350.
- Neyman, J., & Pearson, E. S. (1933). On the Problem of the Most Efficient Tests of Statistical Hypotheses. Philosophical Transactions of the Royal Society A, 231, 289-337.
- Wasserstein, R. L., & Lazar, N. A. (2016). The ASA Statement on p-Values: Context, Process, and Purpose. The American Statistician, 70(2), 129-133.
Embed interactive plots, widgets, and demos using <figure>, <iframe>, or <div class="interactive-embed"> containers. Ensure each embed includes descriptive captions for accessibility.
How to cite
Use the quick export buttons to save citations for reference managers or copy the formatted text directly.
Diogo Ribeiro (2024). P-Values: Sampling Models, Evidence, and Statistical Decisions. Faculty of Media Arts and Design, Technical University of Porto. https://diogoribeiro7.github.io/mathematics/P_value/.


