Theory¶
Population Resemblance Statistic¶
Let
be a fixed reference distribution with \(p_{0j}>0\), and let
be the empirical current distribution.
The Population Resemblance Statistic is
The sample-size-scaled statistic is
Under local alternatives, \(Q_n\) is asymptotically non-central chi-square with \(B-1\) degrees of freedom.
\(\delta\)-resemblance¶
The framework defines a current population \(p\) as \(\delta\)-resemblant to \(p_0\) when
This changes the monitoring question from exact equality to whether the population shift exceeds a pre-specified tolerable amount.
The recommended automatic tolerance is
where \(c>0\) scales the acceptable shift relative to sampling variability.
Least-favourable non-centrality¶
The non-centrality parameter is
Because the true current probabilities are unknown, the framework uses the maximal non-centrality over the \(\delta\)-resemblance region.
For even \(B\),
For odd \(B\),
where \(p_*=\max_j p_{0j}\).
Three decision regions¶
The package implements two nested resemblance hypotheses with tolerances \(\delta\) and \(M\delta\), where \(M>1\).
The PRS decision regions are
and
The critical values are
where \(F^{-1}_{\nu}(\cdot,\lambda)\) is the non-central chi-square quantile function.
The package labels the regions as:
| Region | Package label | Interpretation |
|---|---|---|
| \(R_1\) | acceptable |
Continue using the model or process |
| \(R_2\) | partially discrepant |
Enhanced monitoring |
| \(R_3\) | fully discrepant |
Material discrepancy requiring action |
Structural constraint¶
The widened tolerance must satisfy
This prevents the resemblance region from implying invalid negative probabilities.
Structural sample-size planning¶
For the recommended tolerance,
the structural requirement
can be solved directly for the smallest positive integer sample size.
Use minimum_structural_sample_size() to compute that bound.
What this bound means
The returned value guarantees only that the recommended tolerance satisfies the probability-domain constraint. It is not a power calculation and does not guarantee good finite-sample calibration, sufficient expected cell counts, or asymptotic accuracy.
Calibration¶
Use evaluate_calibration() to inspect the derived tolerance, non-centrality, critical values, and margin to the structural bound before running an assessment.
Use sweep_calibration_parameters() to examine sensitivity across combinations of \(c\), \(M\), \(\alpha_1\), and \(\alpha_2\).
Statistical scope¶
The implementation follows the fixed-reference, one-sample formulation. An empirical reference sample can be converted into a fixed probability vector conditionally, but that does not make the procedure a two-sample test.
This distinction is deliberate. Reference-sample uncertainty requires a different statistical formulation.