Experimental independent two-sample resemblance¶
The package includes an experimental API for comparing two independent categorical samples while propagating sampling uncertainty from both samples.
It is deliberately not exported from the stable package root.
Import¶
Use the experimental namespace explicitly:
from population_resemblance.experimental import (
assess_independent_two_sample_resemblance,
)
result = assess_independent_two_sample_resemblance(
[420, 330, 250],
[390, 360, 250],
samples_are_independent=True,
)
print(result.statistic)
print(result.region)
print(result.effective_sample_size)
print(result.pooled_probabilities)
The explicit samples_are_independent=True argument is intentional. This implementation
must not be used silently for overlapping, paired, repeated, clustered, or otherwise
dependent samples.
Statistic¶
For sample sizes \(n\) and \(m\),
Let \(\widehat p\) and \(\widehat q\) be the sample proportions and let
be the pooled empirical distribution.
The experimental statistic is
At the point null this is exactly Pearson's two-sample homogeneity statistic.
Plug-in resemblance calibration¶
If delta is omitted, the experimental method uses
The least-favourable non-centrality and nested decision boundaries are then calculated from the pooled empirical centre.
This is the plug-in procedure studied in the repository research work. It has finite-sample simulation support, but it does not have a claim of uniform composite-null validity.
Sparse-category policy¶
The experimental API enforces the Phase 2 sparse-category policy.
It rejects the assessment when:
- any pooled category count is zero;
- the minimum pooled expected count is below 5;
- the widened resemblance region violates the structural probability-domain bound.
The method never adds epsilon smoothing, silently drops categories, or merges categories.
Dependence is unsupported¶
The independent effective sample size is invalid in general when samples overlap or contain paired/repeated observations.
The research notes derive an exact correction for one special overlap model, but that correction is not part of this experimental API.
For general dependence, a covariance-aware Wald construction is required.
See:
- Independent asymptotics
- Critical values
- Finite-sample calibration
- Dependence and overlap
- Sparse-category policy
Stability status¶
The experimental namespace is outside the stable package-root compatibility boundary.
Its function names, result objects, validation policy, and statistical calibration may change before promotion to the stable public API.