Two-sample population resemblance: research note¶
This note records the current mathematical research direction for a genuine two-sample extension of the Population Resemblance Statistic (PRS).
It is intentionally not a public statistical API specification. The existing package continues to implement the fixed-reference one-sample framework.
Motivation¶
The current empirical-reference wrapper converts baseline counts into probabilities and then treats those probabilities as fixed. That is a conditional one-sample construction.
A genuine two-sample procedure must propagate uncertainty from both samples.
Let
and
be empirical proportions from two multinomial samples of sizes \(n\) and \(m\).
Independent-sample starting point¶
Assume first that the two samples are independent.
If both samples arise from a common probability vector \(r\), then
where
Define the effective sample size
Then the candidate normalized difference is
Under equality, its limiting covariance is \(\Sigma(r)\), which has rank \(B-1\).
This suggests the candidate quadratic form
If \(r\) were known, the natural conjecture is a limiting \(\chi^2_{B-1}\) distribution under equality.
Because \(r\) is unknown in a true two-sample problem, a practical statistic would need a plug-in estimate, most naturally the pooled empirical distribution
The validity of replacing \(r\) by \(\widehat r\) must be established formally rather than assumed.
Local alternatives¶
A direct analogue of the one-sample local-alternative construction would consider
If the covariance and plug-in arguments hold, the candidate limiting distribution becomes non-central chi-square with
This expression is currently a research target, not yet a package guarantee.
Candidate resemblance definition¶
The simplest two-sample resemblance region would be
That preserves the original category-wise interpretation of resemblance.
However, the two-sample case introduces a nuisance probability vector \(r\), so the least-favourable non-centrality problem is no longer automatically identical to the one-sample derivation.
Questions still to prove include:
- whether pooled weights are asymptotically valid under local alternatives;
- whether the least-favourable configuration retains the same even/odd category structure;
- whether a closed form exists for arbitrary unequal sample sizes;
- how the recommended tolerance should scale with both \(n\) and \(m\);
- whether replacing \(n\) by \(n_{\mathrm{eff}}\) is sufficient.
Dependence and overlapping samples¶
Independence is only the first research case.
In general,
Overlapping individuals, repeated measurements, rolling windows, or nested samples make the cross-covariance non-zero.
A production two-sample method therefore needs one of:
- a model for the dependence structure;
- a consistent covariance estimator;
- a resampling procedure that preserves the dependence;
- or an explicit restriction to independent samples.
The first implementation, if validated, should likely target independent samples only.
Validation requirements before implementation¶
A public two-sample API should not be merged until all of the following are available:
- a complete asymptotic derivation;
- a proof or controlled approximation for the pooled reference weights;
- a well-defined resemblance region;
- least-favourable critical values or a justified numerical optimization;
- finite-sample simulation calibration;
- comparison against the classical multinomial homogeneity test;
- explicit behavior for sparse categories;
- documented handling, or exclusion, of dependent samples.
Research issues¶
The work is split into focused research tasks:
-
44 — independent two-sample asymptotics (derivation);¶
-
45 — two-sample resemblance critical values (derivation);¶
-
46 — finite-sample simulation calibration (study);¶
-
47 — overlapping and dependent samples (derivation).¶
The parent research item remains #32.
Current package policy¶
Until this research is complete, the package should continue to state clearly that:
- the standard PRS API is one-sample with a fixed reference distribution;
- empirical reference counts are treated conditionally as fixed probabilities;
- no existing function should be renamed or reinterpreted as a two-sample test.