Two-sample resemblance critical values¶
This note addresses research issue #45.
The previous derivation established that, for two independent multinomial samples, the pooled Pearson quadratic form is asymptotically non-central chi-square under local alternatives. The question here is whether the one-sample resemblance critical-value construction survives by simply replacing the one-sample size \(n\) with
The answer is only conditionally. If the limiting pooled probability vector is treated as fixed, the one-sample geometry carries over. In a genuine symmetric two-sample problem, however, that pooled vector is itself an unknown nuisance parameter, so a universal closed form does not follow from the one-sample result.
Setup¶
Let the population proportions be \(p\) and \(q\), with independent sample sizes \(n\) and \(m\). Define
and the pooled population centre
Let
Then
Under local alternatives, the two-sample Pearson-type statistic has candidate non-centrality
with
Conditional resemblance region¶
Suppose first that \(r\) is fixed and known.
Define the two-sample resemblance set
If the full symmetric perturbation set is compatible with valid probability vectors \(p\) and \(q\), maximizing \(\lambda\) over \(\mathcal D(\delta)\) is exactly the same weighted convex optimization as in the one-sample PRS framework, with \(n_{\mathrm{eff}}\) replacing \(n\).
Hence, conditionally on fixed \(r\),
where
So the even/odd extreme-point geometry is retained for a fixed centre.
Probability-domain constraint¶
The two samples must both remain valid probability vectors.
Because
and
allowing every coordinate perturbation in \([-\delta,\delta]\) requires
Equivalently,
For the wider resemblance region, replace \(\delta\) by \(M\delta\).
This is stricter than merely requiring \(|p_j-q_j|\le\delta\): it guarantees that the entire symmetric perturbation set around the fixed pooled centre remains inside the probability simplex for both samples.
Conditional critical values¶
For the scaled statistic
a direct conditional analogue of the one-sample nested framework would use
and
The corresponding regions would be
and
If instead an unscaled discrepancy
is reported, the thresholds are \(c_1/n_{\mathrm{eff}}\) and \(c_2/n_{\mathrm{eff}}\).
This conditional construction is mathematically parallel to the one-sample method.
Why the genuine two-sample problem is harder¶
A genuine two-sample test does not know \(r\).
The natural practical statistic uses the pooled empirical distribution
For the point-null problem, replacing \(r\) by \(\widehat r\) is asymptotically valid and yields Pearson's homogeneity statistic.
For resemblance testing, however, the critical values also depend on \(r\) through
Plugging \(\widehat r\) into the critical-value calibration is a stronger step than plugging it into the test statistic.
Pointwise consistency suggests that
for fixed interior \(r\), but that alone does not establish uniform Type I error control over the composite two-sample resemblance null.
The nuisance parameter therefore enters twice:
- in the quadratic-form weights;
- in the least-favourable calibration itself.
Why replacing n by n_eff is not enough¶
The substitution
correctly reproduces the independent two-sample covariance scaling.
But the one-sample PRS has a fixed reference vector \(p_0\). In a symmetric two-sample problem there is no fixed analogue unless one is introduced by design.
Therefore the recipe
replace \(n\) by \(n_{\mathrm{eff}}\) and \(p_0\) by \(\widehat r\)
is a plausible plug-in procedure, but it is not yet a proven resemblance test with uniform error guarantees.
Recommended tolerance: candidate only¶
A natural two-sample analogue of the one-sample recommendation is
With unknown \(r\), the operational version would replace \(r\) by \(\widehat r\).
This has an appealing interpretation: it scales tolerance by the standard error of the difference between two independent proportions.
However, a data-dependent \(\delta\) changes the null region itself. Its calibration must therefore be studied rather than assumed.
Result of issue #45¶
The simple extension succeeds only under a fixed pooled centre:
- \(n_{\mathrm{eff}}\) replaces \(n\);
- the one-sample even/odd least-favourable geometry is preserved;
- conditional critical values follow from the same non-central chi-square quantiles;
- the probability-domain constraint depends on \(\rho\) as well as \(r\).
For the genuine symmetric two-sample problem, the pooled centre is unknown. The simple closed form is therefore conditional on a nuisance parameter, and plugging in \(\widehat r\) requires separate validation.
Consequence for implementation¶
No public two-sample resemblance API should be added yet.
The next required step is #46: simulate the plug-in candidate across sample-size ratios, reference distributions, and resemblance boundaries to determine whether the nominal decision probabilities are preserved in finite samples and whether plug-in calibration is stable enough to justify a practical method.