Independent two-sample PRS asymptotics¶
This note completes the first mathematical research task in #44.
The result below concerns two independent multinomial samples only. It does not yet define a two-sample resemblance decision procedure or critical values for nested \(\delta\)-resemblance hypotheses.
The source PRS paper explicitly keeps \(p_0\) fixed and notes a two-sample formulation as future research. The derivation here is therefore an extension beyond the published one-sample method.
Setup and assumptions¶
For a fixed number of categories \(B\), let
with the two samples independent. Define empirical proportions
Assume:
- \(n\to\infty\) and \(m\to\infty\);
- \(n/(n+m)\to\rho\in(0,1)\);
- \(p_n\to r\) and \(q_m\to r\) for an interior probability vector \(r\), with \(r_j>0\) and \(\sum_j r_j=1\);
- under local alternatives,
[ \sqrt{n_{\mathrm{eff}}}(p_n-q_m)\to\xi, \qquad \mathbf 1^\top\xi=0, ]
where
[ n_{\mathrm{eff}}=\frac{nm}{n+m}. ]
The interior condition is needed because the Pearson weights contain reciprocals of the limiting probabilities.
Difference-of-proportions CLT¶
Write
Using
we obtain
The multinomial central limit theorem gives
and
where
Independence of the samples and \(m/(n+m)\to 1-\rho\), \(n/(n+m)\to\rho\) imply
Under exact equality, \(p_n=q_m=r\), so \(\xi=0\).
Rank and degrees of freedom¶
Let
Then
Because \(u^\top u=1\), the matrix
is symmetric and idempotent:
It is the orthogonal projector onto the \((B-1)\)-dimensional subspace orthogonal to \(u\). Hence
This is the same loss of one degree of freedom caused by the simplex constraint in the one-sample multinomial problem.
Known-weight quadratic form¶
Suppose temporarily that the limiting common probability vector \(r\) is known. Define
Equivalently,
Let
Then
Since \(\mathbf 1^\top\xi=0\),
so the mean vector lies in the range of \(P\). Therefore
with
Hence
Under exact equality, \(\lambda=0\), giving the central \(\chi^2_{B-1}\) limit.
For a local sequence satisfying \(\sqrt{n_{\mathrm{eff}}}(p_n-q_m)\to\xi\), the non-centrality may also be written as
Unknown common probabilities and pooled plug-in¶
A genuine two-sample problem does not know \(r\). Define the pooled empirical distribution
Under the assumptions above,
Consider the plug-in statistic
For each category,
while
Because \(B\) is fixed and every \(r_j>0\),
By Slutsky's theorem,
Thus the pooled plug-in preserves the same limiting distribution under both equality and the stated local alternatives.
Exact connection to Pearson's homogeneity statistic¶
The plug-in statistic is not merely analogous to the classical Pearson two-sample homogeneity statistic: it is exactly the same statistic.
The Pearson statistic for the \(2\times B\) contingency table is
Using
and
the two terms combine to
So the independent two-sample candidate has a classical foundation under the point null.
What remains novel, and unresolved, is the resemblance extension: replacing exact equality with nested composite tolerance regions and deriving valid least-favourable critical values.
What #44 establishes¶
For independent samples and fixed \(B\), with both sample proportions converging to an interior common limit:
- the effective sample size is \(nm/(n+m)\);
- the normalized difference has limiting covariance \(\Sigma(r)\);
- the covariance has rank \(B-1\);
- the known-weight quadratic form has a non-central \(\chi^2_{B-1}\) limit under local alternatives;
- pooled empirical weights are asymptotically valid;
- the pooled quadratic form is exactly Pearson's two-sample homogeneity statistic.
What #44 does not establish¶
This derivation does not yet establish:
- a two-sample \(\delta\)-resemblance hypothesis;
- a recommended two-sample tolerance;
- a least-favourable non-centrality over that tolerance set;
- two nested decision boundaries;
- finite-sample calibration;
- validity under overlapping or dependent samples.
Those questions remain in #45–#47.