Dependence and overlapping two-sample populations¶
This note addresses research issue #47.
All previous two-sample derivations assumed independent multinomial samples. That assumption is not innocuous. In monitoring applications, baseline and current samples may overlap, reuse the same individuals, use rolling windows, or contain repeated observations.
The source PRS paper explicitly notes this complication and therefore keeps the reference population fixed in its published one-sample formulation. The derivation here studies what changes when both samples are random and dependent.
General covariance identity¶
Let
be empirical category-proportion vectors.
For the difference,
the covariance is
Under independence, the cross-covariance terms vanish. Under overlap or repeated sampling, they generally do not.
This means that the independent-sample effective size
is not universally valid.
Exact overlap model¶
Consider a simple case in which:
- sample 1 has size \(n\);
- sample 2 has size \(m\);
- exactly \(k\) observations are shared by both samples;
- shared observations contribute the same category indicator to both samples;
- all non-shared observations are independent draws from the same categorical distribution \(r\).
Let
Then
and the shared observations give
Therefore
Define
Then
has the same first-order covariance \(\Sigma(r)\) as the independent construction.
Special cases¶
If \(k=0\),
recovering the independent result.
If sample 1 is fully contained in sample 2, with \(k=n<m\),
If \(n=m=k\), the two empirical distributions are identical and
The comparison is then degenerate and there is no non-trivial chi-square test.
Consequence for Pearson-style scaling¶
In the exact-overlap model above, dependence changes only the scalar multiplying \(\Sigma(r)\). Therefore the Pearson geometry survives after replacing the independent effective size by
Using the independent effective size when \(k>0\) understates the correct scaling factor. Because the shared observations reduce the variance of the sample difference, the naive independent statistic is systematically too small and becomes conservative under the point null.
This correction is valid only because the cross-covariance has exactly the same \(\Sigma(r)\) structure as the marginal covariance matrices.
Paired repeated measurements¶
A more general repeated-measures setting does not have that scalar structure.
Suppose \(N\) units are observed at two occasions. Let
be the \(B\times B\) joint transition-probability matrix, with marginals \(p\) and \(q\).
For one paired observation, define one-hot vectors \(U\) and \(V\). Then
For \(N\) independent pairs,
Hence
Even when \(p=q=r\), this need not be a scalar multiple of \(\Sigma(r)\).
Therefore a single adjusted effective sample size is generally insufficient.
General covariance-corrected quadratic form¶
For arbitrary dependence, define
Because the category differences sum to zero, \(V\) is singular in the full \(B\)-dimensional space.
Choose any full-row-rank contrast matrix
whose rows span the simplex contrast space. For example,
Let
and
If \(V_C\) is nonsingular, the Wald statistic
has an asymptotic \(\chi^2_{B-1}\) distribution under equality, provided a consistent estimator of \(V_C\) is available.
Under local alternatives with mean contrast \(\mu_C\), the asymptotic non-centrality is
This is the correct general framework for dependent samples.
Why this complicates resemblance calibration¶
The independent resemblance derivation used a diagonal Pearson weighting and a simple closed-form least-favourable non-centrality.
With general dependence, the relevant quadratic form is
not
The least-favourable problem therefore becomes
The one-sample even/odd closed form does not generally survive because the covariance inverse couples categories.
For a known positive-definite \(V_C\), this is a convex quadratic maximization over a polytope. The maximum occurs at an extreme point, but its value depends on the full dependence structure.
Which dependence settings are tractable¶
Known overlap count with identical shared observations¶
This case is analytically tractable.
Required information:
- \(n\);
- \(m\);
- the exact overlap count \(k\).
The scalar effective-size correction above is sufficient under the stated sampling model.
Fully paired repeated observations¶
This case is tractable if the paired category transitions are observed.
The empirical joint transition table estimates \(\Pi\), which gives a direct plug-in estimate of the cross-covariance.
A contrast-space Wald statistic can then be constructed.
Clustered or repeated observations with unit identifiers¶
A sandwich covariance estimator is a natural route when contributions can be grouped by independent clusters or individuals.
The cluster identifier must be available. Aggregate marginal counts alone are not enough.
Rolling windows¶
If the windows share observations and the identities of those shared observations are known, the exact-overlap correction may apply when the repeated records are literally identical.
If observations evolve between windows, a joint transition or cluster-level covariance estimate is required.
Temporal dependence without matched identities¶
Marginal counts from two time points are insufficient to identify the cross-covariance.
A time-series model, block bootstrap, or another dependence model is required.
What aggregate counts cannot identify¶
Suppose only two marginal count vectors are supplied.
Those counts determine \(\widehat p\) and \(\widehat q\), but they do not determine
Many different joint dependence structures can have the same marginals.
Therefore no statistically principled generic dependent-sample correction can be inferred from two count vectors alone.
This is an identifiability limitation, not an implementation limitation.
Bootstrap and resampling¶
When individual-level or cluster-level data are available but an analytic covariance model is undesirable, resampling may be used.
The resampling unit must preserve the dependence structure:
- paired bootstrap for matched observations;
- cluster bootstrap for grouped repeated observations;
- block bootstrap for temporally dependent sequences.
Independent resampling of the two margins would destroy the dependence and is invalid.
Result of issue #47¶
The dependence problem separates into two classes.
-
Scalar-covariance dependence, such as exact shared observations: an adjusted effective sample size can recover the independent Pearson geometry.
-
General dependence, such as paired transitions or temporal dependence: the method requires a full contrast-space covariance estimator and a Wald-type quadratic form.
For the general case, the simple PRS least-favourable closed form is lost.
Implication for the first two-sample implementation¶
The first public two-sample resemblance implementation should be restricted to independent samples.
A later API could support an explicit exact-overlap model if the overlap count is known.
General paired, repeated, clustered, or temporally dependent samples should remain out of scope until a covariance-aware resemblance calibration is separately derived and validated.
This restriction is preferable to silently applying the independent formula where its assumptions do not hold.