Skip to content

API reference

The public API is documented directly from the package docstrings.

Core statistic

population_resemblance.prs

Core Population Resemblance Statistic implementation.

population_resemblance_statistic

population_resemblance_statistic(
    observed: Sequence[float] | npt.NDArray[np.floating],
    reference: Sequence[float] | npt.NDArray[np.floating],
) -> float

Compute the Population Resemblance Statistic (PRS).

The statistic is

.. math::

\mathrm{PRS} = \sum_j \frac{(\hat p_j - p_{0j})^2}{p_{0j}}.

Parameters:

Name Type Description Default
observed Sequence[float] | NDArray[floating]

Current categorical probability distribution.

required
reference Sequence[float] | NDArray[floating]

Reference categorical probability distribution.

required

Returns:

Type Description
float

The PRS value.

Raises:

Type Description
ValueError

If the vectors are invalid, have different lengths, or the reference contains zeros.

PRS decision framework

population_resemblance.framework

Delta-resemblance decision framework.

recommended_delta

recommended_delta(
    reference: Sequence[float] | npt.NDArray[np.floating],
    sample_size: int,
    *,
    c: float = 0.7,
) -> float

Compute the recommended sample-size-aware tolerance delta.

maximum_noncentrality

maximum_noncentrality(
    reference: Sequence[float] | npt.NDArray[np.floating],
    sample_size: int,
    delta: float,
) -> float

Compute the least-favourable non-centrality parameter.

critical_values

critical_values(
    reference: Sequence[float] | npt.NDArray[np.floating],
    sample_size: int,
    delta: float,
    *,
    m: float = 2.0,
    alpha1: float = 0.05,
    alpha2: float = 0.1,
) -> CriticalValues

Compute the lower and upper PRS critical values.

assess_population_resemblance

assess_population_resemblance(
    observed: Sequence[float] | npt.NDArray[np.floating],
    reference: Sequence[float] | npt.NDArray[np.floating],
    sample_size: int,
    *,
    delta: float | None = None,
    c: float = 0.7,
    m: float = 2.0,
    alpha1: float = 0.05,
    alpha2: float = 0.1,
) -> PopulationResemblanceResult

Assess an observed distribution with the PRS decision framework.

Count-based workflows

population_resemblance.counts

Utilities for assessing observed categorical count data.

counts_to_proportions

counts_to_proportions(
    counts: Sequence[int] | npt.NDArray[np.integer],
) -> npt.NDArray[np.float64]

Convert categorical counts to empirical proportions.

Parameters:

Name Type Description Default
counts Sequence[int] | NDArray[integer]

Observed category counts.

required

Returns:

Type Description
ndarray

Empirical category proportions that sum to one.

assess_population_counts

assess_population_counts(
    counts: Sequence[int] | npt.NDArray[np.integer],
    reference: Sequence[float] | npt.NDArray[np.floating],
    *,
    delta: float | None = None,
    c: float = 0.7,
    m: float = 2.0,
    alpha1: float = 0.05,
    alpha2: float = 0.1,
) -> PopulationResemblanceResult

Assess observed counts against a fixed reference distribution.

The current sample size is derived directly from the count vector, which prevents inconsistencies between empirical proportions and the supplied sample size.

Parameters:

Name Type Description Default
counts Sequence[int] | NDArray[integer]

Current observed counts by category.

required
reference Sequence[float] | NDArray[floating]

Fixed reference probability distribution.

required
delta float | None

Optional explicit resemblance tolerance. If omitted, the framework's sample-size-aware recommendation is used.

None
c float

Positive multiplier used when automatically calibrating delta.

0.7
m float

Multiplier defining the wider resemblance region.

2.0
alpha1 float

Error-control parameter for the upper decision boundary.

0.05
alpha2 float

Error-control parameter for the lower decision boundary.

0.1

Returns:

Type Description
PopulationResemblanceResult

Complete PRS assessment.

Sample-size planning

population_resemblance.planning

Structural sample-size planning for recommended PRS tolerances.

minimum_structural_sample_size

minimum_structural_sample_size(
    reference: Sequence[float] | npt.NDArray[np.floating],
    *,
    c: float = 0.7,
    m: float = 2.0,
) -> int

Return the smallest sample size satisfying the PRS probability bound.

For the recommended tolerance delta = c * min_j sqrt(p0_j * (1 - p0_j) / n), the nested PRS framework requires m * delta <= min_j p0_j.

This function solves that inequality for the minimum positive integer sample size. The result is a structural feasibility bound only. It does not guarantee adequate power, finite-sample calibration, or asymptotic accuracy.

Parameters:

Name Type Description Default
reference Sequence[float] | NDArray[floating]

Fixed reference categorical probability distribution.

required
c float

Multiplier used by the recommended sample-size-aware tolerance.

0.7
m float

Multiplier defining the wider resemblance region. Must exceed one.

2.0

Returns:

Type Description
int

Smallest positive integer sample size satisfying the structural bound.

Unified reporting

population_resemblance.reporting

Unified monitoring report across PRS, PSI, and discrete KS benchmarks.

PSIAssessment dataclass

PSIAssessment(statistic: float, status: str)

Population Stability Index result and Lewis benchmark status.

PopulationMonitoringReport dataclass

PopulationMonitoringReport(
    prs: PopulationResemblanceResult,
    psi: PSIAssessment,
    ks: DiscreteKSTestResult,
)

Unified monitoring results for one observed categorical sample.

statuses property

statuses: tuple[str, str, str]

Return PRS, PSI, and KS statuses in a compact tuple.

assess_population_monitoring

assess_population_monitoring(
    counts: Sequence[int] | npt.NDArray[np.integer],
    reference: Sequence[float] | npt.NDArray[np.floating],
    *,
    delta: float | None = None,
    c: float = 0.7,
    m: float = 2.0,
    alpha1: float = 0.05,
    alpha2: float = 0.1,
    ks_simulations: int = 10000,
    ks_seed: int | None = None,
) -> PopulationMonitoringReport

Assess one observed sample with PRS, PSI, and discrete KS.

PRS is the primary tolerance-based framework. PSI and discrete KS are returned as descriptive benchmarks because their decision rules differ from the PRS composite-null formulation.

Parameters:

Name Type Description Default
counts Sequence[int] | NDArray[integer]

Observed category counts.

required
reference Sequence[float] | NDArray[floating]

Fixed reference probability distribution.

required
delta float | None

Optional explicit PRS tolerance.

None
c float

Multiplier used for automatic PRS tolerance calibration.

0.7
m float

Wider PRS resemblance multiplier.

2.0
alpha1 float

Upper PRS error-control parameter.

0.05
alpha2 float

Lower PRS error-control parameter.

0.1
ks_simulations int

Monte Carlo sample count used to calibrate the discrete KS p-value.

10000
ks_seed int | None

Optional random seed used for the discrete KS calibration.

None

Returns:

Type Description
PopulationMonitoringReport

Unified PRS, PSI, and discrete KS results.

Named categories

population_resemblance.named

Named-category convenience API for population monitoring.

NamedPopulationMonitoringReport dataclass

NamedPopulationMonitoringReport(
    categories: tuple[str, ...],
    counts: tuple[int, ...],
    reference_probabilities: tuple[float, ...],
    monitoring: PopulationMonitoringReport,
)

Monitoring report with explicit category labels.

assess_named_population

assess_named_population(
    counts: Mapping[str, int],
    reference: Mapping[str, float],
    *,
    delta: float | None = None,
    c: float = 0.7,
    m: float = 2.0,
    alpha1: float = 0.05,
    alpha2: float = 0.1,
    ks_simulations: int = 10000,
    ks_seed: int | None = None,
) -> NamedPopulationMonitoringReport

Assess a categorical population using named categories.

Category keys are aligned by the insertion order of the reference mapping. The current-count mapping must contain exactly the same category names.

Reusable monitors

population_resemblance.monitor

Reusable configured population monitor.

PopulationMonitor dataclass

PopulationMonitor(
    reference: tuple[float, ...],
    delta: float | None = None,
    c: float = 0.7,
    m: float = 2.0,
    alpha1: float = 0.05,
    alpha2: float = 0.1,
    ks_simulations: int = 10000,
    ks_seed: int | None = None,
)

Reusable monitoring configuration for a fixed reference distribution.

from_reference classmethod

from_reference(
    reference: Sequence[float] | npt.NDArray[np.floating],
    *,
    delta: float | None = None,
    c: float = 0.7,
    m: float = 2.0,
    alpha1: float = 0.05,
    alpha2: float = 0.1,
    ks_simulations: int = 10000,
    ks_seed: int | None = None,
) -> PopulationMonitor

Create a monitor from a fixed reference probability vector.

assess

assess(
    counts: Sequence[int] | npt.NDArray[np.integer],
) -> PopulationMonitoringReport

Assess one current categorical sample.

assess_temporal

assess_temporal(
    counts_by_period: Sequence[Sequence[int]]
    | npt.NDArray[np.integer],
    *,
    labels: Sequence[str] | None = None,
) -> TemporalMonitoringSeries

Assess repeated population snapshots using the stored configuration.

NamedPopulationMonitor dataclass

NamedPopulationMonitor(
    reference: tuple[tuple[str, float], ...],
    delta: float | None = None,
    c: float = 0.7,
    m: float = 2.0,
    alpha1: float = 0.05,
    alpha2: float = 0.1,
    ks_simulations: int = 10000,
    ks_seed: int | None = None,
)

Reusable monitor for a fixed named-category reference distribution.

categories property

categories: tuple[str, ...]

Return category names in canonical monitoring order.

from_reference classmethod

from_reference(
    reference: Mapping[str, float],
    *,
    delta: float | None = None,
    c: float = 0.7,
    m: float = 2.0,
    alpha1: float = 0.05,
    alpha2: float = 0.1,
    ks_simulations: int = 10000,
    ks_seed: int | None = None,
) -> NamedPopulationMonitor

Create a monitor from a named fixed reference distribution.

assess

assess(
    counts: Mapping[str, int],
) -> NamedPopulationMonitoringReport

Assess one named-category current sample.

Diagnostics

population_resemblance.diagnostics

Category-level diagnostics for Population Resemblance Statistic results.

CategoryContribution dataclass

CategoryContribution(
    category: str,
    observed_probability: float,
    reference_probability: float,
    signed_shift: float,
    absolute_shift: float,
    prs_contribution: float,
    contribution_share: float,
)

Contribution of one category to the overall PRS value.

ResemblanceDiagnostics dataclass

ResemblanceDiagnostics(
    statistic: float,
    categories: tuple[CategoryContribution, ...],
)

Decomposition of PRS into category-level contributions.

largest_contributor property

largest_contributor: CategoryContribution

Return the category contributing most to the PRS statistic.

maximum_absolute_shift property

maximum_absolute_shift: float

Return the largest absolute category probability shift.

population_resemblance_diagnostics

population_resemblance_diagnostics(
    observed: Sequence[float] | npt.NDArray[np.floating],
    reference: Sequence[float] | npt.NDArray[np.floating],
    *,
    labels: Sequence[str] | None = None,
) -> ResemblanceDiagnostics

Decompose PRS into interpretable category-level contributions.

Parameters:

Name Type Description Default
observed Sequence[float] | NDArray[floating]

Current categorical probability distribution.

required
reference Sequence[float] | NDArray[floating]

Reference categorical probability distribution.

required
labels Sequence[str] | None

Optional category labels. If omitted, labels are generated as category_1, category_2, and so on.

None

Returns:

Type Description
ResemblanceDiagnostics

Overall PRS statistic and per-category contribution details.

population_count_diagnostics

population_count_diagnostics(
    counts: Sequence[int] | npt.NDArray[np.integer],
    reference: Sequence[float] | npt.NDArray[np.floating],
    *,
    labels: Sequence[str] | None = None,
) -> ResemblanceDiagnostics

Decompose PRS directly from observed category counts.

The empirical distribution is derived from the supplied count vector before delegating to the probability-based diagnostic implementation.

Calibration and sensitivity

population_resemblance.calibration

Calibration diagnostics for PRS decision parameters.

CalibrationDiagnostics dataclass

CalibrationDiagnostics(
    delta: float,
    widened_delta: float,
    lambda_sup: float,
    lower_critical_value: float,
    upper_critical_value: float,
    minimum_reference_probability: float,
    margin_to_probability_bound: float,
    boundaries_overlap: bool,
)

Derived quantities for one PRS calibration configuration.

feasible property

feasible: bool

Return whether the calibration satisfies structural constraints.

evaluate_calibration

evaluate_calibration(
    reference: Sequence[float] | npt.NDArray[np.floating],
    sample_size: int,
    *,
    delta: float | None = None,
    c: float = 0.7,
    m: float = 2.0,
    alpha1: float = 0.05,
    alpha2: float = 0.1,
) -> CalibrationDiagnostics

Evaluate one PRS calibration without performing a population assessment.

This exposes the main derived quantities that determine whether a chosen combination of tolerance and sensitivity parameters is practically valid.

Parameters:

Name Type Description Default
reference Sequence[float] | NDArray[floating]

Fixed reference probability distribution.

required
sample_size int

Current sample size.

required
delta float | None

Optional explicit tolerance. If omitted, the recommended sample-size- aware value is used.

None
c float

Multiplier used for automatic delta calibration.

0.7
m float

Wider resemblance multiplier.

2.0
alpha1 float

Upper decision error-control parameter.

0.05
alpha2 float

Lower decision error-control parameter.

0.1

Returns:

Type Description
CalibrationDiagnostics

Derived tolerance, non-centrality, critical values, and feasibility margins for the supplied configuration.

population_resemblance.sensitivity

Sensitivity grids for PRS calibration parameters.

CalibrationSensitivityPoint dataclass

CalibrationSensitivityPoint(
    c: float,
    m: float,
    alpha1: float,
    alpha2: float,
    diagnostics: CalibrationDiagnostics | None,
    error: str | None,
)

One parameter combination in a PRS calibration sensitivity grid.

feasible property

feasible: bool

Return whether this parameter combination produced a valid calibration.

CalibrationSensitivityGrid dataclass

CalibrationSensitivityGrid(
    points: tuple[CalibrationSensitivityPoint, ...],
)

Collection of PRS calibration results over a parameter grid.

feasible_points property

feasible_points: tuple[CalibrationSensitivityPoint, ...]

Return all valid calibration combinations.

infeasible_points property

infeasible_points: tuple[CalibrationSensitivityPoint, ...]

Return all parameter combinations rejected by the PRS constraints.

sweep_calibration_parameters

sweep_calibration_parameters(
    reference: Sequence[float] | npt.NDArray[np.floating],
    sample_size: int,
    *,
    c_values: Sequence[float] = (0.5, 0.7, 1.0),
    m_values: Sequence[float] = (1.2, 1.5, 2.0),
    alpha1_values: Sequence[float] = (0.05,),
    alpha2_values: Sequence[float] = (0.1,),
) -> CalibrationSensitivityGrid

Evaluate PRS calibration across a Cartesian parameter grid.

Invalid combinations are retained with an explanatory error rather than aborting the entire sweep. This makes structural constraints such as M * delta <= min(p0) visible during sensitivity analysis.

Simulation and operating characteristics

population_resemblance.simulation

Monte Carlo tools for studying PRS decision behavior.

SimulationResult dataclass

SimulationResult(
    simulations: int,
    r1_probability: float,
    r2_probability: float,
    r3_probability: float,
    mean_statistic: float,
)

Empirical PRS decision probabilities from a Monte Carlo experiment.

probabilities property

probabilities: tuple[float, float, float]

Return region probabilities in R1, R2, R3 order.

symmetric_category_shift

symmetric_category_shift(
    reference: Sequence[float] | npt.NDArray[np.floating],
    deviation: float,
) -> npt.NDArray[np.float64]

Construct the symmetric perturbation used in the source simulation study.

The first half of the categories is shifted downward by the supplied deviation and the last half upward by the same amount. For an odd number of categories, the central category remains unchanged.

Parameters:

Name Type Description Default
reference Sequence[float] | NDArray[floating]

Reference categorical probability distribution.

required
deviation float

Absolute probability shift applied to each perturbed category.

required

Returns:

Type Description
ndarray

Perturbed probability distribution.

Raises:

Type Description
ValueError

If the deviation is invalid or would create an invalid probability.

simulate_region_probabilities

simulate_region_probabilities(
    current: Sequence[float] | npt.NDArray[np.floating],
    reference: Sequence[float] | npt.NDArray[np.floating],
    sample_size: int,
    *,
    simulations: int = 10000,
    seed: int | None = None,
    delta: float | None = None,
    c: float = 0.7,
    m: float = 2.0,
    alpha1: float = 0.05,
    alpha2: float = 0.1,
    batch_size: int | None = None,
) -> SimulationResult

Estimate PRS decision probabilities under a specified current population.

Samples are drawn independently from a multinomial distribution with the supplied current probabilities. Each simulated empirical distribution is classified into the PRS regions R1, R2, or R3.

Parameters:

Name Type Description Default
current Sequence[float] | NDArray[floating]

True categorical probabilities from which simulated samples are drawn.

required
reference Sequence[float] | NDArray[floating]

Fixed reference categorical probability distribution.

required
sample_size int

Number of observations in each simulated sample.

required
simulations int

Number of independent Monte Carlo samples.

10000
seed int | None

Optional random seed for reproducibility.

None
delta float | None

Optional explicit resemblance tolerance. If omitted, the recommended sample-size-aware value is used.

None
c float

Tolerance multiplier used for automatic delta calibration.

0.7
m float

Multiplier defining the wider resemblance region.

2.0
alpha1 float

Error-control parameter for the upper decision boundary.

0.05
alpha2 float

Error-control parameter for the lower decision boundary.

0.1
batch_size int | None

Optional maximum number of Monte Carlo samples held in memory at once. If omitted, all simulations are generated in one batch.

None

Returns:

Type Description
SimulationResult

Empirical decision-region probabilities and mean PRS value.

population_resemblance.operating

Operating-characteristic curves for the PRS decision framework.

OperatingCharacteristicPoint dataclass

OperatingCharacteristicPoint(
    delta_multiple: float,
    deviation: float,
    result: SimulationResult,
)

One simulated operating-characteristic point.

OperatingCharacteristicCurve dataclass

OperatingCharacteristicCurve(
    delta: float,
    points: tuple[OperatingCharacteristicPoint, ...],
)

Ordered PRS operating-characteristic results over increasing deviations.

delta_multiples property

delta_multiples: tuple[float, ...]

Return deviation sizes measured in multiples of delta.

deviations property

deviations: tuple[float, ...]

Return absolute category-wise deviations.

r1_probabilities property

r1_probabilities: tuple[float, ...]

Return empirical R1 probabilities along the curve.

r2_probabilities property

r2_probabilities: tuple[float, ...]

Return empirical R2 probabilities along the curve.

r3_probabilities property

r3_probabilities: tuple[float, ...]

Return empirical R3 probabilities along the curve.

source_deviation_grid

source_deviation_grid(
    *, m: float = 2.0, points: int = 30
) -> tuple[float, ...]

Return the deviation grid used by the source simulation design.

The source paper studies deviations from zero to (3M + 2) times delta. This helper returns equally spaced multiples of delta over that interval.

Parameters:

Name Type Description Default
m float

Wider resemblance multiplier.

2.0
points int

Number of equally spaced deviation values.

30

Returns:

Type Description
tuple[float, ...]

Deviation sizes expressed as multiples of delta.

simulate_operating_characteristic_curve

simulate_operating_characteristic_curve(
    reference: Sequence[float] | npt.NDArray[np.floating],
    sample_size: int,
    *,
    delta_multiples: Sequence[float] | None = None,
    simulations: int = 10000,
    seed: int | None = None,
    delta: float | None = None,
    c: float = 0.7,
    m: float = 2.0,
    alpha1: float = 0.05,
    alpha2: float = 0.1,
    batch_size: int | None = None,
) -> OperatingCharacteristicCurve

Simulate PRS decision probabilities over a sequence of shift magnitudes.

Each point uses the symmetric category perturbation employed in the source simulation study. The result is an operating-characteristic curve rather than conventional statistical power, because the PRS framework uses nested composite null hypotheses.

Parameters:

Name Type Description Default
reference Sequence[float] | NDArray[floating]

Fixed reference categorical probability distribution.

required
sample_size int

Number of observations in each Monte Carlo sample.

required
delta_multiples Sequence[float] | None

Shift magnitudes expressed as multiples of delta. If omitted, the source paper's 30-point grid from zero to (3M + 2) delta is used. This grid is not clipped to the probability domain. Supply explicit feasible multiples when its shifts would make any category probability negative or greater than one.

None
simulations int

Monte Carlo samples per curve point.

10000
seed int | None

Optional base random seed. Each curve point receives a deterministic seed offset.

None
delta float | None

Optional explicit resemblance tolerance.

None
c float

Multiplier used for automatic delta calibration.

0.7
m float

Wider resemblance multiplier.

2.0
alpha1 float

Upper PRS error-control parameter.

0.05
alpha2 float

Lower PRS error-control parameter.

0.1
batch_size int | None

Optional Monte Carlo batch size forwarded to each curve-point simulation.

None

Returns:

Type Description
OperatingCharacteristicCurve

Ordered simulation results across the requested deviation magnitudes.

Raises:

Type Description
ValueError

If a requested shift produces invalid category probabilities, or the calibration or simulation parameters are invalid. For example, the default grid is infeasible for a uniform five-category reference with sample size 50 and the default calibration; multiples from 0 to 5 are feasible for that configuration.

population_resemblance.uncertainty

Uncertainty summaries for Monte Carlo classification probabilities.

ProbabilityInterval dataclass

ProbabilityInterval(
    estimate: float, lower: float, upper: float
)

Point estimate and confidence interval for a simulated probability.

SimulationUncertainty dataclass

SimulationUncertainty(
    confidence_level: float,
    r1: ProbabilityInterval,
    r2: ProbabilityInterval,
    r3: ProbabilityInterval,
)

Confidence intervals for PRS region probabilities.

simulation_uncertainty

simulation_uncertainty(
    result: SimulationResult,
    *,
    confidence_level: float = 0.95,
) -> SimulationUncertainty

Attach Wilson confidence intervals to Monte Carlo region probabilities.

These intervals quantify Monte Carlo simulation error only. They are not confidence intervals for the underlying PRS parameters or population shift.

Benchmark methods

population_resemblance.psi

Population Stability Index benchmark implementation.

population_stability_index

population_stability_index(
    observed: Sequence[float] | npt.NDArray[np.floating],
    reference: Sequence[float] | npt.NDArray[np.floating],
) -> float

Compute the Population Stability Index (PSI).

This follows the formulation used in the source paper, including its convention of omitting categories whose observed probability is zero. Reference probabilities must be strictly positive. No smoothing is applied, so zero-count results can differ from PSI implementations that smooth probabilities or return infinity for a zero observed probability.

lewis_psi_status

lewis_psi_status(psi: float) -> str

Classify PSI using the traditional Lewis rule-of-thumb thresholds.

The thresholds are retained only as a benchmark: below 0.10 is green, 0.10 up to 0.25 is amber, and 0.25 or above is red.

population_resemblance.ks

Discrete Kolmogorov-Smirnov benchmark for categorical distributions.

DiscreteKSTestResult dataclass

DiscreteKSTestResult(
    statistic: float,
    p_value: float,
    simulations: int,
    sample_size: int,
)

Monte Carlo calibrated discrete Kolmogorov-Smirnov test result.

status property

status: str

Return the paper-style green/amber/red p-value classification.

discrete_ks_statistic

discrete_ks_statistic(
    observed: Sequence[float] | npt.NDArray[np.floating],
    reference: Sequence[float] | npt.NDArray[np.floating],
) -> float

Compute the discrete Kolmogorov-Smirnov statistic.

The statistic is the maximum absolute difference between the empirical cumulative distribution and the reference cumulative distribution. Cumulative probabilities use the supplied category order. Reordering both vectors together can change the result, so categories should have a meaningful fixed order; nominal labels do not define a unique KS statistic.

Parameters:

Name Type Description Default
observed Sequence[float] | NDArray[floating]

Current categorical probability distribution.

required
reference Sequence[float] | NDArray[floating]

Reference categorical probability distribution.

required

Returns:

Type Description
float

Maximum absolute cumulative-distribution difference.

discrete_ks_test_counts

discrete_ks_test_counts(
    counts: Sequence[int] | npt.NDArray[np.integer],
    reference: Sequence[float] | npt.NDArray[np.floating],
    *,
    simulations: int = 10000,
    seed: int | None = None,
) -> DiscreteKSTestResult

Monte Carlo calibrate the discrete KS statistic under the reference model.

The null distribution is simulated from the multinomial reference because the discrete KS statistic is not distribution-free with respect to the underlying categorical probabilities. Counts and reference probabilities must use the same meaningful category order. Changing that order can change the statistic and calibrated p-value.

Parameters:

Name Type Description Default
counts Sequence[int] | NDArray[integer]

Observed category counts.

required
reference Sequence[float] | NDArray[floating]

Reference categorical probability distribution.

required
simulations int

Number of Monte Carlo samples used for calibration.

10000
seed int | None

Optional random seed for reproducibility.

None

Returns:

Type Description
DiscreteKSTestResult

Statistic, Monte Carlo p-value, simulation count, and sample size.

discrete_ks_status

discrete_ks_status(
    p_value: float,
    *,
    red_alpha: float = 0.01,
    green_alpha: float = 0.1,
) -> str

Classify a discrete KS p-value using the paper's comparison thresholds.

Values below 1% are red, values above 10% are green, and intermediate values are amber by default.

population_resemblance.comparison

Simulation tools for comparing PRS and PSI classifications.

ClassificationProbabilities dataclass

ClassificationProbabilities(
    green: float, amber: float, red: float
)

Empirical green/amber/red classification probabilities.

total property

total: float

Return the total probability mass.

PRSPSIComparisonResult dataclass

PRSPSIComparisonResult(
    simulations: int,
    prs: ClassificationProbabilities,
    psi: ClassificationProbabilities,
)

Monte Carlo comparison between PRS and PSI classifications.

simulate_prs_psi_comparison

simulate_prs_psi_comparison(
    current: Sequence[float] | npt.NDArray[np.floating],
    reference: Sequence[float] | npt.NDArray[np.floating],
    sample_size: int,
    *,
    simulations: int = 10000,
    seed: int | None = None,
    delta: float | None = None,
    c: float = 0.7,
    m: float = 2.0,
    alpha1: float = 0.05,
    alpha2: float = 0.1,
) -> PRSPSIComparisonResult

Estimate PRS and PSI classification frequencies on identical samples.

PRS classifications use the delta-resemblance decision boundaries. PSI classifications use the traditional Lewis thresholds of 0.10 and 0.25.

Parameters:

Name Type Description Default
current Sequence[float] | NDArray[floating]

True categorical probabilities used to generate samples.

required
reference Sequence[float] | NDArray[floating]

Fixed reference categorical distribution.

required
sample_size int

Number of observations per Monte Carlo sample.

required
simulations int

Number of simulated samples.

10000
seed int | None

Optional random seed.

None
delta float | None

Optional explicit PRS tolerance.

None
c float

Multiplier used for automatic delta calibration.

0.7
m float

Wider resemblance multiplier.

2.0
alpha1 float

Upper PRS error-control parameter.

0.05
alpha2 float

Lower PRS error-control parameter.

0.1

Returns:

Type Description
PRSPSIComparisonResult

Empirical classification probabilities for both methods.

Reference and temporal workflows

population_resemblance.reference

Convenience API for empirical reference distributions built from counts.

ReferenceCountMonitoringReport dataclass

ReferenceCountMonitoringReport(
    current_sample_size: int,
    reference_sample_size: int,
    reference_probabilities: tuple[float, ...],
    monitoring: PopulationMonitoringReport,
)

Monitoring report using an empirical reference sample conditionally as fixed.

assess_against_reference_counts

assess_against_reference_counts(
    current_counts: Sequence[int] | npt.NDArray[np.integer],
    reference_counts: Sequence[int]
    | npt.NDArray[np.integer],
    *,
    delta: float | None = None,
    c: float = 0.7,
    m: float = 2.0,
    alpha1: float = 0.05,
    alpha2: float = 0.1,
    ks_simulations: int = 10000,
    ks_seed: int | None = None,
) -> ReferenceCountMonitoringReport

Assess current counts against a reference sample represented by counts.

The reference counts are converted to empirical probabilities and then treated as fixed for the existing one-sample PRS framework. This is a convenience wrapper for conditional inference; it is not a two-sample PRS test and does not propagate reference-sample uncertainty into the critical values.

Parameters:

Name Type Description Default
current_counts Sequence[int] | NDArray[integer]

Current observed counts by category.

required
reference_counts Sequence[int] | NDArray[integer]

Baseline or model-development counts by the same categories.

required
delta float | None

Optional explicit PRS tolerance.

None
c float

Multiplier used for automatic PRS tolerance calibration.

0.7
m float

Wider resemblance multiplier.

2.0
alpha1 float

Upper PRS error-control parameter.

0.05
alpha2 float

Lower PRS error-control parameter.

0.1
ks_simulations int

Monte Carlo sample count used to calibrate the discrete KS benchmark.

10000
ks_seed int | None

Optional random seed used for the discrete KS calibration.

None

Returns:

Type Description
ReferenceCountMonitoringReport

Sample-size metadata, empirical reference probabilities, and the unified PRS/PSI/KS monitoring report.

population_resemblance.temporal

Temporal monitoring utilities for repeated categorical population snapshots.

TemporalMonitoringPoint dataclass

TemporalMonitoringPoint(
    label: str, report: PopulationMonitoringReport
)

Monitoring result for one labeled population snapshot.

TemporalMonitoringSeries dataclass

TemporalMonitoringSeries(
    points: tuple[TemporalMonitoringPoint, ...],
)

Ordered monitoring results across multiple population snapshots.

labels property

labels: tuple[str, ...]

Return snapshot labels in temporal order.

prs_statistics property

prs_statistics: tuple[float, ...]

Return PRS statistics in temporal order.

psi_statistics property

psi_statistics: tuple[float, ...]

Return PSI statistics in temporal order.

ks_statistics property

ks_statistics: tuple[float, ...]

Return discrete KS statistics in temporal order.

assess_temporal_monitoring

assess_temporal_monitoring(
    counts_by_period: Sequence[Sequence[int]]
    | npt.NDArray[np.integer],
    reference: Sequence[float] | npt.NDArray[np.floating],
    *,
    labels: Sequence[str] | None = None,
    delta: float | None = None,
    c: float = 0.7,
    m: float = 2.0,
    alpha1: float = 0.05,
    alpha2: float = 0.1,
    ks_simulations: int = 10000,
    ks_seed: int | None = None,
) -> TemporalMonitoringSeries

Assess repeated categorical samples against one fixed reference distribution.

Optional labels default to period_1, period_2, and so on. If a base KS seed is supplied, each period receives a deterministic seed offset.

Serialization

population_resemblance.serialization

Serialization helpers for monitoring results.

monitoring_report_to_dict

monitoring_report_to_dict(
    report: PopulationMonitoringReport,
) -> dict[str, Any]

Convert a population monitoring report to a JSON-safe dictionary.

temporal_series_to_records

temporal_series_to_records(
    series: TemporalMonitoringSeries,
) -> list[dict[str, Any]]

Convert a temporal monitoring series to flat JSON-safe records.

records_column

records_column(
    records: Sequence[dict[str, Any]], name: str
) -> tuple[Any, ...]

Extract one column from serialized temporal records.

Experimental two-sample API

population_resemblance.experimental.two_sample

Experimental independent two-sample population resemblance assessment.

This module is intentionally outside the package-root stable API. It implements the independent-sample plug-in procedure derived in the repository research notes.

ExperimentalTwoSampleResult dataclass

ExperimentalTwoSampleResult(
    statistic: float,
    effective_sample_size: float,
    pooled_probabilities: tuple[float, ...],
    delta: float,
    lambda_sup: float,
    critical_values: CriticalValues,
    region: DecisionRegion,
    first_sample_size: int,
    second_sample_size: int,
    minimum_expected_count: float,
    structural_margin: float,
)

Result of an experimental independent two-sample resemblance assessment.

label property

label: str

Return the descriptive decision-region label.

assess_independent_two_sample_resemblance

assess_independent_two_sample_resemblance(
    first_counts: Sequence[int] | npt.NDArray[np.integer],
    second_counts: Sequence[int] | npt.NDArray[np.integer],
    *,
    samples_are_independent: bool,
    delta: float | None = None,
    c: float = 0.7,
    m: float = 2.0,
    alpha1: float = 0.05,
    alpha2: float = 0.1,
    minimum_expected_count: float = 5.0,
) -> ExperimentalTwoSampleResult

Assess two independent categorical samples with the experimental method.

This API is experimental and intentionally not exported from the stable package-root public API.