Installation

📊 StatFlow

A modern statistical analysis toolkit for Python

v1.2.3
Python
MIT

Installation

Installation

Terminal
pip install statflow

Requirements

  • Python 3.8 or higher
  • NumPy >= 1.20.0
  • SciPy >= 1.7.0
  • Pandas >= 1.3.0

Optional Dependencies

For advanced visualization features:

1
pip install statflow[viz]

For GPU acceleration:

1
pip install statflow[gpu]

Quick Start

Here's a simple example to get you started with StatFlow:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
import statflow as sf
import numpy as np

# Generate sample data
np.random.seed(42)
x = np.random.normal(100, 15, 1000)
y = np.random.normal(105, 15, 1000)

# Perform t-test
result = sf.ttest(x, y)
print(f"t-statistic: {result.statistic:.4f}")
print(f"p-value: {result.pvalue:.4f}")

# Fit a linear regression
model = sf.LinearRegression()
model.fit(x.reshape(-1, 1), y)
print(f"R-squared: {model.rsquared:.4f}")

Basic Concepts

StatFlow is built around three core concepts:

1. Statistical Tests

All statistical tests in StatFlow return a TestResult object with standardized attributes:

  • statistic: The test statistic value
  • pvalue: The p-value
  • confidence_interval: Confidence intervals (when applicable)
  • effect_size: Standardized effect size measures

2. Models

Statistical models follow the scikit-learn API pattern:

  • fit(X, y): Fit the model to data
  • predict(X): Make predictions
  • summary(): Get detailed model statistics

3. Distributions

Work with probability distributions using a consistent interface:

1
2
3
4
dist = sf.Normal(mu=0, sigma=1)
samples = dist.sample(1000)
pdf_values = dist.pdf(samples)
cdf_values = dist.cdf(samples)

Statistical Tests

StatFlow provides a comprehensive suite of statistical tests for various scenarios.

ttest

statflow.ttest(x, y=None, paired=False, alternative='two-sided', confidence=0.95)

Perform a Student's t-test to compare means.

Parameters

Returns

TestResult

Examples

import statflow as sf
import numpy as np

# Two-sample t-test
group1 = np.random.normal(100, 15, 50)
group2 = np.random.normal(105, 15, 50)
result = sf.ttest(group1, group2)
print(f"p-value: {result.pvalue:.4f}")

# One-sample t-test
sample = np.random.normal(100, 15, 100)
result = sf.ttest(sample, mu=95)
print(f"95% CI: {result.confidence_interval}")

See Also

mannwhitneyu

statflow.mannwhitneyu(x, y, alternative='two-sided')

Perform the Mann-Whitney U test for independent samples.

Parameters

Returns

TestResult

Examples

# Non-parametric alternative to t-test
result = sf.mannwhitneyu(group1, group2)
print(f"U statistic: {result.statistic:.2f}")

See Also

anova

statflow.anova(*groups, post_hoc=None)

Perform one-way analysis of variance (ANOVA).

Parameters

Returns

ANOVAResult

Examples

# Compare multiple groups
group1 = np.random.normal(100, 15, 30)
group2 = np.random.normal(105, 15, 30)
group3 = np.random.normal(110, 15, 30)

result = sf.anova(group1, group2, group3, post_hoc='tukey')
print(result.summary())

See Also

Regression Models

Build and evaluate regression models with ease.

LinearRegression

statflow.LinearRegression(fit_intercept=True, normalize=False)

Ordinary least squares linear regression.

Parameters

Returns

LinearRegression

Examples

# Fit linear regression
X = np.random.randn(100, 3)
y = 2*X[:, 0] + 3*X[:, 1] - X[:, 2] + np.random.randn(100)*0.5

model = sf.LinearRegression()
model.fit(X, y)

print(f"Coefficients: {model.coef_}")
print(f"R-squared: {model.rsquared:.4f}")
print(f"Adjusted R-squared: {model.rsquared_adj:.4f}")

# Get detailed summary
print(model.summary())

# Make predictions
predictions = model.predict(X)

Distributions

Work with probability distributions using a unified interface.

Normal

statflow.Normal(mu=0, sigma=1)

Normal (Gaussian) distribution.

Parameters

Returns

Normal

Examples

# Create a normal distribution
dist = sf.Normal(mu=100, sigma=15)

# Sample from the distribution
samples = dist.sample(1000)

# Compute probability density
x = np.linspace(50, 150, 100)
pdf_values = dist.pdf(x)

# Compute cumulative distribution
cdf_values = dist.cdf(x)

# Get quantiles
median = dist.quantile(0.5)
q95 = dist.quantile(0.95)

Utilities

Helper functions for data analysis and statistics.

bootstrap

statflow.bootstrap(data, statistic, n_resamples=10000, confidence=0.95, random_state=None)

Compute bootstrap confidence intervals for a statistic.

Parameters

Returns

BootstrapResult

Examples

# Estimate confidence interval for median
data = np.random.exponential(scale=2, size=100)
result = sf.bootstrap(data, np.median, random_state=42)

print(f"Median: {np.median(data):.4f}")
print(f"95% CI: {result.confidence_interval}")
print(f"Bootstrap SE: {result.standard_error:.4f}")

Hypothesis Testing

Example: A/B Testing

Compare conversion rates between two groups:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
import statflow as sf
import numpy as np

# Simulate A/B test data
np.random.seed(42)
group_a_conversions = np.random.binomial(1, 0.10, 1000)  # 10% conversion
group_b_conversions = np.random.binomial(1, 0.12, 1000)  # 12% conversion

# Perform proportion test
result = sf.proportions_test(
    [group_a_conversions.sum(), group_b_conversions.sum()],
    [len(group_a_conversions), len(group_b_conversions)]
)

print(f"Group A conversion: {group_a_conversions.mean():.2%}")
print(f"Group B conversion: {group_b_conversions.mean():.2%}")
print(f"p-value: {result.pvalue:.4f}")

if result.pvalue < 0.05:
    print("Significant difference detected!")
else:
    print("No significant difference.")

Example: Multiple Comparisons

When testing multiple hypotheses, control for false discovery rate:

1
2
3
4
5
6
7
8
9
10
11
# Simulate 20 hypothesis tests
p_values = [sf.ttest(
    np.random.normal(0, 1, 100),
    np.random.normal(0.2 if i < 5 else 0, 1, 100)
).pvalue for i in range(20)]

# Apply Benjamini-Hochberg correction
adjusted = sf.multipletests(p_values, method='fdr_bh', alpha=0.05)

print(f"Significant tests (uncorrected): {sum(p < 0.05 for p in p_values)}")
print(f"Significant tests (FDR corrected): {sum(adjusted.reject)}")

Regression Analysis

Example: Multiple Linear Regression

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
import statflow as sf
import pandas as pd
import numpy as np

# Generate synthetic data
np.random.seed(42)
n = 200

data = pd.DataFrame({
    'hours_studied': np.random.uniform(0, 10, n),
    'sleep_hours': np.random.uniform(4, 10, n),
    'previous_score': np.random.uniform(50, 100, n)
})

# Generate exam scores with some noise
data['exam_score'] = (
    3.5 * data['hours_studied'] +
    2.0 * data['sleep_hours'] +
    0.5 * data['previous_score'] +
    np.random.normal(0, 5, n)
)

# Fit model
X = data[['hours_studied', 'sleep_hours', 'previous_score']]
y = data['exam_score']

model = sf.LinearRegression()
model.fit(X, y)

# Print summary
print(model.summary())

# Check assumptions
diagnostics = model.diagnostics()
print(f"Durbin-Watson: {diagnostics.durbin_watson:.2f}")
print(f"Condition Number: {diagnostics.condition_number:.2f}")

# Visualize residuals
model.plot_residuals()

Example: Logistic Regression

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
# Binary classification
X = np.random.randn(500, 5)
z = 1.5*X[:, 0] - 0.8*X[:, 1] + np.random.randn(500)*0.5
y = (z > 0).astype(int)

model = sf.LogisticRegression()
model.fit(X, y)

# Predictions
probabilities = model.predict_proba(X)
predictions = model.predict(X)

# Evaluate
print(f"Accuracy: {(predictions == y).mean():.2%}")
print(f"AUC-ROC: {model.roc_auc_score(X, y):.4f}")

# Confusion matrix
print(model.confusion_matrix(y, predictions))

Time Series Analysis

Example: Autocorrelation Analysis

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
import statflow as sf
import numpy as np

# Generate AR(1) process
np.random.seed(42)
n = 500
rho = 0.7
y = np.zeros(n)
y[0] = np.random.randn()

for t in range(1, n):
    y[t] = rho * y[t-1] + np.random.randn()

# Compute autocorrelation
acf_values = sf.acf(y, nlags=20)
pacf_values = sf.pacf(y, nlags=20)

# Plot ACF/PACF
sf.plot_acf(y, lags=20)
sf.plot_pacf(y, lags=20)

# Test for stationarity
adf_result = sf.adfuller(y)
print(f"ADF Statistic: {adf_result.statistic:.4f}")
print(f"p-value: {adf_result.pvalue:.4f}")

Example: Moving Average Smoothing

1
2
3
4
5
6
7
8
9
# Noisy time series
t = np.linspace(0, 10, 200)
signal = np.sin(t) + 0.3 * np.random.randn(200)

# Apply moving average
smoothed = sf.moving_average(signal, window=10)

# Exponential smoothing
ema = sf.exponential_smoothing(signal, alpha=0.3)

Note

For more advanced time series modeling including ARIMA, SARIMA, and state space models, see a dedicated time series package.

Changelog

Version 1.2.3 (2025-01-15)

Added:

  • New bootstrap() function with percentile and BCa methods
  • Support for weighted statistics in most functions
  • GPU acceleration for large-scale computations (requires statflow[gpu])

Fixed:

  • Improved numerical stability in LinearRegression for ill-conditioned matrices
  • Fixed edge case in mannwhitneyu() with ties
  • Corrected confidence interval calculation in paired t-tests

Changed:

  • Updated minimum NumPy version to 1.20.0 for better type hints
  • Improved performance of anova() by 40% through vectorization

Version 1.1.0 (2024-11-20)

Added:

  • Multiple comparison corrections (multipletests())
  • Proportion tests (proportions_test())
  • Effect size calculations for all major tests

Fixed:

  • Memory leak in bootstrap resampling
  • Incorrect degrees of freedom in ANOVA with unequal sample sizes

Version 1.0.0 (2024-09-01)

Initial release with core functionality:

  • Basic statistical tests (t-test, Mann-Whitney U, ANOVA, etc.)
  • Linear and logistic regression
  • Common probability distributions
  • Time series utilities

Contributing

We welcome contributions! Please see our Contributing Guide for details.

License

StatFlow is released under the MIT License. See LICENSE for details.

Citation

If you use StatFlow in your research, please cite:

1
2
3
4
5
6
7
@software{statflow2024,
  title = {StatFlow: A Modern Statistical Analysis Toolkit},
  author = {StatFlow Contributors},
  year = {2024},
  url = {https://github.com/example/statflow},
  version = {1.2.3}
}
Loading mathematical content