Installation
Installation
pip install statflow
conda install statflow
git clone https://github.com/example/statflow
cd statflow
pip install -e .
Requirements
- Python 3.8 or higher
- NumPy >= 1.20.0
- SciPy >= 1.7.0
- Pandas >= 1.3.0
Optional Dependencies
For advanced visualization features:
1
pip install statflow[viz]
For GPU acceleration:
1
pip install statflow[gpu]
Quick Start
Here's a simple example to get you started with StatFlow:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
import statflow as sf
import numpy as np
# Generate sample data
np.random.seed(42)
x = np.random.normal(100, 15, 1000)
y = np.random.normal(105, 15, 1000)
# Perform t-test
result = sf.ttest(x, y)
print(f"t-statistic: {result.statistic:.4f}")
print(f"p-value: {result.pvalue:.4f}")
# Fit a linear regression
model = sf.LinearRegression()
model.fit(x.reshape(-1, 1), y)
print(f"R-squared: {model.rsquared:.4f}")
Basic Concepts
StatFlow is built around three core concepts:
1. Statistical Tests
All statistical tests in StatFlow return a TestResult object with standardized attributes:
statistic: The test statistic valuepvalue: The p-valueconfidence_interval: Confidence intervals (when applicable)effect_size: Standardized effect size measures
2. Models
Statistical models follow the scikit-learn API pattern:
fit(X, y): Fit the model to datapredict(X): Make predictionssummary(): Get detailed model statistics
3. Distributions
Work with probability distributions using a consistent interface:
1
2
3
4
dist = sf.Normal(mu=0, sigma=1)
samples = dist.sample(1000)
pdf_values = dist.pdf(samples)
cdf_values = dist.cdf(samples)
Statistical Tests
StatFlow provides a comprehensive suite of statistical tests for various scenarios.
ttest
statflow.ttest(x, y=None, paired=False, alternative='two-sided', confidence=0.95)
Perform a Student's t-test to compare means.
Parameters
Returns
TestResult
Examples
import statflow as sf
import numpy as np
# Two-sample t-test
group1 = np.random.normal(100, 15, 50)
group2 = np.random.normal(105, 15, 50)
result = sf.ttest(group1, group2)
print(f"p-value: {result.pvalue:.4f}")
# One-sample t-test
sample = np.random.normal(100, 15, 100)
result = sf.ttest(sample, mu=95)
print(f"95% CI: {result.confidence_interval}")
See Also
mannwhitneyu
statflow.mannwhitneyu(x, y, alternative='two-sided')
Perform the Mann-Whitney U test for independent samples.
Parameters
Returns
TestResult
Examples
# Non-parametric alternative to t-test
result = sf.mannwhitneyu(group1, group2)
print(f"U statistic: {result.statistic:.2f}")
See Also
anova
statflow.anova(*groups, post_hoc=None)
Perform one-way analysis of variance (ANOVA).
Parameters
Returns
ANOVAResult
Examples
# Compare multiple groups
group1 = np.random.normal(100, 15, 30)
group2 = np.random.normal(105, 15, 30)
group3 = np.random.normal(110, 15, 30)
result = sf.anova(group1, group2, group3, post_hoc='tukey')
print(result.summary())
See Also
Regression Models
Build and evaluate regression models with ease.
LinearRegression
statflow.LinearRegression(fit_intercept=True, normalize=False)
Ordinary least squares linear regression.
Parameters
Returns
LinearRegression
Examples
# Fit linear regression
X = np.random.randn(100, 3)
y = 2*X[:, 0] + 3*X[:, 1] - X[:, 2] + np.random.randn(100)*0.5
model = sf.LinearRegression()
model.fit(X, y)
print(f"Coefficients: {model.coef_}")
print(f"R-squared: {model.rsquared:.4f}")
print(f"Adjusted R-squared: {model.rsquared_adj:.4f}")
# Get detailed summary
print(model.summary())
# Make predictions
predictions = model.predict(X)
Distributions
Work with probability distributions using a unified interface.
Normal
statflow.Normal(mu=0, sigma=1)
Normal (Gaussian) distribution.
Parameters
Returns
Normal
Examples
# Create a normal distribution
dist = sf.Normal(mu=100, sigma=15)
# Sample from the distribution
samples = dist.sample(1000)
# Compute probability density
x = np.linspace(50, 150, 100)
pdf_values = dist.pdf(x)
# Compute cumulative distribution
cdf_values = dist.cdf(x)
# Get quantiles
median = dist.quantile(0.5)
q95 = dist.quantile(0.95)
Utilities
Helper functions for data analysis and statistics.
bootstrap
statflow.bootstrap(data, statistic, n_resamples=10000, confidence=0.95, random_state=None)
Compute bootstrap confidence intervals for a statistic.
Parameters
Returns
BootstrapResult
Examples
# Estimate confidence interval for median
data = np.random.exponential(scale=2, size=100)
result = sf.bootstrap(data, np.median, random_state=42)
print(f"Median: {np.median(data):.4f}")
print(f"95% CI: {result.confidence_interval}")
print(f"Bootstrap SE: {result.standard_error:.4f}")
Hypothesis Testing
Example: A/B Testing
Compare conversion rates between two groups:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
import statflow as sf
import numpy as np
# Simulate A/B test data
np.random.seed(42)
group_a_conversions = np.random.binomial(1, 0.10, 1000) # 10% conversion
group_b_conversions = np.random.binomial(1, 0.12, 1000) # 12% conversion
# Perform proportion test
result = sf.proportions_test(
[group_a_conversions.sum(), group_b_conversions.sum()],
[len(group_a_conversions), len(group_b_conversions)]
)
print(f"Group A conversion: {group_a_conversions.mean():.2%}")
print(f"Group B conversion: {group_b_conversions.mean():.2%}")
print(f"p-value: {result.pvalue:.4f}")
if result.pvalue < 0.05:
print("Significant difference detected!")
else:
print("No significant difference.")
Example: Multiple Comparisons
When testing multiple hypotheses, control for false discovery rate:
1
2
3
4
5
6
7
8
9
10
11
# Simulate 20 hypothesis tests
p_values = [sf.ttest(
np.random.normal(0, 1, 100),
np.random.normal(0.2 if i < 5 else 0, 1, 100)
).pvalue for i in range(20)]
# Apply Benjamini-Hochberg correction
adjusted = sf.multipletests(p_values, method='fdr_bh', alpha=0.05)
print(f"Significant tests (uncorrected): {sum(p < 0.05 for p in p_values)}")
print(f"Significant tests (FDR corrected): {sum(adjusted.reject)}")
Regression Analysis
Example: Multiple Linear Regression
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
import statflow as sf
import pandas as pd
import numpy as np
# Generate synthetic data
np.random.seed(42)
n = 200
data = pd.DataFrame({
'hours_studied': np.random.uniform(0, 10, n),
'sleep_hours': np.random.uniform(4, 10, n),
'previous_score': np.random.uniform(50, 100, n)
})
# Generate exam scores with some noise
data['exam_score'] = (
3.5 * data['hours_studied'] +
2.0 * data['sleep_hours'] +
0.5 * data['previous_score'] +
np.random.normal(0, 5, n)
)
# Fit model
X = data[['hours_studied', 'sleep_hours', 'previous_score']]
y = data['exam_score']
model = sf.LinearRegression()
model.fit(X, y)
# Print summary
print(model.summary())
# Check assumptions
diagnostics = model.diagnostics()
print(f"Durbin-Watson: {diagnostics.durbin_watson:.2f}")
print(f"Condition Number: {diagnostics.condition_number:.2f}")
# Visualize residuals
model.plot_residuals()
Example: Logistic Regression
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
# Binary classification
X = np.random.randn(500, 5)
z = 1.5*X[:, 0] - 0.8*X[:, 1] + np.random.randn(500)*0.5
y = (z > 0).astype(int)
model = sf.LogisticRegression()
model.fit(X, y)
# Predictions
probabilities = model.predict_proba(X)
predictions = model.predict(X)
# Evaluate
print(f"Accuracy: {(predictions == y).mean():.2%}")
print(f"AUC-ROC: {model.roc_auc_score(X, y):.4f}")
# Confusion matrix
print(model.confusion_matrix(y, predictions))
Time Series Analysis
Example: Autocorrelation Analysis
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
import statflow as sf
import numpy as np
# Generate AR(1) process
np.random.seed(42)
n = 500
rho = 0.7
y = np.zeros(n)
y[0] = np.random.randn()
for t in range(1, n):
y[t] = rho * y[t-1] + np.random.randn()
# Compute autocorrelation
acf_values = sf.acf(y, nlags=20)
pacf_values = sf.pacf(y, nlags=20)
# Plot ACF/PACF
sf.plot_acf(y, lags=20)
sf.plot_pacf(y, lags=20)
# Test for stationarity
adf_result = sf.adfuller(y)
print(f"ADF Statistic: {adf_result.statistic:.4f}")
print(f"p-value: {adf_result.pvalue:.4f}")
Example: Moving Average Smoothing
1
2
3
4
5
6
7
8
9
# Noisy time series
t = np.linspace(0, 10, 200)
signal = np.sin(t) + 0.3 * np.random.randn(200)
# Apply moving average
smoothed = sf.moving_average(signal, window=10)
# Exponential smoothing
ema = sf.exponential_smoothing(signal, alpha=0.3)
Note
For more advanced time series modeling including ARIMA, SARIMA, and state space models, see a dedicated time series package.
Changelog
Version 1.2.3 (2025-01-15)
Added:
- New
bootstrap()function with percentile and BCa methods - Support for weighted statistics in most functions
- GPU acceleration for large-scale computations (requires
statflow[gpu])
Fixed:
- Improved numerical stability in
LinearRegressionfor ill-conditioned matrices - Fixed edge case in
mannwhitneyu()with ties - Corrected confidence interval calculation in paired t-tests
Changed:
- Updated minimum NumPy version to 1.20.0 for better type hints
- Improved performance of
anova()by 40% through vectorization
Version 1.1.0 (2024-11-20)
Added:
- Multiple comparison corrections (
multipletests()) - Proportion tests (
proportions_test()) - Effect size calculations for all major tests
Fixed:
- Memory leak in bootstrap resampling
- Incorrect degrees of freedom in ANOVA with unequal sample sizes
Version 1.0.0 (2024-09-01)
Initial release with core functionality:
- Basic statistical tests (t-test, Mann-Whitney U, ANOVA, etc.)
- Linear and logistic regression
- Common probability distributions
- Time series utilities
Contributing
We welcome contributions! Please see our Contributing Guide for details.
License
StatFlow is released under the MIT License. See LICENSE for details.
Citation
If you use StatFlow in your research, please cite:
1
2
3
4
5
6
7
@software{statflow2024,
title = {StatFlow: A Modern Statistical Analysis Toolkit},
author = {StatFlow Contributors},
year = {2024},
url = {https://github.com/example/statflow},
version = {1.2.3}
}