Statistics, Machine Learning & Applied Mathematics
Research, technical articles and open-source software on statistical modelling, machine learning, time series, data systems and applied mathematics.
I write for practitioners who need more than a generic tutorial: assumptions, diagnostics, equations, code, failure modes, and the judgment required to use methods responsibly.
About and editorial standards Explore the archive Papers and research Projects and packages
Latest articles
Berkson's Paradox: How Selecting the Cases Worth Looking At Invents Correlations
Among escalated tickets, severity and customer value are correlated at minus 0.55. Across all tickets they are independent. Nothing about...
Staggered Rollouts and Difference-in-Differences: When Two-Way Fixed Effects Get It Wrong
A feature rolls out to regions in three waves. The panel regression with region and month fixed effects says the effect is 0.6. The true ...
Ratio Metrics in A/B Tests: The Session-Level Test Lies and the Delta Method Fixes It
An A/A test on conversion per session, with two thousand users per arm and about three sessions each, comes back significant one time in ...
Information Geometry for Data Science: Curvature, Models, and Learning
Information geometry treats probability models as geometric objects, making it easier to reason about distance, curvature, uncertainty, a...
Discrete Mathematics for Data Science: States, Constraints, and Algorithms
Discrete mathematics is the part of mathematics that explains how data systems make decisions, count possibilities, represent relationshi...
Bayesian Decision Theory for Data Science: From Uncertainty to Action
Bayesian decision theory connects statistical uncertainty to action by asking not only what is likely, but what decision is best under un...
Selected work
These are better entry points than the chronological archive.
How to Detect Data Drift in Machine Learning Models
Production monitoring, statistical tests, drift signals, and the distinction between detectable distribution movement and model failure.
Probability Calibration in Machine Learning
Why probability outputs need calibration, how to score them, and where Venn-Abers style predictors fit.
Forecasting Baselines That Are Hard to Beat
A practical argument for naive and seasonal baselines before treating any forecasting metric as evidence of skill.
Conformal Prediction for Operational Risk Decisions
Prediction intervals connected to decisions, not just uncertainty decoration.
Explore by topic
Statistics & Probability
Inference, modelling, probability, diagnostics, survival analysis, robust methods and uncertainty.
Machine Learning
Model evaluation, monitoring, calibration, drift, tabular learning, data quality and MLOps.
Time Series & Forecasting
Forecasting, baselines, seasonality, anomaly detection, feature engineering and state-space models.
Mathematics
Applied mathematics, stochastic processes, optimization, graph theory and the foundations behind data science.
Data Engineering & Systems
Pipelines, monitoring, real-time processing, software practice and production data systems.
Research Methods & Causal Inference
Experimental design, causal reasoning, preregistration, measurement, ethics and research policy.
Open-source projects
The software pages are technical assets, not just documentation links. The Python portfolio is published on PyPI and spans scientific machine learning, survival simulation, heavy-tailed distributions, design of experiments, QCA, anomaly detection, imbalanced-learning diagnostics, time-series representations, behavioural sensing and WiFi activity recognition.
- PyPI package portfolio — thirteen Python packages with install commands, documentation and source links.
- genSurvPy / gen-surv — survival-data simulation and visualization for statistical research and benchmarking.
- unconfoundedr — R tools for comparing randomized and observational estimands under confounding and transportability concerns.





