Statistics, Machine Learning & Applied Mathematics
Research, technical articles and open-source software on statistical modelling, machine learning, time series, data systems and applied mathematics.
I write for practitioners who need more than a generic tutorial: assumptions, diagnostics, equations, code, failure modes, and the judgment required to use methods responsibly.
About and editorial standards Explore the archive Papers and research Projects and packages
Latest articles
The Trouble With Smooth Curves in Small Simulation Studies
Simulation studies often evaluate a method on a coarse parameter grid and then draw a smooth curve through the results. The plot looks pe...
Why I Don't Automatically Reach for Machine Learning
Machine learning is often treated as the default destination of a data project. I prefer the opposite order: understand the problem, writ...
How Often to Retrain: A Square-Root Rule and Its Limits
Retraining every week spends 15,000 a week to prevent decay worth nothing yet. Retraining twice a year spends nearly nine times as much i...
Proxy Metrics Under Optimisation: Why a Correlation of 0.6 Is Not a Substitute for the Goal
The proxy correlates 0.63 with the metric that matters, so the team optimises it and reports a gain of 3.8. The goal metric moved by 1.9....
Confidence Sets Are Not Just Intervals
We often report uncertainty as a lower and upper bound. That is convenient, but it quietly assumes the accepted parameter values form one...
What Makes Statistical Software Trustworthy?
Statistical software can pass ordinary unit tests and still fail scientifically. The most dangerous bugs often return plausible numbers. ...
Selected work
These are better entry points than the chronological archive.
How to Detect Data Drift in Machine Learning Models
Production monitoring, statistical tests, drift signals, and the distinction between detectable distribution movement and model failure.
Probability Calibration in Machine Learning
Why probability outputs need calibration, how to score them, and where Venn-Abers style predictors fit.
Forecasting Baselines That Are Hard to Beat
A practical argument for naive and seasonal baselines before treating any forecasting metric as evidence of skill.
Conformal Prediction for Operational Risk Decisions
Prediction intervals connected to decisions, not just uncertainty decoration.
Explore by topic
Statistics & Probability
Inference, modelling, probability, diagnostics, survival analysis, robust methods and uncertainty.
Machine Learning
Model evaluation, monitoring, calibration, drift, tabular learning, data quality and MLOps.
Time Series & Forecasting
Forecasting, baselines, seasonality, anomaly detection, feature engineering and state-space models.
Mathematics
Applied mathematics, stochastic processes, optimization, graph theory and the foundations behind data science.
Data Engineering & Systems
Pipelines, monitoring, real-time processing, software practice and production data systems.
Research Methods & Causal Inference
Experimental design, causal reasoning, preregistration, measurement, ethics and research policy.
Open-source projects
The software pages are technical assets, not just documentation links. The Python portfolio is published on PyPI and spans scientific machine learning, survival simulation, heavy-tailed distributions, design of experiments, QCA, anomaly detection, imbalanced-learning diagnostics, time-series representations, behavioural sensing and WiFi activity recognition.
- PyPI package portfolio — thirteen Python packages with install commands, documentation and source links.
- genSurvPy / gen-surv — survival-data simulation and visualization for statistical research and benchmarking.
- unconfoundedr — R tools for comparing randomized and observational estimands under confounding and transportability concerns.





