Statistics, Machine Learning & Applied Mathematics
Research, technical articles and open-source software on statistical modelling, machine learning, time series, data systems and applied mathematics.
I write for practitioners who need more than a generic tutorial: assumptions, diagnostics, equations, code, failure modes, and the judgment required to use methods responsibly.
About and editorial standards Explore the archive Papers and research Projects and packages
Latest articles
When “Science-Based” Becomes a Brand: Ozzi Silva, Fitness Influencing and Scientific Authority
You do not need to be a scientist to communicate science. But once “science-based” becomes part of a public brand, the standard should be...
“Seed Oils Are Toxic”: How Internet Nutrition Turns Real Chemistry into a Health Myth
The seed-oil panic starts from real chemistry: polyunsaturated fats can oxidise, linoleic acid participates in omega-6 metabolism, and re...
Manuel Pinto Coelho, Health Myths, and What the Evidence Actually Says
Some of Manuel Pinto Coelho's public health claims contain a kernel of truth. The problem is what happens next: nuance disappears, uncert...
Inflammation Is Not a Diagnosis: What Social Media Gets Wrong About Inflammatory Processes
“Inflammation” has become a universal explanation on social media for fatigue, bloating, obesity, ageing, anxiety, skin problems and chro...
Why Exact Post-Selection Confidence Intervals Can Be Enormous
A very wide interval is not automatically evidence that an inferential procedure is broken. After selection, weak evidence can force an e...
Writing Statistical Software as Executable Mathematics
In statistical software, many of the strongest tests are not input-output examples. They are equations: test inversion must agree with po...
Selected work
These are better entry points than the chronological archive.
How to Detect Data Drift in Machine Learning Models
Production monitoring, statistical tests, drift signals, and the distinction between detectable distribution movement and model failure.
Probability Calibration in Machine Learning
Why probability outputs need calibration, how to score them, and where Venn-Abers style predictors fit.
Forecasting Baselines That Are Hard to Beat
A practical argument for naive and seasonal baselines before treating any forecasting metric as evidence of skill.
Conformal Prediction for Operational Risk Decisions
Prediction intervals connected to decisions, not just uncertainty decoration.
Explore by topic
Science Communication
Clear explanations of scientific claims, everyday misconceptions, evidence and uncertainty, with worked examples.
Statistics & Probability
Inference, modelling, probability, diagnostics, survival analysis, robust methods and uncertainty.
Machine Learning
Model evaluation, monitoring, calibration, drift, tabular learning, data quality and MLOps.
Time Series & Forecasting
Forecasting, baselines, seasonality, anomaly detection, feature engineering and state-space models.
Mathematics
Applied mathematics, stochastic processes, optimization, graph theory and the foundations behind data science.
Data Engineering & Systems
Pipelines, monitoring, real-time processing, software practice and production data systems.
Research Methods & Causal Inference
Experimental design, causal reasoning, preregistration, measurement, ethics and research policy.
Open-source projects
The software pages are technical assets, not just documentation links. The Python portfolio is published on PyPI and spans scientific machine learning, survival simulation, heavy-tailed distributions, design of experiments, geostatistics, QCA, anomaly detection, imbalanced-learning diagnostics, missing-data imputation, interpretable classification, time-series representations, behavioural sensing and WiFi activity recognition. Two Rust crates on crates.io cover copulas and probabilistic numerics, and the Jekyll theme this site runs on is published on RubyGems.
- PyPI package portfolio — sixteen Python packages with install commands, documentation and source links.
- Rust crates — copula-core for copula modelling and dependence analysis, and uncertain-numerics for Bayesian quadrature and probabilistic linear solvers.
- datalog-theme — the DataLog Jekyll theme for data science and research writing, as a Ruby gem.
- genSurvPy / gen-surv — survival-data simulation and visualization for statistical research and benchmarking.
- unconfoundedr — R tools for comparing randomized and observational estimands under confounding and transportability concerns.

