Diogo Ribeiro (@DiogoRibeiro7)
Lead Data Scientist · Machine Learning Engineer · Statistical ML, Time Series, Causal Inference & Applied AI · Professor
Statistical ML, Time Series, Causal Inference & Applied AI
Research, technical articles and open-source software on statistical modelling, machine learning, time series, causal inference, applied AI, data systems and applied mathematics.
I write for practitioners who need more than a generic tutorial: assumptions, diagnostics, equations, code, failure modes, and the judgment required to use methods responsibly.
About and editorial standards Explore the archive Papers and research Projects and packages
Latest articles
What a Before-and-After Testimonial Can Establish
A genuine improvement does not identify its cause. A worked probability model explains how selected starting measurements, natural variat...
A Database for Analysis: Rows, Columns, Indexes and the Planner
The same five million orders answer an analytical query in 2 milliseconds or in 11 seconds, and a key lookup in 17 microseconds or 570, d...
A Data Lake Is a Directory With Rules
A data lake is files in folders plus the conventions that make them usable. The same six million rows answer a question in 10 millisecond...
Why Exact Post-Selection Confidence Intervals Can Be Enormous
An exact 95% interval of [-71.7, 2.5] for a unit-variance Gaussian mean is not a bug. It is what conditioning on a selection event costs ...
Writing Statistical Software as Executable Mathematics
In statistical software, many of the strongest tests are not input-output examples. They are equations: test inversion must agree with po...
The Trouble With Smooth Curves in Small Simulation Studies
Simulation studies often evaluate a method on a coarse parameter grid and then draw a smooth curve through the results. The plot looks pe...
Selected work
These are better entry points than the chronological archive.
How to Detect Data Drift in Machine Learning Models
Production monitoring, statistical tests, drift signals, and the distinction between detectable distribution movement and model failure.
Probability Calibration in Machine Learning
Why probability outputs need calibration, how to score them, and where Venn-Abers style predictors fit.
Forecasting Baselines That Are Hard to Beat
A practical argument for naive and seasonal baselines before treating any forecasting metric as evidence of skill.
Conformal Prediction for Operational Risk Decisions
Prediction intervals connected to decisions, not just uncertainty decoration.
Explore by topic
Science Communication
Clear explanations of scientific claims, everyday misconceptions, evidence and uncertainty, with worked examples.
Statistics & Probability
Inference, modelling, probability, diagnostics, survival analysis, robust methods and uncertainty.
Machine Learning
Model evaluation, monitoring, calibration, drift, tabular learning, data quality and MLOps.
Time Series & Forecasting
Forecasting, baselines, seasonality, anomaly detection, feature engineering and state-space models.
Mathematics
Applied mathematics, stochastic processes, optimization, graph theory and the foundations behind data science.
Data Engineering & Systems
Pipelines, monitoring, real-time processing, software practice and production data systems.
Research Methods & Causal Inference
Experimental design, causal reasoning, preregistration, measurement, ethics and research policy.
Open-source projects
The software pages are technical assets, not just documentation links. The Python portfolio is published on PyPI and spans scientific machine learning, survival simulation, heavy-tailed distributions, design of experiments, geostatistics, QCA, anomaly detection, imbalanced-learning diagnostics, missing-data imputation, interpretable classification, time-series representations, behavioural sensing and WiFi activity recognition. Two Rust crates on crates.io cover copulas and probabilistic numerics, and the Jekyll theme this site runs on is published on RubyGems.
- PyPI package portfolio — sixteen Python packages with install commands, documentation and source links.
- Rust crates — copula-core for copula modelling and dependence analysis, and uncertain-numerics for Bayesian quadrature and probabilistic linear solvers.
- datalog-theme — the DataLog Jekyll theme for data science and research writing, as a Ruby gem.
- genSurvPy / gen-surv — survival-data simulation and visualization for statistical research and benchmarking.
- unconfoundedr — R tools for comparing randomized and observational estimands under confounding and transportability concerns.



