Data Science

This hub is for the work that sits between a question and a model: exploring a dataset, choosing what to measure, deciding whether an intervention worked, and keeping an analysis honest once it runs in production. It is the second-largest subject on the site, and many of its articles are applied case studies.

Start here

Analysis in Python

Exploratory analysis with pandas, density estimation, entropy, outlier and anomaly detection, and worked examples in the scientific Python stack.

Statistical modelling in practice

Generalised linear models, latent class analysis, count models, missing data in clinical research, and choosing between tests.

Decisions and interventions

Did the change work, and for whom? Synthetic control, uplift modelling and counterfactual evaluation of decision policies.

Models and data in production

Silent data-quality failures, drift detection, validating anomaly detectors, monitoring with wearables and IoT sensors, and what to check before deployment.

Healthcare and ageing

Readmission risk, fall prediction, remote monitoring, chronic disease, clinical text and the ethics of AI in elderly care.

Predictive maintenance

Measuring whether a maintenance programme pays for itself, dashboards, cloud and edge analytics, maintenance text, and the effect on operations.

Latest in Data Science

Every article in this subject is listed under Data Science in the category index. For neighbouring subjects see Statistics, Machine Learning and Research Methods.

Loading mathematical content