Resume where you left off
Online at https://diogoribeiro7.github.io/analytics-blog-jekyll/machine-learning/2024/02/15/machine-learning-walkthrough-churn/
Topics
Business framing
A subscription analytics company wants to proactively engage at-risk customers. The objective is to maximize precision while maintaining acceptable recall so that retention specialists focus on high-quality leads.
Feature engineering pipeline
1
2
3
4
5
6
7
8
9
10
11
12
13
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder, StandardScaler
categorical = ["region", "plan_type", "engagement_cluster"]
numeric = ["tenure_months", "support_tickets", "nps", "usage_minutes"]
preprocessor = ColumnTransformer(
transformers=[
("categorical", OneHotEncoder(handle_unknown="ignore"), categorical),
("numeric", StandardScaler(), numeric),
]
)
Training with cross-validation
1
2
3
4
5
6
7
8
9
10
11
12
13
14
from sklearn.ensemble import GradientBoostingClassifier
from sklearn.model_selection import StratifiedKFold, cross_validate
clf = GradientBoostingClassifier(random_state=42)
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
metrics = cross_validate(
clf,
preprocessor.fit_transform(train_X),
train_y,
cv=cv,
scoring=["precision", "recall", "roc_auc"],
return_estimator=True,
)
Track each run with MLflow to version hyperparameters and feature sets:
1
2
3
4
5
6
7
8
9
10
import mlflow
with mlflow.start_run(run_name="gb-churn-baseline"):
mlflow.log_params(clf.get_params())
mlflow.log_metrics({
"precision_mean": metrics["test_precision"].mean(),
"recall_mean": metrics["test_recall"].mean(),
"roc_auc_mean": metrics["test_roc_auc"].mean(),
})
mlflow.sklearn.log_model(clf.fit(train_X, train_y), "model")
Model evaluation
- Precision@Top500 = 0.71, aligning with business targets.
- Calibration curve indicates slight overconfidence above 0.8 probability.
- SHAP values highlight tenure and NPS as the dominant predictors.
Deployment considerations
- Export the pipeline via
mlflow.sklearn.log_modeland load it in the scoring API. - Schedule batch predictions with Prefect using data quality gates before inference.
- Monitor live metrics with Evidently AI dashboards connected to the feature store.
Download the serving blueprint or open the MLflow dashboard to review the registered model versions.
Your private highlights
Kept in this browser only and never sent anywhere. These are your notes, not comments. With text selected, Alt+Shift+H highlights it and Alt+Shift+N adds a note.
This browser is not letting the site keep data, so saving, progress and highlights are off.
Select a passage of the article to highlight it.
Your reading data
-
Launch the MLflow experiment dashboard
Inspect metrics, parameters, and artifacts captured during training.
© 2024 Diogo Ribeiro. Text and figures under CC BY 4.0.
How to cite
Use the quick export buttons to save citations for reference managers or copy the formatted text directly.
Diogo Ribeiro (2024). Machine Learning Walkthrough: Predicting Customer Churn. DataLog | Data Science & Research Theme. https://diogoribeiro7.github.io/analytics-blog-jekyll/machine-learning/2024/02/15/machine-learning-walkthrough-churn/.
BibTeX
RIS
EndNote
Open science & reproducibility badges
These badges highlight the transparency practices applied to this work. Hover or focus on each badge to learn more about the criteria.
- Open Data Dataset and code repository published with permissive license. Public repository, DOI issued, README with reproduction steps.
- Reproducible Workflow Containerized environment and automated tests provided. Continuous integration pipeline with reproducibility checks.
- Transparent Peer Review Peer review reports archived with DOI and linked to article. Open peer review statement and archived reports on Zenodo.