ML2 project. Predicts whether an ISP customer will churn, using account, billing,
and network-usage features. The entire workflow lives in
churn_predict.ipynb and is built around a single leak-free
scikit-learn / imbalanced-learn pipeline.
All models are tuned for PR-AUC and evaluated on the same untouched held-out test set (25% of the data, stratified). Best model: Random Forest.
| Model | Accuracy | Precision | Recall | F1 | ROC-AUC | PR-AUC |
|---|---|---|---|---|---|---|
| Random Forest | 0.94 | 0.95 | 0.94 | 0.94 | 0.98 | 0.99 |
| Decision Tree | 0.94 | 0.95 | 0.93 | 0.94 | 0.97 | 0.98 |
| Logistic Regression | 0.88 | 0.87 | 0.91 | 0.89 | 0.93 | 0.94 |
| Gaussian NB | 0.74 | 0.69 | 0.96 | 0.80 | 0.91 | 0.92 |
The strongest predictors are remaining_contract, network usage (download_avg /
upload_avg), and whether the customer has a contract at all (the missing-value
indicator) — see "Key modeling decisions" below.
dataset/internet_service_churn.csv — 72,274 customers, 11 columns (the compressed
.7z is the tracked source of truth; the .csv is git-ignored and extracted locally).
| column | description |
|---|---|
id |
customer id (dropped — not predictive) |
is_tv_subscriber, is_movie_package_subscriber |
service flags (0/1) |
subscription_age |
years subscribed |
bill_avg |
avg bill, last 3 months |
remaining_contract |
years left on contract (missing = no active contract) |
service_failure_count |
support calls for service failures |
download_avg, upload_avg |
avg GB down/up (missing = usage never recorded) |
download_over_limit |
times over the download limit |
churn |
target — 1 = churned, 0 = retained |
python -m venv .venv && source .venv/bin/activate # optional but recommended
pip install -r requirements.txt
# extract the dataset (needs the `7z` / p7zip CLI)
cd dataset && 7z x internet_service_churn.7z && cd ..# headless, top-to-bottom (the reproducibility check):
jupyter nbconvert --to notebook --execute --inplace churn_predict.ipynb
# or interactively:
jupyter notebook churn_predict.ipynbRunning the notebook writes the trained pipeline to random_forest_model.joblib
(git-ignored). Because preprocessing is part of the saved pipeline, inference takes
raw data directly — no manual cleaning required, and the artifact reloads in any process:
import joblib, pandas as pd
model = joblib.load("random_forest_model.joblib")
raw = pd.read_csv("dataset/internet_service_churn.csv").drop(columns=["id", "churn"])
preds = model.predict(raw) # handles missing values internally
proba = model.predict_proba(raw)[:, 1]- Missing values are signal, not noise. A missing
remaining_contractmeans the customer has no contract — and those customers churn ~90% of the time. Rather than dropping them (which discarded ~30% of the data and the easiest-to-spot churners), they are kept and the missingness is encoded as a feature. This also keeps the classes naturally balanced (~55/45). - One leak-free pipeline, one code path. A single shared
ColumnTransformer(preprocess) is composed with an optional scaler / SMOTE and the classifier through onemake_pipeline()factory. Imputation, scaling, and any resampling are fit on training folds only (insidePipeline+StratifiedKFold), never on the test set. Every model is tuned the same way (GridSearchCVon PR-AUC) with no special-case branches. - No SMOTE. On the (balanced) full data an explicit cross-validated test showed resampling gives no lift, so it was removed for simplicity and better calibration.
- Tuned for PR-AUC, the appropriate metric for an imbalanced, recall-sensitive churn problem — not raw accuracy.
- Portable artifact. Preprocessing is built from standard scikit-learn transformers, so the saved pipeline loads and predicts in any environment without this notebook's code.
├── churn_predict.ipynb # the full analysis + modeling pipeline
├── dataset/
│ └── internet_service_churn.7z # compressed dataset (extract to .csv locally)
├── requirements.txt
├── .gitignore
└── README.md