Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Internet Service Provider — Customer Churn Prediction

ML2 project. Predicts whether an ISP customer will churn, using account, billing, and network-usage features. The entire workflow lives in churn_predict.ipynb and is built around a single leak-free scikit-learn / imbalanced-learn pipeline.

Results

All models are tuned for PR-AUC and evaluated on the same untouched held-out test set (25% of the data, stratified). Best model: Random Forest.

Model Accuracy Precision Recall F1 ROC-AUC PR-AUC
Random Forest 0.94 0.95 0.94 0.94 0.98 0.99
Decision Tree 0.94 0.95 0.93 0.94 0.97 0.98
Logistic Regression 0.88 0.87 0.91 0.89 0.93 0.94
Gaussian NB 0.74 0.69 0.96 0.80 0.91 0.92

The strongest predictors are remaining_contract, network usage (download_avg / upload_avg), and whether the customer has a contract at all (the missing-value indicator) — see "Key modeling decisions" below.

Dataset

dataset/internet_service_churn.csv — 72,274 customers, 11 columns (the compressed .7z is the tracked source of truth; the .csv is git-ignored and extracted locally).

column description
id customer id (dropped — not predictive)
is_tv_subscriber, is_movie_package_subscriber service flags (0/1)
subscription_age years subscribed
bill_avg avg bill, last 3 months
remaining_contract years left on contract (missing = no active contract)
service_failure_count support calls for service failures
download_avg, upload_avg avg GB down/up (missing = usage never recorded)
download_over_limit times over the download limit
churn target — 1 = churned, 0 = retained

Setup

python -m venv .venv && source .venv/bin/activate   # optional but recommended
pip install -r requirements.txt

# extract the dataset (needs the `7z` / p7zip CLI)
cd dataset && 7z x internet_service_churn.7z && cd ..

Run

# headless, top-to-bottom (the reproducibility check):
jupyter nbconvert --to notebook --execute --inplace churn_predict.ipynb

# or interactively:
jupyter notebook churn_predict.ipynb

Running the notebook writes the trained pipeline to random_forest_model.joblib (git-ignored). Because preprocessing is part of the saved pipeline, inference takes raw data directly — no manual cleaning required, and the artifact reloads in any process:

import joblib, pandas as pd
model = joblib.load("random_forest_model.joblib")
raw = pd.read_csv("dataset/internet_service_churn.csv").drop(columns=["id", "churn"])
preds  = model.predict(raw)          # handles missing values internally
proba  = model.predict_proba(raw)[:, 1]

Key modeling decisions

  • Missing values are signal, not noise. A missing remaining_contract means the customer has no contract — and those customers churn ~90% of the time. Rather than dropping them (which discarded ~30% of the data and the easiest-to-spot churners), they are kept and the missingness is encoded as a feature. This also keeps the classes naturally balanced (~55/45).
  • One leak-free pipeline, one code path. A single shared ColumnTransformer (preprocess) is composed with an optional scaler / SMOTE and the classifier through one make_pipeline() factory. Imputation, scaling, and any resampling are fit on training folds only (inside Pipeline + StratifiedKFold), never on the test set. Every model is tuned the same way (GridSearchCV on PR-AUC) with no special-case branches.
  • No SMOTE. On the (balanced) full data an explicit cross-validated test showed resampling gives no lift, so it was removed for simplicity and better calibration.
  • Tuned for PR-AUC, the appropriate metric for an imbalanced, recall-sensitive churn problem — not raw accuracy.
  • Portable artifact. Preprocessing is built from standard scikit-learn transformers, so the saved pipeline loads and predicts in any environment without this notebook's code.

Repository layout

├── churn_predict.ipynb     # the full analysis + modeling pipeline
├── dataset/
│   └── internet_service_churn.7z   # compressed dataset (extract to .csv locally)
├── requirements.txt
├── .gitignore
└── README.md

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages