A machine learning project that predicts which telecom customers are likely to cancel their subscription, enabling the business to take proactive retention action before it's too late.
Customer churn is one of the most expensive problems in the telecom industry. Acquiring a new customer costs 5–7x more than retaining an existing one. This project builds a classification model to identify at-risk customers before they leave, so the business can intervene with targeted offers or support.
- Source: IBM Telco Customer Churn Dataset
- Size: 7,043 customers × 21 features
- Target:
Churn— Yes (churned) / No (stayed) - Class Distribution: ~73.5% No Churn / ~26.5% Churn (imbalanced)
Key Features Include:
- Demographics — gender, senior citizen, partner, dependents
- Services — phone, internet, streaming, security, tech support
- Account — contract type, payment method, paperless billing
- Charges — monthly charges, total charges, tenure
| Area | Tools |
|---|---|
| Language | Python 3 |
| Data Processing | Pandas, NumPy |
| Visualization | Matplotlib, Seaborn |
| Machine Learning | Scikit-learn |
| Model Saving | Joblib |
| Web App | Streamlit |
- Converted
TotalChargesfrom string to numeric (contained hidden spaces) - Filled 11 missing values in
TotalChargeswith median - Dropped
customerID(irrelevant to prediction)
- Visualized churn distribution and confirmed class imbalance (26.5% churn)
- Found month-to-month contract customers churn ~3x more than two-year customers
- Plotted distributions of Tenure, MonthlyCharges, TotalCharges
- Correlation heatmap showed TotalCharges highly correlated with Tenure
- Applied
pd.get_dummies()for One-Hot Encoding of all categorical features - Used
StandardScalerfor numerical feature normalization - Addressed class imbalance using
class_weight='balanced'— preventing the model from simply predicting "No Churn" every time
| Model | Accuracy | Churn Recall |
|---|---|---|
| Logistic Regression | 74.7% | 82% |
| Random Forest (baseline) | ~79% | ~76% |
| Tuned Random Forest | ~80% | 85% |
Used GridSearchCV with 5-fold cross-validation, optimizing for Recall (not accuracy) because missing a churner costs more than a false alarm.
param_grid = {
'n_estimators': [100, 200, 300],
'max_depth': [5, 10, 15],
'min_samples_split': [2, 5, 10]
}- ROC-AUC Score: 0.86 — excellent class separation
- Plotted ROC Curve and Confusion Matrix
- 5-Fold Cross-Validation for reliable performance estimate
- ✅ Final model: Tuned Random Forest
- ✅ ROC-AUC: 0.86
- ✅ Churn Recall: 85% — model correctly identifies 85% of actual churners
Top 5 Features Driving Churn:
- Tenure (longer = less likely to churn)
- Total Charges
- Monthly Charges
- Contract Type — Month-to-Month customers churn most
- Internet Service — Fiber Optic users show higher churn
- Incentivize long-term contracts — offer discounts to move month-to-month users to 1 or 2-year plans
- High-value customer monitoring — proactively reach out to customers with high monthly charges
- Tenure-based loyalty rewards — implement offers at key tenure milestones where churn risk spikes
- Fiber Optic service audit — investigate whether high churn in Fiber Optic is a pricing or quality issue
- Predictive retention team — use model scores to flag high-risk customers for the customer success team
Built an interactive web app where anyone can enter customer details and get an instant churn prediction with probability score and suggested retention actions.
👉 [Click here to try the app] (https://your-link.streamlit.app)
customer-churn-prediction-ML-Python/
│
├── app.py # Streamlit web application
├── Customer_Churn_Prediction_Improved.ipynb # Full ML notebook
├── customer_churn_model.pkl # Saved trained model
├── scaler.pkl # Saved StandardScaler
├── requirements.txt # Python dependencies
└── README.md # Project documentation
GitHub: Shahnawaz-analytics
Linkedin: Shahnawaz-khan