Data Science Portfolio Project

Telecom Churn Intelligence
Revenue-at-Risk Modeling & Retention Effectiveness

Predicted customer churn, quantified revenue at risk using CLV modeling, and statistically validated a retention campaign via A/B testing — delivering a complete business intelligence system.

Dataset Size
7,032
customers · 21 features
Churn Rate
26.58%
1,869 customers at risk
Best Model ROC-AUC
0.83
Random Forest
Revenue at Risk
₹3.56L
identified portfolio risk
Campaign ROI
1123%
p < 0.0001 · statistically proven

Phase 1
Exploratory Data Analysis
Key patterns identified across contract type, payment method, tenure, and monthly charges.
42.7%
Month-to-month contract customers churn — 15x higher than two-year contracts
45.3%
Electronic check users churn — highest among all payment methods
41.9%
Fiber optic internet users churn — despite paying the highest charges
Churn rate by contract type
Month-to-month: 42.7%, One year: 11.3%, Two year: 2.8%
Churn rate by payment method
Electronic check highest at 45.3%
Tenure distribution by churn status
Churners have much lower tenure
Monthly charges by churn status
Churners avg ₹74.44 vs non-churners ₹61.31

Phase 3
Statistical Analysis
Hypothesis testing confirms all key drivers are statistically significant (p < 0.0001).

Chi-Square Tests — Categorical vs Churn

Contract type vs Churn ✓ p < 0.0001
Electronic check vs Churn ✓ p < 0.0001
Fiber optic vs Churn ✓ p < 0.0001
Fiber optic Chi² statistic 663.36 (strongest)

T-Test & Mann-Whitney U — Numeric vs Churn

MonthlyCharges T-Test ✓ T = 16.48
Churners avg MonthlyCharges ₹74.44
Non-Churners avg MonthlyCharges ₹61.31
Tenure Mann-Whitney U ✓ p < 0.0001

ANOVA — MonthlyCharges across Tenure Groups

New customers (≤12 mo) ₹56.17 avg
Mid customers (13–36 mo) ₹63.25 avg
Loyal customers (36+ mo) ₹72.01 avg
F-Statistic ✓ F = 187.50

Cramér's V — Association with Churn

TenureGroup V = 0.34 ▲ strongest
Fiber optic service V = 0.31
Two-year contract V = 0.30
Electronic check V = 0.30

Phase 4
ML Modeling
SMOTE applied on training set only (4,130 vs 4,130). Two models trained and compared.
Logistic Regression
VIF-cleaned dataset · 10 features · Odds Ratios interpreted
ROC-AUC
0.79
Churn Recall
81%
Churn F1
0.57
Accuracy
67%
Key Odds Ratios
Electronic check: 2.27x more likely to churn
Two-year contract: 0.03x (97% less likely)
Random Forest Best Model
Full 27 features · max_depth=10 · 100 estimators
ROC-AUC
0.83
Churn Recall
74%
Churn F1
0.62
Accuracy
76%
Top features by importance
Tenure 16.9% TotalCharges 9.8% MonthlyCharges 9.3% Electronic check 9.2% Two-year contract 8.9%
ROC-AUC comparison — Logistic Regression vs Random Forest
Random Forest outperforms Logistic Regression

Phase 6
A/B Testing — Retention Campaign
297 high-risk customers (churn prob > 0.7) split into treatment and control groups. Treatment offered a 20% discount.
Campaign effectiveness — statistical validation
27.15%
Churn Reduction
<0.0001
P-Value (Z-Test)
-5.03
Z-Statistic
24.06
Chi² Statistic

H₀: Retention campaign has no effect on churn rate

H₁: Campaign significantly reduces churn rate

✓ Reject H₀ — Campaign effectively reduced churn (p < 0.0001)

95% Confidence Interval: churn reduction between 17.02% and 37.28%

Churn rate — treatment vs control
27% reduction in churn rate
Revenue impact breakdown
1123% ROI on campaign investment

Campaign Return on Investment

Revenue Saved: ₹23,863  |  Campaign Cost: ₹1,951  |  Net Profit: ₹21,912

41 customers retained × avg CLV ₹582

1123%
ROI

Stack
Tech Stack
Tools and libraries used across all phases.
Python · Pandas, NumPy
Statistics · SciPy, Statsmodels, Pingouin
ML · Scikit-learn, Imbalanced-learn
Visualization · Matplotlib, Seaborn
Dashboard · Streamlit
Models · Logistic Regression, Random Forest
Sampling · SMOTE
Testing · Chi-Square, T-Test, ANOVA, Z-Test