← Back to home

Telco Churn Retention: Who to Target, and How to Prove It Worked

Oct 2026 · 7 min read · personal project

Python scikit-learn SciPy SHAP A/B testing uplift modelling Gradio

A churn model tells you who is likely to leave. It does not tell you who a retention offer will keep. A retention team that sends offers to the customers with the highest churn score can spend most of its budget on people who would leave anyway, or on people who would have stayed without an offer.

So the project asks three questions in order: who is likely to leave and how much revenue that puts at risk, who should actually be targeted, and how to design and read the experiment that proves a campaign worked. All data is public sample data: the IBM Telco sample (a fictional company) and the Hillstrom email test (a real randomised retail experiment).

I shipped a logistic regression rather than gradient boosting. Boosting was 0.006 ROC-AUC ahead on the hold-out set, which is less than one standard deviation across cross-validation folds, and the logistic model lets every score be explained exactly. All preprocessing sits inside scikit-learn pipelines, and every customer in the targeting grid is scored by a model that did not see them.

For the experiment I fixed the primary comparison and metric before computing any result, checked that the randomisation held, and reported Holm-adjusted p-values across the six tests. The one finding I only spotted after seeing the results is labelled as a hypothesis for the next test, not a result.

Churn drivers: Contract type, tenure and internet service dominate. Month-to-month customers churn at 43% against 3% on two-year contracts, and 53% of customers leave in their first six months.

Model and targeting: Hold-out ROC-AUC 0.84. Contacting the top 20% of customers by score reaches 49.7% of all churners, 2.5 times a random list of the same size.

Value x risk grid: Combining churn risk with what each customer pays shows where the money is. The "Save first" segment holds 18% of customers but 53% of the revenue at risk.

A/B test and uplift: On the 64,000-customer Hillstrom test, the women's email lifted the visit rate by 4.5 points (95% CI 3.9 to 5.2). Emailing the 30% with the highest predicted uplift gave 8.4 points of lift, against 3.4 for a "most likely to respond" list, which did no better than emailing everyone.

Retention test plan: Sample size and power for testing an offer on the "Save first" segment, plus a clearly labelled simulation that shows when uplift targeting beats churn-risk targeting and when it does not.

Live app: A Gradio app with a what-if risk scorer with plain-word reasons, the driver charts, a targeting grid with a "contact the top X%" slider, and the experiment results with a sample size calculator.

Customers scored

7,043

Experiment size

64,000

Uplift vs propensity

8.4 vs 3.4 pts

The experiment passed its validity checks: the arm sizes matched the design (sample ratio p = 0.90) and the largest covariate imbalance was 0.014, well under the usual 0.1 limit. The visit lift was more than five times the smallest effect the test could detect.

The most practical result came from the test plan. With 634 customers per arm, a retention test on "Save first" alone can only detect a churn drop of 7.6 points or more. Detecting a realistic 2 to 5 point effect would need between 1,457 and 8,945 customers per arm. Knowing that before running a campaign saves a team from an inconclusive test.

Reaching a churner is not the same as saving one. The propensity list looked sensible and did no better than random, because it picked people who would visit anyway. Ranking by predicted uplift is the step most churn projects skip.

A model choice should be justified against the noise. A 0.006 AUC gain inside the cross-validation spread is not a reason to give up exact explanations.

Simulations need clear labels. The simulated part shows a mechanism, not a fact, and in one scenario the simple risk list wins. Saying so openly is what makes the real-data results credible.