Enhancing Telecom Customer Churn Prediction Using Voting Ensemble Learning with SMOTE and ADASYN
Main Article Content
Abstract
Customer churn prediction is an important task in the telecommunications sector, particularly when class imbalance affects the ability of machine learning models to identify churned customers. This study evaluates the performance of several machine learning algorithms using the Telco Customer Churn dataset, which contains 7,043 customer records and 21 features. Five classification algorithms, namely Support Vector Machine (SVM), Decision Tree, Random Forest, XGBoost, and CatBoost, were evaluated. A Voting Ensemble model combining Random Forest, XGBoost, and CatBoost was also developed. To address class imbalance, SMOTE and ADASYN oversampling techniques were applied to the training data. The results showed that ensemble-based models generally outperformed individual classifiers. The Voting Ensemble model combined with SMOTE achieved the best performance, obtaining 97% accuracy, 0.95 precision, 0.93 recall, 0.95 F1-score, and 0.96 ROC-AUC. The findings indicate that integrating ensemble learning with oversampling techniques can improve churn prediction performance in imbalanced telecom datasets.
Article Details
Section

This work is licensed under a Creative Commons Attribution 4.0 International License.
How to Cite
References
[1] P. Wachwanakijkul, S. Junsiritrakhoon, N. Kantanantha, G. Narayanamurthy, and P. Jarumaneeroj, “Data-driven approaches to predicting customer churn in a non-contractual car-sharing company,” Transportation Research Interdisciplinary Perspectives, vol. 33, p. 101600, 2025, doi: 10.1016/j.trip.2025.101600.
[2] H. Shoja, E. Sabet, and S. S. Sadat, “Customer churn prediction system using machine learning: A case study ROSHAN Telecom-Afghanistan,” International Journal of Integrated Science and Technology, vol. 4, pp. 123–137, 2026, doi: 10.59890/ijist.v4i2.287.
[3] T. Zhang, S. Moro, and R. Ramos, “A data-driven approach to improve customer churn prediction based on telecom customer segmentation,” Future Internet, vol. 14, pp. 1–19, 2022, doi: 10.3390/fi14030094.
[4] L. Theodorakopoulos, A. Theodoropoulou, and C. Klavdianos, “Big data analytics and AI for consumer behavior in digital marketing: Applications, synthetic and dark data, and future directions,” Big Data and Cognitive Computing, vol. 10, p. 46, 2026, doi: 10.3390/bdcc10020046.
[5] M. Martinović, K. Dokic, and D. Pudić, “Comparative analysis of machine learning models for predicting innovation outcomes: An applied AI approach,” Applied Sciences, vol. 15, p. 3636, 2025, doi: 10.3390/app15073636.
[6] Y. Rimal, N. Sharma, A. Alsadoon, and S. K. Abbas, “A comparative analysis of ensemble AutoML machine learning prediction accuracy of STEM student grade prediction: A multi-class classification perspective,” Multimedia Tools and Applications, vol. 84, pp. 38343–38369, 2025, doi: 10.1007/s11042-024-20554-8.
[7] N. M. AbdelAziz et al., “A comprehensive evaluation of machine learning and deep learning models for churn prediction,” Information, vol. 16, no. 7, p. 537, 2025, doi: 10.3390/info16070537.
[8] C. Kaope and Y. Pristyanto, “The effect of class imbalance handling on datasets toward classification algorithm performance,” MATRIK: Jurnal Manajemen, Teknik Informatika dan Rekayasa Komputer, vol. 22, pp. 227–238, 2023, doi: 10.30812/matrik.v22i2.2515.
[9] P. Magadum and M. A. Shaikhsurab, “Enhancing customer churn prediction in telecommunications: An adaptive ensemble learning approach,” 2024, doi: 10.13140/RG.2.2.24573.78569.
[10] A. Bhatnagar and S. Srivastava, “Customer churn prediction: A machine learning approach with data balancing for telecom industry,” International Journal of Computing, pp. 9–18, 2025, doi: 10.47839/ijc.24.1.3873.
[11] T. Xu, Y. Ma, and K. Kim, “Telecom churn prediction system based on ensemble learning using feature grouping,” Applied Sciences, vol. 11, p. 4742, 2021, doi: 10.3390/app11114742.
[12] X. Chen et al., “A comprehensive analysis of churn prediction in telecommunications using machine learning,” arXiv preprint arXiv:2509.22654, 2025.
[13] A. Bhatnagar, “Customer churn prediction using machine learning approach: A comprehensive study,” Journal of Information Systems Engineering and Management, vol. 10, pp. 80–92, 2025, doi: 10.52783/jisem.v10i25s.3944.
[14] Y. Wei, “Telco customer churn prediction,” Highlights in Science, Engineering and Technology, vol. 92, pp. 218–226, 2024, doi: 10.54097/84bmrd32.
[15] D. Gupta, N. Patel, and A. Shri, “Enhancing telecom customer retention through data mining-based churn prediction,” International Journal of Scientific Research in Science and Technology, vol. 11, pp. 566–572, 2024, doi: 10.32628/IJSRST24122223.
[16] K. Peng and Y. Peng, “Research on telecom customer churn prediction based on GA-XGBoost and SHAP,” Journal of Computer and Communications, vol. 10, pp. 107–120, 2022, doi: 10.4236/jcc.2022.1011008.
[17] A. Kumar and D. Mishra, “Improving support vector machine using modified kernel function,” International Journal of Scientific Research and Modern Technology, vol. 4, pp. 1–5, 2025, doi: 10.38124/ijsrmt.v4i5.501.
[18] I. Mienye and N. Jere, “A survey of decision trees: Concepts, algorithms, and applications,” IEEE Access, 2024, doi: 10.1109/ACCESS.2024.3416838.
[19] H. Salman, A. Kalakech, and A. Steiti, “Random forest algorithm overview,” Babylonian Journal of Machine Learning, pp. 69–79, 2024, doi: 10.58496/BJML/2024/007.
[20] K. İleri, “Comparative analysis of CatBoost, LightGBM, XGBoost, RF, and DT methods optimized with PSO,” International Journal of Machine Learning and Cybernetics, vol. 16, pp. 6937–6956, 2025, doi: 10.1007/s13042-025-02654-5.
[21] S. Ganie, P. D. Pramanik, and Z. Zhao, “Ensemble learning with explainable AI for improved heart disease prediction,” Scientific Reports, vol. 15, 2025, doi: 10.1038/s41598-025-97547-6.
[22] H. Weko and H. Suparwito, “SVM and ensemble majority voting algorithm on sentiment analysis of using ChatGPT in education,” International Journal of Applied Sciences and Smart Technologies, vol. 7, pp. 451–468, 2025, doi: 10.24071/ijasst.v7i2.12680.
[23] S. Ganie, P. D. Pramanik, and Z. Zhao, “Ensemble learning with explainable AI for improved heart disease prediction,” Scientific Reports, vol. 15, 2025, doi: 10.1038/s41598-025-97547-6.
[24] A. Omari et al., “A predictive analytics approach to improve telecom’s customer retention,” Frontiers in Artificial Intelligence, vol. 8, 2025, doi: 10.3389/frai.2025.1600357.
[25] S. J. Haddadi et al., “Customer churn prediction in imbalanced datasets with resampling methods: A comparative study,” Expert Systems with Applications, vol. 246, p. 123086, 2024, doi: 10.1016/j.eswa.2023.123086.
[26] N. M. S. Neethu and V. C. S. Vinod, “A review of unlabeled and imbalanced data challenges in machine learning: Strategies and solutions,” Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, vol. 15, 2025, doi: 10.1002/widm.70043.
[27] D. Elreedy, A. Atiya, and F. Kamalov, “A theoretical distribution analysis of SMOTE for imbalanced learning,” Machine Learning, vol. 113, 2023, doi: 10.1007/s10994-022-06296-4.
[28] Y. Li et al., “An improved SMOTE algorithm for enhanced imbalanced data classification,” Scientific Reports, vol. 15, 2025, doi: 10.1038/s41598-025-09506-w.
[29] S. F. Taskiran et al., “A comprehensive evaluation of oversampling techniques for enhancing text classification performance,” Scientific Reports, vol. 15, p. 21631, 2025, doi: 10.1038/s41598-025-05791-7.
[30] A. Taskeen, S. U. R. Khan, and A. Mashkoor, “An adaptive synthetic sampling and batch generation-oriented hybrid approach,” Soft Computing, vol. 28, pp. 13595–13614, 2024, doi: 10.1007/s00500-024-10378-x.
[31] D. Salifu et al., “Data augmentation and machine learning algorithms for multi-class imbalanced morphometrics data,” Heliyon, vol. 11, p. e42214, 2025, doi: 10.1016/j.heliyon.2025.e42214.
[32] X. Gao et al., “A comprehensive survey on imbalanced data learning,” Frontiers of Computer Science, vol. 20, 2026, doi: 10.1007/s11704-025-50274-7.
[33] O. Rainio, J. Teuho, and R. Klén, “Evaluation metrics and statistical tests for machine learning,” Scientific Reports, vol. 14, 2024, doi: 10.1038/s41598-024-56706-x.
[34] K. Sujon et al., “Accuracy, precision, recall, F1-score, or MCC? Empirical evidence from advanced statistics, ML, and XAI,” Journal of Big Data, vol. 12, 2025, doi: 10.1186/s40537-025-01313-4.