Learning Under Extreme Class Imbalance: A Comparative Study of Algorithmic and Data-Level Solutions
DOI:
https://doi.org/10.61453/jods.v20260210Keywords:
Class Imbalance, Imbalanced Learning, Cost-Sensitive Learning, Data Preprocessing, Model EvaluationAbstract
Extreme class imbalance remains a persistent challenge in machine learning, particularly in high-impact domains such as fraud detection, medical diagnosis, and risk analysis, where minority classes represent critical outcomes. Conventional models often fail in such settings due to their bias toward majority classes, resulting in poor minority detection despite high overall accuracy. Although various data-level and algorithm-level techniques have been proposed, existing studies typically evaluate them in isolation and lack a comprehensive understanding of their effectiveness across different imbalance conditions. To address this gap, this study proposes a systematic comparative framework that integrates data-level resampling, cost-sensitive learning, and hybrid approaches to evaluate their performance under varying imbalance ratios and noise levels. Multiple benchmark datasets are utilized, and experiments are conducted using standardized preprocessing, controlled imbalance simulation, and repeated trials to ensure robustness. Performance is assessed using imbalance-aware metrics, including precision, recall, F1-score, and ROC-AUC. The results indicate that hybrid approaches consistently outperform standalone methods, achieving the most stable and balanced performance across all scenarios. In particular, hybrid models demonstrate superior minority class recall and F1-score while maintaining competitive precision, especially under extreme imbalance conditions. The primary goal of this research is to provide a comprehensive evaluation of imbalance-handling strategies and offer practical guidance for selecting appropriate techniques based on dataset characteristics. The findings highlight the importance of combining data-centric and model-centric approaches to enhance robustness and reliability in imbalanced learning environments. The results demonstrate up to a 32% improvement in recall compared to baseline models.
References
Adnan, M., Alarood, A. A. S., Uddin, M. I., & Rehman, I. ur. (2022). Utilizing grid search cross-validation with adaptive boosting for augmenting performance of machine learning models. PeerJ Computer Science, 8, e803. https://doi.org/10.7717/PEERJ-CS.803/SUPP-7
Aguiar, G., Krawczyk, B., & Cano, A. (2024). A survey on learning from imbalanced data streams: taxonomy, challenges, empirical study, and reproducible experimental framework. Machine Learning, 113(7), 4165–4243. https://doi.org/10.1007/s10994-023-06353-6
Altalhan, M., Algarni, A., & Turki-Hadj Alouane, M. (2025). Imbalanced Data Problem in Machine Learning: A Review. IEEE Access, 13, 13686–13699. https://doi.org/10.1109/ACCESS.2025.3531662
Araf, I., Idri, · Ali, Chairi, I., & Idri, A. (2024). Cost-sensitive learning for imbalanced medical data: a review. Artificial Intelligence Review, 57, 80. https://doi.org/10.1007/s10462-023-10652-8
Carvalho, M., Pinho, A. J., & Brás, S. (2025). Resampling approaches to handle class imbalance: a review from a data perspective. Journal of Big Data 2025 12:1, 12(1), 71-. https://doi.org/10.1186/S40537-025-01119-4
Chen, W., Yang, K., Yu, Z., Shi, Y., & Chen, C. L. P. (2024). A survey on imbalanced learning: latest research, applications and future directions. Artificial Intelligence Review, 57(6), 137-. https://doi.org/10.1007/S10462-024-10759-6
Das, S., Mullick, S. S., & Zelinka, I. (2022). On Supervised Class-Imbalanced Learning: An Updated Perspective and Some Key Challenges. IEEE Transactions on Artificial Intelligence, 3(6), 973–993. https://doi.org/10.1109/TAI.2022.3160658
Farhadpour, S., Warner, T. A., & Maxwell, A. E. (2024). Selecting and Interpreting Multiclass Loss and Accuracy Assessment Metrics for Classifications with Class Imbalance: Guidance and Best Practices. Remote Sensing 2024, Vol. 16, 16(3). https://doi.org/10.3390/RS16030533
Ghosh, K., Bellinger, C., Corizzo, R., Branco, P., Krawczyk, B., & Japkowicz, N. (2022). The class imbalance problem in deep learning. Machine Learning 2022 113:7, 113(7), 4845–4901. https://doi.org/10.1007/S10994-022-06268-8
Hancock, J. T., Khoshgoftaar, T. M., & Johnson, J. M. (2023). Evaluating classifier performance with highly imbalanced Big Data. Journal of Big Data 2023 10:1, 10(1), 42-. https://doi.org/10.1186/S40537-023-00724-5
Hemmatian, J., Hajizadeh, R., & Nazari, F. (2025). Addressing imbalanced data classification with Cluster-Based Reduced Noise SMOTE. PLOS ONE, 20(2), e0317396. https://doi.org/10.1371/JOURNAL.PONE.0317396
Imani, M., Beikmohammadi, A., & Arabnia, H. R. (2025). Comprehensive Analysis of Random Forest and XGBoost Performance with SMOTE, ADASYN, and GNUS Under Varying Imbalance Levels. Technologies 2025, Vol. 13, 13(3). https://doi.org/10.3390/TECHNOLOGIES13030088
Joshi, I., Grimmer, M., Rathgeb, C., Busch, C., Bremond, F., & Dantcheva, A. (2024). Synthetic Data in Human Analysis: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(7), 4957–4976. https://doi.org/10.1109/TPAMI.2024.3362821
Li, J., Sun, H., Li, J., Han, B., Liu, T., Yao, Q., Gong, M., Niu, G., Tsang, I. W., Sugiyama, M., & Li jyli, J. (2022). Beyond confusion matrix: learning from multiple annotators with awareness of instance features. Machine Learning 2022 112:3, 112(3), 1053–1075. https://doi.org/10.1007/S10994-022-06211-X
Li, P., Wu, Z., Wu, Y., Liang, H., Guo, X., & Li, Y. (2026). A hybrid higher-order graph convolutional network for fault diagnosis based on small sample label propagation. Measurement, 268, 120680. https://doi.org/10.1016/J.MEASUREMENT.2026.120680
Lima, F. T., & Souza, V. M. A. (2023). A Large Comparison of Normalization Methods on Time Series. Big Data Research, 34, 100407. https://doi.org/10.1016/J.BDR.2023.100407
Lin, C., Tsai, C. F., & Lin, W. C. (2022). Towards hybrid over- and under-sampling combination methods for class imbalanced datasets: an experimental study. Artificial Intelligence Review 2022 56:2, 56(2), 845–863. https://doi.org/10.1007/S10462-022-10186-5
Mastour, H., Dehghani, T., Moradi, E., & Eslami, S. (2023). Early prediction of medical students’ performance in high-stakes examinations using machine learning approaches. Heliyon, 9(7), e18248. https://doi.org/10.1016/j.heliyon.2023.e18248
Pan, I., Mason, L. R., & Matar, O. K. (2022). Data-centric Engineering: integrating simulation, machine learning and statistics. Challenges and opportunities. Chemical Engineering Science, 249, 117271. https://doi.org/10.1016/J.CES.2021.117271
Salmi, M., Atif, D., Oliva, D., Abraham, A., & Ventura, S. (2024). Handling imbalanced medical datasets: review of a decade of research. Artificial Intelligence Review 2024 57:10, 57(10), 273-. https://doi.org/10.1007/S10462-024-10884-2
Sraitih, M., Jabrane, Y., & Hajjam El Hassani, A. (2022). A Robustness Evaluation of Machine Learning Algorithms for ECG Myocardial Infarction Detection. Journal of Clinical Medicine 2022, Vol. 11, 11(17). https://doi.org/10.3390/JCM11174935
Sujon, K. M., Hassan, R., Choi, K., & Samad, M. A. (2025). Accuracy, precision, recall, f1-score, or MCC? empirical evidence from advanced statistics, ML, and XAI for evaluating business predictive models. Journal of Big Data 2025 12:1, 12(1), 268-. https://doi.org/10.1186/S40537-025-01313-4
Thiyagalingam, J., Shankar, M., Fox, G., & Hey, T. (2022). Scientific machine learning benchmarks. Nature Reviews Physics 2022 4:6, 4(6), 413–420. https://doi.org/10.1038/s42254-022-00441-7
Zhao, L., Han, F., Ling, Q., Han, H., Yao, Z., Liu, W., & Zhou, Z. (2025). A Survey on Class Imbalance Learning Algorithms in Complex Scenarios. IEEE Access, 13, 180799–180833. https://doi.org/10.1109/ACCESS.2025.3618909
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Journal of Data Science

This work is licensed under a Creative Commons Attribution 4.0 International License.