Optimization of Support Vector Machine for Imbalanced Credit Risk Classification

Authors

  • Andre Pratama Adiwijaya Gunadarma University, Indonesia

DOI:

https://doi.org/10.56127/ijst.v5i2.3038

Keywords:

credit risk, credit default, machine learning, Support Vector Machine, classification

Abstract

Credit risk classification is an important component of financial decision-making because inaccurate classification may increase payment-default exposure and reduce lending quality. Credit datasets commonly contain numerical variables with different measurement scales and imbalanced distributions between default and non-default clients, which can affect the reliability of machine-learning models. Objective: This study aims to develop and evaluate a Support Vector Machine model for classifying credit card clients into default and non-default categories. The study also examines the influence of numerical standardization, kernel selection, hyperparameter optimization, and balanced class weighting on classification performance. Methodology: A quantitative experimental approach was applied using the Default of Credit Card Clients dataset from the UCI Machine Learning Repository. The dataset consisted of 30,000 observations and 23 predictor variables. Data were divided into training and testing subsets using a stratified 80:20 ratio. Categorical variables were encoded, numerical variables were standardized, and several SVM kernels were evaluated. Hyperparameter selection was conducted using five-fold cross-validation. Model performance was assessed using accuracy, precision, recall, F1-score, specificity, balanced accuracy, and ROC–AUC. Findings: The SVM model trained without standardization failed to identify default clients effectively. Numerical standardization substantially improved classification performance, while the radial basis function kernel produced the strongest validation results. The selected balanced RBF-SVM achieved 77.13% accuracy, 48.54% precision, 56.22% recall, 52.09% F1-score, 83.07% specificity, 69.65% balanced accuracy, and 75.10% ROC–AUC. Balanced class weighting improved default detection but increased false-positive predictions. Implications: The model can support financial institutions as an initial credit-risk screening tool. Its predictions should be combined with document verification, repayment-capacity analysis, and manual assessment rather than being used as the sole basis for credit approval. Originality: This study provides a controlled evaluation of SVM performance by integrating feature standardization, kernel selection, hyperparameter optimization, class-imbalance treatment, and class-sensitive performance metrics. The study demonstrates that credit-risk models should be selected based on balanced default detection rather than overall accuracy alone.

References

Bao, W., Ning, L., & Yue, K. (2019). Integration of unsupervised and supervised machine learning algorithms for credit risk assessment. Expert Systems with Applications, 128, 301-315. https://doi.org/10.1016/j.eswa.2019.02.033

Bhatore, S., Mohan, L., & Reddy, Y. R. (2020). Machine learning techniques for credit risk evaluation: A systematic literature review. Journal of Banking and Financial Technology, 4(1), 111-138. https://doi.org/10.1007/s42786-020-00020-3

Bücker, M., Szepannek, G., Gosiewska, A., & Biecek, P. (2022). Transparency, auditability, and explainability of machine learning models in credit scoring. Journal of the Operational Research Society, 73(1), 70-90. https://doi.org/10.1080/01605682.2021.1922098

Bussmann, N., Giudici, P., Marinelli, D., & Papenbrock, J. (2021). Explainable machine learning in credit risk management. Computational Economics, 57(1), 203-216. https://doi.org/10.1007/s10614-020-10042-0

Cervantes, J., Garcia-Lamont, F., Rodríguez-Mazahua, L., & Lopez, A. (2020). A comprehensive survey on support vector machine classification: Applications, challenges and trends. Neurocomputing, 408, 189-215. https://doi.org/10.1016/j.neucom.2019.10.118

Chang, V., Sivakulasingam, S., Wang, H., Wong, S. T., Ganatra, M. A., & Luo, J. (2024). Credit risk prediction using machine learning and deep learning: A study on credit card customers. Risks, 12(11), 174. https://doi.org/10.3390/risks12110174

Dastile, X., Celik, T., & Potsane, M. (2020). Statistical and machine learning models in credit scoring: A systematic literature survey. Applied Soft Computing, 91, 106263. https://doi.org/10.1016/j.asoc.2020.106263

Dumitrescu, E. I., Hué, S., Hurlin, C., & Tokpavi, S. (2022). Machine learning for credit scoring: Improving logistic regression with non-linear decision-tree effects. European Journal of Operational Research, 297(3), 1178-1192. https://doi.org/10.1016/j.ejor.2021.06.053

Emmanuel, I., Sun, Y., & Wang, Z. (2024). A machine learning-based credit risk prediction engine system using a stacked classifier and a filter-based feature selection method. Journal of Big Data, 11, 23. https://doi.org/10.1186/s40537-024-00882-0

Laborda, J., & Ryoo, S. (2021). Feature selection in a credit scoring model. Mathematics, 9(7), 746. https://doi.org/10.3390/math9070746

Markov, A., Seleznyova, Z., & Lapshin, V. (2022). Credit scoring methods: Latest trends and points to consider. The Journal of Finance and Data Science, 8, 180-201. https://doi.org/10.1016/j.jfds.2022.07.002

Moscato, V., Picariello, A., & Sperlì, G. (2021). A benchmark of machine learning approaches for credit score prediction. Expert Systems with Applications, 165, 113986. https://doi.org/10.1016/j.eswa.2020.113986

Noriega, J. P., Rivera, L. A., & Herrera, J. A. (2023). Machine learning for credit risk prediction: A systematic literature review. Data, 8(11), 169. https://doi.org/10.3390/data8110169

Suhadolnik, N., Ueyama, J., & Da Silva, S. (2023). Machine learning for enhanced credit risk assessment: An empirical approach. Journal of Risk and Financial Management, 16(12), 496. https://doi.org/10.3390/jrfm16120496

Trivedi, S. K. (2020). A study on credit scoring modeling with different feature selection and machine learning approaches. Technology in Society, 63, 101413. https://doi.org/10.1016/j.techsoc.2020.101413

Wang, T., & Li, J. (2019). An improved support vector machine and its application in P2P lending personal credit scoring. IOP Conference Series: Materials Science and Engineering, 490(6), 062041. https://doi.org/10.1088/1757-899X/490/6/062041

Yeh, I. C. (2009). Default of credit card clients (UCI Machine Learning Repository. https://doi.org/10.24432/C55S3H

Downloads

Published

2026-07-31

How to Cite

Andre Pratama Adiwijaya. (2026). Optimization of Support Vector Machine for Imbalanced Credit Risk Classification. International Journal Science and Technology, 5(2), 274–295. https://doi.org/10.56127/ijst.v5i2.3038

Citation Check

Similar Articles

1 2 3 4 5 6 > >> 

You may also start an advanced similarity search for this article.