Identifikasi Pesan Penipuan Berkedok Hadiah Digital Menggunakan Machine Learning

Authors

  • Muhammad Shafari Rahmat Universitas Gunadarma, Indonesia
  • Syahrizal Andhika Universitas Gunadarma, Indonesia

DOI:

https://doi.org/10.56127/jts.v5i2.2948

Keywords:

digital prize scam, machine learning, text classification, TF-IDF, XGBoost

Abstract

Digital prize scams are increasingly distributed through text messages, social media, and instant messaging platforms using persuasive expressions, suspicious links, and urgent instructions. This study aims to develop a machine learning-based approach for identifying Indonesian-language scam messages disguised as digital prize notifications. The dataset consisted of 6,000 text messages, comprising 3,000 scam messages and 3,000 non-scam messages. The data were manually labeled by two independent annotators and processed through case folding, text cleaning, tokenization, stopword removal, and stemming. Text features were transformed into numerical representations using Term Frequency–Inverse Document Frequency (TF-IDF). Five classification algorithms were evaluated, including Multinomial Naive Bayes, Support Vector Machine, Logistic Regression, Random Forest, and XGBoost. Model performance was assessed using stratified five-fold cross-validation based on accuracy, precision, recall, F1-score, and ROC-AUC. The results showed that XGBoost achieved the best performance, with an accuracy of 97.0%, precision of 97.2%, recall of 96.8%, F1-score of 97.0%, and ROC-AUC of 99.3%. Scam messages were commonly characterized by words related to prizes, winning notifications, free offers, claims, and urgency. These findings indicate that TF-IDF combined with XGBoost can effectively support the automatic detection of digital prize scam messages. Future studies should incorporate URL characteristics, sender metadata, and contextual language models to improve detection performance.

References

Abayomi-Alli, O. O., Misra, S., & Abayomi-Alli, A. (2022). A deep learning method for automatic SMS spam classification: Performance of learning algorithms on indigenous dataset. Concurrency and Computation: Practice and Experience, 34(17), e6989. doi:10.1002/cpe.6989.

Altunay, H. C., & Albayrak, Z. (2024). SMS spam detection system based on deep learning architectures for Turkish and English messages. Applied Sciences, 14(24), 11804. doi:10.3390/app142411804.

Carroll, F., Adejobi, J. A., & Montasari, R. (2022). How good are we at detecting a phishing attack? Investigating the evolving phishing attack email and why it continues to successfully deceive society. SN Computer Science, 3, 170. doi:10.1007/s42979-022-01069-1.

Ejaz, A., Mian, A. N., & Manzoor, S. (2023). Life-long phishing attack detection using continual learning. Scientific Reports, 13, 11488. doi:10.1038/s41598-023-37552-9.

Goel, D., & Jain, A. K. (2018). Mobile phishing attacks and defence mechanisms: State of art and open research challenges. Computers & Security, 73, 519–544. doi:10.1016/j.cose.2017.12.006.

Jain, A. K., Yadav, S. K., & Choudhary, N. (2020). A novel approach to detect spam and smishing SMS using machine learning techniques. International Journal of E-Services and Mobile Applications, 12(1), 21–38. doi:10.4018/IJESMA.2020010102.

Liu, X., Lu, H., & Nayak, A. (2021). A spam transformer model for SMS spam detection. IEEE Access, 9, 80253–80263. doi:10.1109/ACCESS.2021.3081479.

Mahmud, T., Prince, M. A. H., Ali, M. H., Hossain, M. S., & Andersson, K. (2024). Enhancing cybersecurity: Hybrid deep learning approaches to smishing attack detection. Systems, 12(11), 490. doi:10.3390/systems12110490.

Otoritas Jasa Keuangan. (2024, March 26). Wajib tahu, OJK tidak pernah meminta dana untuk pencairan hadiah.

Otoritas Jasa Keuangan. (2025a, November 15). Satgas PASTI imbau masyarakat waspadai penipuan menggunakan artificial intelligence.

Otoritas Jasa Keuangan. (2025b, March 5). Waspada penipuan mengatasnamakan bank melalui SMS.

Pramakrisna, F. D., Adhinata, F. D., & Tanjung, N. A. F. (2022). Aplikasi klasifikasi SMS berbasis web menggunakan algoritma logistic regression. Teknika, 11(2), 90–97. doi:10.34148/teknika.v11i2.466.

Roy, P. K., Singh, J. P., & Banerjee, S. (2020). Deep learning to filter SMS spam. Future Generation Computer Systems, 102, 524–533. doi:10.1016/j.future.2019.09.001.

Shaaban, M. A., Hassan, Y. F., & Guirguis, S. K. (2022). Deep convolutional forest: A dynamic deep ensemble approach for spam detection in text. Complex & Intelligent Systems, 8(6), 4897–4909. doi:10.1007/s40747-022-00741-6.

Shinde, A., Shahra, E. Q., Basurra, S., Saeed, F., AlSewari, A. A., & Jabbar, W. A. (2024). SMS scam detection application based on optical character recognition for image data using unsupervised and deep semi-supervised learning. Sensors, 24(18), 6084. doi:10.3390/s24186084.

Downloads

Published

2026-07-19

How to Cite

Muhammad Shafari Rahmat, & Syahrizal Andhika. (2026). Identifikasi Pesan Penipuan Berkedok Hadiah Digital Menggunakan Machine Learning. Jurnal Teknik Dan Science, 5(2), 88–99. https://doi.org/10.56127/jts.v5i2.2948

Citation Check

Similar Articles

<< < 1 2 3 

You may also start an advanced similarity search for this article.