Comparative Analysis Of Classification Algorithm On Problem Data Of ATM Machines With Rapid

Authors

DOI:

https://doi.org/10.22303/csrid-.16.2.2024.188-200

Keywords:

Data Mining, Klasifikasi, Naïve Byaes, Rapidminer

Abstract

The aim of the proposed research is to compare and test the accuracy of data mining classification algorithms. Comparing algorithms that depend on different parameters of a given data set. There are learning and classification algorithms that are used to analyze, study and classify the available data. However, the problem is finding the best algorithm and the desired results with the highest level of accuracy in predicting future values ​​or events from a data set. Where the classification models used are the C4.5 and Naïve Bayes algorithms. Testing and validation using k-fold Cross Validation as well as evaluating the performance of the prediction model using the ROC-AUC graph with graphic visualization. The data used as samples were taken from ATM machine problem data with a total of approximately 250 samples. Testing was carried out with the help of the Rapidminer tool with operators and parameters used in creating models of the algorithms being compared. The tests that have been carried out prove that the C4.5 algorithm has the best performance with an average accuracy value of 96.00%, a recall value of 97.78% and a precision value of 92.14%, while the naïve Bayes algorithm produces an accuracy value of 83. 00%, the recall value is 76.40% and the precision value is 84.82%. Apart from that, evaluation and validation in this test is also seen based on the ROC curve called AUC (Area Under the ROC Curve) where for the C4.5 algorithm the value is 0.931 while naïve Bayes is 0.894 so the C4.5 algorithm is categorized as Very Good Classification because it has a value between 0.90-1.00. These results show that the C4.5 algorithm is proven to be a potentially effective and efficient classification algorithm.

References

D. Y. Hakim Tanjung, “Penerapan Algoritma Naïve Bayes Untuk Klasifikasi Data Pengisian ATM,” Infosys (Information Syst. J., vol. 7, no. 1, p. 12, 2022, doi: 10.22303/infosys.7.1.2022.12-24.

K. and S. A. Rajesh, “Analysis of SEER Dataset for Breast Cancer Diagnosis using C4.5 Classification Algorithm,” Int. J. Adv. Res. Comput. Commun. Eng., vol. 1, no. 2, pp. 72–77, 2012, [Online]. Available: www.ijarcce.com

S. Hendrian, “Algoritma Klasifikasi Data Mining Untuk Memprediksi Siswa Dalam Memperoleh Bantuan Dana Pendidikan,” Fakt. Exacta, vol. 11, no. 3, pp. 266–274, 2018, doi: 10.30998/faktorexacta.v11i3.2777.

D. Huchon, N. Crozet, N. Cantenot, and R. Ozon, “Germinal vesicle breakdown in the Xenopus laevis oocyte: Description of a transient microtubular structure,” Reprod. Nutr. Dev., vol. 21, no. 1, pp. 135–148, 1981, doi: 10.1051/rnd:19810112.

S. Masripah, “Komparasi Algoritma Klasifikasi Data Mining untuk Evaluasi Pemberian Kredit,” Bina Insa. ICT J., vol. 3, no. 1, p. 234336, 2016.

A. Saleh, M. Maryam, and K. Puspita, “Determination of Corn Quality using the Decision Tree of C 4.5 Algorithm,” 2019 7th Int. Conf. Cyber IT Serv. Manag. CITSM 2019, pp. 5–8, 2019, doi: 10.1109/CITSM47753.2019.8965334.

P. Subarkah, E. P. Pambudi, and S. O. N. Hidayah, “Perbandingan Metode Klasifikasi Data Mining untuk Nasabah Bank Telemarketing,” MATRIK J. Manajemen, Tek. Inform. dan Rekayasa Komput., vol. 20, no. 1, pp. 139–148, 2020, doi: 10.30812/matrik.v20i1.826.

M. Kamil and W. Cholil, “Analisis Perbandingan Algoritma C4.5 dan Naive Bayes pada Lulusan Tepat Waktu Mahasiswa di Universitas Islam Negeri Raden Fatah Palembang,” J. Inform., vol. 7, no. 2, pp. 97–106, 2020, doi: 10.31294/ji.v7i2.7723.

S. Hulu and P. Sihombing, “Analysis of Performance Cross Validation Method and K-Nearest Neighbor in Classification Data,” Int. J. Res. Rev., vol. 7, no. 4, pp. 69–73, 2020.

H. Azis, P. Purnawansyah, F. Fattah, and I. P. Putri, “Performa Klasifikasi K-NN dan Cross Validation Pada Data Pasien Pengidap Penyakit Jantung,” Ilk. J. Ilm., vol. 12, no. 2, pp. 81–86, 2020, doi: 10.33096/ilkom.v12i2.507.81-86.

F. Tempola, M. Muhammad, and A. Khairan, “Perbandingan Klasifikasi Antara KNN dan Naive Bayes pada Penentuan Status Gunung Berapi dengan K-Fold Cross Validation,” J. Teknol. Inf. dan Ilmu Komput., vol. 5, no. 5, p. 577, 2018, doi: 10.25126/jtiik.201855983.

P. Mochamad Rizki Ilham, “Implementasi Data Mining Menggunakan Algoritma C4.5 Untuk Prediksi Kepuasan Pelanggan Taksi Kosti,” Simplementasi Data Min. Menggunakan Algoritm. C4.5 Untuk Prediksi Kepuasan Pelangg. Tak. Kosti, vol. Vol. 4, No, no. 5, p. 11, 2016.

D. L. Arisandy, “Analisis Perbandingan Algoritma Naive Bayes dan Algoritma C4.5 untuk Klasifikasi Multi Data,” no. 1310651061, 2017.

J. Keilwagen, I. Grosse, and J. Grau, “Area under precision-recall curves for weighted and unweighted data,” PLoS One, vol. 9, no. 3, pp. 1–13, 2014, doi: 10.1371/journal.pone.0092209.

T. Saito and M. Rehmsmeier, “The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets,” PLoS One, vol. 10, no. 3, pp. 1–21, 2015, doi: 10.1371/journal.pone.0118432.

M. Buckland and F. Gey, “The relationship between Recall and Precision,” J. Am. Soc. Inf. Sci., vol. 45, no. 1, pp. 12–19, 1994, doi: 10.1002/(SICI)1097-4571(199401)45:1<12::AID-ASI2>3.0.CO;2-L.

T. Fawcett, “An introduction to ROC analysis,” Pattern Recognit. Lett., vol. 27, no. 8, pp. 861–874, 2006, doi: 10.1016/j.patrec.2005.10.010.

T. k and M. Wadhawa, “Analysis and Comparison Study of Data Mining Algorithms Using Rapid Miner,” Int. J. Comput. Sci. Eng. Appl., vol. 6, no. 1, pp. 9–21, 2016, doi: 10.5121/ijcsea.2016.6102.

M. A. Syahab and Z. Purnama, “Analisis Rapid Miner Terhadap Tujuan Pendekatan Keahlian Pada keinginan Untuk Keluar Dengan Motivasi Intrinsik Sebagai Variabel Moderator,” vol. 12, pp. 1593–1602, 2023.

Y. Mardi, “Data Mining : Klasifikasi Menggunakan Algoritma C4.5,” J. Edik Inform., vol. 2, no. 2, pp. 213–219, 2017.

F. Harahap, “Penerapan data Mining dalam Pemilihan Mobil Menggunakan Algoritma C4.5,” J. VOI (Voice Informatics), vol. 7, no. x, 2018.

A. Lestari, “Increasing Accuracy of C4 . 5 Algorithm Using Information Gain Ratio and Adaboost for Classification of Chronic Kidney Disease,” pp. 32–38, 2020.

K. S. Raju, M. R. Murty, M. V. Rao, and S. C. Satapathy, “Support Vector Machine with K-fold Cross Validation Model for Software Fault Prediction,” Int. J. Pure Appl. Math., vol. 118, no. 20, pp. 321–334, 2018, [Online]. Available: https://www.researchgate.net/publication/329414359_Support_Vector_Machine_with_K-fold_Cross_Validation_Model_for_Software_Fault_Prediction

H. Zhang and J. Su, “Learning probabilistic decision trees for AUC,” Pattern Recognit. Lett., vol. 27, no. 8, pp. 892–899, 2006, doi: 10.1016/j.patrec.2005.10.013.

F. Vilariño, L. I. Kuncheva, and P. Radeva, “ROC curves and video analysis optimization in intestinal capsule endoscopy,” Pattern Recognit. Lett., vol. 27, no. 8, pp. 875–881, 2006, doi: 10.1016/j.patrec.2005.10.011.

Downloads

Published

2024-06-15

Issue

Section

Articles

How to Cite

Comparative Analysis Of Classification Algorithm On Problem Data Of ATM Machines With Rapid . (2024). CSRID (Computer Science Research and Its Development Journal), 16(2), 188-200. https://doi.org/10.22303/csrid-.16.2.2024.188-200

Similar Articles

11-20 of 86

You may also start an advanced similarity search for this article.