Pengembangan Model Hybrid Long Short-Term Memory – Extreme Gradient Boosting untuk Prediksi Diabetes dan Kolesterol
Abstract
Diabetes and high cholesterol are two health conditions that can increase the risk of various chronic diseases if not detected and treated early. Therefore, a prediction method is needed that can help identify disease risks based on patient health data. This study aims to develop a diabetes and cholesterol prediction model using a hybrid Long Short-Term Memory (LSTM) and Extreme Gradient Boosting (XGBoost) method. The dataset used was obtained from the Kaggle platform and includes various health parameters, such as age, gender, body mass index (BMI), blood glucose levels, HbA1c levels, blood pressure, as well as the patient's medical history and lifestyle. The research stages include data collection, data pre-processing, training and testing data division, data balancing using the Synthetic Minority Over-sampling Technique (SMOTE), feature extraction using LSTM, and the classification process using XGBoost. Based on the test results, the diabetes prediction model obtained an accuracy value of 74.42%, precision of 20.20%, recall of 68.06%, and an F1-score of 31.15%. Meanwhile, the cholesterol prediction model achieved an accuracy of 71.73%, a precision of 45.16%, a recall of 60.73%, and an F1-score of 51.80%. These results demonstrate that the hybrid LSTM–XGBoost method is capable of identifying patterns in patient health data and producing quite good classification performance to support early detection of diabetes and high cholesterol risks.
References
Afsaneh, E., Sharifdini, A., Ghazzaghi, H., & Zarei Ghobadi, M. (2022). Recent applications of machine learning and deep learning models in the prediction, diagnosis, and management of diabetes: a comprehensive review. Diabetology & Metabolic Syndrome, 14, 196. https://dmsjournal.biomedcentral.com/articles/10.1186/s13098-022-00969-9
Araf, I., Idri, A., & Chairi, I. (2024). Cost-sensitive learning for imbalanced medical data: A review. Artificial Intelligence Review, 57, 80. https://doi.org/10.1007/s10462-023-10652-8
Chawla, N. V, Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic Minority Over-sampling Technique. Journal of Artificial Intelligence Research, 16, 321–357. https://doi.org/10.1613/jair.953
Firthous, J. J., & Murugeswari, G. (2026). Diabetes Prediction using PPG Signal and Clinical Data with XGBoost Classification. The Bioscan, 21(2), 764–779. https://thebioscan.com/index.php/pub/article/view/5765
Fregoso-Aparicio, L., Noguez, J., Montesinos, L., & García-García, J. A. (2021). Machine learning and deep learning predictive models for type 2 diabetes: a systematic review. Diabetology & Metabolic Syndrome, 13, 148. https://dmsjournal.biomedcentral.com/articles/10.1186/s13098-021-00767-9
Goodfellow, I., Bengio, Y., & Courville, A. (2021). Deep Learning. MIT Press.
Hayashi, T., Shimizu, T., & Fukami, Y. (2021). Collaborative Problem Solving on a Data Platform Kaggle. ArXiv.
Hochreiter, S., & Schmidhuber, J. (1997). Long Short-Term Memory. Neural Computation, 9(8), 1735–1780. https://doi.org/10.1162/neco.1997.9.8.1735
Kavitha, V. (2021). Design of intelligent diabetes mellitus detection system using hybrid feature selection based XGBoost classifier. Computers in Biology and Medicine, 136, 104664. https://www.sciencedirect.com/science/article/pii/S0010482521004583
Khanam, J., & Foo, S. (2022). A Comparison of Machine Learning Algorithms for Diabetes Prediction.
Khanam, J. J., & Foo, S. Y. (2022). Diabetes mellitus prediction and diagnosis from a data preprocessing and machine learning perspective. Computer Methods and Programs in Biomedicine, 220, 106773. https://www.sciencedirect.com/science/article/pii/S0169260722001596
Kumar, V., Lalotra, G. S., & Sasikala, P. (2022). Addressing binary classification over class imbalanced clinical datasets using computationally intelligent techniques. Healthcare, 10(7), 1293. https://doi.org/10.3390/healthcare10071293
Mufidah, S. I., & Handayani, I. (2025). Deteksi Risiko Diabetes Berdasarkan Faktor Kesehatan Menggunakan Long Short-Term Memory (LSTM) dan XGBoost. INTECOMS: Journal of Information Technology and Computer Science. https://journal.ipm2kpe.or.id/index.php/INTECOM/article/view/17291
Pudjihartono, N., Fadason, T., Kempa-Liehr, A. W., & O’Sullivan, J. M. (2022). A review of feature selection methods for machine learning-based disease risk prediction. Frontiers in Bioinformatics, 2, 927312. https://doi.org/10.3389/fbinf.2022.927312
Rabby, M. F. F. (2021). Blood Glucose Prediction Using LSTM-Based Deep Learning Model.
Rabby, M. F., Tu, Y., & Hossen, M. I. (2021). Stacked LSTM based deep recurrent neural network with Kalman smoothing for blood glucose prediction. BMC Medical Informatics and Decision Making, 21, 101. https://bmcmedinformdecismak.biomedcentral.com/articles/10.1186/s12911-021-01462-5
Sathyanarayanan, S., & Tantri, B. R. (2024). Confusion Matrix-Based Performance Evaluation Metrics. 27(4).
Schiborn, C., & Schulze, M. B. (2022). Precision prognostics for the development of complications in diabetes. Diabetologia, 65, 1867–1882. https://link.springer.com/article/10.1007/s00125-022-05731-4
Sharma, A., Kumar, R., & Singh, P. (2023). XGBoost-based machine learning models for healthcare prediction: A review. Healthcare Analytics, 4, 100215.
Sowjanya, A. M., & Mrudula, O. (2023). Effective treatment of imbalanced datasets in health care using modified SMOTE coupled with stacked deep learning algorithms. Applied Nanoscience, 13(3), 1829–1840. https://doi.org/10.1007/s13204-021-02063-4
Verma, A., Gupta, D., & Kaur, M. (2022). Performance analysis of XGBoost algorithm for classification and prediction tasks. Expert Systems with Applications, 198, 116805.
Zanelli, S., Ammi, M., Hallab, M., & El Yacoubi, M. A. (2022). Diabetes Detection and Management through Photoplethysmographic and Electrocardiographic Signals Analysis: A Systematic Review. Sensors, 22(13), 4890. https://www.mdpi.com/1424-8220/22/13/4890
Buku:
Goodfellow, I., Bengio, Y., & Courville, A. (2021). Deep Learning. MIT Press.


