Comparative analysis of machine learning algorithms for tuberculosis classification based on symptom data

Authors

  • Ihsan Fathoni Amri Departement of Data Science, Universitas Muhammadiyah Semarang, Central Java, Indonesia
  • Muhammad Ivan Ardiansyah Departement of Data Science, Universitas Muhammadiyah Semarang, Central Java, Indonesia
  • Wikanastri Hersoelistyorini Department of Food Technology, Universitas Muhammadiyah Semarang, Central Java, Indonesia

DOI:

https://doi.org/10.30762/f_m.v9i1.8357

Keywords:

Tuberculosis Classification, Machine Learning, K-Nearest Neighbor, XGBoost, Symptom-Based Prediction

Abstract

Tuberculosis remains a major global health issue due to its high transmission rate and delays in early detection. Early identification of suspected tuberculosis cases based on patient symptoms is important for timely screening and reducing disease transmission. This study aims to compare the performance of Logistic Regression, Support Vector Machine, K-Nearest Neighbor, and Extreme Gradient Boosting in classifying suspected tuberculosis cases using symptom-based data. The dataset, consisting of clinical indicators such as cough, fever, shortness of breath, and other tuberculosis-related symptoms, was preprocessed and divided into training and testing sets. Model performance was evaluated using accuracy, confusion matrix, sensitivity/recall, specificity, precision, F1-score, and balanced accuracy. The results show that K-Nearest Neighbor achieved the highest accuracy of 87%, compared with Support Vector Machine at 80%, Logistic Regression at 72%, and XGBoost at 71%. However, the confusion matrix showed that KNN produced 110 false-negative cases and only 19 true-positive TB cases, resulting in a TB recall of approximately 14.7%. These findings indicate that high accuracy does not necessarily reflect good screening performance, especially when sensitivity is low. Therefore, although KNN showed the highest accuracy, it cannot yet be considered adequate as a standalone tuberculosis screening model. Further improvement and validation using larger, balanced, and clinically confirmed datasets are required.

Downloads

Download data is not yet available.

References

Ahadiyah, K., Dewi, A. F., & Romadewanti, S. H. (2024). Application of binary logistic regression analysis to factors that influence participation in 2024 presidential election. Journal Focus Action of Research Mathematic (Factor M), 7(2), 224–238. https://doi.org/10.30762/f_m.v7i2.3703

Alqaissi, E. Y., Alotaibi, F. S., & Ramzan, M. S. (2022). Modern machine-learning predictive models for diagnosing infectious diseases. Computational and Mathematical Methods in Medicine, 2022(1), 1–13. https://doi.org/10.1155/2022/6902321

Amri, I. F., Rohim, F. H. N., Ardiansyah, M. I., Saputra, F. S., Supriyanto, N. A. F., & Nakib, A. M. (2025). Analysis of suspected factors associated with tuberculosis cases in Semarang City using a logistic regression model. Scientific Journal of Computer Science, 1(1), 23–34. https://doi.org/10.64539/sjcs.v1i1.2025.32

Amri, I. F., Rohim, F. H. N., Azka, M. I. N., & Rakhmawati, M. S. (2025). Analysis of gallstone incidence factors using a binary logistic regression model. EKSAKTA Journal of Science and Data Analysis, 6(2), 72–83. https://doi.org/10.20885/EKSAKTA.vol6.iss2.art3

B, H. M. K., Jose, S. A., Jirawattanapanit, A., & Mathew, K. (2025). A comprehensive study on tuberculosis prediction models: Integrating machine learning into epidemiological analysis. Journal of Theoretical Biology, 597, Article 111988. https://doi.org/10.1016/j.jtbi.2024.111988

Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785–794. https://doi.org/10.1145/2939672.2939785

Chinagudaba, S. N., Gera, D., Dasu, K. K. V., S, U. S., K, K., Singarajpure, A., U, S., N, S., Chadda, V. K., & N, S. B. (2023). Predictive analysis of tuberculosis treatment outcomes using machine learning: A Karnataka TB data study at a scale. arXiv. https://arxiv.org/abs/2403.08834

Dasarwar, P., Yadav, U., Morris, K., Chavhan, N., Bondre, S., & Kalamkar, S. (2025). Revolutionizing tuberculosis prediction: A cutting-edge approach. Engineering, Technology & Applied Science Research, 15(3), 22929–22936. https://doi.org/10.48084/etasr.10449

Dunias, Z. S., Calster, B. V., Timmerman, D., Boulesteix, A. L., & Smeden, M. V. (2024). A comparison of hyperparameter tuning procedures for clinical prediction models: A simulation study. Statistics in Medicine, 43(6), 1119–1134. https://doi.org/10.1002/sim.9932

Esteva, A., Robicquet, A., Ramsundar, B., Kuleshov, V., Depristo, M., Chou, K., Cui, C., Corrado, G., Thrun, S., & Dean, J. (2019). A guide to deep learning in healthcare. Nature Medicine, 25(1), 24–29. https://doi.org/10.1038/s41591-018-0316-z

Fisher, A., Rudin, C., & Dominici, F. (2019). All models are wrong, but many are useful: Learning a variable's importance by studying an entire class of prediction models simultaneously. Journal of Machine Learning Research, 20(177), 1–81. http://jmlr.org/papers/v20/18-760.html

Hindustani, B. D., Hindustani, S., & Nguyen, P. (2007). Tackling tuberculosis: A comparative dive into machine learning for tuberculosis detection. arXiv. https://doi.org/10.48550/arXiv.2512.02364

Hussain, F., Elmedany, W., & Hamad, M. (2019). Evaluating risk factors in cardiovascular disease diagnosis using machine learning approaches. SSRN. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4896719

Izzati, N., Yannuansa, N., Ummah, I., Wati, D. A. R., Indahwati, E., Purwasih, S. M., & Kholis, N. (2024). Dynamical analysis of a mathematical model on the spread of diphtheria disease with vaccination completeness factor. Journal Focus Action of Research Mathematic (Factor M), 7(2), 206–223. https://doi.org/10.30762/f_m.v7i2.3901

Lundberg, S. M., Erion, G., Chen, H., Degrave, A., Prutkin, J. M., Nair, B., Katz, R., Himmelfarb, J., Bansal, N., & Lee, S. (2019). Explainable AI for trees: From local explanations to global understanding. arXiv. https://doi.org/10.48550/arXiv.1905.04610

Molnar, C., Casalicchio, G., & Bischl, B. (2020). Interpretable machine learning – a brief history, state-of-the-art, and challenges. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases (pp. 417–431). Springer International Publishing. https://link.springer.com/chapter/10.1007/978-3-030-65965-3_28

Multiyaningrum, R., Dawi, H. R., Hartanto, R. N. M., Al Haris, M., & Amri, I. F. (2025). Forecasting the rupiah exchange rate against the US Dollar using the LSTM algorithm. Journal Focus Action of Research Mathematic (Factor M), 8(2), 201–216. https://doi.org/10.30762/f_m.v8i2.6530

Ningrum, A. F., & Amri, I. F. (2024). Sentiment analysis of public opinion on handling stunting in Indonesia using random forest. J Statistika: Jurnal Ilmiah Teori dan Aplikasi Statistika, 17(1), 645–654. https://doi.org/10.36456/jstat.vol17.no1.a9088

Olojido, J. B., & Oladimeji, B. J. (2025). Prediction of tuberculosis using machine learning techniques. Faculty of Natural and Applied Sciences Journal of Mathematical and Statistical Computing, 2(2), 8–14. https://fnasjournals.com/index.php/FNAS-JMSC/article/view/723

Pagan, M., Zarlis, M., & Candra, A. (2023). Investigating the impact of data scaling on the k-nearest neighbor algorithm. Computer Science and Information Technologies, 4(2), 135–142. https://doi.org/10.11591/csit.v4i2.p135-142

Priyono, E. (2024). Prediction of tuberculosis patients with machine learning algorithms. IPI (Jurnal Ilmiah Penelitian dan Pembelajaran Informatika), 9(4), 2334–2341. https://doi.org/10.29100/jipi.v9i4.5486

Santosa, Y. T., Maulida, D. W., Sukowati, B. A., & Mahmudah, M. H. (2025). Comparative analysis of first-order linear differential equations and arithmetic methods in projecting Surakarta City's population. Journal Focus Action of Research Mathematic (Factor M), 8(1), 35–51. https://doi.org/10.30762/f_m.v8i1.5026

Sari, K. N. (2025). Panel data regression analysis utilizing CEM and FEM methods in relation to the profitability of Sharia commercial banks (2015–2024). Journal Focus Action of Research Mathematic (Factor M), 8(2), 242–261. https://doi.org/10.30762/f_m.v8i2.5633

Saudjhana, A., Budiman, A., Fernando, H., Juliantio, J., Junianto, K., Venessa, K., Salim, S., & Tomy, T. (2024). Implementasi big data terhadap pengecekan medis dan konsultasi kesehatan di Indonesia [Implementation of big data on medical check-ups and health consultations in Indonesia]. Journal of Information Systems and Technology, 5(1), 1–6. https://doi.org/10.37253/joint.v5i1.4323

Shahzad, M. F., Xu, S., Lim, W. M., Yang, X., & Khan, Q. R. (2024). Artificial intelligence and social media on academic performance and mental well-being: Student perceptions of positive impact in the age of smart learning. Heliyon, 10(8). https://doi.org/10.1016/j.heliyon.2024.e29523

Sim, J. Z. T., Fong, Q. W., Huang, W., & Tan, C. H. (2021). Machine learning in medicine: What clinicians should know. Singapore Medical Journal, 64(2), 91–97. https://doi.org/10.11622/smedj.2021054

Sun, D., Wu, X., Zhang, Y., Wang, W., He, M., & Diao, L. (2026). Machine learning models for predicting latent tuberculosis infection risk in close contacts of patients with pulmonary tuberculosis — Henan Province, China, 2024. China CDC Weekly, 8(3), 71–79. https://weekly.chinacdc.cn/en/article/doi/10.46234/ccdcw2026.012?viewType=HTML

Topol, E. J. (2019). High-performance medicine: The convergence of human and artificial intelligence. Nature Medicine, 25(1), 44–56. https://doi.org/10.1038/s41591-018-0300-7

Yang, Y., Zhao, C., & Lu, H. (2025). Leveraging machine learning for food waste reduction: An analysis of predictive models. Applied and Computational Engineering, 112, 154–160. https://doi.org/10.54254/2755-2721/112/2025.18115

Yudistira, I., Rohman, N., Mardianto, M. F. F., & Kuzairi. (2023). Comparison modeling of the number population who have been vaccinated in East Java using the biresponse Fourier series estimator method with the trend function. Journal Focus Action of Research Mathematic (Factor M), 6(1), 37–47. https://doi.org/10.30762/factor_m.v6i1.685

Downloads

Published

29-06-2026

How to Cite

Amri, I. F., Ardiansyah, M. I., & Hersoelistyorini, W. (2026). Comparative analysis of machine learning algorithms for tuberculosis classification based on symptom data. Journal Focus Action of Research Mathematic (Factor M), 9(1), 155–173. https://doi.org/10.30762/f_m.v9i1.8357

Issue

Section

Articles