International Research Journal of Engineering and Technology (IRJET)
e-ISSN: 2395-0056
Volume: 12 Issue: 01 | Jan 2025
p-ISSN: 2395-0072
www.irjet.net
Predictive Modeling for Chronic Kidney Disease Using Machine Learning Hemangi Patil1, Gaurav Acharya2 1Independent Researcher, Mumbai, India
2Department of master’s in computer science, IIT Chicago, Illinois
-------------------------------------------------------------------------***----------------------------------------------------------------------1.2. OBJECTIVES Abstract - Chronic Kidney Disease (CKD) is a global health problem since it starts asymptomatically and progresses in an irreversible manner. Therefore, the earlier and accurate diagnosis plays a significant role in improving patient outcomes. This paper presents a machine learning (ML) framework that handles data problems such as missing values and evaluated different ML-models for addressing the problem of CKD diagnosis. K-Nearest Neighbors (KNN) data imputation and classifier evaluation on six classifiers show Random Forest (RF) out of all trained classifiers performs best with the accuracy of 99.75%. A new hybrid method with Logistic Regression (LOG) and RF gives an additional improvement in accuracy (99.83%). This approach is scalable and adaptable in a clinical environment.
The main objective of this study is to develop an early diagnostic model for Chronic Kidney Disease (CKD) that minimizes testing and cost while achieving high accuracy. The study's objectives are to use K-Nearest Neighbors (KNN) imputation to manage missing data in the CKD dataset and feature selection with information gain to determine which features are most crucial for CKD identification. To find the best accurate model, a variety of machine learning methods will be used and compared, such as Feedforward Neural Networks, Random Forest, Support Vector Machine, K-Nearest Neighbor, Naive Bayes, and Logistic Regression. The objective is to use 24 important predictors to predict the presence of CKD and evaluate the accuracy and other pertinent metrics of various algorithms.
Key Words: Chronic Kidney Disease, Machine Learning, Data Imputation, Random Forest, Logistic Regression, Predictive Modeling
2. LITERATURE SURVEY
1. INTRODUCTION
Different papers and articles have been reviewed for this project. Also, their conclusions are summarized in this section. The section present documents that were studied prior and post project development. The mentioned articles provide with a better understanding about structure of the system and how various algorithms could be combined together so as to build a system with higher efficiency.
Chronic Kidney Disease (CKD) is a progressive condition characterized by the gradual decline of kidney function, often remaining undetected until it reaches advanced stages. Early detection is critical for preventing further complications and improving patient outcomes. However, traditional diagnostic methods are time-consuming and may not always be practical in clinical settings. In recent years, machine learning (ML) has emerged as a promising solution for automating CKD diagnosis, offering faster and more accurate predictions. Despite its potential, existing ML models face challenges such as incomplete datasets and limited generalization. This paper proposes a novel MLbased framework for CKD diagnosis that addresses these challenges. Specifically, it utilizes K-Nearest Neighbors (KNN) imputation to handle missing data, making the model applicable even when diagnostic categories are unknown or incomplete. The study evaluates several ML algorithms, including Logistic Regression (LOG), Random Forest (RF), Support Vector Machine (SVM), K-Nearest Neighbor (KNN), Naive Bayes (NB), and Feedforward Neural Networks (FNN) to establish CKD diagnostic models. Furthermore, a hybrid model combining Logistic Regression and Random Forest is introduced to improve the accuracy of the predictions. This hybrid model achieves an impressive accuracy of 99.83%, highlighting its effectiveness and potential for clinical adoption in CKD diagnosis.
© 2025, IRJET
|
Impact Factor value: 8.315
Table -1: Publications Cited: Title
Year
Diagnosis of 2016 patients with (Chemometr. chronic kidney Intell. Lab.) disease by using two fuzzy classifiers Diagnosis of chronic kidney disease by using random forest
|
Author
Summary
Z. Chen et Used fuzzy al. classifiers to diagnose CKD, handling incomplete datasets for better accuracy.
2017
A. Subasi, Applied Random (Int. Conf. E. Medical and Alickovic, J. Forest for CKD Kevric prediction, Biological with emphasis Engineering) on feature selection.
ISO 9001:2008 Certified Journal
|
Page 220