Skip to main content

Diabetes Detection Using Machine Learning

Page 1


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

Diabetes Detection Using Machine Learning

Dr. Sri Hari Nallamala1 , Supriya makineni2 , Sonti Kusuma3, Shaik Khaja Muzeer4

1Professor, Dept. of Computer Science and Engineering, VVIT, Guntur, India

234Student, Dept. of Computer Science and Engineering, VVIT, Guntur, India

Abstract - Diabetes is a chronic disease that requires early detection to prevent severe health complications such as heart disease, kidney failure, and vision loss. This paper presents a machine learning-based system for diabetes prediction using the XGBoost algorithm. The model is trained using medical parameters such as glucose level, BMI, blood pressure, insulin level, and age. Data preprocessing techniques including normalization and handling missing values are applied to improve model performance. The model is evaluated using accuracy, precision, recall, and F1-score. Experimental results show that XGBoost outperforms traditional algorithms and provides reliable predictions for early diagnosis.

Key Words: Diabetes Detection, XGBoost, Machine Learning, Healthcare, Prediction

1. INTRODUCTION

Diabetes is a chronic disease characterized by high blood sugarlevels,whichcanleadtoserioushealthcomplications such as heart disease, kidney failure, nerve damage, and vision problems if not detected early. According to global health reports, the number of diabetes cases has been increasingrapidlyduetochangesinlifestyle,unhealthydiet, andlackofphysicalactivity.Earlypredictionanddiagnosis of diabetes play a crucial role in reducing its impact and preventinglong-termcomplications. Traditional methods of diabetes diagnosis often rely on laboratory tests and clinical expertise, which can be timeconsuming and may not always provide early-stage detection.Theselimitationshighlighttheneedforintelligent systemsthatcanassistinfasterandmoreaccuratediagnosis usingavailablemedicaldata.

Machine learning techniques have emerged as powerful toolsinthehealthcaredomain,enablingtheanalysisoflarge datasetstoidentifyhiddenpatternsandrelationships.These techniques can be used to build predictive models that support medical decision-making.In recentyears,various machinelearningalgorithmshavebeenappliedfordisease prediction,showingpromisingresultsintermsofaccuracy andefficiency.

In this project, a diabetes detection system is developed using the XGBoost algorithm implemented in Python. XGBoostisanadvancedensemblelearningtechniqueknown foritshighperformanceandscalability.Themodelanalyzes importantmedicalparameterssuchasglucoselevel,blood pressure, bodymassindex(BMI),insulinlevel,and age to

predict whether a person is diabetic or not. The model analyzesseveralimportantmedicalparametersthatplaya keyroleindeterminingthelikelihoodofdiabetes:

Themodelanalyzesseveralimportantmedicalandlifestyle parameterstopredictthelikelihoodofdiabetes:

1.1 Gender

Gender indicates whetherthepatientis maleor female.It caninfluencediabetesriskduetodifferencesinhormonal levels,bodycomposition,andlifestylepatterns.

1.2 Age

Ageisasignificantfactor,astheriskofdevelopingdiabetes increaseswithage.OlderindividualsaremorepronetoType 2 diabetes due to reduced metabolic activity and lifestyle factors.

1.3 Hypertension

Hypertension refers to high blood pressure. It is closely associated with diabetes and increases the risk of cardiovascular complications. Patients with hypertension aremorelikelytodevelopdiabetes.

1.4 Heart Disease

Thisparameterindicateswhetherthepatienthasahistoryof heart disease. Diabetes and heart disease are strongly related, and the presence of heart disease can increase diabetesrisk.

1.5 Smoking History

Smokingisalifestylefactor thataffectsoverallhealthand increasestheriskofdiabetes.Itcanleadtoinsulinresistance andothermetabolicdisorders.

1.6 Body Mass Index(BMI)

BMI is a measure of body fat based on height and weight. Higher BMI values indicate obesity, which is a major risk factorfordiabetes.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

1.7 HbA1c Level(%)

HbA1c represents the average blood sugar level over the past2–3months.Itisoneofthemostimportantindicators fordiagnosingdiabetes.HigherHbA1cvaluesindicatepoor bloodsugarcontrol.

1.8 Blood Glucose Level(mg/dL)

Bloodglucoselevelmeasuresthecurrentamountofsugarin theblood.Elevatedglucoselevelsareaprimaryindicatorof diabetes.

These parameters collectively provide comprehensive information about the patient’s health condition. By analyzing these features, the XGBoost model identifies patternsandrelationshipstoaccuratelypredicttheriskof diabetes.

Theproposedsystemaimstoprovideareliableandefficient tool for early diabetes prediction. It can assist healthcare professionals in making informed decisions and enable timely intervention. By integrating machine learning into healthcaresystems,thisapproachcontributestoimproved patient outcomes and supports the development of smart healthcaresolutions.

2. PROBLEM STATEMENT

Diabetesisamajorglobalhealthconcernthatrequiresearly detection to prevent serious complications such as heart disease,kidneyfailure,andvisionloss.Traditionaldiagnostic methods can be time-consuming and may not always provideaccuratepredictionsatanearlystage.

Thereisaneedforanefficientandautomatedsystemthat canpredictdiabetesbasedonmedicalparameterswithhigh accuracy.Thechallengeliesinselectingappropriatemachine learningalgorithmsandhandlingmedical data effectively. Thisprojectaimstoaddressthesechallengesbydeveloping adiabetesdetectionsystemusingtheXGBoostalgorithm.

3. EXISTING SYSTEM

Theexistingsystemsfordiabetespredictionprimarilyrely ontraditionalmachinelearningalgorithmssuchasLogistic Regression,K-NearestNeighbors,andDecisionTrees.While these methods provide basic prediction capabilities, they often suffer from limitations such as lower accuracy and overfitting.

Additionally,manualanalysisofmedicaldataistimeconsumingandrequiresexpertknowledge.Thesesystems lack scalability and may not perform well with large datasets. Therefore, there is a need for more advanced approaches that can handle complex data and provide reliablepredictions.

4. PROPOSED SYSTEM

The proposed system uses the XGBoost algorithm for accurate and efficient diabetes prediction. XGBoost is an advancedensemblelearningtechniquethatenhancesmodel performance by combining multiple decision trees in a sequentialmanner.Eachtreeisbuilttocorrecttheerrorsof thepreviousone,resultinginahighlyoptimizedandrobust predictionmodel.

The system takes important medical parameters such as glucoselevel,bloodpressure,bodymassindex(BMI),age, andotherhealth-relatedfactorsasinput.Theseparameters arefirstprocessedthroughdatapreprocessingtechniques, which include handling missing values, removing inconsistencies,andnormalizingthedatatoimprovemodel performance.

Afterpreprocessing,thecleaneddatasetisusedtotrainthe XGBoostmodel.Themodellearnspatternsandrelationships betweentheinputfeaturesandthetargetvariable,enabling ittoaccuratelyclassifywhetherapersonisdiabeticornondiabetic.

The proposed system offers several advantages, including improved prediction accuracy, faster computation, and efficient handling of missing data. Additionally, XGBoost includes regularization techniques that help reduce overfittingandimprovegeneralization.Thesefeaturesmake the system more reliable and suitable for real-world healthcareapplications.

5. LITERATURE SURVEY

Several studies have been conducted in recent years to predictdiabetesusingmachinelearningtechniques,aiming to improve early diagnosis and reduce the risk of severe health complications. With the growing availability of healthcare data, researchers have explored various datadriven approaches to enhance prediction accuracy and supportclinicaldecision-making.

Traditional machine learning algorithms such as Logistic Regression and Decision Tree have been widely used for diabetes prediction. Logistic Regression is a statistical methodthatmodelstheprobabilityofabinaryoutcomeand issimpletoimplementandinterpret.DecisionTreesprovide ahierarchicalstructureofdecisionrules,makingthemeasy to understand. However, these methods often produce moderateaccuracyandaresensitivetonoiseandoverfitting, especiallywhendealingwithcomplexdatasets.

OthertechniquessuchasK-NearestNeighbors(KNN)and SupportVectorMachines(SVM)havealsobeenapplied in diabetespredictionsystems.KNNclassifiesinstancesbased on similarity measures, but it becomes computationally expensive for large datasets and is sensitive to irrelevant

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

features.SVMiseffectiveinhandlinghigh-dimensionaldata and can achieve good accuracy, but it requires careful parameter tuning and kernel selection, which increases complexity.

Toovercomethelimitations ofsinglemodels,researchers haveshiftedtowardsensemblelearningtechniquessuchas Random Forest and Gradient Boosting. Random Forest constructsmultipledecisiontreesusingdifferentsubsetsof data and combines their outputs to improve prediction accuracy and reduce overfitting. Gradient Boosting builds models sequentially, where each new model focuses on correcting the errors made by previous ones. These techniques have demonstrated improved performance comparedtotraditionalapproaches.

Amongtheensemblemethods,XGBoost(ExtremeGradient Boosting) has emerged as one of the most powerful algorithms for classification tasks. XGBoost enhances the gradient boosting framework by incorporating regularization techniques, which help prevent overfitting and improve generalization. It also supports parallel processing,makingitcomputationallyefficientandsuitable for large-scale datasets. Additionally, XGBoost can handle missingdata effectivelyandprovidesbuilt-inmechanisms forfeatureimportanceanalysis.

Several research works have reported that XGBoost outperformsothermachinelearningalgorithmsintermsof accuracy,precision,andoverallperformanceinhealthcare applications. It has been successfully applied in various diseasepredictionsystems,includingdiabetes,heartdisease, and cancer detection. The ability of XGBoost to capture complex relationships among features makes it highly suitableformedicaldataanalysis.

Furthermore, recent studies emphasize the importance of proper data preprocessing and feature selection in improving model performance. Techniques such as normalization, handling missing values, and selecting relevantfeaturessignificantlycontributetobetterprediction accuracy. Combining these preprocessing steps with advancedalgorithmslikeXGBoostresultsinmorereliable andefficientsystems.

Basedonthefindingsfrompreviousresearch,itisevident that ensemble learning techniques, particularly XGBoost, provide superior performance compared to traditional machinelearningmodels.Therefore,thisprojectutilizesthe XGBoost algorithm to develop an efficient and accurate diabetes detection system that can assist healthcare professionalsinearlydiagnosisanddecision-making.

6. METHODOLOGY

The proposed system uses a machine learning-based approachtodetectdiabetesbasedonmedicalparameters.

The implementation is carried out using Python and the XGBoostalgorithm,whichisknownforitshighperformance andefficiencyinclassificationproblems.Themethodology consists of several stages including data collection, preprocessing, feature selection, model training, and prediction.

6.1 Dataset

The dataset used in this project is obtained from a standard medical dataset repository. It contains various health-related attributes that are important for diabetes prediction.Thekeyfeaturesinthedatasetincludeglucose level,bloodpressure,bodymassindex(BMI),insulinlevel, diabetes pedigree function, skin thickness, and age. Each recordinthedatasetrepresentsapatient’smedicaldetails along with an outcome indicating whether the person is diabeticornot.

Thedatasetplaysacrucialroleintrainingthemodel,asthe quality and relevance of the data directly affect the predictionaccuracy.Asufficientnumberofsamplesareused toensurereliablemodelperformance.

6.2 Data Preprocessing

Data preprocessing is an essential step in machine learning,asrawdata often containsmissingvalues,noise, andinconsistencies.Inthisstage,thedatasetiscleanedand transformedtoimprovethequalityofinputdata.

Missing values in attributes such as insulin and skin thicknessarehandledusingappropriatetechniquessuchas replacing them with mean or median values. Data normalizationisappliedtoscalethefeaturesintoauniform range,whichhelpsimprovetheperformanceofthemodel.

Thedatasetisthendividedintotrainingandtestingsets.The trainingsetisusedtotrainthemodel,whilethetestingsetis used to evaluate its performance. This ensures that the modeliscapableofgeneralizingtonewandunseendata.

6.3 Feature Selection

Feature selection is performed to identify the most relevantattributesthatcontributetodiabetesprediction.By selecting important features and removing irrelevant or redundantones,themodelbecomesmoreefficientandless complex. This also helps in reducing overfitting and improvingpredictionaccuracy.

6.4 Model Implementation

TheXGBoost(ExtremeGradientBoosting)algorithmisused for classification due to its high accuracy, speed, and scalability.XGBoostisanensemblelearningtechniquethat builds multiple decision trees sequentially. Each new tree

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

focusesoncorrectingtheerrorsmadebytheprevioustrees, therebyimprovingoverallperformance.

XGBoost uses gradient boosting techniques along with regularizationtopreventoverfitting.Italsosupportsparallel processing, making it computationally efficient. The algorithmautomaticallyhandlesmissingdataandprovides featureimportancescores,whichhelpinunderstandingthe contributionofeachparameter.

The model is trained using the training dataset and optimizedusingappropriatehyperparameters.Oncetrained, itiscapableofmakingpredictionsonnewinputdata.

6.5 System Workflow

Theworkflowoftheproposedsystemfollowsastructured sequenceofsteps:

1. Input medical data such as glucose level, blood pressure,BMI,insulinlevel,andage.

2. Perform data preprocessing including cleaning, normalization,andhandlingmissingvalues.

3. Selectrelevantfeaturesformodeltraining.

4. Train the XGBoost model using the processed dataset.

5. Predict the diabetes outcome based on input parameters.

6. Displaytheresultasdiabeticornon-diabetic.

Thismethodologyensuresefficientdataprocessing,accurate prediction, and reliable performance, making the system suitableforreal-worldhealthcareapplications.

7. SYSTEM ARCHITECTURE

Thesystemarchitectureoftheproposeddiabetesdetection systemisdesignedtoprocessmedical data efficientlyand generateaccuratepredictions.Itconsistsofmultiplestages includingdatainput,preprocessing,featureselection,model training,andpredictionoutput.

In the first stage, patient data is collected in the form of medical parameters such as glucose level, blood pressure, body mass index (BMI), insulin level, and age. This data serves as the input to the system and forms the basis for prediction.

The next stage involves data preprocessing, where the collected data is cleaned and prepared for analysis. This includeshandlingmissingvalues,removinginconsistencies, and normalizing the data to ensure uniformity. Proper preprocessing improves the quality of the dataset and enhancestheperformanceofthemodel.

After preprocessing, feature selection is performed to identify the most relevant attributes that contribute to

diabetes prediction. Selecting important features helps in reducingmodelcomplexityandimprovingaccuracy.

Inthemodeltrainingstage,theprocesseddataisfedintothe XGBoostalgorithm.XGBoostbuildsmultipledecisiontreesin asequentialmanner,whereeachtreecorrectstheerrorsof thepreviousone.Thisresultsinastrongpredictivemodel withhighaccuracyandreducedoverfitting.

Oncethemodelistrained,itisusedtomakepredictionson newinputdata.Thesystemanalyzesthegivenparameters andclassifiestheresultaseitherdiabeticornon-diabetic.

Finally, the output is displayed in a user-friendly format, providing a clear and understandable result. This system architecture ensures efficient data processing, accurate prediction, and easy interpretation, making it suitable for real-worldhealthcareapplications.

Fig -1:SystemArchitectureofDiabetesDetectionSystem

Fig -1 illustrates the system architecture of the proposed diabetes detection system. The process begins with the collectionofadatasetcontainingmedicalparameters.The dataisthenpreprocessedtohandlemissingvalues,remove noise,andnormalizethefeatures.

Afterpreprocessing,featureanalysisisperformedtoidentify importantattributesthatinfluencediabetesprediction.The processed data is then used to train the XGBoost model, whichbuildsanefficientandaccuratepredictionmodel.

The trained model is used to predict whether a person is diabeticornon-diabeticbasedoninputparameters.Finally, performanceanalysisiscarriedoutusingevaluationmetrics suchasaccuracy,precision,recall,andF1-scoretoassessthe effectivenessofthemodel.

8. PERFORMANCE ANALYSIS

Theperformanceoftheproposeddiabetesdetectionsystem is evaluated using standard classification metrics such as accuracy, precision, recall, and F1-score. These metrics provide a comprehensive understanding of the model’s effectivenessinpredictingdiabeticandnon-diabeticcases.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

Accuracyrepresentstheoverallcorrectnessofthemodelby measuringtheratioofcorrectlypredictedinstancestothe totalnumberofinstances.Itgivesageneralideaofhowwell themodelperformsonthedataset.

Precision measures the proportion of correctly predicted positivecasesamongallpredictedpositivecases.Itindicates how reliable the model is when it predicts a patient as diabetic.Ahigherprecisionvaluemeansfewerfalsepositive predictions.

Recall,alsoknownassensitivity,measurestheabilityofthe model to correctly identify actual positive cases. It is an important metric in healthcare applications, as missing a diabetic case (false negative) can lead to serious consequences.

TheF1-scoreistheharmonicmeanofprecisionandrecall, providingabalancedmeasureofthemodel’sperformance.It is particularly useful when there is an imbalance in the dataset.

Inadditiontothesemetrics,aconfusionmatrixcanbeused to visualize the performance of the model. The confusion matrix consists of four components: true positives, true negatives,falsepositives,andfalsenegatives.Thishelpsin understanding how well the model distinguishes between diabeticandnon-diabeticcases.

The XGBoost model achieved high accuracy and demonstratedsuperiorperformancecomparedtotraditional machinelearningalgorithmssuchasLogisticRegressionand Decision Tree. Its ability to handle complex relationships, performregularization,andreduceoverfittingcontributesto itsimprovedaccuracy.

Overall, the performance analysis indicates that the proposed system is reliable and effective for diabetes prediction, making it suitable for real-world healthcare applications.

9. RESULTS AND DISCUSSION

Theproposeddiabetesdetectionsystemwasimplemented using Python and the XGBoost algorithm. The model was trainedandtestedonamedicaldatasetcontainingvarious healthparameters.

AsshowninTable-I,theXGBoostalgorithmachieveshigher accuracy compared to other machine learning algorithms such as Logistic Regression, Decision Tree, and Random Forest. This improvement is due to its ability to handle complexrelationshipsamongfeaturesandreduceoverfitting through ensemble learning techniques. In addition to accuracy,XGBoostalsoprovidesbetterprecision,recall,and F1-score,indicatingitsoveralleffectivenessandreliabilityin predictingdiabetes.TheseresultsdemonstratethatXGBoost is a suitable and efficient model for healthcare prediction systems.

Theperformanceofthemodelwasevaluatedusingaccuracy as the primary metric. The XGBoost model achieved high prediction accuracy compared to traditional machine learningalgorithmssuchasLogisticRegressionandDecision Tree. This is due to its ability to handle complex relationships and reduce overfitting through ensemble learning.

Theresultsindicatethatthe model caneffectivelyclassify patientsasdiabeticornon-diabeticbasedoninputfeatures. Thesystemprovidesfastandreliablepredictions,makingit suitableforreal-timehealthcareapplications.

DiabetesDetectionWebApplication

Fig. 2 illustrates the home page of the proposed diabetes detection web application. The interface is designed to be user-friendlyandintuitive,allowinguserstoeasilynavigate throughtheplatform.Thehomepageprovidesanoverview of the system’s functionality, highlighting its ability to perform early diabetes risk assessment using machine learningtechniques.

Theapplicationdisplayskeyfeaturessuchasaccuracyrate, numberofassessments,andavailability,whichenhanceuser confidenceinthesystem.Italsoincludesnavigationoptions likefeatures,workingprocess,andaboutsections,enabling userstounderstandthesystemindetail.

Theplatformallowsuserstoinputtheirhealth-relateddata such as age, body mass index (BMI), and glucose levels to assesstheirdiabetesrisk.Thecleanlayoutandresponsive

Table -1: ComparingMachineLearningAlgorithms
Fig -2:HomePageof

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

designensureaccessibilityacrossdifferentdevices,making itsuitableforreal-timeusage.

Overall, the home page serves as the entry point of the system,providingessentialinformationandguidingusersto utilizethediabetespredictionserviceefficiently.

Fig -3:DiabetesPredictionResult

Fig. 3 illustrates the prediction result generated by the proposed diabetes detection system based on the input medical parameters provided by the user. The system processes the input data, such as glucose level, blood pressure,bodymassindex(BMI),insulinlevel,andage,and feedsitintothetrainedXGBoostmodel.

The model analyzes the relationships between these parameters and evaluates the probability of the patient beingdiabetic.Basedonthisanalysis,thesystemclassifies theresultaseitherdiabetic ornon-diabetic.Theoutput is displayed clearly on the user interface, making it easy for userstounderstandtheirhealthstatus.

The prediction result is generated in real time, demonstratingtheefficiencyandspeedofthesystem.The use of the XGBoost algorithm ensures high accuracy and reliabilityinthepredictionprocess.Additionally,thesystem provides a simple and interactive output, which enhances userexperienceandaccessibility.

Thisresultinterfaceplaysacrucialroleinthesystem,asit directly communicates the outcome of the analysis to the user.Itcanassisthealthcareprofessionalsandindividualsin making informed decisions regarding further medical consultationandpreventivemeasures.

Overall,Fig.3highlightsthepracticalimplementationofthe proposed system and demonstrates its effectiveness in deliveringaccurateanduser-friendlydiabetespredictions.

Fig -4:DiabetesPredictionResult

Fig.4showsthepredictionresultgeneratedbythesystem basedontheinputmedicalparameters.

10. ADVANTAGES AND DISADVANTAGES

10.1 Advantages:

 HighAccuracy:

The system uses the XGBoost algorithm, which provides high prediction accuracy by capturing complexrelationshipsamongmedicalparameters.

 FastPrediction:

Themodelgeneratesresultsquickly,enablingrealtimediabetespredictionandmakingitsuitablefor practicalhealthcareapplications.

 ReducedHumanEffort:

Theautomatedsystemminimizesmanualanalysis byhealthcareprofessionals,savingtimeandeffort indiagnosis.

 ImprovedDecision-Making:

Thesystemassistsdoctorsandusersbyproviding reliablepredictions,helpinginearlydiagnosisand bettertreatmentplanning.

 EfficientHandlingofData:

XGBoost effectively handles missing values and noisy data, improving the overall performance of themodel.

 Scalability:

Thesystemcanbeeasilyextendedtoincludemore data and additional features without significant changestothemodel.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

 User-FriendlyInterface:

The web-based application provides an intuitive interface, allowing users to easily input data and obtainresults.

 EarlyDetection:

Thesystemenablesearlyidentificationofdiabetes risk, which helps in preventing severe health complications.

 Cost-EffectiveSolution:

It reduces the need for frequent medical tests by providing an initial prediction based on input parameters.

 Adaptability:

The model can be adapted for predicting other diseasesbymodifyingthedatasetandfeatures.

10.2 Limitations:

The model depends on the quality of the dataset. It may require large datasets for better performance. The system currentlyfocusesonlimitedparametersandcanbeexpanded further.

11. CONCLUSIONS

Thispaperpresentsamachinelearning-basedapproachfor diabetes detection using the XGBoost algorithm implemented in Python. The proposed system effectively analyzes important medical parameters such as glucose level,bloodpressure,bodymassindex(BMI),insulinlevel, andagetopredictwhetherapatientisdiabeticornot.

The use of XGBoost significantly improves prediction accuracyduetoitsabilitytohandlecomplexrelationships among features, perform regularization, and reduce overfitting. Compared to traditional machine learning algorithmssuchasLogistic Regression and Decision Tree, theXGBoostmodeldemonstratessuperiorperformancein termsofaccuracy,efficiency,andreliability.

Theexperimentalresultsconfirmthattheproposedsystem can accurately classify patients and provide consistent predictions. This makes the system a valuable tool for assisting healthcare professionals in early diagnosis and decision-making. By enabling early detection, the system helpsinreducingtheriskofseverecomplicationsassociated withdiabetes.

Furthermore,theintegrationofmachinelearningtechniques intohealthcaresystemscontributestothedevelopmentof intelligent and automated diagnostic tools. The proposed

approachissimple,scalable,andcanbeextendedtoother diseasepredictionsystems.

Infuturework,themodelcanbeenhancedbyusinglarger and more diverse datasets to further improve accuracy. Advanced techniques such as deep learning can also be explored.Additionally,thesystemcanbeintegratedintoa real-time web or mobile application, making it more accessible and useful for both patients and healthcare providers.

Infuture,themodelcanbeimprovedbyusingmoredata from different sources to increase accuracy. Advanced methods like deep learning can also be used to get better results.Thesystemcanbedevelopedintoawebormobile applicationsothatitcanbeusedinrealtime.Moremedical features can also be added to improve prediction performance.

11. CONCLUSIONS

12.1 Additional Features:

More medical parameters can be included to improve the accuracyandeffectivenessoftheprediction.

12.2 Integration with Healthcare Systems:

The model can be integrated with hospital or healthcare systemstoassistdoctorsindiagnosisandmonitoring.

12.3 Continuous Model Improvement:

The model can be updated regularly with new data to improveitsperformanceovertime.

12.4 User-Friendly Enhancements:

Theinterfacecanbefurtherimprovedtomakethesystem moreinteractiveandeasiertouseforallusers.

12.5 Multi-Disease Prediction:

Thesystemcanbeextendedtopredictotherdiseasesusing similarmachinelearningtechniques.

ACKNOWLEDGEMENT

Theauthorswouldliketoexpresstheirsinceregratitude to Dr. N. Sri Hari for his valuable guidance, continuous support,andencouragementthroughoutthedevelopmentof thisproject.Hisinsightsandsuggestionsgreatlycontributed tothesuccessfulcompletionofthiswork.

We also thank the Head of the Department, faculty members, and the management of Vasireddy Venkatadri InstituteofTechnologyforprovidingthenecessaryresources andsupport.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

REFERENCES

[1]InternationalDiabetesFederation,"IDFDiabetesAtlas," 2021.

[2] UCI Machine Learning Repository, "Pima Indians DiabetesDataset,"Available:https://archive.ics.uci.edu.

[3] T. Chen and C. Guestrin, "XGBoost: A Scalable Tree BoostingSystem,"inProc.22ndACMSIGKDDInt.Conf. KnowledgeDiscoveryandDataMining,2016.

[4] F. Pedregosa et al., "Scikit-learn: Machine Learning in Python,"JournalofMachineLearningResearch,vol.12, pp.2825–2830,2011.

[5]WorldHealthOrganization,"GlobalReportonDiabetes," 2021.

2026, IRJET | Impact Factor value: 8.315 | ISO 9001:2008

Turn static files into dynamic content formats.

Create a flipbook
Diabetes Detection Using Machine Learning by IRJET Journal - Issuu