
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Sayera Yousufa1 , Jakkula Jayanthi2 , Racha Siri Varsha3 , Omkar Radhika4
123Computer Science Engineering Student, JNTUH, Telangana, India
4Omkar Radhika, Guest lecturer, Dept. of Computer Science and Engineering, JNTUH,Telangana,India ***
Abstract - Machine learning models often achieve high predictive accuracy but lack interpretability, limiting their applicability in critical domains. This paper proposes an Explainability-Guided Feature Selection Model (XGFS-ML), which integrates SHAP-based feature importance with supervised learning algorithms. A dynamic threshold-based featureselectionmechanismisintroducedtoadaptivelyselect relevant features based on their contribution to model predictions. The framework is evaluated using two datasets: the Breast Cancer Wisconsin dataset and a Telecom Churn dataset, ensuring improved generalization across domains. Performance evaluation is conducted using 5-fold crossvalidation, reporting mean accuracy and standard deviation for robustness. Experimental results demonstrate that the proposed approach reduces feature dimensionality while maintaining or improving classification performance, particularly enhancing results in high-dimensional datasets. The framework provides a balanced solution for building accurate, interpretable, and efficient machine learning systems.
Key Words: Machine Learning, Feature Selection, XGFSML, Random Forest, Logistic Regression, Explainable AI, Classification.
MachineLearning(ML)hasbecomeanessentialtechnology for solving complex real-world problems across multiple domains including healthcare, finance, agriculture, cybersecurity,andintelligentautomation[1],[9].Withthe rapidgrowthofdataavailability,machinelearningmodels arecapableoflearningcomplexpatternsandrelationships betweeninputvariablesandtargetoutputs.
However,manyhigh-performancemachinelearningmodels operateasblack-boxsystems,makingtheirdecision-making processdifficulttointerpret[10].Thislackoftransparency reducesusertrust,especiallyincriticalapplicationswhere understandingmodelbehaviorisessential.
Featureselectionisacrucialstepinmachinelearningthat aimstoidentifythemostrelevantfeaturesfromadataset.By eliminating irrelevant or redundant features, feature selectionimprovesmodelaccuracy,reducescomputational complexity,and enhances generalization performance [4]. However, traditional feature selection techniques such as filter,wrapper,andembeddedmethodsprimarilyfocuson improvingperformancewhileoftenignoringinterpretability.
To address this limitation, this paper proposes an Explainability-Guided Feature Selection Model (XGFS-ML) that integrates SHAP-based feature importance with supervised machine learning algorithms. A dynamic threshold-based mechanism is introduced to adaptively select important features based on their contribution to modelpredictions.
Furthermore,toensurerobustnessandgeneralization,the proposed framework is evaluated on multiple datasets, including a breast cancer dataset [6] and a telecom churn dataset.Model performance isassessedusingstratified 5foldcross-validation,withresultsreportedasmeanaccuracy andstandarddeviation.
Theproposedapproachenhancesbothinterpretabilityand performance, providing an effective solution for building transparentandefficientmachinelearningmodels.
Many machine learning models achieve high predictive accuracybutlackinterpretability.Insensitivedomainssuch ashealthcareandfinance,itisimportanttounderstandhow inputfeaturesinfluencemodeldecisions.Traditionalfeature selectionapproachesdonotincorporateexplainabilityinto the selection process, which limits their usefulness in transparentdecision-makingsystems.
Themainobjectivesofthisresearchare:
1. To develop an explainability-guided feature selectionframework.
2. TointegrateSHAP-basedfeatureimportancewith machinelearningalgorithms.
3. Toreducefeaturedimensionalitywhilemaintaining orimprovingpredictionaccuracy.
4. To improve transparency and interpretability of machinelearningmodels.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Theproposedframeworkconsistsofthefollowingstages:
Dataset Collection → Data Preprocessing → SHAP-Based Feature Selection → Model Training → Performance Evaluation→ModelExplanation
Initially,datasetsarecollectedandpreprocessedtohandle missingvalues,encodecategoricalfeatures,andstandardize thedata.FeatureselectionisthenperformedusingSHAPbasedimportancescores.Adynamicthresholdmechanismis appliedtoselectthemostrelevantfeaturesbasedontheir contributiontomodelpredictions.
The selected features are used to train multiple machine learning models, including Logistic Regression, Random Forest,andXGBoost.Modelperformanceisevaluatedusing stratified 5-fold cross-validation, and results are reported using mean accuracy and standard deviation to ensure robustness.
SHAP explainability is further used to analyze feature contributions,providinginsightsintomodelbehavior. The pipeline enhances interpretability, reduces feature dimensionality,andmaintainsstrongperformance,makingit suitableforreal-worlduse.


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Two datasets were used in this study to evaluate the generalizationcapabilityoftheproposedmodel.
1.BreastCancerWisconsinDataset:
Source:UCIMachineLearningRepository[6]
Instances:569
Features:30numericalfeatures
Task:Binaryclassification(0=Benign/1=Malignant)
2.TelecomChurnDataset:
Source:Publictelecomcustomerdataset(Kaggle)
Features: Customer-related attributes such as tenure, monthly charges, contract type, payment method, and internetservices
Task:Predictcustomerchurn(Yes/No)
Data preprocessing is an essential step in preparing the dataset for machine learning algorithms. The following preprocessingtechniqueswereapplied:
Handlingmissingvaluesbyreplacingthemwithappropriate statisticalmeasures
Encoding categorical variables using one-hot encoding Featurescalingusingstandardization
Unliketraditionalapproaches,thedatasetwasnotsplitusing a fixed train-test ratio. Instead, stratified 5-fold crossvalidationwasusedduringmodeltrainingandevaluationto ensurerobustandunbiasedperformanceestimation.
Thesepreprocessingstepsensurethatthemachinelearning models can effectively learn patterns while maintaining consistencyandreliabilityinresults[9].
2.4
FeatureselectionisperformedusingSHAP(ShapleyAdditive Explanations)[5].Adynamicthresholdiscomputedbased on the mean SHAP importance values, and features exceeding this threshold are selected. This adaptive approachensuresthatfeatureselectionisdata-drivenand flexible across different datasets. A minimum feature constraint is also applied to prevent excessive dimensionalityreduction.
The following machine learning algorithms are used for performanceevaluation:
LogisticRegression-
Logistic Regression is a statistical model used for binary classification tasks. It estimates the probability that an instancebelongstoaspecificclass[1].
RandomForest-
Random Forest is an ensemble learning algorithm that constructs multiple decision trees and combines their predictionstoimproveaccuracyandreduceoverfitting[2]
XGBoost-
XGBoost is an advanced gradient boosting algorithm designedforefficiencyandhighperformanceinlarge-scale machinelearningtasks[3]
The proposed XGFS-ML framework was evaluated using multiple machine learning models, including Logistic Regression,RandomForest,andXGBoost,ontwodatasets: Breast Cancer and Telecom Churn. The experiments were conducted to analyze the effectiveness of explainabilityguided feature selection in improving classification performancewhilereducingfeaturedimensionality.
Themodelswereinitiallytrainedusingthefullfeatureset. SHAP-based feature importance was then used to rank features,andadynamicthresholdmechanismwasappliedto selectthemostrelevantfeatures.Themodelswereretrained usingtheoptimizedfeaturesubset Performanceevaluation wascarriedoutusingstratified5-foldcross-validation,and resultsarereportedasmeanaccuracyalongwithstandard deviationto ensure robustnessand reliability.The results demonstratethat the proposedapproachimprovesmodel performancewhileenhancinginterpretabilityandreducing complexity.
Thefollowingevaluationmetricwasusedtomeasuremodel performance:
•Accuracy–proportionofcorrectlyclassifiedsamples.
To ensure reliable evaluation, stratified 5-fold crossvalidation was used, and results are reported as mean accuracy±standarddeviation.Thisprovidesamorerobust assessmentcomparedtoasingletrain-testsplit.
Theclassificationperformanceofdifferentmachinelearning modelsanddatasetsissummarizedinTable1and2

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Table -1: BreastCancerDataset
MODEL BEFOREXGFS AFTERXGFS
LOGISTIC REGRESSION 0.9737±0.0166
±0.0131 RANDOM FOREST
±0.0123
±0.0153 XGBOOST
±0.0163
Table -2: TelecomDataset
±0.0151
MODEL BEFOREXGFS AFTERXGFS
LOGISTIC REGRESSION
±0.0162
RANDOM FOREST 0.7346±0.0002
XGBOOST
±0.0110
±0.0117
±0.0070
±0.0113
The results show that the proposed XGFS framework maintainsstableperformanceonthebreastcancerdataset while improving performance in the telecom dataset. Overall, the approach enhances efficiency and interpretability across different datasets, demonstrating goodgeneralization.
Theaccuracycomparisongraphillustratestheperformance ofthemachinelearningmodelsonthebreastcancerdataset before and after applying the proposed XGFS feature selectionapproach.
The results indicate that Logistic Regression achieved an accuracy of 0.9737 before XGFS and improved to 0.9789 after feature selection, showing a noticeable performance gain. Random Forest achieved 0.9561 before XGFS and 0.9596afterXGFS,indicatingaslightimprovement.XGBoost achieved 0.9684 before XGFS and 0.9631 after XGFS, showing a minor decrease while still maintaining competitiveperformance.
Overall, the results demonstrate that the proposed XGFS approach effectively reducesfeature dimensionality while maintaining or slightly improving model performance. Logistic Regression showed the most improvement, indicating that feature selection helped enhance model efficiencywithoutcompromisingaccuracy.

Tounderstandhowindividualfeaturesinfluenceprediction outcomes, SHAP (Shapley Additive Explanations) [5] was usedtocomputefeatureimportancescores.
The SHAP summary plot illustrates the impact of each feature on model predictions, where features are ranked based on their importance and the spread indicates their influence. Features such as area worst, concave_points_worst, concave_points_mean, perimeter worst,andconcavityworstareobservedtohavethehighest impactonclassification.
Higherfeaturevalues(showninred)generallycontribute positivelytowardtheprediction,whilelowervalues(shown inblue)haveanegativeimpact.Thisindicateshowfeature variationsaffectthemodeloutput.
AdynamicthresholdbasedonmeanSHAPimportancewas appliedtoselectthemostrelevantfeatures,ensuringadatadrivenandadaptivefeatureselectionprocess.Thisimproves interpretabilitywhilemaintainingmodelperformance.


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
The feature importance graph ranks the most influential features based on their SHAP importance values. It highlights the relative contribution of each feature to the predictionoutput.
The analysis shows that features such as area worst, concave_points_worst, concave_points_mean, perimeter worst, and concavity worst have the highest importance, indicating their strong influence on classification. These features consistently appear at the top of the ranking and playakeyroleinmodeldecision-making.
By applying SHAP-based feature ranking along with a dynamic threshold, the proposed XGFS-ML framework effectivelyremoveslesssignificantfeatureswhileretaining the most informative ones. This leads to improved interpretability and reduced dimensionality while maintainingstrongclassificationperformance.

4. CONCLUSIONS
This paper proposed an Explainability-Guided Feature Selectionframework(XGFS-ML)thatintegratesSHAP-based feature importance with machinelearning modelsusing a dynamic threshold-based selection mechanism. The approach identifies the most influential features and removes redundant attributes, reducing dataset dimensionalitywhilemaintainingmodelperformance[4]
ExperimentalevaluationonmultipledatasetsusingLogistic Regression,RandomForest,andXGBoostdemonstratesthat the proposed method maintains or improves predictive accuracy while enhancing model efficiency. Notable improvements were observed in the telecom dataset, highlighting the effectiveness of the approach in handling high-dimensionalreal-worlddata.
SHAP analysis further provided insights into feature contributions,improvingtransparencyandinterpretability of the models. Overall, the proposed framework offers a
reliable and efficient solution for building interpretable machinelearningsystems.
Futureworkwillfocusonapplyingtheframeworktolargescaledatasetsandexploringadvancedlearningtechniques.
[1]T.Mitchell,“MachineLearning,”McGraw-Hill,1997.
[2]L.Breiman,“RandomForests,”MachineLearningJournal, vol.45,no.1,pp.5–32,2001.
[3] T. Chen and C. Guestrin, “XGBoost: A Scalable Tree Boosting System,” Proceedings of the 22nd ACM SIGKDD InternationalConferenceonKnowledgeDiscoveryandData Mining,2016.
[4]I.GuyonandA.Elisseeff,“AnIntroductiontoVariableand Feature Selection,” Journal of MachineLearning Research, vol.3,pp.1157–1182,2003.
[5] S. Lundberg and S. Lee, “A Unified Approach to Interpreting Model Predictions,” Advances in Neural InformationProcessingSystems(NeurIPS),2017.
[6]D.DuaandC.Graff,“UCIMachineLearningRepository,” UniversityofCalifornia,Irvine,2017.
[7]J.Han,M.Kamber,andJ.Pei,“DataMining:Conceptsand Techniques,”MorganKaufmannPublishers,2011.
[8]A.Géron,“Hands-OnMachineLearningwithScikit-Learn, KerasandTensorFlow,”O’ReillyMedia,2019.
[9] P. Domingos, “A Few Useful Things to Know About MachineLearning,”CommunicationsoftheACM,vol.55,no. 10,pp.78–87,2012