
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Vikrant Singh
¹Department of Artificial Intelligence and Data Science Pune Institute of Computer Technology Pune, Maharashtra, India
Abstract - This study presents a comparative analysis of Multinomial Naive Bayes and Logistic Regression for spam classification using the SMS Spam Collection Dataset. The primary objective is to evaluate the effect of Laplace smoothing on Naïve Bayes performance and analyse the impactoftheindependenceassumption.Experimentalresults showthatoptimalsmoothing(α = 0.1) significantly improves classification performance. Naïve Bayes achieved 98.39% accuracyand92%recall,outperformingLogisticRegression, which achieved 95.52% accuracy and 68.67% recall. The results demonstrate that proper smoothing enhances generative classifier performance and that Naive Bayes is particularly effective for high-dimensional sparse text classification tasks.
Key Words: Spam Detection, Naive Bayes, Logistic Regression, Laplace Smoothing, Machine Learning, Text Classification
1. INTRODUCTION
Spam detection is a fundamental problem in text classification and machine learning. Among various classification algorithms, Multinomial Naive Bayes and Logistic Regression are widely used due to their computational efficiency and effectiveness. Naive Bayes assumesconditionalindependencebetweenfeatures,which isoftenviolatedinreal-worldtextdata.LogisticRegression, ontheotherhand,isadiscriminativemodelthatdoesnot rely on independence assumptions. This research aims to experimentally evaluate the sensitivity of the Multinomial NaiveBayesclassifiertoLaplacesmoothingandtocompare itsperformancewithLogisticRegressioninspamdetection tasks. Specifically, the study investigates how varying the smoothing parameter influences classification accuracy, precision,recall,andF1score.Theobjectiveistoanalyzethe practical implications of the independence assumption in NaiveBayesandtodetermineoptimalsmoothingconditions forimprovedclassificationperformance.
Themaincontributionsofthisstudyareasfollows:
•ExperimentalsensitivityanalysisofLaplacesmoothing ParameteronMultinomialNaiveBayesperformance.
•ComparativeperformanceevaluationofNaiveBayesand LogisticRegressionusingTF-IDFfeatures.
•Empirical demonstration that moderate smoothing significantlyimprovesclassificationperformance.
•AnalysisshowingthatNaiveBayesachieveshigherrecall inspamdetection,makingitmoresuitableforreal-world spamfilteringapplications.
NaiveBayeshasbeenwidelyusedinspamdetectiondue to its simplicity, efficiency, and relatively strong performance in high-dimensional text classification problems.LogisticRegressionhasalsodemonstratedrobust performanceinclassificationtasksduetoitsdiscriminative nature and ability to model feature relationships directly. PreviousresearchhasshownthatLaplacesmoothingplays an important role in improving probability estimation in Naïve Bayes classifiers by preventing zero probabilities. However, excessive smoothing can negatively impact classificationaccuracy.WhilebothNaiveBayesandLogistic Regression have been extensively studied, experimental analysisfocusingonsmoothingsensitivityandindependence assumptioneffectsremainslimited.NaiveBayesisawidely used probabilistic classifier in text classification tasks [1]. LogisticRegressionisapowerfuldiscriminativemodelfor classification problems [2], [3]. The experiments in this studyusetheSMSSpamCollectionDataset[4].
A. Naive Bayes Classification
NaiveBayesisaprobabilisticclassifierbasedonBayes’ theorem:
P(Y|X)=(P(X|Y)×P(Y))/P(X) (1)
whereYrepresentstheclasslabelandXrepresentsthe featurevector.
Laplace smoothing modifies probability estimation as follows:
P(xi|Y)=(count(xi,Y)+α)/(count(Y)+αn)(2)
whereαisthesmoothingparameterandnisthenumber offeatures.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
LogisticRegressionisadiscriminativeclassifierthatmodels theprobabilityofaclasslabelusingthesigmoidfunction:
P(y=1|x)=1/(1+e^−(wTx+b)) (3)
where:
•xisthefeaturevectorrepresentingtheinputdocument,
•wistheweightvectorthatrepresentstheimportanceof eachfeature,
•bisthebiasterm,
• w^T x represents the dot product between the weight vectorandfeaturevector.
Theweightvectorwandbiastermbarelearnedduring trainingtominimizeclassificationerror.UnlikeNaiveBayes, LogisticRegressiondoesnotassumeindependencebetween features and directly learns decision boundaries from the data.
Thedatasetwasdividedintotrainingandtestingsetsusing an80:20split.TextmessagesweretransformedusingTF-IDF Vectorization,whichassignsweightstowordsbasedontheir importance across documents. This helps reduce the influence of commonly occurring words and improves classificationperformance.
TheMultinomialNaiveBayesclassifierwastrainedusing multiplesmoothingparametervaluesrangingfrom0.001to 100. Logistic Regression was trained using default parameters.
AllexperimentswereconductedusingPython3.11and scikit-learnonastandardcomputingenvironment.
5.1 Dataset
The spam dataset consists of labelled text messages categorized as spam or non-spam. The text data was preprocessedusingstandardtechniquesincludingtokenization andvectorization.
The experiments were conducted using the SMS Spam Collection Dataset, which contains 5,572 labeled SMS messages categorized as spam or ham (non-spam). The dataset consists of 747 spam messages and 4,825 ham messages and is widely used as a benchmark for spam classificationresearch.
5.2 Feature Extraction
TextmessagesweretransformedusingTF-IDFVectorization. TF-IDFassignsweightstowordsbasedontheirfrequency within a document and their inverse frequency across all documents. This representation improves classification performance by emphasizing informative words and reducingtheimpactofcommonwords.
•MultinomialNaiveBayes
•LogisticRegression.
5.4
Thesmoothingparameterαwasvariedacrossmultiple valuesrangingfrom0.001to100toanalyzeitsimpacton classificationperformance.
5.5
Thefollowingevaluationmetricswereused:
•Accuracy
•Precision
•Recall
•F1Score
Themodelswereimplementedusingthescikit-learnmachine learning library in Python. TF-IDF vectorization was performedusingTfidfVectorizer.MultinomialNaiveBayes was implemented using MultinomialNB, and Logistic RegressionwasimplementedusingLogisticRegression.
Modelperformancewasevaluatedonthetestdatasetusing accuracy,precision,recall,andF1scoremetrics.
Table -1: PERFORMANCECOMPARISONOFMODELS
Theoptimalsmoothingparameterwasfoundtobeα=0.1, whichproducedthehighestF1score.

Chart -1:EffectofLaplacesmoothingonAccuracy

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072



The experimental results reveal several important observations regarding classifier behavior and smoothing sensitivity.
First, Multinomial Naive Bayes demonstrated strong performancewhenmoderatesmoothingvalueswereused. Optimal smoothing improved probability estimation and prevented zero-probability issues without excessively distortingfeaturelikelihoods.However,excessivesmoothing resultedinperformancedegradation.Largesmoothingvalues reducedtheinfluenceofdiscriminativewords,causingthe modeltoassignoverlyuniformprobabilitiesacrossfeatures.
Second, Logistic Regression exhibited lower recall compared to Naive Bayes. This indicates that Logistic Regression failed to identify a significant portion of spam messages. This behavior may be attributed to its linear decisionboundaryandsensitivitytoclassimbalance.Naive Bayes achieved higher recall, making it more effective in spamdetectionscenarioswhereidentifyingspammessagesis morecriticalthanavoidingfalsepositives.Thesuperiorrecall achieved by Naive Bayes may also be attributed to its probabilistic nature,whichallowsit to bettercapture rare wordoccurrencescommonlyassociatedwithspammessages. LogisticRegression,asadiscriminativemodel,maybemore conservativeinclassification,resultinginlowerrecall. The dataset exhibits class imbalance, with spam messages representing only 13.41% of the total dataset. Class imbalance can significantly affect classifier performance, particularlyrecall.GenerativeclassifierssuchasNaiveBayes may handle such imbalance more effectively due to probabilistic modeling of feature distributions, which may explainthesuperiorrecallobservedinthisstudy.
These results suggest that generative models such asNaive Bayes can outperform discriminative models in highdimensionalsparsetextclassificationtasks,particularly whenpropersmoothingisapplied.Thefindingsconfirmthat Laplace smoothing plays a crucial role in improving Naive Bayes performance and that careful parameter tuning is essentialforoptimalclassification.
Thisstudywasconductedusingasinglespamdatasetand default Logistic Regression parameters. Further research may include cross-validation, hyperparameter tuning for LogisticRegression,andevaluationonadditionaldatasetsto improvegeneralizability.
Thisstudyconductedadetailedexperimentalanalysisof LaplacesmoothingeffectsonMultinomialNaiveBayesand compareditsperformancewithLogisticRegressionforspam classification.
The results demonstrate that Laplace smoothing significantlyinfluencesNaiveBayesperformance.Moderate smoothing values improve probability estimation and enhanceclassificationaccuracy,whileexcessivesmoothing reducesmodeleffectiveness.
Furthermore,NaiveBayesachievedhigherrecallcompared toLogisticRegression,makingitmoresuitableforspam detectiontaskswhereidentifyingspammessagesiscritical. Thesefindingsprovideempiricalevidencethatproper smoothingparameterselectionisessentialforoptimalNaive

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Bayes performance and confirm its effectiveness in highdimensionaltextclassificationproblems. Futureworkmayextendthisresearchtolargerdatasets, cross-domainclassificationtasks,anddeeplearning-based approaches.
ACKNOWLEDGMENT
Theauthorwouldliketothanktheacademicmentorsand peerswhoprovidedguidanceduringthisresearch.
[1] A. McCallum and K. Nigam,“A comparison of event models for naive bayes text classification,” in AAAI WorkshoponLearningforTextCategorization,1998.
[2] C.M.Bishop,PatternRecognitionandMachineLearning Springer,2006.
[3] T.Hastie,R.Tibshirani,andJ.Friedman,TheElements ofStatisticalLearning.Springer,2009.
[4] T.A.Almeida,J.M.G.Hidalgo,andA.Yamakami,“Sms spamcollectiondataset,”UCIMachineLearning Repository,2011.