Skip to main content

COMPARATIVE STUDY OF COST-SENSITIVE LEARNING AND DATA- LEVEL BALANCING STRATEGIES FOR HIGHLY SKEWED

Page 1


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | A pr 2026 www.irjet.net p-ISSN: 2395-0072

COMPARATIVE STUDY OF COST-SENSITIVE LEARNING AND DATALEVEL BALANCING STRATEGIES FOR HIGHLY SKEWED SPAM EMAIL DATASETS

1Master of Technology, Computer Science and Engineering, Lucknow Institute of Technology, Lucknow, India

2Assistant Professor, Department of Computer Science and Engineering, Lucknow Institute of Technology, Lucknow, India ***

Abstract - Spam email detection has become a critical research area due to the increasing volume of unsolicited emails and their role in phishing and malware delivery. A major challenge inspamclassificationisthepresenceofhighly imbalanceddatasets, where legitimate emailsdominatespam samples, leading to biased learning behavior in traditional machine learning classifiers. This study presents a comparative analysis of data-level balancing strategies and cost-sensitive learning methods for spam classification under extreme skew conditions. Three benchmark datasets, namely SpamAssassin, Ling-Spam, and CSDMC2010, are utilized and further modified to generate highly skewed variants with imbalance ratios of 90:10, 95:5, and 98:2. Text preprocessing is applied, followed byTF-IDF feature extraction.Experiments are performed usingMultinomial NaïveBayes,SupportVector Machine, and Random Forest classifiers. Data-level strategies such as RandomUndersampling(RUS),RandomOversampling (ROS), and SMOTE are compared against cost-sensitive techniques including class-weighted SVM, MetaCost, and decision threshold tuning. Results demonstrate that SMOTE combined with Random Forest yields strong improvements in recall and F2-score, whereas threshold-tuned SVM achieves better precision and reduces false positives. The findings confirm that cost-sensitive learningoffers stableperformance under severe skewness, while resampling methods provide significant recall gains but may increase computational overhead. This study contributes a structured evaluation framework for selecting imbalance-handling techniques in real-world spam filtering systems.

Key Words: Spam Detection, Class Imbalance, CostSensitive Learning, SMOTE, Threshold Tuning, TF-IDF, Imbalanced Classification

1. INTRODUCTION

Spamemaildetectionremainsacrucialresearchdomainin cybersecurityandmachinelearningduetotherapidgrowth ofelectroniccommunicationsystems.Emailiswidelyused forprofessionalcommunication,academiccorrespondence, e-commercetransactions,andofficialnotifications,makingit an essential component of modern digital infrastructure. However,thispopularityalsoattractsmaliciousactorswho exploit email as a low-cost and high-reach medium to distributeunsolicitedandharmfulmessages.Spamemails

notonlycauseinconveniencebyoverwhelminginboxesbut alsocreateseveresecuritythreatsbyservingasacarrierfor phishingattacks,identitytheft,andmalwaredistribution.As spamtechniquescontinuetoevolve,thereisanincreasing needforrobustandintelligentspamfilteringmechanisms thatcanadapttochangingattackstrategiesandeffectively detectmaliciouscontent.

1.1 Background

Email communication has become a foundational tool for global interaction, supporting both organizational operationsandpersonalcommunication.Withtheexpansion of internet connectivity and cloud-based messaging platforms,emailhasachievedwidespreadadoptionbecause of its efficiency, asynchronous communication model, and abilitytotransferlarge-scaleinformationinstantly.Despite its advantages, email is frequently targeted by spammers due to its open architecture and massive user base. Spam emailsrepresentunsolicitedbulkmessagesthataretypically sent for advertising, scams, or malicious purposes. The growingvolumeofspamhasmadeautomatedspamfiltering anessentialrequirementforensuringsecureandproductive emailcommunication(HeandGarcia,2009).

1.1.1 Importance of Email Communication and Spam Threats

Email is considered one of the most reliable digital communicationtoolsduetoitsformalnature,accessibility, and ability to support secure document exchange. Governments, businesses, educational institutions, and individuals rely heavily on email for official notifications, legalcommunication,andfinancialtransactions.However, spamemailshavebecomeamajorchallenge,contributingto wasted bandwidth, reduced productivity, and increased cybersecurity risks. Many spam messages are designed to manipulate recipients into revealing confidential information,downloadingmaliciousattachments,orvisiting fraudulentwebsites.Consequently,spamfilteringisnotonly amatterofconveniencebutalsoasignificantcybersecurity requirement in modern information systems (Metsis, AndroutsopoulosandPaliouras,2006).

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | A pr 2026 www.irjet.net p-ISSN: 2395-0072

1.2 Problem of Highly Skewed Spam Datasets

Amajortechnical challengeinspamemail classificationis thehighlyimbalancednatureofreal-worlddatasets.Inmost practicalenvironments,legitimateemails(ham)significantly outnumber spam emails. This imbalance creates biased learningconditionswheremachinelearningmodelstendto prioritize the majority class. As a result, classifiers may achievehighoverallaccuracybutfailtodetectspamemails effectively. This issue becomes even more severe in corporate and enterprise environments, where spam may represent only a very small fraction of the overall email traffic.Therefore,handlingdatasetskewnessisessentialfor designingeffectivespamdetectionsystems(Japkowiczand Shah,2011).

1.2.1 Imbalance Ratios such as 95:5 and 98:2

Datasetimbalanceisoftenexpressedasaratioofmajority classsamplestominorityclasssamples.Inspamdetection,a datasetwitha95:5ratiomeansthat95%ofemailsareham whileonly5%arespam.Similarly,a98:2ratiorepresentsan extremescenariowherespamemailsconstituteonly2%of thetotaldataset.Suchextremeskewnesscausesclassifiers tobecomeoverlyconservativeinpredictingspambecause predicting ham most of the time results in fewer classificationerrors overall.Consequently, the model may fail to identify rare spam emails, making it ineffective for real-worlddeployment.Theseimbalanceconditionsrequire specializedstrategiessuchasresamplingmethodsorcostsensitive learning approaches to improve spam detection performance(Chawlaetal.,2002).

1.3 Cost Asymmetry in Spam Filtering

Spamfilteringisnotonlyaffectedbyclassimbalancebutalso by cost asymmetry, meaning that different types of misclassificationerrorshaveunequalconsequences.Afalse positive occurs when a legitimate email is incorrectly classifiedasspam,whileafalsenegativeoccurswhenaspam email is incorrectly classified as legitimate. These errors havedifferentimpactsonusersandorganizations.Inrealworldemailsystems,falsepositivesaregenerallyconsidered more critical because they may lead to missed business opportunities,lossoflegaldocuments,orfailuretoreceive important personal communication. Therefore, spam filteringsystemsmustbedesignedtominimizesuchcostly errors(DrummondandHolte,2006).

1.3.1 False Positives vs False Negatives Impact

Falsepositivescanhaveseriousconsequences,particularly incorporateandinstitutionalenvironments.Ifanimportant emailsuchasajoboffer,legalnotice,orfinancialtransaction confirmation is incorrectly blocked, it may result in significant harm. On the other hand, false negatives allow spammessagestoreachtheinbox,causingannoyanceand potential exposure to phishing or malware. While both

errors are undesirable, the cost of false positives is often higher because legitimate communication may be permanentlylost.Thismakesspamfilteringfundamentally different from many other classification problems, as the objective is not only high accuracy but also balanced decision-makingbasedonreal-worldconsequences(Powers, 2011).

1.4 Research Objectives

The primary objective of this research is to conduct a systematiccomparisonofdata-levelbalancingstrategiesand cost-sensitivelearningtechniquesforhighlyskewedspam emaildatasets.Thestudyaimstodeterminewhichapproach provides the most reliable performance under increasing skewness while maintaining acceptable computational efficiency. Since spam detection systems must operate in real-time and handle large-scale datasets, identifying a method that balances accuracy, recall, precision, and computationalcostisessential.

1.4.1 Comparative Analysis of Data-Level vs CostSensitive Approaches

Thisresearchevaluatesdata-levelmethodssuchasRandom Undersampling,RandomOversampling,andSMOTE,which modify the dataset distribution to improve minority class learning.Thesearecomparedwithcost-sensitivemethods such as class-weighted learning, MetaCost, and threshold tuning, which modify the classifier’s decision-making process.Thegoalistoquantifyhoweachstrategyimpacts spam detection performance, particularly under severe imbalanceratios(Chawlaetal.,2002).

1.4.2

Identifying the Most Robust Strategy Under Increasing Skewness

A major objective of this research is to identify which technique remains stable when imbalance increases from moderate levels to extreme skewness. This involves evaluating classifier performance across multiple dataset configurations such as 90:10, 95:5, and 98:2. The study focuses on determining whether data-level balancing methods remain effective or whether cost-sensitive strategiesprovidebetterstabilityunderextremeconditions (SaitoandRehmsmeier,2015).

2. RELATED WORK

Spam filtering has been widely studied as a classical text classification problem in machine learning and natural languageprocessing.Overthepasttwodecades,researchers have proposed various statistical and computational techniquestoimprovespamdetectionaccuracy,reducefalse positives, and enhance model robustness against evolving spamstrategies.However,thepresenceofhighlyimbalanced datasets remains a major challenge, particularly because real-worldemailenvironmentsoftencontainfarfewerspam

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | A pr 2026 www.irjet.net p-ISSN: 2395-0072

messages compared to legitimate emails. Consequently, modern research increasingly focuses on imbalancehandling strategies such as resampling, synthetic data generation, and cost-sensitive learning approaches. This section reviews existing literature on spam filtering algorithms,classimbalancechallenges,balancingstrategies, cost-sensitivemethods,andevaluationmetricsrelevantto highlyskewedspamdatasets.

2.1 Spam Filtering and Traditional ML Approaches

Spamemaildetectionhastraditionallybeenaddressedusing supervised machine learning models trained on labeled email datasets. Early machine learning-based spam filters replaced rule-based approaches by learning statistical patterns directly from email content. Among these, Naïve Bayesclassifiersbecameoneoftheearliestandmostwidely adopted methods due to their simplicity, probabilistic foundation,andstrongperformanceinhigh-dimensionaltext spaces. The success of Naïve Bayes in spam filtering is largelyattributedtoitsabilitytoefficientlycomputeclass probabilitiesevenwhenthenumberoffeaturesisextremely large,suchaswhenBag-of-WordsorTF-IDFrepresentations areused(Metsis,AndroutsopoulosandPaliouras,2006).

2.2

Class Imbalance in Text Classification

Classimbalancereferstoadatasetconditionwhereoneclass containssignificantlymoresamplesthantheother.Inspam filtering, legitimate emails (ham) typically dominate the dataset, while spam messages form a minority class. This imbalancenegativelyaffectsclassifiertrainingbecausemost machine learning algorithms are designed to minimize overall error. As a result, classifiers may become biased toward the majority classandfail to detect minorityclass instances effectively. In highly skewed spam datasets, a classifier may incorrectly classify most spam messages as ham while still achieving high accuracy due to the overwhelmingmajorityoflegitimateemails(HeandGarcia, 2009).

2.3 Data-Level Balancing Techniques

Data-level balancing strategies attempt to reduce class imbalance by modifying the dataset distribution before modeltraining.Thesetechniquesaimtoeitherincreasethe numberofminorityclasssamplesorreducethenumberof majorityclasssamplessothattheclassifiercanlearnspam patterns more effectively. Data-level methods are widely used because they are classifier-independent and can be applied as preprocessing steps without changing the learning algorithm itself. The most common balancing techniquesincludeRandomUndersampling(RUS),Random Oversampling(ROS),andsyntheticsamplingmethodssuch asSMOTE(Chawlaetal.,2002).

2.3.1 Random Undersampling (RUS)

Random Undersampling is a technique that reduces the majorityclasssizebyrandomlyremovinghamemailsuntila morebalancedclassdistributionisachieved.Thisapproach iscomputationallyefficient becauseit reducesthedataset size, leading to faster training time. However, the main disadvantageofRUSisthat itmaydiscard importantham samples that contain useful information for classification. Removingsuchsamplescanreducetheclassifier’sabilityto distinguish between legitimate and spam messages, potentiallyincreasingfalsepositives.Despitethislimitation, RUS is often used in large-scale applications where computationalefficiencyiscritical(HeandGarcia,2009).

2.4 Cost-Sensitive Learning Techniques

Unlikedata-levelbalancingmethods,cost-sensitivelearning approaches handle imbalance by modifying the learning algorithm rather than the dataset distribution. These methods assign higher penalties to misclassifications involvingtheminorityclassorcriticalerrorssuchasfalse positives.Cost-sensitiveapproachesareespeciallyimportant inspamfilteringbecausefalsepositives(hamclassifiedas spam) may cause serious consequences such as missed business communication. Cost-sensitive learning allows classifierstoincorporatereal-worldmisclassificationcosts, making them more practical for deployment in real email filteringsystems(Elkan,2001).

2.4.1 Class Weighting in SVM

Classweightingisacommonlyusedcost-sensitivetechnique in Support Vector Machines. In this approach, the optimization function assigns higher penalty weights to minorityclasserrors.Thisforcestheclassifiertofocusmore oncorrectlydetectingspammessages.Class-weightedSVM iseffectivebecauseitdoesnotrequirealteringthedataset distribution and works naturally within margin-based learningframeworks.Byincreasingthecostofmisclassifying spam emails, the classifier shifts its decision boundary toward the majority class, improving recall for spam detection(Veropoulos,CampbellandCristianini,1999).

2.5 Evaluation Metrics for Imbalanced Spam Classification

Evaluation of spam detection models under imbalance conditionsrequirescarefulselectionofperformancemetrics. Accuracyisoftenmisleadingbecauseitcanremainhigheven when the classifier fails to detect spam. Therefore, researchersemphasizeconfusion-matrix-basedmetricssuch asprecision, recall,andF-scores. Precisionmeasureshow manypredictedspamemailsareactuallyspam,reflectingthe falsepositiverate.Recallmeasureshowmanyactualspam emails are correctly identified, representing detection capability. The F1-score provides a harmonic balance between precision and recall, while the F2-score gives

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | A pr 2026 www.irjet.net p-ISSN: 2395-0072

greater weight to recall, making it more suitable when detecting spam is more critical than avoiding false alarms (Powers,2011).

3. MATERIALS AND METHODS

This section presents the methodology adopted for conducting a comparative study of cost-sensitive learning anddata-levelbalancingstrategiesonhighlyskewedspam email datasets. The overall experimental design follows a structured workflow involving dataset preparation, preprocessing, feature extraction, imbalance-handling, model training, and performance evaluation. Since spam emailclassificationisatextminingtask,theresearchuses standard natural language processing (NLP) techniques combinedwithmachinelearningclassifiers.Additionally,the studyemphasizesimbalance-awarelearningbysimulating real-world skewness conditions and evaluating multiple strategies to address minority class under-representation (HeandGarcia,2009).

3.1 Research Workflow

Theresearchworkflowisdesignedasasystematicpipeline to ensure fair comparison among different imbalancehandling techniques. The workflow begins with dataset collection,followedbytextpreprocessingtoremovenoise and standardize email content. After preprocessing, the textualemailsareconvertedintonumericalrepresentations usingTF-IDFvectorization,whichtransformstheemailtext intohigh-dimensionalfeaturevectorssuitableformachine learning. Once features are extracted, imbalance-handling strategiesareappliedintwocategories:data-levelbalancing methods, such as oversampling and undersampling, and cost-sensitivelearningtechniques,suchasclassweighting and threshold tuning. After applying these strategies, machinelearningclassifiersaretrainedusingtheprocessed trainingdataset.Finally,modelevaluationisconductedusing imbalance-aware performance metrics such as precision, recall, F1-score, F2-score, and AUC-PR to ensure reliable assessment under skewed distributions (Saito and Rehmsmeier,2015).

3.1.1 Pipeline: Preprocessing → TF-IDF → Balancing/Cost-Sensitive→Training→Evaluation

The experimental pipeline follows a sequential process where each stage contributes to improving classification reliability. First, preprocessing ensures that irrelevant elementssuchasHTMLtags,punctuation,andstopwordsdo not distort the feature space. Next, TF-IDF converts the cleaned text into weighted vectors representing term importance.Afterfeaturetransformation,balancingmethods suchasSMOTEorundersamplingareappliedtothetraining datatoimprovespamrepresentation,whilecost-sensitive methodsmodifythelearningbehaviorbyassigningdifferent penalties to misclassification errors. Following this, classifiers are trained on the prepared dataset. Finally,

evaluation is conducted on the test set using suitable metrics, ensuring that results are not biased by the imbalanceproblemandthatminorityclassperformanceis accuratelycaptured(Chawlaetal.,2002).

3.2 Dataset Description

This research uses three publicly available benchmark datasetstoensuregeneralizabilityandreproducibility.The datasetsincludeSpamAssassin,Ling-Spam,andCSDMC2010, which are widely used in spam filtering research. These datasets contain labeled email messages categorized into spam and ham. The use of multiple datasets is important because spam content varies across sources, and a model thatperformswellononedatasetmaynotgeneralizewellto others.Moreover,usingmultipledatasetshelpsevaluatethe stability of imbalance-handling techniques across diverse spam distributions and linguistic patterns (Metsis, AndroutsopoulosandPaliouras,2006).

3.2.1 SpamAssassin Dataset

TheSpamAssassinPublicCorpusisoneofthemostwidely useddatasetsforspamemailresearch.Itcontainsreal-world spamandlegitimateemailscollectedfrommultiplesources. The dataset includes a mixture of commercial spam messages, phishing-like emails, and normal personal or professional emails. SpamAssassin is valuable for experimentation because it contains realistic spam characteristicssuchasembeddedURLs,advertisingcontent, and obfuscated words. The dataset is frequently used for benchmarkingmachinelearningspamclassifiersduetoits open accessibility and reliable labeling (Metsis, AndroutsopoulosandPaliouras,2006).

3.3 Creation of Highly Skewed Dataset Variants

To simulate real-world spam filtering conditions, highly skeweddatasetvariantswerecreatedfromeachbenchmark dataset.Manyrealemailenvironmentscontainaverysmall proportion of spam compared to legitimate emails, often reaching imbalance ratios such as 95:5 or even 98:2. Therefore, this research generated controlled dataset variantswithimbalanceratiosof90:10,95:5,and98:2.This was achieved by randomly reducing the number of spam sampleswhilekeepingthehamclassunchanged,ensuring thatthedatasetsizeremainsrealisticandtheimbalanceis systematically increased. Such dataset simulation allows evaluation of classifier performance under increasing skewness and provides insight into which imbalancehandling strategies remain stable when spam becomes extremelyrare(HeandGarcia,2009).

3.4 Text Preprocessing

Preprocessingisacrucialstepinspamemailclassification because raw email data often contains noise such as punctuation,HTMLtags,URLs,andirrelevanttokens.Ifsuch

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | A pr 2026 www.irjet.net p-ISSN: 2395-0072

elements are not removed, they can distort feature representation and reduce classifier performance. This researchappliedastandardNLPpreprocessingpipelineto normalize the email text and reduce vocabulary size. PreprocessingalsohelpsimprovetheeffectivenessofTF-IDF featureextractionbyensuringthatonlymeaningfultextual components are included in the feature space. Proper preprocessingenhancesmodelgeneralizationbyreducing overfittingtoirrelevantemailformattingpatterns(Manning, RaghavanandSchütze,2008).

3.4.1 Tokenization, Stop word Removal, and Stemming/Lemmatization

Tokenizationistheprocessofsplittingtextintoindividual wordsortokens.Aftertokenization,stopwordssuchas“the”, “is”, “and”, and “of” are removed because they occur frequentlybutprovideminimaldiscriminatoryinformation for spam classification. Following this, stemming or lemmatizationisappliedtoreducewordstotheirrootforms. Forinstance,“running”,“runs”,and“ran”maybereducedto “run”, improving vocabulary consistency. These preprocessingstepsreducefeaturedimensionalityandhelp classifierslearnmeaningfulspampatternsmoreeffectively (Manning,RaghavanandSchütze,2008).

3.5 Feature Extraction

Machine learning classifiers require numerical input, but emailmessagesareinherentlyunstructuredtext.Therefore, feature extraction is required to convert emails into numerical vectors.In thisstudy,TF-IDF vectorizationwas usedbecauseiteffectivelyrepresentstextbycapturingterm importance relative to the document corpus. TF-IDF is widely used in spam filtering because spam emails often containdistinctkeywordsandrepeatedpromotionalphrases thatreceivehigherweightsinTF-IDFrepresentation(Metsis, AndroutsopoulosandPaliouras,2006).

3.6 Base Classifiers Used

To ensure a comprehensive comparison, this study employed three diverse machine learning classifiers: Multinomial Naïve Bayes, Support Vector Machine, and Random Forest. The selection of these classifiers was motivatedbytheirstrongperformanceintextclassification tasksandtheirdistinctlearningprinciples.Usingclassifiers withdifferentlearningbehaviorsprovidesafairevaluation ofwhetherimbalance-handlingtechniquesgeneralizeacross algorithms. Additionally, these models represent probabilistic, margin-based, and ensemble-based learning paradigms,makingthemsuitableforcomparativeanalysisin spamfilteringresearch(Breiman,2001).

4. IMBALANCE HANDLING STRATEGIES

Class imbalance is a dominant challenge in spam email classification because legitimate emails usually far

outnumber spam emails in real-world inbox traffic. This imbalance causes traditional machine learning models to become biased toward the majority class, leading to poor spam detection performance, especially under extreme skewnesssuchas95:5or98:2distributions.Toaddressthis issue, imbalance-handling strategies are generally categorized into two major groups: data-level balancing methods, which modify the dataset distribution, and costsensitive learning methods, which modify the classifier’s learningobjective.Thisstudyevaluatesbothcategoriesto identify the most effective approach for improving spam detection under highly imbalanced conditions (He and Garcia,2009).

4.1 Data-Level Balancing Methods

Data-levelbalancingstrategiesattempttoimproveminority class learning by altering the dataset composition before training. These methods either increase minority class samplesorreducemajorityclasssamplestocreateamore balanceddistribution.Theprimaryadvantageofdata-level approaches is that they are independent of the classifier, meaning they can be applied to any machine learning algorithm. However, these methods may introduce limitationssuchasinformationlossoroverfitting,depending onthechosenbalancingtechnique.Inspamfiltering,datalevel balancing is particularly useful because it allows the classifier to observe more spam patterns during training, improvingitsabilitytogeneralizeminorityclassbehavior (Chawlaetal.,2002).

4.2 Cost-Sensitive Learning Methods

Cost-sensitive learning methods handle imbalance by modifying the classifier training objective rather than changing the dataset distribution. These approaches incorporate misclassification costs into learning, ensuring thaterrorsonminorityclassinstancesorhigh-costerrors are penalized more heavily. Cost-sensitive learning is particularly important in spam filtering because false positives,wherelegitimateemailsarewronglyblocked,can bemoredamagingthanfalsenegatives.Byintroducingcost sensitivity, classifiers can be trained to reflect real-world priorities,improvingpracticalusability.Unlikeresampling methods,cost-sensitiveapproachestypicallyavoidartificial changestodatasetdistributionandmayreducetheriskof overfittingcausedbyoversampling(Elkan,2001).

4.2.1 Class-Weighted SVM

Class-weightedSupportVectorMachineisacost-sensitive methodwheredifferentpenaltyweightsareassignedtothe spamandhamclassesduringtraining.Inthisapproach,the optimization function applies a higher penalty when the minorityclass(spam)ismisclassified,forcingtheclassifier to shift the decision boundary in favor of spam detection. Thekeyadvantageofclass-weightedSVMisthatitimproves recall for spam without altering the dataset composition.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | A pr 2026 www.irjet.net p-ISSN: 2395-0072

This makes it computationally efficient and suitable for large-scalespamfilteringsystems.However,thelimitationis that assigning excessively high weights may lead to an increaseinfalsepositives,astheclassifierbecomesoverly aggressive in labeling emails as spam. Therefore, weight tuning is essential to achieve an appropriate balance between precision and recall (Veropoulos, Campbell and Cristianini,1999).

5. EXPERIMENTAL SETUP

Theexperimentalsetupisdesignedtoensureacontrolled andfaircomparisonbetweendifferentimbalance-handling strategies.Sincetheobjectiveofthisstudyiscomparative analysis,itisessentialthatallmodelsareevaluatedunder identical conditions, including consistent datasets, preprocessingsteps,featureextraction,andclassifiers.This controlledexperimentalframeworkensuresthatdifferences in results can be attributed to the imbalance-handling methodsratherthanvariationsindatapreparationormodel configuration.Furthermore,evaluationisperformedusing imbalance-awaremetricstoprovidemeaningfulinsightsinto spam detection performance under skewed distributions (SaitoandRehmsmeier,2015).

5.1

Experimental Design

The study follows a controlled comparison methodology where the same datasets and the same feature extraction approach are used across all experiments. The base classifiers, namely Naïve Bayes, SVM, and Random Forest, are trained under multiple conditions: baseline (no balancing), data-level balancing (RUS, ROS, SMOTE), and cost-sensitivelearning(classweighting,MetaCost,threshold tuning).Bykeepingtheexperimentalpipelineconstantand varying only the imbalance-handling strategy, the study ensuresthatperformance differencesarefairly evaluated. Thisdesignprovidesaclearunderstandingofwhichmethod offers the best robustness under increasing skewness conditions(HeandGarcia,2009).

5.2 Train-Test Split and Cross Validation

To ensure reliable performance evaluation, the dataset is divided into training and testing subsets using stratified splitting.Stratificationensuresthattheproportionofspam and ham emails remains consistent across training and testing datasets. This is important in imbalanced learning becauserandomsplittingmaycreatesubsetswithveryfew spamsamples,resultinginunreliableevaluation.Inaddition, cross-validationisappliedtoensurethatresultsarestable andnotdependentonaparticularsplit.Acriticalaspectof this setup is avoiding data leakage: balancing techniques such as SMOTE and oversampling are applied only on the trainingdataset,whilethetestdatasetremainsuntouched. Thispreventsartificialinflationofperformancemetricsand ensures realistic model evaluation (Japkowicz and Shah, 2011).

5.3 Hyper parameter Tuning

Hyperparametertuningisperformedtooptimizeclassifier performance and ensure fair comparison. This study uses GridSearchCV to explore multiple hyperparameter combinations systematically. For Naïve Bayes, smoothing parameteralphaistunedtocontrolprobabilityestimation stability. For SVM, the penalty parameter C is optimized, alongwithkernelselectionifrequired.ForRandomForest, keyparameterssuchasthenumberoftrees(n_estimators), maximumtreedepth(max_depth),andminimumsamples split are tuned. GridSearchCV is performed using crossvalidation on the training set, ensuring that the selected hyperparameters generalize effectively. Hyperparameter tuning is essential because imbalance-handling strategies mayinfluenceoptimalmodelsettingsdifferently,especially inhighlyskeweddatasets(BergstraandBengio,2012).

5.4 Evaluation Metrics

Evaluationmetricsplayacentralroleinimbalancedspam classification because traditional accuracy fails to capture minority class detection capability. Therefore, multiple metrics are used to evaluate classifier performance comprehensively. The metrics include confusion matrix, precision, recall, F1-score, F2-score, and AUC-PR. Each metricprovidesadifferentperspectiveonspamdetection performance,ensuringabalancedevaluationframework.In particular, recall and F2-score are emphasized because missing spam emails can allow phishing and malware attackstoreachusers(Powers,2011).

5.4.1 Confusion Matrix

The confusion matrix is used to summarize classification outcomesbyreportingtruepositives,truenegatives,false positives, and false negatives. In spam filtering, true positives represent correctly detected spam emails, while falsepositivesrepresentlegitimateemailswronglylabeled as spam. The confusion matrix is essential because it providesaclearunderstandingoferrordistribution,which is critical in cost-sensitive spam filtering environments. It alsoformsthebasisforcalculatingotherevaluationmetrics suchasprecisionandrecall(Fawcett,2006).

5.5 Computational Environment

AllexperimentsareconductedusingPythonprogramming languageinastandardmachinelearningenvironment.The implementation uses widely adopted libraries such as NumPy and Pandas for data handling, Scikit-learn for machinelearningmodelsandevaluation,andImbalancedlearn for resampling techniques such as SMOTE. The computationalsetupincludesasystemwithatleastanIntel i5/i7 processor (or equivalent), 8GB or higher RAM, and sufficient storage for dataset processing. Since TF-IDF featurematricesaresparseandhigh-dimensional,memory efficiencyisanimportantconsideration.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | A pr 2026 www.irjet.net p-ISSN: 2395-0072

6. RESULTS

This section presents a detailed analysis of experimental findings obtained from applying baseline classifiers, datalevel balancing methods, and cost-sensitive learning strategies to highly skewed spam email datasets. The evaluation focuses on imbalance-aware performance metrics,particularlyrecall,F2-score,precision,andAUC-PR. Resultsareanalyzedacrossdifferentskewnesslevels(90:10, 95:5,and98:2)toassessmethodrobustness.

6.1 Dataset Skewness Statistics

Before applying any imbalance-handling strategy, the datasetswereanalyzedtoconfirmthedistributionofham andspamemailsunderdifferentsimulatedskewconditions. The imbalance ratio significantly influences classifier behavior,particularlyunderextremeskewnesssuchas98:2. As the proportion of spam decreases, minority class detection becomes more challenging due to limited representationduringtraining(HeandGarcia,2009).

6.2 Baseline Results Without Any Strategy

Baselineexperimentswereconductedwithoutapplyingany balancing or cost-sensitive techniques. The results demonstrate that classifier performance deteriorates significantlyasimbalanceincreases,especiallyat the98:2 ratio.Althoughoverallaccuracyremainshigh,recallandF2score decline sharply, indicating failure to detect spam effectively.

6.2.1 Performance Collapse Under 98:2 Imbalance

Under extreme skewness (98:2), classifiers such as Naïve Bayes and SVM predict the majority class in most cases. Whileaccuracymayexceed97%,recalldropssubstantially because many spam emails are misclassified as ham. This behavior confirms that traditional classifiers are not inherentlyrobusttosevereimbalanceconditions(Japkowicz andShah,2011).

6.2.2 Accuracy Paradox Highlight

The baseline results clearly demonstrate the accuracy paradox, where high accuracy masks poor minority class performance.Forexample,aclassifierpredictingallemails as ham achieves 98% accuracy in a 98:2 dataset but fails entirelyatspamdetection.Thisreinforcestheimportanceof using recall, F2-score, and AUC-PR rather than accuracy alone when evaluating spam classifiers under imbalance (SaitoandRehmsmeier,2015).

6.3 Results of Data-Level Strategies

Data-level balancing strategies significantly improved minority class detection. By modifying the dataset distribution,classifierswereexposedtomorespamsamples duringtraining,resultinginimprovedrecallandF2-score.

6.3.1 RUS Results

Random Undersampling improved recall compared to baseline by reducing majority class dominance. However, because many legitimate emails were removed, some classifiers experienced reduced precision. While RUS enhanced spam detection sensitivity, its information loss sometimes negatively affected overall generalization. F2scoreimprovedmoderatelybutwasnotconsistentlystable acrossdatasets(HeandGarcia,2009).

6.4 Results of Cost-Sensitive Strategies

Cost-sensitive strategies improved performance without modifyingdatasetdistribution.Theseapproachesprimarily enhancedrecallwhilemaintainingbettercontroloverfalse positivescomparedtodata-levelmethods.

7. CONCLUSION

Thisresearchpresentedacomparativeevaluationofdatalevel balancing techniques and cost-sensitive learning strategies for spam email classification under highly imbalanceddatasetconditions.Experimentswereconducted onthreebenchmarkdatasets,namelySpamAssassin,LingSpam,andCSDMC2010,withsimulatedimbalanceratiosof 90:10, 95:5, and 98:2 to reflect real-world spam filtering environments. The baseline classifiers showed strong performancedegradationunderextremeskewness,where accuracy remained high but recall and F2-score dropped significantly, confirming the presence of the accuracy paradoxinimbalancedspamdatasets.Data-levelmethods suchasRandomUndersamplingandRandomOversampling improvedminorityclassdetectionbutexhibitedlimitations suchasinformationlossandoverfitting.SMOTEachievedthe mostconsistentimprovementamongresamplingtechniques by generating synthetic spam samples, leading to higher recall and F2-score values. Cost-sensitive strategies, including class-weighted SVM, MetaCost, and threshold tuning,demonstratedstableperformancewithoutmodifying dataset distribution. Among these, threshold optimization provedhighlyeffectiveinimprovingF2-scorebyincreasing spam recall while maintaining acceptable precision. The comparativeresultsindicatethatSMOTE-basedresampling iseffectivewhenmaximizingspamdetectionistheprimary objective, whereas threshold tuning and class-weighted learningaremoresuitablewhenminimizingfalsepositives is critical. Overall, this study provides a structured frameworkforselectingimbalance-handlingtechniquesin spamfilteringsystems,highlightingtheimportanceofusing imbalance-awaremetricssuchasF2-scoreandAUC-PRfor realisticevaluation.

8. FUTURE SCOPE

Futureworkcanextendthisresearchbyincorporatingdeep learning-basedarchitecturessuchasCNN-LSTM,Bi-LSTM, and Transformer models, which can capture contextual

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | A pr 2026 www.irjet.net p-ISSN: 2395-0072

relationships beyond TF-IDF features. Advanced oversamplingtechniquessuchasADASYNandBorderlineSMOTE may also be explored to generate more realistic minority samples and reduce class overlap. Additionally, concept drift handling should be investigated since spam patterns evolve continuously over time, making static models less reliable in real-world deployment. Future studies may also develop adaptive cost-sensitive frameworks where misclassification costs dynamically changebasedonuserfeedbackandemailpriority.Finally, hybrid models combining resampling with cost-sensitive learning could be designed to achieve better trade-offs between spam recall and false positive minimization in highlyskewedenvironments.

REFERENCES

1. Androutsopoulos,I.,Koutsias,J.,Chandrinos,K.V.and Spyropoulos, C.D. (2000) ‘An evaluation of naive Bayesian anti-spam filtering’, in Proceedings of the WorkshoponMachineLearningintheNewInformation Age.Barcelona,Spain,pp.9–17.

2. Bergstra,J.andBengio,Y.(2012)‘Randomsearchfor hyper-parameter optimization’, Journal of Machine LearningResearch,13,pp.281–305.

3. Breiman,L.(2001)‘Randomforests’,MachineLearning, 45(1),pp.5–32.

4. Chawla,N.V.,Bowyer,K.W.,Hall,L.O.andKegelmeyer, W.P.(2002)‘SMOTE:SyntheticMinorityOver-sampling Technique’,JournalofArtificialIntelligenceResearch, 16,pp.321–357.

5. Domingos,P.(1999)‘MetaCost:Ageneralmethodfor makingclassifierscost-sensitive’,inProceedingsofthe Fifth ACM SIGKDD International Conference on KnowledgeDiscoveryandDataMining(KDD’99).San Diego,CA:ACM,pp.155–164.

6. Drummond,C.andHolte,R.C.(2006)‘Costcurves:An improved method for visualizing classifier performance’,MachineLearning,65(1),pp.95–130.

7. Elkan, C. (2001) ‘The foundations of cost-sensitive learning’, in Proceedings of the Seventeenth InternationalJointConferenceonArtificialIntelligence (IJCAI).Seattle,WA,USA,pp.973–978.

8. Fawcett, T. (2006) ‘An introduction to ROC analysis’, PatternRecognitionLetters,27(8),pp.861–874.

9. Han,H.,Wang,W.Y.andMao,B.H.(2005)‘BorderlineSMOTE: A new over-sampling method in imbalanced datasetslearning’,inProceedingsoftheInternational Conference on Intelligent Computing (ICIC). Berlin: Springer,pp.878–887.

10. He, H. and Garcia, E.A. (2009) ‘Learning from imbalanceddata’,IEEETransactionsonKnowledgeand DataEngineering,21(9),pp.1263–1284.

11. Japkowicz,N.andShah,M.(2011)EvaluatingLearning Algorithms: A Classification Perspective. Cambridge: CambridgeUniversityPress.

12. Joachims,T.(1998)‘Textcategorizationwithsupport vector machines: Learning with many relevant features’, in Proceedings of the 10th European Conference on Machine Learning (ECML). Berlin: Springer,pp.137–142.

13. Manning, C.D., Raghavan, P. and Schütze, H. (2008) Introduction to Information Retrieval. Cambridge: CambridgeUniversityPress.

14. Metsis,V.,Androutsopoulos,I.andPaliouras,G.(2006) ‘SpamfilteringwithNaiveBayes–WhichNaiveBayes?’, inProceedingsof the Third ConferenceonEmail and Anti-Spam (CEAS 2006). Mountain View, California, USA.

15. Pedregosa, F. et al. (2011) ‘Scikit-learn: Machine learning in Python’, Journal of Machine Learning Research,12,pp.2825–2830.

16. Powers, D.M.W. (2011) ‘Evaluation: From precision, recall and F-measure to ROC, informedness, markedness and correlation’, Journal of Machine LearningTechnologies,2(1),pp.37–63.

17. Saito, T. and Rehmsmeier, M. (2015) ‘The precisionrecallplotismoreinformativethantheROCplotwhen evaluatingbinaryclassifiersonimbalanceddatasets’, PLOSONE,10(3),e0118432.

18. Sculley, D., O’Connor, M., McCallum, A. and CorradaEmmanuel, A. (2011) ‘Spam filtering using machine learningtechniques’,inProceedingsoftheConference onEmailandAnti-Spam(CEAS).

19. Veropoulos,K.,Campbell,C.andCristianini,N.(1999) ‘Controllingthesensitivityofsupportvectormachines’, inProceedingsoftheInternationalJointConferenceon Artificial Intelligence (IJCAI). Stockholm, Sweden, pp. 55–60.

Turn static files into dynamic content formats.

Create a flipbook