Skip to main content

FAKE NEWS DETECTION USING NATURAL LANGUAGE PROCESSING AND MACHINE LEARNING

Page 1


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

FAKE NEWS DETECTION USING NATURAL LANGUAGE PROCESSING AND MACHINE LEARNING

Khushboo Yadav1 , Dr. J. B. Singh2

1Master of Technology, Computer Science and Engineering, Sagar Institute of Technology and Management, Barabanki, India

2Professor, Department of Computer Science and Engineering, Sagar Institute of Technology and Management, Barabanki, India ***

Abstract - The rapid proliferation of digital media platforms has significantly increased the spread of fake news, posing serious challenges to societal trust, political stability, and information integrity. Traditional manual verification methods are inadequate to handle the scale and speed at which misinformation propagates online. This study proposes an automated fake news detection framework using Natural Language Processing (NLP) and Machine Learning (ML) techniques. The approachfocusesonanalyzingtextualcontent to distinguish between genuine and fabricated news by leveraging linguistic, statistical, and contextual features. Benchmark datasets, including LIAR and FakeNewsNet, are utilized to ensure experimental reliability and comparative evaluation. The methodologyincorporatescomprehensivetext preprocessing, followed by feature extraction using TF-IDF, word embeddings (Word2Vec, GloVe), and transformer-based contextual models such as BERT. Multiple classification algorithms, including Logistic Regression, Support Vector Machine, Random Forest, ConvolutionalNeuralNetworks,and transformer-based models, are implemented and compared. Performance is evaluatedusingaccuracy,precision,recall,and F1-score. Experimental results demonstratethattransformerbased models outperform traditional machine learning approaches by effectively capturing semantic and contextual nuances in text. The proposed framework achieves improved classification accuracy androbustnessacrossdiversedatasets. This research contributes to the development of scalable and efficient fake news detection systems, offering practical applications for social media platforms, news organizations, and policymakers in combating misinformation.

Key Words: Fake News Detection, Natural Language Processing, Machine Learning, Deep Learning, BERT, Text Classification

1. INTRODUCTION

Therapidevolutionofdigitalcommunicationtechnologies has fundamentally transformed the way information is created,disseminated,andconsumed.Onlineplatformshave enabled instantaccessto newsandglobal events,thereby enhancinginformationavailabilityandpublicengagement. However,thistransformationhasalsointroducedsignificant challenges,particularlythe widespreadcirculationof fake news. Fake news, defined as deliberately fabricated or

misleading information presented as legitimate news, has emerged as a major threat to societal trust, democratic processes, and public decision-making. The increasing reliance on digital platforms for news consumption has amplifiedtheimpactofmisinformation,makingitessential todevelopautomatedandscalabledetectionmechanisms.In this context, Natural Language Processing (NLP) and MachineLearning(ML)havegainedprominenceaseffective tools for analyzing textual data and identifying deceptive contentpatterns(Shuetal.,2017).

1.1 Background

1.1.1

Rise of Digital Media and Social Platforms

Overthepastdecade,digital mediaandsocial networking platformshavebecometheprimarysourcesofinformation for a vast global audience. Platforms such as Facebook, Twitter(nowX),Instagram,andWhatsAppenableusersnot onlytoconsumenewsbutalsotocreateandsharecontentin real time. This shift from traditional journalism to userdriven content generation has democratized information dissemination, allowing diverse voices to participate in public discourse. However, the absence of strict editorial oversightontheseplatformshasreducedthereliabilityof sharedinformation.Unliketraditionalmediaoutlets,which follow rigorous verification processes, social media platforms often lack mechanisms to ensure content authenticitybeforepublication.Asaresult,misleadingand falseinformationcaneasilyreachalargeaudiencewithout priorvalidation(Castilloetal.,2011).

1.1.2 Rapid Spread of Misinformation

The structural and algorithmic design of social media platforms significantly contributes to the rapid spread of misinformation. Content that is emotionally engaging, sensational,orcontroversialtendstoreceivehighervisibility duetouserinteractionssuchaslikes,shares,andcomments. This creates an environment where fake news can spread faster than factual information, often going viral within a short period. Studies have shown that false information propagatesmorerapidlyandwidelythantrueinformation due to its novelty and emotional appeal (Vosoughi et al., 2018). Additionally, the interconnected nature of social networks allows misinformation to cascade across

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

communities, making it difficult to control once it begins spreading. This rapid dissemination not only misleads individuals but also influences public opinion, potentially leadingtoserioussocialandpoliticalconsequences.

1.1.3

Need for Automated Detection Systems

Given the sheer volume of content generated on digital platforms, manual fact-checking is no longer sufficient to address the problem of fake news. Human verification processes are time-intensive, resource-dependent, and incapableofkeepingpacewithreal-timeinformationflow. Consequently, there is a growing need for automated detection systems that can efficiently process large-scale textualdataandidentifymisleadingcontent.NLPtechniques enable machines to understand and analyze human language, while ML algorithms provide predictive capabilities by learning patterns from labeled data. The integrationofthesetechnologiesallowsforthedevelopment of scalable and efficient systems capable of detecting fake news with high accuracy, thereby supporting content moderationandinformationverificationefforts(Shuetal., 2017).

1.2 Problem Statement

1.2.1

Linguistic Ambiguity and Misleading Semantics

Oneoftheprimarychallengesinfakenewsdetectionliesin thecomplexityofhumanlanguage.Fakenewsoftenemploys subtle linguistic manipulations, including ambiguity, exaggeration,andselectiveframing,toappearcrediblewhile conveyingmisleadinginformation.Wordsandphrasesmay have multiple meanings depending on context, making it difficult for detection systems to interpret the intended message accurately. Additionally, fake news creators frequentlyuseemotionallychargedorpersuasivelanguage to influence readers’ perceptions. These semantic complexitiesposesignificantchallengesfortraditionaltext classification methods, which may struggle to capture nuanceddifferencesbetweengenuineanddeceptivecontent (Rubinetal.,2016).

1.2.2 Dynamic Nature of Fake News

Fakenewsisnotstatic;itevolvescontinuouslyinresponse to current events, technological advancements, and user behavior.Newnarratives,terminologies,anddissemination strategiesemergefrequently,makingitdifficultfordetection systems to remain effective over time. Models trained on historical data may fail to generalize to new types of misinformation, leading to reduced performance in realworld scenarios. Furthermore, fake news spans multiple domains, including politics, healthcare, finance, and entertainment,eachwithitsownlinguisticcharacteristics. This dynamic and domain-specific nature of fake news necessitatesthedevelopmentofadaptivemodelscapableof learning from diverse datasets and generalizing across differentcontexts(ZhouandZafarani,2019).

1.2.3 Limitations of Manual and Rule-Based Detection

Traditional approaches to fake news detection, such as manualfact-checkingandrule-basedsystems,areinherently limited in scalability and adaptability. Manual verification reliesonhumanexpertiseandisunabletohandlethevast andcontinuouslygrowingvolumeofonlinecontent.Onthe other hand, rule-based systems depend on predefined patternsorkeywords,whichmaynotcapturethecomplexity andvariabilityoffakenews.Thesesystemsoftenfailwhen facedwithneworsophisticatedmisinformationstrategies thatdonotconformtopredefinedrules.Asaresult,thereis asignificantgapbetweentherateoffakenewsgeneration and the capacity of traditional detection methods, highlighting the need for more advanced, data-driven approaches(KumarandShah,2018).

1.3 Research Objectives

1.3.1

Comparative Analysis of ML/DL Models

A key objective of this research is to conduct a comprehensive comparative analysis of various machine learninganddeeplearningmodelsforfakenewsdetection. ByevaluatingmodelssuchasLogisticRegression,Support Vector Machines, Random Forests, Convolutional Neural Networks,andtransformer-basedarchitectures,thestudy aims to identify the strengths and limitations of each approach.Thiscomparisonprovidesvaluableinsightsinto modelperformance,computationalefficiency,andsuitability fordifferenttypesoftextualdata.

1.3.2 Improve Detection Using Semantic Features

Another important objective is to enhance detection accuracy by incorporating advanced semantic features. Traditional methods such as TF-IDF focus on word frequency but often fail to capture contextual meaning. In contrast,moderntechniquessuchaswordembeddingsand transformer-basedmodelsenableadeeperunderstandingof semanticrelationshipsandcontextualdependencieswithin text. By leveraging these advanced representations, the researchaimstoimprovethemodel’sabilitytodetectsubtle andcomplexformsofmisinformation(Devlinetal.,2019).

1.3.3 Develop a Robust and Scalable Detection System

Thestudyalsoaims to developa robust andscalable fake newsdetectionsystemcapableofhandlinglargevolumesof data and adapting to evolving misinformation patterns. Scalability is essential for real-world applications, where systems must process continuous streams of data from multiple sources. Robustness ensures that the model performsconsistentlyacrossdifferentdatasetsanddomains, makingitsuitablefordeploymentinpracticalenvironments such as social media platforms and news verification systems.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

1.4 Contributions of the Paper

1.4.1 Hybrid NLP + ML/DL Framework

Thisresearchintroducesahybridframeworkthatintegrates NaturalLanguageProcessingtechniqueswithbothmachine learninganddeeplearningmodels.Bycombininglinguistic analysis with computational intelligence, the proposed approachenhancesthecapabilitytodetectfakenewsmore accurately and efficiently. The hybrid nature of the framework allows it to leverage the strengths of different methodologies,resultinginimprovedperformance.

1.4.2 Comparative Evaluation (TF-IDF, Embeddings, Transformers)

Thestudyprovidesacomprehensiveevaluationofvarious featureextractiontechniques,includingtraditionalstatistical methodssuchasTF-IDF,distributedwordembeddingslike Word2Vec and GloVe, and advanced transformer-based models such as BERT. This comparison highlights the effectiveness of different representations in capturing linguisticandsemanticpatterns,therebycontributingtothe understandingoffeatureimportanceinfakenewsdetection.

1.4.3 Performance Benchmarking on Standard Datasets

To ensure reliability and reproducibility, the proposed modelsareevaluatedonstandardbenchmarkdatasetssuch as LIAR and FakeNewsNet. Benchmarking against widely used datasets allows for meaningful comparison with existingstudies and demonstratesthe effectiveness ofthe proposedapproach.Theresultsprovideempiricalevidence supportingthesuperiorityofadvancedmodelsindetecting fake news with higher accuracy and robustness (Wang, 2017).

2. RELATED WORK (LITERATURE REVIEW)

Theproblemoffakenewsdetectionhasattractedsignificant attentionfromresearchersacrossdomainssuchasNatural Language Processing, data mining, and social network analysis. Over time, the field has evolved from basic linguistic analysis to sophisticated deep learning and transformer-based models. This section reviews the progressionofmethodologies,highlightingkeycontributions andexistinglimitations.

2.1 Early Approaches

2.1.1 Linguistic and Credibility-Based Detection

Initialresearchinfakenewsdetectionprimarilyfocusedon analyzinglinguisticfeaturesandcredibilityindicatorswithin textual content. Early studies explored how deceptive information could be identified through stylistic patterns, includingwordfrequency,sentimentpolarity,andsyntactic structures.Theseapproachesreliedontheassumptionthat fakenewsexhibitsdistinguishablelinguisticcharacteristics,

suchasexaggeratedlanguage,inconsistency,oremotional tone. Researchers also incorporated credibility-based features,suchassourcereliabilityandauthorreputation,to enhance detection accuracy. For instance, studies demonstratedthatcombiningtextualanalysiswithmetadata such as publisher credibility significantly improves classification performance (Mihalcea and Strapparava, 2010). While these methods laid the foundation for fake news detection, they were limited in capturing complex contextualrelationshipswithintext.

2.1.2 Social Media Misinformation Studies

Withtheriseofsocialmediaplatforms,researchersbeganto investigatemisinformationfromanetworkandpropagation perspective.Unliketraditionaltext-basedapproaches,these studies analyzed how information spreads across social networks,focusingonuser behavior,interaction patterns, and temporal dynamics. It was observed that fake news often follows distinct propagation patterns, such as rapid initialdiffusionfollowedbysuddendecline.Byincorporating features such as retweet patterns, user credibility, and engagementmetrics,earlymodelswereabletodistinguish between credible and misleading information more effectively. This shift highlighted the importance of combiningtextualfeatureswithsocialcontextforimproved detection(Castilloetal.,2011).

2.2 Machine Learning-Based Approaches

2.2.1

Logistic Regression, SVM, Random Forest

Asthefieldprogressed,machinelearningalgorithmsbecame widelyadoptedforfakenewsdetectionduetotheirabilityto learn patterns from labeled data. Models such as Logistic Regression, Support Vector Machines (SVM), and Random Forestswereextensivelyusedforbinaryclassificationtasks. Logistic Regression provided a simple yet effective probabilistic framework, while SVM demonstrated strong performanceinhigh-dimensionalfeaturespacestypicalof textual data. Random Forest, as an ensemble method, improvedrobustnessbycombiningmultipledecisiontrees. Thesemodelsofferedadvantagesintermsofinterpretability and computational efficiency, making them suitable for baselineevaluationsandlarge-scaleapplications.

2.2.2 Feature Engineering (TF-IDF, BoW)

The effectiveness of machine learning models largely depended on feature engineering techniques used to representtextualdata.TraditionalapproachessuchasBagof-Words (BoW) and Term Frequency–Inverse Document Frequency(TF-IDF)transformedtextintonumericalvectors basedonwordfrequencyandimportance.Thesemethods enabled models to capture discriminative keywords and phrasesassociatedwithfakenews.However,theylackedthe ability to capture semantic meaning and contextual relationshipsbetweenwords.Despitethislimitation,feature engineeringplayedacrucialroleinearlydetectionsystems

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

and provided a strong foundation for subsequent advancements(Wang,2017).

2.3 Deep Learning Approaches

2.3.1

CNN, RNN, LSTM Models

The introduction of deep learning marked a significant advancementinfakenewsdetectionbyenablingautomatic feature extraction and representation learning. Convolutional Neural Networks (CNNs) were applied to capturelocalpatternsintext,suchasphrasesandn-grams, whichareindicativeofdeceptivecontent.RecurrentNeural Networks (RNNs) and their variants, particularly Long Short-Term Memory (LSTM) networks, were designed to modelsequentialdependenciesintext,capturingcontextual relationshipsacrosssentences.Thesemodelseliminatedthe need for manual feature engineering and demonstrated improved performance compared to traditional machine learningapproaches.

2.3.2

Semantic Learning Improvements

Deep learning models further enhanced semantic understanding by learning distributed representations of wordsandsentences.Wordembeddingsallowedmodelsto capturesimilaritiesbetweenwordsbasedontheircontext, improvingtheabilitytodetectsubtlelinguisticcuesinfake news.Theseadvancementsenabledsystemstomovebeyond surface-leveltextanalysisandincorporatedeepersemantic insights.However,deeplearningmodelsoftenrequiredlarge datasets and significant computational resources, which posedchallengesforpracticalimplementation(Kaliyaretal., 2020).

2.4 Transformer-Based Models

2.4.1 BERT, RoBERTa, DistilBERT

Transformer-basedmodelsrepresentthestate-of-the-artin fakenewsdetection,offeringsignificantimprovementsover traditional and deep learning approaches. Models such as BERT (Bidirectional Encoder Representations from Transformers), RoBERTa, and DistilBERT utilize attention mechanismstocapturecontextualrelationshipswithintext. Unlikeearliermodels,transformersconsiderbothleftand right context simultaneously, enabling a deeper understanding of language semantics. RoBERTa improves upon BERT through optimized training strategies, while DistilBERTprovidesalightweightalternativewithreduced computationalrequirements.

2.4.2

Context-Aware Fake News Detection

Theprimaryadvantageoftransformer-basedmodelsliesin theirabilitytogeneratecontext-awareembeddings,where the meaning of a word is influenced by its surrounding context. This capability is particularly important for fake news detection, where subtle differences in phrasing and

contextcansignificantlyimpactinterpretation.Transformer models have demonstrated superior performance in capturingnuancedlinguisticpatternsandachievinghigher classification accuracy. Their ability to generalize across differentdatasetsanddomainsmakesthemhighlyeffective forreal-worldapplications(Devlinetal.,2019).

2.5 Research Gaps

Despite significant advancements, existing models often struggle to generalize across different domains. A model trained on political news may not perform well on healthcareorfinancialmisinformationduetovariationsin languageandcontext.Thislimitationhighlightstheneedfor moreadaptableanddomain-independentapproaches.

Mostcurrentfakenewsdetectionsystemsoperateinoffline settings using pre-collected datasets. Real-time detection remains a challenge due to the high computational requirementsandtheneedforcontinuousdataprocessing. Addressingthislimitationiscrucialforpracticaldeployment indynamiconlineenvironments.

Anothermajorchallengeisdatasetbias,asmanybenchmark datasets are limited in scope and may not represent realworlddiversity.Additionally,moststudiesfocusonEnglishlanguage data, restricting the applicability of models in multilingualcontexts.Theselimitationsemphasizetheneed for more diverse datasets and cross-lingual detection techniques(ZhouandZafarani,2019).

3. METHODOLOGY

Thissectionpresentsthesystematicapproachadoptedfor developing an effective fake news detection system. The methodologyintegratesNaturalLanguageProcessing(NLP) and Machine Learning (ML) techniques into a structured pipeline,ensuringreproducibility,scalability,andaccuracy indetectingmisinformation.

3.1 Overall Framework

3.1.1 Detection Pipeline

The proposed framework follows a multi-stage pipeline consisting of data collection, preprocessing, feature extraction,modeltraining,andevaluation.Initially,datasets arecollectedfromreliablesourcesandpreparedforanalysis. Thepreprocessingstagetransformsrawtextualdataintoa cleanandstructuredformat.Featureextractiontechniques are then applied to convert text into numerical representations suitable for machine learning algorithms. Subsequently, multiple models are trained and evaluated using standard performance metrics. This structured pipeline ensures a logical flow of operations and enables systematiccomparisonofdifferentapproaches(Shuetal., 2017).

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

3.2 Dataset Description

3.2.1 LIAR

Dataset

TheLIARdatasetisawidelyusedbenchmarkforfakenews detection, consisting of short political statements labeled acrossmultipletruthfulnesscategories.Forthisstudy,the datasetistransformedintoabinaryclassificationproblem by grouping labels into “real” and “fake” categories. Its balanced distribution and structured annotations make it suitableforevaluatingclassificationmodels(Wang,2017).

3.2.2 FakeNewsNet Dataset

FakeNewsNet provides a more comprehensive dataset containingfull-lengthnewsarticlesalongwithsocialcontext featuressuchasuserengagementandpublisherinformation. It includes subsets like PolitiFact and GossipCop, offering diversetextualandcontextualdata.Thisdatasetenablesthe evaluation of models on real-world news content and enhancesgeneralizability.

3.2.3 Data Distribution and Splitting Strategy

To ensure fair evaluation, the datasets are divided into training,validation,andtestingsetsusinga70–15–15ratio. Stratified sampling is applied to maintain class balance acrossallsubsets.

Table 1: Data Splitting Strategy

Dataset

Portion Percentage Purpose

Training 70% Modellearning

Validation 15% Hyperparameter tuning

Testing 15% Finalevaluation

3.3 Data Preprocessing

3.3.1 Text Cleaning

Textpreprocessingbeginswithcleaningoperationssuchas convertingalltexttolowercaseandremovingpunctuation, specialcharacters,andURLs.Thesestepsreducenoiseand ensure uniformity in textual data representation, thereby improvingmodelperformance(KumarandShah,2018).

3.3.2 Tokenization and Lemmatization

Tokenizationdividestextintoindividual wordsortokens, formingthebasicunitsforanalysis.Lemmatizationisthen applied to reduce words to their base forms, ensuring consistencyacrossvariationsofthesameword.Thisprocess

enhances semantic clarity and reduces feature dimensionality.

3.3.3

Stop-word Removal

Commonwordssuchas“the,”“is,”and“and”areremovedas they do not contribute significantly to classification. Eliminating these stop-words improves computational efficiency and allows the model to focus on meaningful features.

3.3.4

Handling Missing and Noisy Data

Missingvaluesandirrelevantentriesarehandledthrough datacleaningtechniques,ensuringdatasetintegrity.Noisy data,includingredundantorinconsistenttext,isfilteredto improvethequalityofinputdataformodeltraining.

3.4

Feature Extraction

Feature extraction transforms textual data into numerical representationsthatmachinelearningmodelscanprocess effectively.

3.4.1

Statistical Features

TF-IDF

Term Frequency–Inverse Document Frequency (TF-IDF) assignsweightstowordsbasedontheirimportancewithina document relative to the entire dataset. It highlights discriminative terms while reducing the influence of commonlyoccurringwords.

Bag-of-Words

The Bag-of-Words (BoW) model represents text as a collection of word frequencies without considering order. Although simple, it provides a strong baseline for text classificationtasks.

3.4.2 Word Embeddings

Word2Vec

Word2Vecgeneratesdensevectorrepresentationsofwords by capturing semantic relationships based on context. It enablesmodelstounderstandsimilaritiesbetween words beyondsimplefrequencycounts.

GloVe

GloVe (Global Vectors) utilizes global co-occurrence statistics to create word embeddings, providing a balance betweenlocalcontextandglobalsemanticrelationships.

FastText

FastText extends Word2Vec by incorporating subword information,allowingittohandlerareandmisspelledwords

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

effectively,whichisparticularlyusefulinsocialmediatext (Bojanowskietal.,2017).

Table 2: Comparison of Feature Extraction Techniques

Technique Type Key Advantage Limitation

TF-IDF Statistical Simple, interpretable Nocontext awareness

BoW Statistical Easy implementation Ignores wordorder

Word2Vec Embedding Captures semantics Static embeddings

GloVe Embedding Globalcontext Requires large corpus

FastText Embedding Handlesrare words Increased complexity

3.4.3 Contextual Embeddings

BERT

BERTusesbidirectionaltransformerstounderstandcontext frombothdirectionsinasentence,enablingdeepsemantic understanding.

RoBERTa

RoBERTaimprovesuponBERTthroughoptimizedtraining strategies and larger datasets, resulting in enhanced performance.

DistilBERT

DistilBERT is a lightweight version of BERT that reduces computationalcostwhilemaintaininghighaccuracy,making itsuitableforreal-timeapplications(Devlinetal.,2019)

3.5 Handling Imbalanced Data

3.5.1 SMOTE Technique

Classimbalanceisaddressedusingthe SyntheticMinority Oversampling Technique (SMOTE), which generates synthetic samples for the minority class. This approach improvesmodelperformancebyensuringbalancedlearning andreducingbiastowardthemajorityclass(Chawlaetal., 2002).

3.6 Model Development

3.6.1 Machine Learning Models

Classical machine learning models are implemented as baselineclassifiers.

 Logistic Regression: Provides probabilistic classificationandinterpretability.

 Support Vector Machine (SVM): Effective in highdimensionalfeaturespaces.

 Random Forest: Ensemble method that improves accuracyandreducesoverfitting.

 Naive Bayes: Efficientprobabilisticclassifiersuitable fortextdata.

3.6.2

Deep Learning Models

Deep learning models automatically learn hierarchical featurerepresentations.

CNN: Captureslocaltextualpatternssuchasphrasesandngrams.

LSTM: Models sequential dependencies and long-term contextintext.

These models outperform traditional approaches by capturingcomplexlinguisticstructures(Kaliyaretal.,2020).

3.6.3 Transformer Models

Transformer-basedmodelsrepresentthestate-of-the-artin fakenewsdetection.

BERT Fine-tuning: Adapts pre-trained BERT for classificationtasks.

RoBERTa: Offersimprovedperformancethroughenhanced training.

DistilBERT: Providesefficiencywithreducedcomputational cost.

These models excel in capturing contextual nuances and semanticrelationshipswithintext.

3.7 Experimental Setup

3.7.1 Software

Tools

TheimplementationiscarriedoutusingPython,leveraging librariessuchasScikit-learnforclassicalmachinelearning, TensorFlow and PyTorch for deep learning, and NLP librarieslikeNLTKandspaCyforpreprocessing.Thesetools provideflexibilityandscalabilityformodeldevelopment.

3.7.2 Hardware Environment

TheexperimentsareconductedonaGPU-enabledsystemto handlecomputationallyintensivetaskssuchastrainingdeep learning and transformer models. High-performance hardware significantly reduces training time and enables experimentationwithcomplexarchitectures.

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net

Table 3: Experimental Setup

4.1.4 F1-Score

TheF1-scoreistheharmonicmeanofprecisionandrecall, providinga balancedmeasureofmodel performance.It is especiallyusefulwhendealingwithimbalanceddatasets,as it considers both false positives and false negatives in its calculation.

4.1.5 ROC-AUC

The Receiver Operating Characteristic–Area Under Curve (ROC-AUC) metric evaluates the model’s ability to distinguish between classes across different threshold values. A higher ROC-AUC score indicates better classificationperformanceandrobustness(Fawcett,2006).

Table 4: Performance Metrics

4. RESULTS AND ANALYSIS

This section presents the experimental outcomes of the proposed fake news detection framework. The evaluation focuses on measuring model performance using standard classificationmetrics,comparingdifferentapproaches,and analyzingtheeffectivenessoffeatureextractiontechniques. The results provide insights into the strengths and limitations of machine learning, deep learning, and transformer-basedmodels.

4.1 Performance Metrics

4.1.1

Accuracy

Accuracy measures the proportion of correctly classified instancesamongthetotalnumberofsamples.Itprovidesa generalindicationofmodelperformance;however,itmaybe misleadinginthepresenceofimbalanceddatasets.Despite thislimitation,accuracyremainsacommonlyusedbaseline metric for evaluating classification models (Sokolova and Lapalme,2009).

4.1.2

Precision

Precision evaluates the proportion of correctly predicted positive instances out of all predicted positives. In the contextoffakenewsdetection,itreflectshowmanyarticles identifiedasfakeareactuallyfake.Highprecisionindicatesa low false positive rate, which is crucial in avoiding the misclassificationofgenuinenews.

4.1.3 Recall

Recall measures the proportion of correctly identified positive instances out of all actual positive samples. It is particularlyimportantinfakenewsdetection,asfailingto detectfakecontentcanhaveseriousconsequences.Ahigh recall value ensures that most fake news instances are successfullyidentified.

Metric Description Importance in Fake News Detection

Accuracy Overall correctness General performance indicator

Precision Truepositives/ predicted positives Reducesfalse alarms

Recall Truepositives/ actualpositives Detectsmaximum fakenews

F1-score

ROC-AUC

Balanceof precision&recall Handlesimbalance

Classification separability Robustevaluation

4.2 Comparative Analysis

4.2.1 ML vs DL vs Transformer

Models

A comparative evaluation of machine learning (ML), deep learning (DL), and transformer-based models reveals significant differences in performance. Traditional ML models such as Logistic Regression and Support Vector Machines provide stable baseline results with lower computational requirements. However, their reliance on manual feature engineering limits their ability to capture complexlinguisticpatterns.

Deep learning models, including CNN and LSTM, demonstrate improved performance by automatically learninghierarchicalrepresentationsoftext.Theyeffectively capture sequential and contextual information, leading to betterclassificationaccuracycomparedtoMLmodels.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

Transformer-based models, such as BERT and RoBERTa, outperformbothMLandDLapproachesduetotheirability to understand bidirectional context and semantic relationships.Thesemodelsachievesuperiorperformance across multiple evaluation metrics, making them the preferredchoiceforfakenewsdetectiontasks(Devlinetal., 2019).

Table 5: Model Performance Comparison

4.3 Results Discussion

4.3.1 Best-Performing

Model (BERT/RoBERTa)

Among all evaluated models, transformer-based architecturessuchasBERTandRoBERTademonstratethe highest performance in fake news detection tasks. Their ability to capture contextual dependencies and semantic nuances enablesthem tooutperformtraditional and deep learning models. Fine-tuning these pre-trained models on domain-specific datasets further enhances their accuracy andgeneralizationcapabilities.

4.3.2 Trade-off: Accuracy vs Computational Cost

4.2.2 Feature Comparison (TF-IDF vs Embeddings)

Featureextractionplaysacriticalroleindeterminingmodel performance.TraditionalmethodssuchasTF-IDFprovide sparserepresentationsbasedonwordfrequency,whichare effective for capturing keyword-level patterns. However, they fail to represent semantic meaning and contextual relationships.

Incontrast,wordembeddingssuchasWord2VecandGloVe providedensevectorrepresentationsthatcapturesemantic similaritiesbetweenwords.Contextualembeddings,suchas thosegeneratedbyBERT,furtherenhancethiscapabilityby considering the surrounding context of each word. As a result, models using embeddings consistently outperform those relying on TF-IDF features, particularly in detecting subtlelinguisticcuesassociatedwithfakenews(Mikolovet al.,2013).

Table 6: Feature Comparison Feature Type

Despite their superior performance, transformer models requiresubstantialcomputationalresources,includinghigh memory and processing power. In contrast, machine learning models offer faster training and lower resource consumption but at the cost of reduced accuracy. Deep learningmodelsprovideabalancebetweentheseextremes, offeringmoderateaccuracywithmanageablecomputational requirements. This trade-off must be carefully considered whenselectingmodelsforreal-worlddeployment.

4.3.3 Impact of Feature Engineering

The choice of feature extraction technique significantly influences model performance. While traditional methods such as TF-IDF provide a strong baseline, advanced embedding techniques enhance the model’s ability to understandsemanticrelationships.Contextualembeddings, in particular, contribute to substantial improvements in classification accuracy by capturing nuanced linguistic patterns. This highlights the importance of feature engineering in developing effective fake news detection systems(Penningtonetal.,2014).

5.

DISCUSSION

Thissectioninterpretstheexperimentalfindingsandplaces them within the broader context of fake news detection research.Ithighlightskeyobservations,discussespractical applications, and identifies limitations that influence the effectivenessandgeneralizabilityoftheproposedapproach.

5.1 Key Findings

5.1.1

Transformer Models Outperform Traditional ML

Theexperimentalresultsclearlyindicatethattransformerbased models, such as BERT and RoBERTa, outperform traditionalmachinelearninganddeeplearningapproaches acrossallevaluationmetrics.Thisperformancegaincanbe

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

attributedtotheattentionmechanism,whichenablesthese modelstocapturelong-rangedependenciesandcontextual relationshipswithin text.Unlikeclassical modelsthatrely heavilyonhandcraftedfeatures,transformerarchitectures learn rich semantic representations directly from data. Consequently, they demonstrate superior accuracy, precision,andF1-scoresindistinguishingbetweenrealand fake news. These findings are consistent with recent advancements in NLP, where transformer models have achieved state-of-the-art performance across various text classificationtasks(Devlinetal.,2019).

5.1.2

Contextual Understanding Improves Detection

Another significant observation is the importance of contextualunderstandinginimprovingfakenewsdetection. Models utilizing contextual embeddings outperform those relying on static representations such as TF-IDF or traditionalwordembeddings.Contextualmodelsinterpret the meaning of words based on their surrounding text, allowingthemtodetectsubtlelinguisticcues,sarcasm,and misleadingnarratives.Thiscapabilityisparticularlyvaluable in fake news detection, where deceptive content often mimics legitimate writing styles. The results demonstrate that incorporating context-aware features significantly enhancesclassificationperformanceandrobustness(Peters etal.,2018).

Table 7: Conclusion of Key Findings

Aspect Observation Impact

Model Performance Transformers outperform ML/DL Higheraccuracy andreliability

Feature Representation Contextual embeddings superior Improved semantic understanding

Detection Capability Betterhandling ofcomplextext Reduced misclassification

5.2 Practical Implications

5.2.1 Use in Social Media Moderation

Theproposedfakenewsdetectionframeworkhassignificant implications for social media platforms, where misinformation spreads rapidly. Automated detection systemscanbeintegratedintocontentmoderationpipelines to identify and flag suspicious content in real time. By leveragingNLPandmachinelearningtechniques,platforms can reduce the spread of harmful misinformation and improvetheoverallqualityofinformationsharedbyusers. Additionally,suchsystemscanassistinprioritizingcontent

for manual review, thereby optimizing moderation efforts andreducinghumanworkload(Shuetal.,2017).

5.2.2 Support for Journalists and Policymakers

Beyond social media, fake news detection systems can support journalists and policymakers in verifying informationandmakinginformeddecisions.Journalistscan usethesetoolstocross-checkfactsandidentifypotentially misleadingsources,whilepolicymakerscanleveragethemto monitor misinformation trends and design effective countermeasures. The ability to analyze large volumes of textual data in real time provides valuable insights into information ecosystems, enabling proactive responses to emergingthreats.

Table 8: Practical Applications

Domain Application Benefit

SocialMedia Content moderation Reduces misinformation spread

Journalism

Fact-checking support Enhancescredibility

Governance Policy monitoring Informeddecisionmaking

6. CONCLUSION

This research presents a comprehensive approach to fake newsdetectionbyintegratingNaturalLanguageProcessing (NLP)techniqueswithmachinelearning,deeplearning,and transformer-based models. The study systematically evaluatesmultiplefeatureextractionmethods,includingTFIDF,wordembeddings,andcontextualembeddings,toassess their effectiveness in capturing linguistic and semantic patterns. Experimental results demonstrate that transformer-basedmodels,particularlyBERTandRoBERTa, significantly outperform traditional machine learning and deep learning approaches in terms of accuracy, precision, recall, and F1-score. Their ability to capture contextual dependenciesandnuancedsemanticrelationshipsenables morereliableidentificationofmisleadingcontent.

Furthermore,thecomparativeanalysishighlightsthecritical role of feature representation in improving classification performance.Whiletraditionalstatisticalmethodsprovidea strong baseline, contextual embeddings offer superior performance by incorporating dynamic language understanding.Theproposedframeworkalsodemonstrates robustness across benchmark datasets, indicating its potentialapplicabilityinreal-worldscenarios.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

Overall, this research contributes to the advancement of automated fake news detection systems by providing a scalableandeffectivesolution.Thefindingsemphasizethe importanceofadoptingadvancedNLPtechniquestocombat misinformation.Futureenhancementscanfurtherimprove system efficiency and broaden its applicability across domains and languages, supporting efforts to maintain information integrity in digital ecosystems (Devlin et al., 2019).

7. LIMITATIONS OF THE RESEARCH

Despiteitscontributions,thisstudyhasseverallimitations. First,therelianceonbenchmarkdatasetssuchasLIARand FakeNewsNetintroducespotentialdatasetbias,whichmay affectthegeneralizabilityofthemodelstodiversereal-world scenarios(ZhouandZafarani,2019). Second, the focus on English-language data restricts the applicability of the proposedframeworkinmultilingualenvironments,limiting its global usability. Third, transformer-based models, althoughhighlyaccurate,requiresubstantialcomputational resources, making real-time deployment challenging Additionally,thestudyprimarilyfocusesontextualfeatures and does not consider multimodal data such as images or videos, which are increasingly used in misinformation campaigns. Addressing these limitations is essential for developing more robust and scalable fake news detection systems.

REFERENCES

1. Bojanowski, P., Grave, E., Joulin, A. and Mikolov, T., 2017. Enriching word vectors with subword information. Transactions of the Association for ComputationalLinguistics,5,pp.135–146.

2. Castillo, C., Mendoza, M. and Poblete, B., 2011. InformationcredibilityonTwitter.In:Proceedingsof the20thInternationalConferenceonWorldWideWeb (WWW2011).pp.675–684.

3. Chawla,N.V.,Bowyer,K.W.,Hall,L.O.andKegelmeyer, W.P.,2002.SMOTE:Syntheticminorityover-sampling technique. Journal of Artificial Intelligence Research, 16,pp.321–357.

4. Devlin,J.,Chang,M.W.,Lee,K.andToutanova,K.,2019. BERT:Pre-trainingofdeepbidirectionaltransformers forlanguageunderstanding.In:ProceedingsofNAACLHLT2019.pp.4171–4186.

5. Fawcett, T., 2006. An introduction to ROC analysis. PatternRecognitionLetters,27(8),pp.861–874.

6. Kaliyar, R.K., Goswami, A. and Narang, P., 2020. FakeBERT:Fakenewsdetectioninsocialmediawitha BERT-baseddeeplearningapproach.MultimediaTools andApplications,79,pp.1–19.

7. Kumar,S.andShah,N.,2018.Falseinformationonweb and social media: A survey. arXiv preprint arXiv:1804.08559.

8. Mikolov, T., Chen, K., Corrado, G. and Dean, J., 2013. Efficientestimationofwordrepresentationsinvector space.arXivpreprintarXiv:1301.3781.

9. Mihalcea,R.andStrapparava,C.,2010.Theliedetector: Explorationsintheautomaticrecognitionofdeceptive language.In:ProceedingsoftheACL2010Conference ShortPapers.pp.309–312.

10. Pennington, J., Socher, R. and Manning, C.D., 2014. GloVe: Global vectors for word representation. In: ProceedingsofEMNLP2014.pp.1532–1543.

11. Peters,M.E.,Neumann,M.,Iyyer,M.,Gardner,M.,Clark, C., Lee, K. and Zettlemoyer, L., 2018. Deep contextualizedwordrepresentations.In:Proceedings ofNAACL-HLT2018.pp.2227–2237.

12. Shu, K., Sliva, A., Wang, S., Tang, J. and Liu, H., 2017. Fake news detection on social media: A data mining perspective. ACM SIGKDD Explorations Newsletter, 19(1),pp.22–36.

13. Sokolova, M. and Lapalme, G., 2009. A systematic analysis of performance measures for classification tasks. Information Processing & Management, 45(4), pp.427–437.

14. Vosoughi,S.,Roy,D.andAral,S.,2018.Thespreadof true and false news online. Science, 359(6380), pp.1146–1151.

15. Wang, W.Y., 2017. “Liar, liar pants on fire”: A new benchmark dataset for fake news detection. In: ProceedingsofACL2017.pp.422–426.

16. Zhou,X.andZafarani,R.,2019.Fakenews:Asurveyof research,detectionmethods,andopportunities.ACM ComputingSurveys,53(5),pp.1–36.

17. Aphiwongsophon, S. and Chongstitvatana, P., 2018. Detectingfakenewswithmachinelearningmethod.In: Proceedings of the 15th International Conference on Electrical Engineering/Electronics, Computer, TelecommunicationsandInformationTechnology.

18. Capuano, N., Fenza, G., Loia, V. and Nota, F.D., 2023. Content-basedfakenewsdetectionwithmachineand deeplearning:Asystematicreview.Neurocomputing, 530,pp.91–103.

19. Rastogi,S.andBansal,D.,2023.Areviewonfakenews detection:Typology,timeofdetection,andtaxonomies. International Journal of Information Security, 22(1), pp.177–212.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

20. Kang,M.,Seo,J.,Park,C.andLim,H.,2022.Utilization strategyofuserengagementsinfakenewsdetection. IEEEAccess,10,pp.79516–79525.

21. Rohera, D., Jain, A., Mishra, A. and Soni, R., 2022. A taxonomyoffakenewsclassificationtechniques.IEEE Access,10,pp.30367–30394.

22. Saleh,H.,Alharbi,A.andAlsamhi,S.H.,2021.OPCNNFAKE:Optimizedconvolutionalneuralnetworkforfake newsdetection.IEEEAccess,9,pp.129471–129489.

23. Puri,R.,2024.Fakenewsdetection:Acomprehensive study on modern classification techniques. ResearchGatePreprint.

24. Sharma, U., Saran, S. and Patil, S.M., 2021. Fake news detection using machine learning algorithms. International Journal of Engineering Research & Technology.

25. Nigam, A. and Dhruv, A., 2021. Fake news detection usingNLP.InternationalJournalofAdvanceResearch, IdeasandInnovationsinTechnology.

26. Sharma, P.K., Divakar, M.S. and Lodhi, R., 2026. DetectionandclassificationoffakenewsusingNLPand deeplearning.IJRASET.

27. Liu, Y. et al., 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692.

28. Sanh, V., Debut, L., Chaumond, J. and Wolf, T., 2019. DistilBERT:AdistilledversionofBERT.arXivpreprint arXiv:1910.01108.

29. Yang,Z.etal.,2019.XLNet:Generalizedautoregressive pretrainingforlanguageunderstanding.NeurIPS.

30. Qian,F.,Gong,C.,Sharma,K.andLiu,Y.,2018.Neural userresponsegeneratorforfakenewsdetection.ACL.

31. Gong, S., Sinnott, R.O., Qi, J. and Paris, C., 2023. Fake newsdetectionthroughgraph-basedneuralnetworks: Asurvey.arXiv.

32. Mayank, M., Sharma, S. and Sharma, R., 2021. DEAPFAKED:Knowledgegraph-basedfakenewsdetection. arXiv.

33. Monti, F., Frasca, F., Eynard, D., Mannion, D. and Bronstein, M., 2019. Fake news detection on social media using geometric deep learning. arXiv preprint arXiv:1902.06673.

34. Oshikawa,R.,Qian,J.andWang,W.Y.,2018.Asurveyon natural language processing for fake news detection. arXiv.

35. Khan,J.Y.etal.,2019.Abenchmarkstudyofmachine learningmodelsforfakenewsdetection.arXiv.

36. Bondielli,A.andMarcelloni,F.,2019.Asurveyonfake news and rumour detection techniques. Information Sciences,497,pp.38–55.

37. Thorne, J. and Vlachos, A., 2018. Automated fact checking: Task formulations, methods and future directions.COLING.

38. Shu, K., Mahudeswaran, D. and Liu, H., 2019. FakeNewsTracker:Atoolforfakenewscollectionand detection.SIGKDDExplorations.

39. Zhou,X.,Wu,J.andZafarani,R.,2020.SAFE:Similarityawarefakenewsdetection.PAKDD.

40. Baly, R. et al., 2018. Integrating stance detection and factchecking.NAACL.

41. Singhal, S. et al., 2021. SpotFake+: Multimodal frameworkforfakenewsdetection.IEEETransactions onMultimedia.

2026, IRJET | Impact Factor value: 8.226 | ISO 9001:2008

Turn static files into dynamic content formats.

Create a flipbook
FAKE NEWS DETECTION USING NATURAL LANGUAGE PROCESSING AND MACHINE LEARNING by IRJET Journal - Issuu