
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Chaitali Charandas Daware1 , Prof. V. K. Shandilya2 , Dr. N. P. Mohod3
1Final Year M. Tech Student, Department of Computer Sciences and Engineering, Sipna College of Engineering and Technology Amravati, Maharashtra, India
2Professor, Department of Computer Sciences and Engineering, Sipna College of Engineering and Technology Amravati, Maharashtra, India
3Assistant Professor, Department of Computer Sciences and Engineering, Sipna College of Engineering and Technology Amravati, Maharashtra, India
Abstract - The rapid growth of artificial intelligence and deep learning has led to the creation of hyper-realistic synthetic media known as deepfakes. These pose serious challenges to security, trust, and digital integrity. While deepfakes can be used positively in areas like entertainment, education, and virtual reality, their misuse in spreading misinformation, committing identity theft, and engaging in cybercrimes has raised global concerns. Therefore, detecting manipulated images, especially of human faces, has become an urgent research priority. Early detection methods relied on handcrafted features, such as inconsistencies in facial landmarks and textures. However, these approaches often struggle against more sophisticated generation models. Convolutional Neural Networks (CNN) have become the most effective solution. They offer automated feature extraction and hierarchical learning capabilities that capture subtle spatial and frequency-domain artifacts. This review paper provides an overview of CNNbased deepfake detection techniques. It highlights benchmark datasets like Face Forensics++, Celeb-DF, and the Deepfake Detection Challenge (DFDC), as well as performance metrics commonly used to evaluate models. Recent advancements, including lightweight CNN, hybrid deep learning frameworks, and ensemble architectures, are discussed, along with challenges such as dataset bias, generalization, adversarial robustness, and real-time deployment. Furthermore, the paper outlines promising directions for future research. These include multimodal approaches, interpretable AI, and efficient models that can be deployed on edge devices. By synthesizing existing literature and experimental findings, this review emphasizes the strengths and limitations of CNN-based methods. It aims to guide researchers toward more robust, explainable, and scalable detection systems that can counter the evolving threat of deep-fake technologies.
Key Words: Deepfake detection, Convolutional Neural Networks (CNN), Generative Adversarial Networks (GAN), Autoencoders, Image forensics, Explainable AI, Adversarial robustness, and Benchmark dataset.
Intoday’sdigitalage,thetermdeepfakehasbecomeahottopicinartificialintelligenceandcybersecurity.Theword“deepfake” combinestwoterms:deepandfake.Theprefix“deep”referstodeeplearning,whichisabranchofartificialintelligence.Ituses layered neural networksto learn patterns fromlargeamounts of data automatically.“Fake” refersto the manipulatedor artificiallycreatedcontentthatismeanttolookreal.Together,deepfakesrepresenthighlyrealisticbutfabricatedimages, videos, or audio recordings generated by deep learning algorithms, like Generative Adversarial Networks (GAN) and autoencoders.Tograsptheseriousnessofdeepfakes,itiscrucialtotellapartrealandfake.Realimagesorvideosgenuinely capturepeople,events,orobjectsthroughcamerasorrecordingdeviceswithoutanymanipulation.Ontheotherhand,fake imagesorvideosaresyntheticcreationswherefacialfeatures,movements,orvoicesarealteredorreplaceddigitally.Awellmadedeepfakecanmakeitalmostimpossibleforthehumaneyetotellauthenticcontentfromfakes.Thiscreatesarisky environment where misinformation, identity theft, political manipulation, and cybercrimes can thrive. Early methods for detecting fake content used traditional machine learning (ML) techniques. These approaches relied on manually created features,suchasunusualeyeblinking,faciallandmarks,lightinginconsistencies,orcolormismatches.Whilemachinelearning madesomeprogress,itseffectivenesswaslimited.Handcraftedfeaturesoftendidnotworkwellagainstmoreadvancedforgery techniques. Deepfake generators became better at correcting simple inconsistencies. Additionally, traditional ML models neededexplicitfeaturedesignandcouldnotautomaticallylearncomplexpatternsinlargedatasets.
Thisgapwaseffectivelyfilledbydeeplearning,especiallythroughConvolutionalNeuralNetworks(CNN).Unliketraditional machine learning, deep learning does not require manual feature extraction. CNN automatically identify fine details and structures,suchaspixel-levelanomalies,textureinconsistencies,andsubtlefrequencyartifactsthatarehardforpeopletosee. Throughmultiplelayersofconvolutions,pooling,andactivationfunctions,CNNcancapturebothlocalandglobalfeaturesof images.Thismakesthemveryeffectiveattellingrealthingsfromfakefacialimages.Infact,CNN-basedmodelshaveshown muchbetterperformancethantraditionalMLmethodsinaccuracyandrobustness.Thegrowthofdeepfaketechnologyhasalso

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
beendrivenbypowerfulgenerativemodelsandlargedatasets.GenerativeAdversarialNetworks,introducedin2014,markeda significantadvancementincreatingsyntheticmedia.GANworkbyhavingtwonetworkscompete:ageneratorcreatesfake images,whileadiscriminatortriestotellthefakefromtherealones.

Fig -1: TheGenerativeAdversarialNetwork(GAN)andtheDeepfakeDetectionChallenge
Thisfigureillustratestheprimarychallengeindeepfakedetection:thedifficultyinvisuallydistinguishingbetweenanauthentic imageanditsdeepfakecounterpart.Thecomparisonhighlightshowadvanceddeeplearningmodels,likeGANs,generate hyper-realisticsyntheticcontent.Whilethehumaneyemayfailtonoticetheforgery,ConvolutionalNeuralNetworks(CNNs) aredesignedtoautomaticallyextractandclassifythesubtle,non-visualartifactsandinconsistenciesleftbythegenerative process. Over time, the generator improves and produces increasingly realistic images that challenge both humans and machines. This complicates detection, requiring advanced CNN-based systems that can keep up with new manipulation techniques.Beyondthetechnicalside,thesocietalandethicalimplicationsofdeepfakesareserious.Theythreatenprivacy, politicalstability,digitaltrust,andevennationalsecurity.Fakevideosofleaders,falseevidenceincourtcases,andmanipulated celebritycontentshowthepotentialharmwhendetectionsystemsfail.Thishighlightstheurgentneedforstrong,scalable,and real-timedetectionmethods.Inconclusion,whilemachinelearninglaidthegroundworkforinitialdetectionefforts,itisdeep learning, particularly CNN, that offers the most reliable way to combat deepfake images. By automatically learning representationsfromlargedatasets,CNNprovideunmatchedaccuracyandadaptabilityindigitalforensics.Thispaperfocuses onCNN-basedmethods,datasets,challenges,andfuturedirections,offeringathoroughreviewofdeepfakeimagedetection strategies.
Deepfakesarehighlyrealistic syntheticmediacreatedbymachinelearningmodelslikeGenerativeAdversarial Networks (GAN),autoencoders,anddiffusion-basedframeworks.Thesemethodsallowthemanipulationoffaces,speech,andexpressions withalmostperfectrealism.Thisraiseconcernsinareassuchaspolitics,journalism,andpersonalsafety.Unliketraditional videoediting,whichleavesobvioussigns,moderndeepfakegeneratorsproduceoutputsthatlookveryconvincingandarehard tospotwiththenakedeye.Theincreasingavailabilityofdeepfakecreationtoolsmakesresearchondetectionatoppriorityfor maintainingdigitaltrustandethicalmediause[1].
Researchers have studied the unique features of deepfakes to create detection methods. For instance, studies show that synthetic content often has small inconsistencies in facial areas, blinking patterns, or head movements. Additionally, compressionartifactsandblendingissuesfromthecreationprocessoffercluesfordetectionmodels.Manystudiesstressthe needtocapturebothspatialandtemporaldifferencestoensurestrengthagainsthigh-qualitymanipulations[2],[3].

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Deeplearningmethods,particularlyConvolutionalNeuralNetworks(CNN),havebecomethemainapproachfordetecting alteredmedia.CNNcanautomaticallylearndistinguishingfeaturesthatsetapartfakefacesfromrealones,doingbetterthan oldermethodsbasedonfrequencyanalysisortexturedetails.Usinglarge-scalepre-trainedmodelslikeVGG16,ResNet,and EfficientNethasbeenespeciallyhelpful,allowingresearcherstouseknowledgefromnaturalimagedatasetsfor deepfake detectiontasks[4].
Severalresearchersnotethatdeeplearningmodelsperformatahighlevelwhentrainedonstandarddatasetsbutstruggleto adapttonewmanipulations.HybridsolutionsthatmixCNNwithextracomponents,likeattentionmechanismsoradversarial traineddiscriminators,havebeensuggested.Thesemodelsnotonlyimprovedetectionaccuracybutalsomakeiteasierto understandtheresultsbyfocusingonspecificproblemsinmanipulatedfacialareas[5],[6].
CNNarethefoundationofmostdetectionframeworksbecausetheycapturefineimagedetailswell.Earlierworkshowedthat standardCNNdesignscouldspotvisualartifactsfromfaceswappingandreenactment.Newermethodsuseresidualnetworks and deeper convolutional layers to gather layered representations of facial features, boosting resistance to subtle manipulations.ResearchershavealsoinvestigatedlightweightCNNarchitecturestocutdownoncomputationaldemandsand enablereal-timeuse[7].
Furtherprogresscentresonadversarialrobustness.Sincedeepfakecreationmethodsadvancequickly,modelstrainedonone typeoffakecanoftenstruggleagainstnewtechniques.SomeresearchershavesuggestedusingCNNensemblesorprediction fusionstrategiestoimproveresilienceagainstadversarialattacks.Othershaveaddedspatiotemporalconvolutions,allowing CNNtocapturemotiondynamicsinvideosinsteadofjustrelyingonstillframes.TheseadvancementsshowhowflexibleCNNbasedframeworksareforreal-worldapplication[8],[9].
While CNN aregreatat extractingspatial features,they oftenfail tocapturethesequential relationshipsfound invideos. ResearchershavedevelopedhybridmodelsthatcombineCNNwithrecurrentnetworkslikeLongShort-TermMemory(LSTM) units.Thiscombinationhelpsdetectinconsistenciesacrossmultipleframes,especiallyregardinglipsynchronizationandfacial movements.HybridCNN-LSTMmodelshavebeenshowntoperformbetterthanonesusingonlyCNN,particularlyinvideo detectiontasks[10].
In addition to LSTMs, recent research has investigated transformer-based architectures. By mixing CNN with Vision Transformers(ViT),thesemodelscanlearnbothlocalandglobalrepresentationsoffaces.Theyshowbetterperformancein cross-datasettests,indicatinggreateradaptabilitytonewdeepfaketechniques.Researchersarguethathybridframeworksare morefutureproofsincetheybringtogethermultiplefeaturelearningmethodsintoasingledetectionsystem[11],[12].
Creatingreliabledeepfakedetectorsdependsgreatlyonhigh-qualitydatasets.PublicresourceslikeFaceForensics++andthe DeepfakeDetectionChallenge(DFDC)datasetprovidestandardbenchmarksfortrainingandevaluation.Thesedatasetsinclude millionsofmanipulatedandgenuinesamples,allowingresearcherstodevelopandcomparevariousdetectionmodelsunder controlledconditions.However, dataset biasisstill a concern because modelstrainedonone dataset maystrugglewhen appliedtoanotherduetovariationsinmanipulationmethodsorcompressionlevels[13].
Severalstudiesstresstheimportanceofcross-datasetgeneralizationforreal-worldapplications.Forexample,whileCNNbaseddetectorsperformwellonFaceForensics++,theiraccuracydropssharplyondatasetslikeDFDCthattheyhavenotseen before.Researchersbelievethatthevarietyandqualityoftrainingdataareasimportantasthedetectionmodelitself.Totackle thisissue,newdatasetsthatincludedifferentmanipulationtypes,ethnicdiversity,andreal-worldnoisefactorsarebeing proposed.Theseeffortsarecrucialforensuringscalability,strength,andfairnessindeepfakedetectionsystems[14],[15] this section (Section II) has thoroughly reviewed the foundational and contemporary techniques for deepfake detection, rangingfromtheearlysuccessesofhandcraftedfeaturestothedominanceofadvancedCNNarchitectures.Theliterature reveals a consistent evolutionary path toward models that integrate spatial and frequency domain analysis to improve robustnessandgeneralizability.Buildingonthesefindings,thesubsequentsectionsofthisreviewwillproceedasfollows: SectionIIItransitionsfromdetectionmethodstothecriticalresourcethatpowersthembydiscussingandcomparingthe differentbenchmarkdatasetsusedindeepfakeresearch.SectionIVsynthesizesthebestpracticesidentifiedinthereviewed

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
literaturetoproposeanovel,highlyeffectiveCNN-baseddetectionmethodology.Finally,SectionVshiftsthefocustoacritical analysisofthefield'scurrentdrawbacksandlimitations,pavingthewayforfutureresearchdirections.Thisstructureensuresa logicalprogressionfromthecurrentstateofthearttothecontributionsandfutureoutlookofthisreview.
Deepfakedetectionresearchdependsondatasetsthatofferlargeamountsofrealandfakesamples.Realmediacomesfrom actual recordings, while fake data is created using face-swap algorithms, GAN, or autoencoders. For CNN-based models, datasetsareessentialbecausetheyinfluencehowwellmodelscanlearnspecificsignsofforgeriesandapplythatknowledgeto newmanipulations.Datasetsvaryinsize,type,resolution,andrealism.Somefocusoncontrolled,lab-createdforgeries,like FaceForensics++,whileotherspresentdifficult,real-worldfakesgatheredfromonlineplatforms,suchasWildDeepfake.The authenticityofthefakesalsoaffectsthelevelofchallenge;earlierdatasetshadnoticeableartifacts,butmodernbenchmarks likeCeleb-DFandDFDCincludehighlyrealisticmanipulations.
Table -1: MajorDatasetsforDeepfakeImageDetection
Dataset
UADFV(2018) Video
49 real + 49 fake videos ~300 frames/video, 720p
Face Forensics++ (2019)
Video+Images
Celeb-DF (2019, v2) Video
1,000 real + 4,000 fakes ~1.8Mframes, 256×256
Autoencoder faceswaps Low
590 real + 5,639 fakes ~563,000 frames, up to 720p
Deepfakes, Face2Face, Neural Textures Medium–High
First deepfake dataset; small scale, used for early CNN testing.
Most widely used benchmark; includes multiple forgery methods and compression levels.
GAN-based faceswaps High
DF-TIMIT(2018) Video
320 real + 640 fake videos 720p
Deepfake Detection Challenge (DFDC, 2019)
Video
23,654 real + 104,500 fake ~3,000 subjects, 300 GB
Lip-sync & faceswap Medium
High-quality celebrityfakes; difficult even for human detection.
Focused on expressionand mouth manipulation; useful for lipsyncstudies.
Multiple GAN methods VeryHigh
Industry-scale dataset with diverse actors, lighting, and ethnicity coverage.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Wild Deepfake (2020) Video (in-thewild)
DeeperForensics1.0(2020) Video
707 real + 3,509 fakes Variable (collected online)
60,000 real + 60,000 fakes
Internetsourced face swaps High
>50M frames, 1080p GAN swaps + real-world perturbations VeryHigh
Real-world fakes with noise,blur,and uncontrolled conditions.
Includes environmental noise (blur, occlusion) to mimic social media.
TheproposeddeepfakeimagedetectionframeworkisbasedonConvolutionalNeuralNetworks(CNN)andisdesignedto overcomekeylimitationsreportedinexistingliterature,includinglimitedcross-datasetgeneralization,reducedrobustness underreal-worldconditions,andvulnerabilitytoevolvingmanipulationtechniques.Althoughpriorstudiesdemonstratehigh detection accuracy on controlled datasets, their performance often degrades when exposed to unseen data or advanced deepfakegenerationmodels.Toaddressthesechallenges,theproposedapproachadoptsamulti-stagedetectionpipelinethat integratesdatasetpreparation,preprocessing,multi-domainfeatureextraction,ensemble-basedclassification,interpretability, andcontinuouslearning.
TheoverallarchitectureoftheproposedsystemisillustratedinFig.2,whilethedetailedoperationalworkflowispresentedin Fig.3.

Fig.2illustratestheproposedmulti-stagedeepfakedetectionpipeline,highlightingthetransitionfromtraditionalhandcrafted feature-basedapproachestoCNN-drivenautomaticfeaturelearning.ThediagramemphasizestheabilityofCNNtocapture complexspatialandfrequency-domainartifactsthataredifficulttodetectvisually.
Thefirststageinvolvesdatasetpreparationandpreprocessing.CNN-basedmodelsrequirelarge-scale,well-labeleddatasetsfor effectivetraining.Therefore,theproposedframeworkutilizesbenchmarkdatasetssuchasFaceForensics++,Celeb-DF,andthe DeepfakeDetectionChallenge(DFDC)dataset.Toenhancegeneralizationandreal-worldapplicability,in-the-wilddatasets
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072 © 2026, IRJET | Impact Factor value: 8.315 | ISO 9001:2008

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
includingWildDeepfakeandDeeperForensics-1.0arealsoincorporated.Alldatasetsaredividedintotraining,validation,and testingsubsetswithbalanceddistributionsofrealandmanipulatedsamples.
Preprocessing isperformed toensurethat only relevant facial informationisanalyzed.Face detectionandalignmentare carried out using established tools such as OpenCV, Dlib, or RetinaFace, enabling accurate localization of facial regions. Detectedfacesareresizedtoafixedresolutionandnormalizedtominimizevariationscausedbypose,illumination,andcamera conditions.Additionalpreprocessingtechniques,includinghistogramequalizationandilluminationcorrection,areappliedto improve visual consistency. To improve robustness against real-world distortions, data augmentation strategies such as rotation,flipping,blurring,cropping,andcompressionareemployed,simulatingtransformationscommonlyintroducedduring socialmediasharing.
ThesecondstagefocusesonfacialfeatureextractionusingCNNacrossmultipledomains.ConventionalCNN-basedmethods primarilyrelyonspatialfeaturestodetectpixel-levelanomalies,blendingartifacts,andboundaryinconsistencies.However, moderndeepfakegenerationtechniquessignificantlyreducevisibledistortions,necessitatingricherandmorediscriminative featurerepresentations.
Intheproposedframework,CNNareemployedtoextractmulti-domainfacialfeatures.Spatialfeaturesarelearneddirectly fromRGBimagesthroughhierarchicalconvolutionallayers.Earlylayerscapturelow-levelfeaturessuchasedges,textures,and contours, while deeperlayersextracthigh-level semanticrepresentationsandglobal facial inconsistenciesintroduced by generative models. In addition to spatial features, frequency-domain features are extracted using Fourier Transform or DiscreteCosineTransformtechniquestoidentifyabnormalspectraldistributionsandGAN-specificfingerprints.Furthermore, noise residual features are obtained using filtering techniques such as Spatial Rich Models (SRM), which capture inconsistenciesinsensornoisepatternscommonlydisruptedinsyntheticimages.
Theintegrationofspatial,frequency,andnoise-basedfeaturesenablestheCNNtodetectbothvisibleandhiddenmanipulation artifacts,therebyimprovingdetectionaccuracyandrobustness.
ThethirdstageinvolvesclassificationusingCNNandensemblelearningstrategies.MultipledeepCNNbackbones,including XceptionNet,ResNet,andEfficientNet,areemployedtolearncomplementaryrepresentationsofmanipulatedcontent.Each architecturefocusesondistinctcharacteristicsofdeepfakeartifacts,suchastextureirregularities,frequencyanomalies,or globalstructuralinconsistencies.
Toenhancereliability,predictionsfromindividualmodelsarecombinedusingensembletechniquessuchasmajorityvoting, weightedaveraging,orstacking.Thisensemblestrategyreducesdependenceonasinglearchitecture,improvesgeneralization acrossdatasets,andincreasesrobustnessagainstunseenmanipulationtechniques. Forvideo-basedextensions,temporal modeling may be incorporated using 3D CNNs or recurrent neural networks, enabling the detection of frame-level inconsistenciessuchasunnaturaleyeblinkingorlipsynchronizationerrors.
Thefourthstagefocusesondecision-makingandinterpretability.Ratherthanprovidingonlybinaryclassificationoutputs,the systemgenerates probabilityscores representing predictionconfidence. To improvetransparency and forensicusability, interpretabilitytechniquessuchasGradient-weightedClassActivationMapping(Grad-CAM)areemployedtovisualizefacial regionscontributingtothemodel’sdecision.Theseheatmapshelpidentifymanipulatedareas,enhanceusertrust,andsupport applicationsindigitalforensicsandmediaverification.Confidencethresholdsarealsoappliedtoreducefalsepositivesin sensitivescenarios.
The final stage addresses continuous learning and deployment. As deepfake generation techniques evolve rapidly, static detectionmodelsmaybecomeoutdated.Tomitigatethisissue,theproposedframeworkincorporatesanincrementallearning mechanismthatallowsnewmanipulationsamplestobeintegratedthroughfine-tuningwithoutretrainingtheentiremodel. Thisapproachensuresadaptabilitywhileminimizingcatastrophicforgetting.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Forreal-worlddeployment,thesystemcanbeimplementedasascalableweb-basedapplicationusingframeworkssuchas FlaskorFastAPI.Inferenceoptimizationtechniques,includingONNXRuntimeorTensorRT,canbeemployedtoaccelerate processing.CloudplatformssuchasAWSorGoogleCloudenablelarge-scaledeployment,whilelightweightCNNvariantscan beoptimizedformobileandedgedevicestosupportreal-timedetection.
Fig.3presentsadetailedflowchartofthevisualdeepfakedetectionprocess,depictingthesequentialstagesfromdatainput andpreprocessingtomulti-domainfeatureextraction,classification,decision-making,andcontinuouslearning.

5. DRAWBACKS AND LIMITATIONS
5.1 Dataset Dependency
ManydetectionmodelsachievehighaccuracywhentrainedandtestedonbenchmarkdatasetslikeFaceForensics++[15]or DFDC[14].However,performanceoftendecreaseswhenappliedtounseenorreal-worlddeepfakesduetodatasetbias.Several studiesoverlookthisgeneralizationissue,whichlimitsthepracticaluseoftheirmodels[3],[5],[10].
5.2 Limited Robustness to Evolving Deepfakes
As deepfake generation techniques improve, existing detection models often struggle to identify high-quality forgeries. Methods relyingheavilyon spatial artifactsorspecific manipulation patternsmay become outdatedwithnewerGANsor diffusionmodels[2],[7],[13].Adversarialtrainedmodelscanimproverobustness[2],buttheyintroduceextracomputational demands.
5.3 Computational Complexity
DeepCNNs,3Dspatiotemporalnetworks,andhybridCNN-LSTM/Transformerarchitecturesachievebetteraccuracy,butthey require significant computational power. Lightweight models [6] lower resource usage but often compromise detection accuracy, creating a trade-off between speed and performance. This issue limits real-time deployment and edge-device applications.

International Research
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net
WhileCNNsexcelatextractingspatialfeatures,manymodelsdonotadequatelyaddresstemporalinconsistenciesinvideobaseddeepfakes.ApproacheslackingLSTMorattentionmechanismsoftenmisssubtleframe-to-framemanipulations[3],[8], [9].
Severalstudiesshowthatmodelscanbetrickedbyminorchanges,compressionartifacts,oradversarialattacks.Fewworks providewaystoensurestrongdefenseagainstcraftedattacks,leavingmodelsexposedinreal-worldsituations[2],[11].
Mostmodelsoperateasblackboxes,makingithardtounderstandtheirdecision-makingprocesses.Onlyafewstudiesinclude attentionmechanismsorvisualizationtoolstohighlightmanipulatedareas,whichlimitsforensicandregulatoryapplications [11],[12].
5.7
Mostresearchfocusesonlyonvisualcues,ignoringadditionalaudioorphysiologicalsignalsthatcouldimprovedetection. Multimodalapproachesremainunderexplored,leavingpotentialaccuracyimprovementsunexploited[4],[9].
Table -2: ComparativeAnalysisofDetectionTechniques

13 Issue: 03 | Mar 2026 www.irjet.net
Deep-faketechnologyisoneofthebiggestchallengesofthedigitalage.Itofferscreativeopportunitiesbutalsoposesserious riskstosecurity,privacy,andtrust.Detectingalteredimageshasbecomeessential,andConvolutionalNeuralNetworks(CNN) haveproventobethemosteffectivesolution.Unliketraditionalmachinelearningmethodsthatdependonhandmadefeatures, CNNlearnpatternsautomatically.Theycanspotsubtleflawsthatareoftenhardforpeopletosee.Thispaperlooksatexisting research,datasets,methods,andtechnicalprogressinCNN-baseddeepfakedetection.KeydatasetslikeFaceForensics++,CelebDF,andDFDChavebeencrucialfortrainingandevaluatingmodels.Despiteadvancements,challengespersist,suchasdataset bias,limiteddiversity,andpoorperformanceacrossdifferentareas.AnalyzingCNNarchitectureshowsthatdeeperandhybrid modelsaremorerobustbutusuallyrequiremorecomputingpower.Akeypointisinterpretability;CNNshouldnotonlytellif contentisrealorfakebutalsoexplainhowtheycometothatconclusion.Techniqueslikeheatmapscanshowalteredareas, whichbuildtrustandsupportspracticalusesinforensics,journalism,anddigitalmediaverification.Additionally,theriseof adversarialattacksandnewforgerytechniqueshighlightstheneedforsystemsthatcanadapt.Staticmodelscanquicklybecome outdated,makingongoinglearningimportant.Theproposedapproachcombinesvariousdatasets,preprocessingsteps,multidomainfeatureextraction,ensembleCNN models,andcleardecision-makinglayers.Thisbroadframeworkaimstocreate detectionsystemsthataredependable,scalable,generalizable,andtransparent.Thiswillhelpensuretrustworthinessinthefastchangingworldofsyntheticmedia.
[1] A.Badale,L.Castelino,C.Darekar,andJ.Gomes,“DeepFakeDetectionUsingNeuralNetworks,”Int.J.EngineeringResearch &Technology(IJERT),NTASU-2020Conf.Proc.,vol.9,no.3(SpecialIssue),pp.349–354,2021
[2] A.Heidari,N.J.Navimipour,H.Dag,andM.Unal,“DeepfakeDetectionUsingDeepLearningMethods:ASystematicand Comprehensive Review,” WIREs Data Mining and Knowledge Discovery, vol. 14, no. 2, e1520, Feb. 2024, doi: 10.1002/widm.1520.
[3] A.H.Soudy,O.Sayed,H.Tag-Elser,R.Ragab,S.Mohsen,T.Mostafa,A.A.Abohany,andS.O.Slim,“DeepfakeDetectionUsing ConvolutionalVisionTransformersandConvolutionalNeuralNetworks,”NeuralComputingandApplications,vol.36,pp. 19759–19775,2024,doi:10.1007/s00521-024-10181-7.
[4] A. Malik, M. Kuribayashi, S. M. Abdullahi, and A. N. Khan, “Deepfake Detection for Human Face Images and Videos: A Survey,”IEEEAccess,vol.10,pp.18757–18791,Feb.2022,doi:10.1109/ACCESS.2022.3151186.
[5] D.Samal,P.Agrawal,andV.Madaan,“DeepfakeImageDetection&ClassificationUsingConv2DNeuralNetworks,”inProc. ACI’23:WorkshoponAdvancesinComputationalIntelligenceatICAIDS2023,Hyderabad,India,Dec.2023,pp.113–122.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
[6] H.H.Nguyen,F.Fang,J.Yamagishi,andI.Echizen,“Multi-taskLearningforDetectingandSegmentingManipulatedFacial Images and Videos,” in Proc. IEEE Int. Conf. on Biometrics (ICB), Crete, Greece, Jun. 2019, pp. 1–8, doi: 10.1109/ICB45273.2019.8987362.
[7] J.C.Neves,R.Tolosana,R.Vera-Rodriguez,V.Lopes,H.Proença,andJ.Fierrez,“GANprintR:Improvedfakesandevaluation ofthestateoftheartinfacemanipulationdetection,”IEEEJournalofSelectedTopicsinSignalProcessing,earlyaccess, 2020,doi:10.1109/JSTSP.2020.3007250.
[8] K.K.R.,I.Maji,A.K.Kumar,A.N.S.,andV.Mekali,“DeepfakeImageDetectionUsingConvolutionalNeuralNetworks:A Web-BasedApproach,”Int.J.CreativeRes.Thoughts(IJCRT),vol.13,no.7,pp.771–776,Jul.2025.
[9] M.Amerini,L.Galteri,R.Caldelli,andA.DelBimbo,“DeepfakeVideoDetectionThroughOpticalFlowBasedCNN,”inProc. IEEE Int. Conf. on Computer Vision Workshops (ICCVW), Seoul, South Korea, Oct. 2019, pp. 1205–1207, doi: 10.1109/ICCVW.2019.00156.
[10] M.S.Rana,M.N.Nobi,B.Murali,andA.H.Sung,“DeepfakeDetection:ASystematicLiteratureReview,”IEEEAccess,vol. 10,pp.25493–25518,Mar.2022,doi:10.1109/ACCESS.2022.3154404.
[11] M.S.RanaandA.H.Sung,‘‘DeepfakeStack:Adeepensemble-basedlearningtechniquefordeepfakedetection,’’inProc.7th IEEEInt.Conf.CyberSecur.CloudComput.(CSCloud)/6thIEEEInt.Conf.EdgeComput.ScalableCloud(EdgeCom),New York,NY,USA,Aug.2020,pp.70–75,doi:10.1109/CSCloud-EdgeCom49738.2020.00021.
[12] M.TaebandH.Chi,“ComparisonofDeepfakeDetectionTechniquesthroughDeepLearning,”JournalofCybersecurityand Privacy,vol.2,no.1,pp.89–106,Mar.2022,doi:10.3390/jcp2010007.
[13] M.T.Jafar,M.Ababneh,M.Al-Zoube,andA.Elhassan,‘‘Forensicsandanalysisofdeepfakevideos,’’inProc.11thInt.Conf. Inf.Commun.Syst.(ICICS),Irbid,Jordan,Apr.2020,pp.053–058,doi:10.1109/ICICS49469.2020.239493.
[14] R.Durall,M.Keuper,andJ.Keuper,‘‘Watchyourup-convolution:CNNbasedgenerativedeepneuralnetworksarefailing toreproducespectraldistributions,’’inProc.IEEE/CVFConf.Comput.Vis.PatternRecognit.(CVPR),Seattle,WA,USA,Jun. 2020,pp.7887–7896,doi:10.1109/CVPR42600.2020.00791.
[15] X.Chang,J.Wu,T.Yang,andG.Feng,“DeepFakeFaceImageDetectionbasedonImprovedVGGConvolutionalNeural Network,”inProc.39thChineseControlConf.(CCC),Jul.2020,pp.7252–7257,doi:10.23919/CCC50068.2020.9189596.
[16] X.Ding,Z.Raziei,E.C.Larson,E.V.Olinick,P.Krueger,andM.Hahsler,‘‘Swappedfacedetectionusingdeeplearningand subjectiveassessment,’’EURASIPJ.Inf.Secur.,vol.2020,no.1,pp.1–12,Dec.2020,doi:10.1186/s13635-020-00109-8.
[17] X. Wang, T. Yao, S. Ding, and L. Ma,‘‘Face manipulation detection via auxiliary supervision,’’ in Neural Information Processing(ICONIP)(LectureNotesinComputerScience),vol.12532,H.Yang,K.Pasupa,A.C.Leung,J.T.Kwok,J.H.Chan, I.King,Eds.Cham,Switzerland:Springer,2020,pp.313–324,doi:10.1007/978-3-030-63830-6_27.
[18] X.Zhu,H.Wang,H.Fei,Z.Lei,andS.Z.Li,‘‘Faceforgerydetectionby3Ddecomposition,’’2020,arXiv:2011.09737.
[19] Y.Li,X.Yang,P.Sun,H.Qi,andS.Lyu,“ExposingDeepFakeVideosbyDetectingFaceWarpingArtifacts,”inProc.IEEEConf. on Computer Vision and Pattern Recognition Workshops (CVPRW), Long Beach, CA, USA, Jun. 2019, pp. 46–52, doi: 10.1109/CVPRW.2019.00012.
[20] Y.Patel,S.Tanwar,P.Bhattacharya,R.Gupta,T.Alsuwian,I.E.Davidson,andT.F.Mazibuko,“AnImprovedDenseCNN Architecture for Deepfake Image Detection,” IEEE Access, vol. 11, pp. 22081–22099, 2023, doi: 10.1109/ACCESS.2023.3251417.