
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Nagla Elhaj Babiker1, Khalid Hamid Bilal2, Magdi B. M. Amien3
1Department. of Electronics Engineering (Communications and Control), University of Gezira, College of Engineering and Technology.
2 Professor, Department of Electrical Engineering, Omdurman Islamic University, Omdurman, Sudan.
3Department of Electrical and Electronics Engineering, University of Khartoum, Khartoum, Sudan ***
Abstract - Signal classification in noisy environments is a common issue for wireless communication systems, particularly while working with modulation strategies. The complexity of signal data makes it difficult to use traditional methods that rely on simple convolutional neural network (CNN) architectures and conventional machine-learning models, particularly when operating in environments with high signal-to-noise ratios (SNRs). In addition, for real-world heterogeneous datasets, these methods often lack required generalizability. This work proposes a hybrid CNN transformer model to overcome these constraints. The proposed model performs better at classification when the SNR changes because it combines the sequential modelling capabilities of the transformer architecture with the feature extraction capabilities of the convolutional layers. It is trained using a dataset of 11 modulation schemes under varying SNRs to ensure the resilient performance in various noise settings of the model. The model performance is assessed using its accuracy and bit error rate (BER). The model outperforms standard methods by both accuracy and generalization for classification. The model performance under various settings was examined employing a confusion matrix visualization.
Key Words: Automatic modulation recognition (AMR), Deep-learning neural networks
Throughout the growth of new wireless communication technologies, the limited availability of radio spectra leads to a considerable drawback. This drawback is not a genuine deficiency of spectrum resources but rather ineffective regulatory frameworksthatassignfrequenciesin a strictandunyieldingfashion[1]. A wide range oforganizations,suchasthecivilian, government, commercial, and military organizations, share the electromagnetic spectrum. Therefore, modern wireless communicationenvironmentsrequireadynamicspectrumallocationpolicyinsteadofthefixedpolicythatiscommonlyadopted today,leadingtolowspectrumutilizationdifficulties.Cognitiveradio(CR)technologycreationletsradioschangeanduseunused frequencyresources,itiscalled“spectrumholes”or“whitespaces”.Thisworkindicatestheconceptofdynamicaccesstotheradio spectrum.Bythis,secondaryusers(SUs)canopportunisticallyaccessthefrequencyallocationtoprimaryusers(PUs)whenthey are not in use. This approach aims to enhance spectral efficiency by enabling transmissions on detected free bands, which addressestheissueofspectrumshortage[2].
Because spectrum sensing and detection are becoming more critical in spectrum monitoring, management, and secure communicationssuchas5Gcommunicationsandbeyond,IoTnetworks,andotherservices,CRhasessentialrolesandcapabilities indetectingactivePUtransmissionsovertheband.DecidingtotransmitthesensingoutcomesindicatesthateachPUtransmitters isinactiveatthisbandwithahighprobability,andCRbecomesthepromotionsolutionforscarcityandunderutilizationproblems [3].AMRisanessentialitemindigitalcommunicationsystemsandservesasacriticalelementfortheeffectivenessofCR.Itenables dynamicradioresourcemanagementbyusingreconfigurablesoftware-definedtransceivers.Thesetransceiverscanreconfigure theirtransmissionparametersbyusingtheaccessiblecommunicationresourcesintheelectromagneticenvironmenttorecognize modulationtypesofunknownsignalswithoutpreviousdata[4].ByusingAMR,CRisexpectedtoaccuratelyrecognizeorclassify themodulationstructureofthereceivedsignalrapidly,withoutanylatency.ThisprocesshelpstoidentifywhetherthereisaPUin the channel, as all PUs use a single modulation technique for transmission over the frequency channel. This allows the correspondingdemodulationprocesstobedoneonthereceivingside[5],[6].
AMRisatransitionalstagebetweensignalmodulationanddemodulation.Becauseitischallengingtodistinguishbetween numerousmodulationschemesowingtoseveralfactors,includingmultipathfading,noise,centerfrequencyoffset,andsignal

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
structure distortion caused by inadequate hardware design or crystal oscillator drifting, AMR technology is crucial before demodulatingsignalsatthereceiverandavitallinkinwirelesscommunication.Thetechniquecanautomaticallydeterminethe modulationtypeofthesignalandextracttheinformationitcontainswithoutbeingacquaintedwiththesystemspecifications.Its automationcanadvancethesignalmodulationdetectionaccuracywhilesignificantlyreducingtheneedforhumanresources, therebyenhancingtheaccuracyofsignalmodulationrecognition.Manualmodulationrecognition(MMR)hasbeenreplacedbya newautomaticmethod(AMR)inwhichMMRrequiresmanualobservationwithanoscilloscope.TheMMRmethodwillbelimited byfourtypesofinformationrequiredforthesearchoperator:theintermediatefrequency(IF)timewaveform,theaverageaswell asinstantspectrumofthesignal,sound,immediateamplitude,andbothlongrecognitionanderrorrecognitiontimes.
AMRplaysamajorroleinCR,whichenablesdynamicradioresourcemanagementbyemployingreconfigurablesoftwaredefined transceivers. These transceivers can reconfigure their transmission parameters by using obtainable communication resourcesinanelectromagneticenvironmenttoidentifytheunknownsignalsmodulationtypeswithoutpriorknowledge[7].
As in Figure 1, the AMR block consists of two components, such as signal preprocessing and classifier modules. The preprocessingmoduleestimatessynchronizationparameters,includingthefrequencyoffsetofthetimingrecovery,receivedsignal, andpower.Inthesecondpart,thereiseliminationofsignaldisturbancessuchasinterferenceidentification,resultinginenhanced performance
AMRisclassifiedintotwoclasses.ThefirstcomprisestheexistingAMRtechniques,consistingofdecision-theoreticapproachesand feature-based(FB)methods.Thesemethodsdependontheformofthesignalinput,suchassignalstatisticalapproachesand image-based methods. The likelihood-based (LB) approach is the initial decision-theoretic approach. The probability density function(PDF)coverstheAMC-LBoftheobservedwaveformalsolearnsfrommodulatedsignalsandFBmethods[8].LBrequires priorknowledgeofthePDFandthecalculationofthemaximumlikelihoodvalueforalltheproposedmodulationschemes,which addshighcomputationalcomplexity,lacksrobustnessagainstmodelmismatch,andposesachallengeforreal-timesystems.FB approachesconcentratelowcomputationalcomplexityonfeaturesextractionfromthesignaldirectly,removingtheneedforextra channelorsignaldataand,thus,loweringcomputationaldemands.Featureextractionandfeatureclassificationcanbeapplied,in whichinthefeatureextractionmodule,suddenfeatures,highordercumulant(HOC)features,andwaveletfeaturescanbeused[9]. Mostexistingwell-developedAMCmethodsareFB-based,particularlywiththeriseofmachinelearning(ML)approachesusingthe deep neural network (DNN), where different ML approaches can be exploited. The support vector machine (SVM), k-nearest neighbour(KNN),etc.areregularlyassumed.Thesemethodssignificantlydependontheextractionandanalysisofsignalfeatures [10],[11],[12],[13].
ThesecondandmostmodernclassisadvancedMLanddeeplearning(DL)methods,whichhavedrawnincreasingcarefor improved spectrum-sensing detection, precision, and accuracy. DL with artificial neurons organized in a stacked multilayer architectureintroducesinnovativetechniquesthatsignificantlyimprovetheefficiencyofmodernwirelesscommunicationsystems. SeveralDLalgorithms,plusconvolutionalneuralnetworks(CNNs),DNN,alongwithlongshort-termmemory(LSTM)[14],[15], [16],[17],aswellastheirhybrids,demonstratesignificantadvantagesinaddressingchallengeslikespectrumsensingandAMRin CR.Thisimprovementwasdonebyaccuratelyevaluatingcriticalfeaturesofthereceivedsignals,suchasmodulationtype,which canbeenhancedtoenableprecisepredictionsofchannelavailability.Recently,DLtransformer-basedmodelshavebeendeveloped, introducinganewrecognitionalgorithmforreal-timedecision-makingandleveragingself-attentiontoreducecomplexityconcerns in the AMR of signals, thereby achieving accurate modulation classification under practical channel conditions. Recently, transformer-basedmodelsexploreforreal-timedecision-making,leveragingself-attentiontohandlecomplexspectrum-allocation tasks.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
SpectrumsensingisacorefunctionofCRNs;however,severallimitations,suchasnoiseandtherequirementforpreceding informationofsignalfeatures,causesittofailinenvironmentswithlowsignal-to-noiseratio(SNR).Inaddition,recognizingthe type of modulation is critical to intelligent spectrum access and interference management. Some existing spectrum-sensing methods,forinstanceenergydetectionandmatchedfiltering,combineAMRwithspectrumsensing.AMRinvolvesclassifyingthe modulationschemes(e.g.,quadraturephase-shiftkeying,binaryphase-shiftkeying(BPSK),16-QAM,frequencymodulation,and amplitude modulation (AM)) used by the transmitted signal. These schemes can significantly improve the performance and robustnessofspectrumsensing,particularlyunderreal-worldconditionswithvaryingnoiselevels,fading,andsignaldistortions. Becausethenecessityforautonomousmodulationrecognitionrisesinwirelesssystems,wheremodulationschemesarelikelyto alterregularlyastheenvironmentchanges,DNNareconsideredasanewandpowerfulmethodforAMR,overcomingthereliance onexpertanalysisandtheinabilitytoscalewiththeincreasingsignalcomplexityexhibitedbytraditionalmethods.Thecoreissue thisworkaddressisthedesignanddevelopmentofanaccurate,noise-resilient,andcomputationallyefficientDL-basedframework forAMRtoenhancethespectrumsensingperformanceinCRnetworks.
1. Thispaperpresentsadata-driven,adaptive,andscalablesolutionthatenhancestheCRNsintelligenceandreliability.Thework is done by combining signal classification and spectrum detection capabilities in a CNN-transformer model to improve modulationrecognition,mainlywhileworkinginlow-SNRenvironments.
2. Theproposedframeworktrainsandvalidates11modulationschemesundervariousSNRconditionstoachieverobusttraining acrossnoiselevels,whilecomparingthemwithpreviousmethodsthatsignificantlydeclineatlowSNRs.
3. Aspertheexperimentalevaluationsusingaccuracy,BER,andconfusionmatrixanalysis,theproposedhybridmodelperforms betterinbothclassificationaccuracyandgeneralizationabilitythantraditionalCNN-basedandclassicalMLapproachesto improvedgeneralizationandaccuracy.
Thisstudyoffersausefulpathforcreatingnoise-resilientmodulationrecognitionsystemsapplicabletoreal-worldsystemsand contemporarywirelessnetworksbyfusingsequentialmodellingwithfeatureextractioninalightweighthybridarchitecture-based CR. Thoughtheexistingworkshavesomeadvantages,theseDL-basedAMRmethodsoftenfailunderlowSNRenvironmentsowing tolimitedglobalfeaturemodelingandover-dependenceonlocalpatterns.Thisstudypurposestoworkwiththeseexistinggapsby proposingahybridCNN-Transformermodel,whichenhancesrobustness.
2. Literature Review
AMRisanessentialtechniqueinCRnetworksandpossessestheorskilltoclassifythereceivedsignalmodulationwithoutpast informationofthemodulationscheme.Itisablindorsemi-blindmethodthatenablesSUsorCRreceiverstosense,classify,and adaptwithoutpriorknowledgeorcoordinationtoidentifythemodulationtypesofunknownsignals.Suchanapproachhelpswith intelligentspectrumsensing,robustness,andadaptivity,andreducesthesignallingoverheadrequiredinCRnetworks[18].
DLnetworkarchitecturestargetmodulationrecognitionalgorithmstoenhancethereliability,simplicity,andeffectivenessofAIbasedAMCmodelsinwirelesscommunicationapplications.TheresultsrevealthatDLarchitecturescansignificantlyoutperform traditionalmethods,providingastrongfoundationforfutureresearchinthisarea[19],[20].
Several DL techniques based on CNN models are presented for recognition of modulation techniques. In [21], Mohsen 2024 designedtwoCNNmodelsfortheautomaticrecognitionofmodulationmethods:onebasedontheRadioML2016.10adatasetand theotherusinganimagedataset.Thehyperparametermodificationenabledbothmodelstoachievehighvalidationandtesting accuracies. DL models,particularlyCNNsandRNNs,canlearndiscriminative representationsdirectlyfromrawI/Qsamples establishedtoavoidtheneedformanualorhandcraftedsignalfeatures,aresuitableforextractinglocalwaveformpatterns,andare extensivelyusedtoimprovetheaccuracyofAMR,especiallyinlowSNRenvironments.However,capturingglobalfeaturesand mitigatingirrelevantfactorsremainchallengingforimprovingtheAMRperformance.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Hybrid networks have been introduced to improve AMRs by effectively capturing global features and enhancing the model generalization. Several studies [22] have suggested hybrid designs that take advantage of the complementary capabilities of transformers(globalcontextmodeling)andCNNs(localfeatureextraction)andtheimportanceofintegratingtransformersand convolutionalnetworkstoincreaseclassificationaccuracyinwirelesscommunicationsatdifferentSNRs.Byutilizingbothdesigns advantages,thismethodprovidesreliableperformanceinsettingswithbothlow-andhigh-SNRsignals.
BasedonthecombinationbetweenthetransformerandLSTM(TLDNN)framework[23],QuimprovedthemainchallengeofAMR, whichiscapturingglobalfeaturesandenhancingmodelgeneralization,byproposingadataaugmentationstrategycalledsegment substitution(SS).TheSSenhancesrobustnessbyalteringpartsofsignalstoforcethenetworktorelymoreonglobalconstellation patternsandlessonspuriousfeaturesorchannelartifactsandmitigatetheimpactofirrelevantinherentvariables,suchasRF fingerprintcharacteristicsandchannelcharacteristics,andsolvetheprimaryproblemofAMR.Ahybridconvolutionaltransformer classifier(HCTC)hasbeen designedtoclassifyunknown signals[24].TheHCTC model employs a three-stageframework for featuresextractionfromin-phase/quadraturesignalsusingaconvolutionallayer,atransformerlayer,aswellasfeaturemapping. Whencomparing,theHCTCmodelgivessuperiorperformancebymaintaininghighaverageaccuraciesacrosstheSNRrange.This demonstrates complete robustness and reliability across various practical noise environments. This model shows good classificationaccuracyandworksathighSNRlevels.However,itslowsdownatverylowSNRlevelsfrom0dBandlower,and henceitislimitedinnoisycommunication.
ThelightweightradiotransformermethodforAMCwasusedin[25],leveragingbothlarge-scaleandsmaller-scaleRadioML 2018.01AdatasetstocomprehensivelyevaluatetheperformanceofMobileRaT,consideringvariousmodulationschemesand realisticcommunicationconditions.Althoughthemodelisuseful,itdoesnotdirectlyaddresstheadvantagesoftheCNNfrontends formaintainingthelocalwaveformstructure.Becauseitismostlytransformer-centricratherthanacompleteCNN–transformer hybrid,whichlimitsthefront-endadvantagesoftheCNNarchitecture.Theresearchersin[26]generatedasolutionforAMCnoise, NMformer,basedonavisiontransformer(ViT),whichissuitableforcomplexfeaturerelationshipsandimagereconstruction. Constellationdiagramshavebeengeneratedfrommodulatedsignals,convertingthesignalinformationintoa2-Drepresentation, whichachieveshighaccuracyacrossvariousSNRsanddemonstratesstrongresiliencetoout-of-distributiondata,outperforming baselineclassifiers,comparedwithtraditionalCNN-basedapproachesthatfocusonlocalfeatures.
TraditionalCNN-basedapproachesin[21]and[22]workwellincapturinglocalfeaturesbutoftenfailtogeneralizeinlowSNR environments owing to their limited capability to model long-range dependencies. Alternatively, transformer-based models [24][25][26], while are effective at global feature extraction, commonly overlook localized signal characteristics, which are essentialformodulationrecognition.RecenthybriddesignslikeHCTC[24]andNMformer[26]workwiththisgapbuteitherlack effective local feature encoding or exhibit instability at very low SNRs. The model proposed in this work addresses these shortcomingsbyjoiningconvolutionallayersforrobustfeatureextractionwithTransformerlayersforsequencemodeling.This strategyhelpsinachievinghighclassificationaccuracyandnoiseresilienceacrossabroadSNRspectrum.Thisdual-capability frameworkenhancesmodulationrecognitionperformancewherebothlocalandglobalsignalattributesareimportant
3. Proposed Methodology

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
ThisworkutilizestheRML2016.10adataset,forverifyingtheresults.TheRML2016.10adatasetcomprises20SNRsthatranges from-20to18dBfor22,000samplesintotalacross11modulationschemes.TheseschemesincludeWBFM,4PAM,AM-SSB,BPSK, AM-DSB,16QAM,CPFSK,GFSK,QPSK,64QAM,and8PSK.ThesamplescountofeverymodulationmodeunderallSNRconditionis 1000.Also,thedataformatofallthesamplesis128×2,where128signifiesthelengthofthesignal.Theproposedarchitectureis showninfigure2.
Datasplittingistheprocessofseparatingtheinputdatasetintotwodistinctsubsets,suchasthetraining,andtestsets.Thegoal istoletthemodelbetrainedonasubsetofthedata,checkedonanothersubsettochangehyperparameters,andestimatedona finalsubset(testset)toguesshowwellitwilldoondatathatithasn'tseenbefore.Itiscommonpracticetouse80%ofthedataset fortrainingpurposes.Themodelisfitted,anditsweightsareadjustedusingthisdata.Byadjustinghyperparametersincludingthe batchsizeandlearningrate,thevalidationsetfacilitatesmodelselection.Thisworkuses20%ofthedatafortesting.Thereisno biasinthetestfindingssincethemodeldoesnotviewthisdatawhentraining.
Normalizationandone-hotencodingareimportantprocessestopreparethedataforDLtrainingandencoding.Scalingthe featuressuchthattheyhavecomparableranges,usuallybetween0and1or-1and1,iswhatnormalizationdoes.Thisisespecially importantforneuralnetworksbecauseitmakessurethatthemodeltreatsallfeaturesthesame.Thisway,problemsdon'thappen wheresomefeaturesdominatethelearningprocessbecausetheyarebiggerinequation(1). (1)

Where isthemeanand isthestandarddeviationofthefeature.Abinarymatrixmaybegeneratedfromcategoricallabels usingtheone-hotencodingprocess.EachoftheNclassesinthedatasetisrepresentedbyavectoroflengthN,exceptfortheclass's index,whichissetto1.Allothervaluesare0.
Ifthedataissortedbyclassortime,forexample,shufflingisconductedtopreventanypossiblebiasesinthetrainingprocess. To avoid the model from learning from data order effects, the dataset may be rearranged at random. Data augmentation approachesartificiallyenlargethedatasetbycreatingvariantsofthealready-existingdata.Thismightincludedatamodifications likeflipping,rotating,orintroducingnoise,whichenhancesthemodel'sgeneralizationinequation(2)
Tasksinvolvingsignalprocessingmayincludedoingthingslikeintroducingrandomnoisetotheinputsignalsoradjustingthe signalintensity.
Forsomesignalprocessingwork,thehybridCNN-transformermodelworkswithDLarchitectures.Thisworkmakesuseofthe bestfeaturesoftwodifferentnetworktypes,suchasthetransformermodelsandtheCNNs.
Toextractlocalfeaturesfrominputdata,CNNsmakesuseoftheconvolutionallayers.CNNsareconsideredaviableoptionfor signal detectionsince they detectedgesandotherhigh-frequencypartsthatshow importantlocal patternsinthewaysignal representationisdoneintermsoftimeandfrequency.Inaconvolutionallayer,theinputdatapassesthroughafilters,orkernels

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
set.Eachfilterlooksforvariousfeatures,includingtextures,edges,orfrequencypatternsinsignalprocessing.A2Dconvolution operation’smathematicalformulationisinequation(3):


(3)
Where istheoutput, istheinputdata, isthekernel/filter,and and arethekerneldimensions.Forclassifying themodulations,itisnecessarytoremovethefrequencycomponentsandphaseshiftsfromthesignals.Additionally,theselayers alsoworkwellwithothersignalprocessingtasksrelatedtothiswork.
Transformersarewell-suitedfortime-seriesdata,suchassignaldata,sincetheyareusedtodescribesequentialrelationships in data. Transformer layers in the hybrid CNN-transformer model can identify long-range relationships in signal data by concentrating on diverse input sequence parts. Mechanisms for self-attention are the basis of the transformer architecture. Prioritizing one part of a sequence above another is the main point. Here is the definition of the self-attention method as in equation(4) (4)


Where isthequerymatrix, isthekeymatrix, isthevaluematrix,and isthekeyvectordimension.Self-attention mechanismsmayhelpmodelsbyfocusingonsignalfeatures,suchascertainfrequencycomponents,thatmaybeimportantfor classifying
Figure2depictsthelayerarrangementoftheproposedhybridCNN-transformermodel.Thearchitecturecontainsnumerous vitalanddiverselayers,eachperformingauniquerole.Theinputtotheinputlayerisanissue-specific1Dor2Dsignal,suchas time-seriesorfrequency-domaindata.Theformofasignalconsistsofitslengthandthenumberofchannels(e.g.,1forunivariate dataandmoreformultivariatedata),whicharerepresentedasbatchsize,sequencelength,andNumchannels,respectively.Onthe otherhand,convolutionallayersextractlocalizedfeaturesandpatternsfromtheinputdata.
Theconvolutionlayerlearnstheimportantfeatures,likefrequencycomponents,edges,orshiftsfromtheinputsignal.Asingle convolutionalblockcomprisesvariouslayers,whichincludetheconvolutional,max-pooling,anddropoutlayers.
TheCNNnetworkutilizesfilterstogeneratefeaturemaps.Themaximumpoolingmethodhelpsintheidentificationofthe highestvalueintheareaofinterest.Thismethodkeepsthefeaturesthatareconsideredmostusefulalongwithdimensionality reduction.Bychangingaunit'sportionoftheinputto0,thedropoutlayerpreventsoverfittingduringthetrainingprocess
FollowingtheCNN layer, thiswork usesa flattenlayer, whichconverts the outputtoa 1Dvectorandpasses it onto the transformerlayers.ThischangeturnsthefeaturemapsfromtheCNNlayersintoasequencethatthetransformeraccepts.Thefour layersinthisnetworkincludeafeedforwardnetwork,multi-headself-attention,normalization,andadropoutlayer.Toidentify attentionscoresandhighlighttheimportantfeatures,multi-headself-attentionusesself-attentionontheinputsequence.Twofully connectedlayersinafeedforwardnetworkapplyaReLUactivationfunctiontotheattentionoutputs. Thedropoutlayerensures constanttrainingbystandardizingtheoutputofeverytransformerblock.Byrandomlydeactivatingneurons,thedropoutlayer preventsoverfitting.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Input Layer
[batch_size, sequence length, num channels]
Extract local features from signal
Flatten Layer
[batch_size, flattened length]
Block
(Attention, Feed-Forward, Norm, Drop)
Captures long-range dependencies
Global Average Pool
[batch_size, num features]
Dense Layer
[batch_size, num classes]
Output Layer
Predicted class probabilities
Aglobalaveragepoolinglayerhelpssummarizethelearningfeaturesinthesequence.Thisprocedureisdoneafterpassing throughthetransformerlayersbyreducingthesequencedimensionandoutputtingasinglevalueperfeature.Theoutputshapeof thistransformerlayeris(batch_size,num_features). Adenselayerisinsertedafterthetransformerblocktomapthelearned featurestotheoutputclasses.Thislayerisresponsibleforcarryingoutthelastchange.Inmostcases,thisworkintroducesnonlinearityusingReLUoranequivalentactivationfunction.Asalaststep,theoutputlayerusessoftmaxasanactivationfunction The finalresultforclassificationisgotfromthislayer.
3.3.
TheMLmodellearnsfromthedatathroughthemodeltrainingphase.Thisworkusesmultiplemethodstoconfirmthatthe modeltrainsandgeneralizeswell.

3.3.1.
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Early stopping is meant to avoid the overfitting issue. By this method the training process ends once the validation loss improvesfora prearranged epochcount. After the model hasachieved excellentgeneralization,this preventsit from further learningnoisefromthetrainingdata.
3.3.2.
ThisworkusesatoolforlearningrateschedulingcalledReduceLROnPlateau,whichchangesthelearningrateduringtraining. Thismethodologyensuresthatthemodelisfine-tunedasitgrowsclosertoanoptimalstatebydecreasingthelearningratewhen thevalidationperformanceplateausinequation(5).
(5)
Thefactorisusuallylessthan1,andinthiswork,itis0.5.
3.3.3.
Duringtraining,theproposedmodelkeepsanoteofitscheckpointsforsomeregularintervals,usuallywhenthevalidationloss islow.Thetrainingprocessmayresumefromtheoptimalconditionifoverfittingoccursorinterruptionisnecessary.
4. Results
Thissectiondiscussestheadvantagesoftheproposedapproachanditsimplementation.Thisstudyimplementationisdonein Pythonbyusingthewell-knownlibrarieslikescikit-learn,TensorFlow,andKeras.Theselibrariesserveasbuildingblocksfor creation,training,andassessmentoftheDLmodels.TheproposedhybridCNN-transformermodelcombinesthebestfeaturesof convolutionallayersforfeatureextractionalsotransformerlayersforsequencemodelling.Thismodelisalsosuitableforrealworldapplicationsincommunicationsystems,asitworksbetterthanexistingworksingeneralizationandaccuracy.Furthermore, theproposedworkmanagesnoisydataandlong-rangerelationshipsinsignals.
Table – 1: ModelArchitectureTable Layer
InputLayer (None,128,2,1)
Conv2D (None,128,2, 64)
MaxPooling2D (None,64,2,64)
Conv2D (None,64,2, 128)
MaxPooling2D (None,32,2, 128)
Flatten (None,8192)
Reshape (None,128,64)
MultiHeadAttention (None,128,64) 33,216
LayerNormalization (None,128,64)
Flatten (None,8192)

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Dense (None,128)
Dropout (None,128) 0
Dense (None,11) 1,419
Table1providestheoutputshapeandtheparametersusedinthiswork.Eachlayerdeterminesthefinalproduct'ssize.The firstparameter,None,intheinputlayerrepresentsthedynamicbatchsize.Theoutputtensor'swidth,height,andchanneldepth aretheremainingvaluesintheinputlayer.Therequiredparametersneededforlayeroperations,suchasweightsandbiases,add uptothetotalparameters.Thismodelhandlesfeatureextraction,sequencemodelling,andclassificationtasks.Theinputlayer analyzes the data first, then the CNN layers extract spatial features, and finally the transformer layers capture long-range relationships.
Fromthetransformerprocess,afterflatteningtheoutput,isamulti-headattentionlayerthatreshapesthedataandpassesit on.Thenthenormalizationlayerstabilizesthelearningprocess.Lastly,decision-makingandclassificationaredonebydenselayers, which include dropout layers to avoid overfitting. The final output layer predicts the class labels using a softmax activation function.
4.1. Testing and Evaluation
Oncetrained,itiscrucialtomeasurethemodel'sperformanceonunseentestdatatodetermineitsgeneralizationcapabilities. Theproportionofaccuratepredictionstoallpredictionsisknownasaccuracy.Theformulaisinequation(6): (6)
BERcalculatestherateofmisclassifiedbitsasfollowsinequation(7): (7)

Bycomparingthepredictedlabelswiththetruelabels,aconfusionmatrixdisplaystheclassificationmodelperformance.It revealsthekindsoferrorsthemodelcommits.
4.2. Comparison
TheCNNmodel,DNN,andtheexistingCLDNNmodelarethethreeadditionalmodels.Tables2,3,and4describethestructure oflayersintheseexistingmodels.
Table -2: ExistingCLDNNArchitecture
Layers
Convolution(Conv2d)
Input Parameters
Customactivation
MaxPooling2D poolsize=(1,2)
Dropout(rate) 0.3
Conv2D
Customactivation
MaxPooling2D poolsize=(1,2)

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Dropout(rate)
Dense
128hiddenlayer,Custom activation
Dropout(rate) 0.3
Reshape (2,4096) LSTM Customactivation
Dropout(rate)
Dense
Dropout(rate)
128hiddenlayer,Custom activation
Dense 11hiddenlayer,softmax
Table -3: ExistingCNNArchitecture Layers Input Parameters
Convolution(Conv2d)
Relu
BatchNormalization -
MaxPooling2D pool_size=(1,2)
Dropout(rate)
Conv2D
0.3
Relu
BatchNormalization -
MaxPooling2D poolsize=(1,2)
Dropout(rate) 0.3
Convolution(Conv2d)
Relu
BatchNormalization -
MaxPooling2D pool_size=(1,2)
Dropout(rate) 0.3
Convolution(Conv2d)
Relu
BatchNormalization -
MaxPooling2D poolsize=(1,2)
Dropout(rate) 0.3
2026, IRJET | Impact Factor value: 8.315 | ISO 9001:2008

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Flatten -
Dense 128hiddenlayer,Relu
BatchNormalization -
Dense 11hiddenlayers,SoftMax
Table – 4: ExistingDNNArchitecture
Layers
Dense
Input Parameters
256hiddenlayers,Relu
Dense 128hiddenlayers,Relu
Dense 250hiddenlayers,Relu
Dense 64hiddenlayer,Relu
Flatten -
Dense 128hiddenlayer,Relu
Dense 11hiddenlayer,softmax
Table - 5: Proposedaccuracy,loss,validationlossandvalidationaccuracy

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072


Chart – 4: ValidationLossComparison
Table – 6: Proposedotherperformancemeasures

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

– 6: BERcomparison

– 7: PrecisionComparison

– 8: RecallComparison

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Theabovefiguresfrom3to9showtheevaluationresultsofseveralMLmodelsusingfiveimportantmetrics,whichinclude precision, recall, accuracy, F1 score, and BER. All the metric calculations are done for diverse SNR values in the dataset. The proposedmodelhasthebestaccuracyofallthemodelsatSNRsof18,16,and14dB.Theproposedmodelgets0.7830accuracyat SNR=18dBandkeepsthataccuracyprettyhighevenasthenoiselevelgoesup,whichismuchbetterthanCLDNN,CNN,andDNN. However,sinceithassuchachallengingtimedealingwithnoise,theCNNmodelconsistentlyshowsverypooraccuracy,hovering around0.1.Theresultshowsthattheproposedmodelisbetteratwithstandingnoiseandalsokeepsitsaccuracyhighevenin difficultenvironments.
Thegeneralizabilityoftheproposedmodelisshownbyusingthevalidationlossmetric.WhencomparingtheexistingCNNand DNNmodels,theproposedmodelshowsalowandconstantloss.WhileexistingmodelsshowsharplinesinlossathigherSNR levels,theproposedmodel'slossisconstant.At18dBSNR,thelossfortheproposedmodelis0.4861;however,CNN'slossis 4.7325.ThissuggeststhatCNNstruggleswithhighSNRdata,mostlikelybecauseofitslackofgeneralizability.Thus,theproposed modeloutperformsotherexistingworks,especiallyundersettingswithgreaterSNRs.Theproposedmodelidentifiessignalswitha probabilityof0.7645at18dBSNR.WithlowerSNRvalues,everymodelshowsareductionindetectionlikelihoodasthereisarise innoiselevels.Buttheproposedmodelmaintainshigherdetectionrates,whichindicateshighnoiseflexibility.
WithalowerBER,theproposedmodelperformsbettersinceitassessestherateoferrorbitsinitspredictionsincomparison totheexistingmodels.TheproposedmodelhasaBERof0.235forSNR=18dB,whileCNNhasaveryhighBERof5.8593.This showstheproposedmodel'sabilitytoconveymoretrustworthydata,withoutconsiderationofthenoiselevels.Theperformance metricsshowthetrade-offbetweencorrectlyidentifyingpositivesamples(precision)andgettingasmanytruepositives.Thisis shownbytheproposedmodel'shighF1scoresincomparisontotheexistingCNNalongwithDNN.Theproposedmodelismuch betterthanCNNandDNNatdealingwithunbalancedclassesornoisydata.ItsF1scoreisaround0.77atSNR=18dB,whichis muchhigherthantheirscoresofalmostzero.
4.3.
Theproposedmodelshowsbetteroutcomesthantheexistingmodels(CLDNN,CNN,andDNN)onallimportantcriteria.This advantagemakesitthecleargoodinsituationswithdifferentamountsofnoise.EvenatlowerSNRlevels,theproposedmodel keeps validation accuracy high, loss low, and signal detection efficient. The proposed model maintains a more constant performanceatlowerSNRlevels,whileCNN'sperformancedeclinessignificantly.Thisstudyshowsthattheproposedmodelhas beenmodifiedtobemoreflexibletonoiseandisperfectforreal-worldapplicationswheresignaldeteriorationisprevalent.Asthe validationlossoftheproposedmodelislowandconstantoverarangeofSNRvalues,itappearstobeabletogeneralizeverywell withnewdata.Alternatively,CNNandDNNmodelsshowalotoflosswhenthere'salotofnoise,whichmeanstheyeitheroverfitor don'tgeneralizewell.Theproposedmodelisthepreferredalternativewhenadaptingtodifferentnoisesituations.
TheproposedmodeldoesanimprovedjobthantheCNNandDNNmodelsbecauseitcanbeusedrepeatedlyandisaccurateat findingpositivecases while reducingfalsepositivesand negatives.Foractivitieswheretheproperidentificationofsignalsis

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
critical,suchasincommunicationsystemsordetectiontasksinnoisysettings,thisisessential.Incommunicationsystemswhere theintegrityofthesentdataiscrucial,theproposedmodelconsistentlyshowsreducedbiterrorrates(BER).Whencomparedto CNNandDNN,whichhavemuchhigherBERvaluesandarethereforelessreliable,theproposedmodelhastheuniqueabilityto lowerBERevenwhenthereisalotofnoise.
Theproposedmodeldoesagoodjobwithdatasetsthataren'tbalancedornoisysituationswherebothfalsepositivesaswellas falsenegativesareimportant,asintheF1score,whichcomparesaccuracyandrecall.Becauseofitssuperioroverallperformance, theproposedmodelismoresuitedforreal-worldapplicationsthatneedprecisionanddependability,asshownbyitshighF1score. Duetoitstolerancetonoise,superiorgeneralization,increasedaccuracyandrecall,anddecreasederrorrates,theproposedmodel surpassestheexistingmodels.Itsabilitytohandledifferentamountsofnoisewithoutsacrificingperformancemakesitamore dependableoptionforreal-worldapplications.
Thehybridmodel’ssuperiorresultsqualifiedtoitsabilitytojointlyexploitlocalpatternsviaCNNsandcontextualdependenciesvia Transformerlayers.Thisallowsmorerobustmodulationclassificationevenunderseverenoise,unlikeclassicalmodelswhich strugglewithlong-rangedependencies.





International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072


5. Conclusion
ThisstudyintroducedahybridCNN-TransformermodelforAMRinwirelesscommunicationsystemsoperatingundervarying noise conditions. The proposed model successfully overcomes the drawbacks of conventional CNNs and classical ML techniques, particularly in high SNR amplitude settings, by combining the long-range dependency modelling abilities of transformer architectures with the local feature extraction capabilities of convolutional layers. Its robustness is further demonstratedbytheconfusionmatrixanalysis,particularlyinaccuratelyclassifyingcomplexmodulationtypeseveninnoisy environments.Insummary,thisproposedhybridmodel,evaluatedonadatasetcomprisingelevenmodulationschemesacross awideSNRrange,demonstratesimprovedclassificationaccuracy,generalization,andnoiserobustness.Performancemetrics, includingconfusionmatricesandBER,confirmitssuperiorityoverbaselinemethods,particularlyindifferentiatingclosely spacedmodulationtypes.Makingitsuitableforreal-worldapplicationssuchasCR,software-definedradios,andadaptive communicationsystems.
[1] A.S.JamwalandG.Kaur,“CognitiveRadio:AnEmergingTrendforBetterSpectrumUtilization,”IJCATR,vol.2,no.3,pp. 229–231,May2013.
[2] R.Zhang,Y.-C.Liang,andS.Cui,“DynamicResourceAllocationinCognitiveRadioNetworks:AConvexOptimization Perspective,”IEEESignalProcessingMagazine,vol.27,no.3,pp.102–114,May2010.
[3] M.U.MuzaffarandR.Sharqi,“AReviewofSpectrumSensinginModernCognitiveRadioNetworks,”Telecommunication Systems,vol.85,no.2,pp.347–363,Feb.2024.
[4] E.M.Ali,G.M.Salama,K.A.A.,andM.Ezz-Eldin,“AutomaticModulationClassificationforEnhancedCognitiveRadiofor IoTSystemsBasedonDeepLearning,”JournalofAdvancedEngineeringTrends,vol.44,no.1,pp.0–0,Jan.2025.
[5] B.Tang,Y.Tu,Z.Zhang,andY.Lin,“DigitalSignalModulationClassificationwithDataAugmentationUsingGenerative AdversarialNetsinCognitiveRadioNetworks,”IEEEAccess,vol.6,pp.15713–15722,2018.
[6] Q. Zheng, X. Tian, L. Yu, A. Elhanashi, and S. Saponara, “Recent Advances in Automatic Modulation Classification Technology:Methods,Results,andProspects,”InternationalJournalofIntelligentSystems,vol.2025,no.1,p.4067323, Jan.2025.
[7] B.Jdid,K.Hassan,I.Dayoub,W.H.Lim,andM.Mokayef,“MachineLearning-BasedAutomaticModulationRecognitionfor WirelessCommunications:AComprehensiveSurvey,”IEEEAccess,vol.9,pp.57851–57873,2021.
[8] X. Liu, C. J. Li, C. T. Jin, and P. H. W. Leong, “Wireless Signal Representation Techniques for Automatic Modulation Classification,”IEEEAccess,vol.10,pp.84166–84187,2022.
[9] O. A. Dobre, A. Abdi, Y. Bar-Ness, and W. Su, “A Survey of Automatic Modulation Classification Techniques: Classical ApproachesandNewTrends.”

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
[10] A.R.,L.C.,C.C.,J.C.W.A.,andA.B.R.,“ModulationClassificationinCognitiveRadio,”inFoundationsofCognitiveRadio Systems,S.Cheng,Ed.InTech,2012.
[11] K.Tekbıyık,A.R.Ekti,A.Görçin,G.K.Kurt,andC.Keçeci,“RobustandFastAutomaticModulationClassificationwithCNN underMultipathFadingChannels,”inProc.IEEE91stVehicularTechnologyConference(VTC-Spring),May2020,pp.1–6.
[12] S.Huangetal.,“GeneralizedAutomaticModulationClassificationforOFDMSystemsunderUnseenSyntheticChannels,” IEEETransactionsonWirelessCommunications,vol.23,no.9,pp.11931–11941,Sep.2024.
[13] H. Zhang, F. Zhou, H. Du, Q. Wu, and C. Yuen, “Revolution of Wireless Signal Recognition for 6G: Recent Advances, ChallengesandFutureDirections,”arXivpreprint,arXiv:2503.08091,Mar.2025.
[14] M. C. Park and D. S. Han, “Deep Learning-Based Automatic Modulation Classification with Blind OFDM Parameter Estimation,”IEEEAccess,vol.9,pp.108305–108317,2021.
[15] S.K.Jagatheesaperumal,I.Ahmad,M.Höyhtyä,S.Khan,andA.Gurtov,“DeepLearningFrameworksforCognitiveRadio Networks:ReviewandOpenResearchChallenges,”arXivpreprint,arXiv:2410.23949,Oct.2024.
[16] E.VijayandK.Aparna,“RNN-BIRNN-LSTMBasedSpectrumSensingforProficientDataTransmissioninCognitiveRadio,” e-Prime–AdvancesinElectricalEngineering,ElectronicsandEnergy,vol.6,p.100378,Dec.2023.
[17] T.XuandY.Ma,“SignalAutomaticModulationClassificationandRecognitioninViewofDeepLearning,”IEEEAccess,vol. 11,pp.114623–114637,2023.
[18] T. Zhang, C. Shuai, and Y. Zhou, “Deep Learning for Robust Automatic Modulation Recognition Method for IoT Applications,”IEEEAccess,vol.8,pp.117689–117697,2020.
[19] B.Xuetal.,“TowardsExplainabilityforAI-BasedEdgeWirelessSignalAutomaticModulationClassification,”Journalof CloudComputing,vol.13,no.1,p.10,Jan.2024.
[20] Y. Wang et al., “An Improved Modulation Recognition Algorithm Based on Fine-Tuning and Feature Re-Extraction,” Electronics,vol.12,no.9,p.2134,May2023,doi:10.3390/electronics12092134.
[21] S.Mohsen,A.M.Ali,andA.Emam,“AutomaticmodulationrecognitionusingCNNdeeplearningmodels,”MultimedTools Appl,vol.83,no.3,pp.7035–7056,Jan.2024,doi:10.1007/s11042-023-15814-y.
[22] Z.Elkhatib,F.Kamalov,S.Moussa,A.B.Mnaouer,M.C.E.Yagoub,andH.Yanikomeroglu,“RadioModulationClassification Optimization Using Combinatorial Deep Learning Technique,” IEEE Access, vol. 12, pp. 17552–17570, 2024, doi: 10.1109/ACCESS.2024.3357628.
[23] Y.Qu,Z.Lu,R.Zeng,J.Wang,andJ.Wang,“EnhancingAutomaticModulationRecognitionthroughRobustGlobalFeature Extraction,”arXivpreprint,arXiv:2401.01056,Jan.2024.
[24] J.D.Ruikar,D.-H.Park,S.-Y.Kwon,andH.-N.Kim,“HCTC:HybridConvolutionalTransformerClassifierforAutomatic ModulationRecognition,”Electronics,vol.13,no.19,p.3969,Oct.2024,doi:10.3390/electronics13193969.
[25] Q.Zhengetal.,“MobileRaT:ALightweightRadioTransformerMethodforAutomaticModulationClassificationinDrone CommunicationSystems,”Drones,vol.7,no.10,p.596,Sep.2023,doi:10.3390/drones7100596.
[26] A. Faysal, M. Rostami, R. G. Roshan, H. Wang, and N. Muralidhar, “NM former: A Transformer for Noisy Modulation ClassificationinWirelessCommunication,”Oct.30,2024,arXiv:arXiv:2411.02428.doi:10.48550/arXiv.2411.02428.