Skip to main content

Hybrid CNN-Transformer Architecture for Enhanced Signal Classification in Wireless Communications

Page 1


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Hybrid CNN-Transformer Architecture for Enhanced Signal Classification in Wireless Communications

1Department. of Electronics Engineering (Communications and Control), University of Gezira, College of Engineering and Technology.

2 Professor, Department of Electrical Engineering, Omdurman Islamic University, Omdurman, Sudan.

3Department of Electrical and Electronics Engineering, University of Khartoum, Khartoum, Sudan ***

Abstract - Signal classification in noisy environments is a common issue for wireless communication systems, particularly while working with modulation strategies. The complexity of signal data makes it difficult to use traditional methods that rely on simple convolutional neural network (CNN) architectures and conventional machine-learning models, particularly when operating in environments with high signal-to-noise ratios (SNRs). In addition, for real-world heterogeneous datasets, these methods often lack required generalizability. This work proposes a hybrid CNN transformer model to overcome these constraints. The proposed model performs better at classification when the SNR changes because it combines the sequential modelling capabilities of the transformer architecture with the feature extraction capabilities of the convolutional layers. It is trained using a dataset of 11 modulation schemes under varying SNRs to ensure the resilient performance in various noise settings of the model. The model performance is assessed using its accuracy and bit error rate (BER). The model outperforms standard methods by both accuracy and generalization for classification. The model performance under various settings was examined employing a confusion matrix visualization.

Key Words: Automatic modulation recognition (AMR), Deep-learning neural networks

1.I NTRODUCTION

Throughout the growth of new wireless communication technologies, the limited availability of radio spectra leads to a considerable drawback. This drawback is not a genuine deficiency of spectrum resources but rather ineffective regulatory frameworksthatassignfrequenciesin a strictandunyieldingfashion[1]. A wide range oforganizations,suchasthecivilian, government, commercial, and military organizations, share the electromagnetic spectrum. Therefore, modern wireless communicationenvironmentsrequireadynamicspectrumallocationpolicyinsteadofthefixedpolicythatiscommonlyadopted today,leadingtolowspectrumutilizationdifficulties.Cognitiveradio(CR)technologycreationletsradioschangeanduseunused frequencyresources,itiscalled“spectrumholes”or“whitespaces”.Thisworkindicatestheconceptofdynamicaccesstotheradio spectrum.Bythis,secondaryusers(SUs)canopportunisticallyaccessthefrequencyallocationtoprimaryusers(PUs)whenthey are not in use. This approach aims to enhance spectral efficiency by enabling transmissions on detected free bands, which addressestheissueofspectrumshortage[2].

Because spectrum sensing and detection are becoming more critical in spectrum monitoring, management, and secure communicationssuchas5Gcommunicationsandbeyond,IoTnetworks,andotherservices,CRhasessentialrolesandcapabilities indetectingactivePUtransmissionsovertheband.DecidingtotransmitthesensingoutcomesindicatesthateachPUtransmitters isinactiveatthisbandwithahighprobability,andCRbecomesthepromotionsolutionforscarcityandunderutilizationproblems [3].AMRisanessentialitemindigitalcommunicationsystemsandservesasacriticalelementfortheeffectivenessofCR.Itenables dynamicradioresourcemanagementbyusingreconfigurablesoftware-definedtransceivers.Thesetransceiverscanreconfigure theirtransmissionparametersbyusingtheaccessiblecommunicationresourcesintheelectromagneticenvironmenttorecognize modulationtypesofunknownsignalswithoutpreviousdata[4].ByusingAMR,CRisexpectedtoaccuratelyrecognizeorclassify themodulationstructureofthereceivedsignalrapidly,withoutanylatency.ThisprocesshelpstoidentifywhetherthereisaPUin the channel, as all PUs use a single modulation technique for transmission over the frequency channel. This allows the correspondingdemodulationprocesstobedoneonthereceivingside[5],[6].

AMRisatransitionalstagebetweensignalmodulationanddemodulation.Becauseitischallengingtodistinguishbetween numerousmodulationschemesowingtoseveralfactors,includingmultipathfading,noise,centerfrequencyoffset,andsignal

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

structure distortion caused by inadequate hardware design or crystal oscillator drifting, AMR technology is crucial before demodulatingsignalsatthereceiverandavitallinkinwirelesscommunication.Thetechniquecanautomaticallydeterminethe modulationtypeofthesignalandextracttheinformationitcontainswithoutbeingacquaintedwiththesystemspecifications.Its automationcanadvancethesignalmodulationdetectionaccuracywhilesignificantlyreducingtheneedforhumanresources, therebyenhancingtheaccuracyofsignalmodulationrecognition.Manualmodulationrecognition(MMR)hasbeenreplacedbya newautomaticmethod(AMR)inwhichMMRrequiresmanualobservationwithanoscilloscope.TheMMRmethodwillbelimited byfourtypesofinformationrequiredforthesearchoperator:theintermediatefrequency(IF)timewaveform,theaverageaswell asinstantspectrumofthesignal,sound,immediateamplitude,andbothlongrecognitionanderrorrecognitiontimes.

AMRplaysamajorroleinCR,whichenablesdynamicradioresourcemanagementbyemployingreconfigurablesoftwaredefined transceivers. These transceivers can reconfigure their transmission parameters by using obtainable communication resourcesinanelectromagneticenvironmenttoidentifytheunknownsignalsmodulationtypeswithoutpriorknowledge[7].

As in Figure 1, the AMR block consists of two components, such as signal preprocessing and classifier modules. The preprocessingmoduleestimatessynchronizationparameters,includingthefrequencyoffsetofthetimingrecovery,receivedsignal, andpower.Inthesecondpart,thereiseliminationofsignaldisturbancessuchasinterferenceidentification,resultinginenhanced performance

Chart – 1: AMRinthereceiverpartofcommunicationsystems

AMRisclassifiedintotwoclasses.ThefirstcomprisestheexistingAMRtechniques,consistingofdecision-theoreticapproachesand feature-based(FB)methods.Thesemethodsdependontheformofthesignalinput,suchassignalstatisticalapproachesand image-based methods. The likelihood-based (LB) approach is the initial decision-theoretic approach. The probability density function(PDF)coverstheAMC-LBoftheobservedwaveformalsolearnsfrommodulatedsignalsandFBmethods[8].LBrequires priorknowledgeofthePDFandthecalculationofthemaximumlikelihoodvalueforalltheproposedmodulationschemes,which addshighcomputationalcomplexity,lacksrobustnessagainstmodelmismatch,andposesachallengeforreal-timesystems.FB approachesconcentratelowcomputationalcomplexityonfeaturesextractionfromthesignaldirectly,removingtheneedforextra channelorsignaldataand,thus,loweringcomputationaldemands.Featureextractionandfeatureclassificationcanbeapplied,in whichinthefeatureextractionmodule,suddenfeatures,highordercumulant(HOC)features,andwaveletfeaturescanbeused[9]. Mostexistingwell-developedAMCmethodsareFB-based,particularlywiththeriseofmachinelearning(ML)approachesusingthe deep neural network (DNN), where different ML approaches can be exploited. The support vector machine (SVM), k-nearest neighbour(KNN),etc.areregularlyassumed.Thesemethodssignificantlydependontheextractionandanalysisofsignalfeatures [10],[11],[12],[13].

ThesecondandmostmodernclassisadvancedMLanddeeplearning(DL)methods,whichhavedrawnincreasingcarefor improved spectrum-sensing detection, precision, and accuracy. DL with artificial neurons organized in a stacked multilayer architectureintroducesinnovativetechniquesthatsignificantlyimprovetheefficiencyofmodernwirelesscommunicationsystems. SeveralDLalgorithms,plusconvolutionalneuralnetworks(CNNs),DNN,alongwithlongshort-termmemory(LSTM)[14],[15], [16],[17],aswellastheirhybrids,demonstratesignificantadvantagesinaddressingchallengeslikespectrumsensingandAMRin CR.Thisimprovementwasdonebyaccuratelyevaluatingcriticalfeaturesofthereceivedsignals,suchasmodulationtype,which canbeenhancedtoenableprecisepredictionsofchannelavailability.Recently,DLtransformer-basedmodelshavebeendeveloped, introducinganewrecognitionalgorithmforreal-timedecision-makingandleveragingself-attentiontoreducecomplexityconcerns in the AMR of signals, thereby achieving accurate modulation classification under practical channel conditions. Recently, transformer-basedmodelsexploreforreal-timedecision-making,leveragingself-attentiontohandlecomplexspectrum-allocation tasks.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

1.1 Problem Formulation

SpectrumsensingisacorefunctionofCRNs;however,severallimitations,suchasnoiseandtherequirementforpreceding informationofsignalfeatures,causesittofailinenvironmentswithlowsignal-to-noiseratio(SNR).Inaddition,recognizingthe type of modulation is critical to intelligent spectrum access and interference management. Some existing spectrum-sensing methods,forinstanceenergydetectionandmatchedfiltering,combineAMRwithspectrumsensing.AMRinvolvesclassifyingthe modulationschemes(e.g.,quadraturephase-shiftkeying,binaryphase-shiftkeying(BPSK),16-QAM,frequencymodulation,and amplitude modulation (AM)) used by the transmitted signal. These schemes can significantly improve the performance and robustnessofspectrumsensing,particularlyunderreal-worldconditionswithvaryingnoiselevels,fading,andsignaldistortions. Becausethenecessityforautonomousmodulationrecognitionrisesinwirelesssystems,wheremodulationschemesarelikelyto alterregularlyastheenvironmentchanges,DNNareconsideredasanewandpowerfulmethodforAMR,overcomingthereliance onexpertanalysisandtheinabilitytoscalewiththeincreasingsignalcomplexityexhibitedbytraditionalmethods.Thecoreissue thisworkaddressisthedesignanddevelopmentofanaccurate,noise-resilient,andcomputationallyefficientDL-basedframework forAMRtoenhancethespectrumsensingperformanceinCRnetworks.

1.2 Contribution

1. Thispaperpresentsadata-driven,adaptive,andscalablesolutionthatenhancestheCRNsintelligenceandreliability.Thework is done by combining signal classification and spectrum detection capabilities in a CNN-transformer model to improve modulationrecognition,mainlywhileworkinginlow-SNRenvironments.

2. Theproposedframeworktrainsandvalidates11modulationschemesundervariousSNRconditionstoachieverobusttraining acrossnoiselevels,whilecomparingthemwithpreviousmethodsthatsignificantlydeclineatlowSNRs.

3. Aspertheexperimentalevaluationsusingaccuracy,BER,andconfusionmatrixanalysis,theproposedhybridmodelperforms betterinbothclassificationaccuracyandgeneralizationabilitythantraditionalCNN-basedandclassicalMLapproachesto improvedgeneralizationandaccuracy.

Thisstudyoffersausefulpathforcreatingnoise-resilientmodulationrecognitionsystemsapplicabletoreal-worldsystemsand contemporarywirelessnetworksbyfusingsequentialmodellingwithfeatureextractioninalightweighthybridarchitecture-based CR. Thoughtheexistingworkshavesomeadvantages,theseDL-basedAMRmethodsoftenfailunderlowSNRenvironmentsowing tolimitedglobalfeaturemodelingandover-dependenceonlocalpatterns.Thisstudypurposestoworkwiththeseexistinggapsby proposingahybridCNN-Transformermodel,whichenhancesrobustness.

2. Literature Review

AMRisanessentialtechniqueinCRnetworksandpossessestheorskilltoclassifythereceivedsignalmodulationwithoutpast informationofthemodulationscheme.Itisablindorsemi-blindmethodthatenablesSUsorCRreceiverstosense,classify,and adaptwithoutpriorknowledgeorcoordinationtoidentifythemodulationtypesofunknownsignals.Suchanapproachhelpswith intelligentspectrumsensing,robustness,andadaptivity,andreducesthesignallingoverheadrequiredinCRnetworks[18].

DLnetworkarchitecturestargetmodulationrecognitionalgorithmstoenhancethereliability,simplicity,andeffectivenessofAIbasedAMCmodelsinwirelesscommunicationapplications.TheresultsrevealthatDLarchitecturescansignificantlyoutperform traditionalmethods,providingastrongfoundationforfutureresearchinthisarea[19],[20].

Several DL techniques based on CNN models are presented for recognition of modulation techniques. In [21], Mohsen 2024 designedtwoCNNmodelsfortheautomaticrecognitionofmodulationmethods:onebasedontheRadioML2016.10adatasetand theotherusinganimagedataset.Thehyperparametermodificationenabledbothmodelstoachievehighvalidationandtesting accuracies. DL models,particularlyCNNsandRNNs,canlearndiscriminative representationsdirectlyfromrawI/Qsamples establishedtoavoidtheneedformanualorhandcraftedsignalfeatures,aresuitableforextractinglocalwaveformpatterns,andare extensivelyusedtoimprovetheaccuracyofAMR,especiallyinlowSNRenvironments.However,capturingglobalfeaturesand mitigatingirrelevantfactorsremainchallengingforimprovingtheAMRperformance.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Hybrid networks have been introduced to improve AMRs by effectively capturing global features and enhancing the model generalization. Several studies [22] have suggested hybrid designs that take advantage of the complementary capabilities of transformers(globalcontextmodeling)andCNNs(localfeatureextraction)andtheimportanceofintegratingtransformersand convolutionalnetworkstoincreaseclassificationaccuracyinwirelesscommunicationsatdifferentSNRs.Byutilizingbothdesigns advantages,thismethodprovidesreliableperformanceinsettingswithbothlow-andhigh-SNRsignals.

BasedonthecombinationbetweenthetransformerandLSTM(TLDNN)framework[23],QuimprovedthemainchallengeofAMR, whichiscapturingglobalfeaturesandenhancingmodelgeneralization,byproposingadataaugmentationstrategycalledsegment substitution(SS).TheSSenhancesrobustnessbyalteringpartsofsignalstoforcethenetworktorelymoreonglobalconstellation patternsandlessonspuriousfeaturesorchannelartifactsandmitigatetheimpactofirrelevantinherentvariables,suchasRF fingerprintcharacteristicsandchannelcharacteristics,andsolvetheprimaryproblemofAMR.Ahybridconvolutionaltransformer classifier(HCTC)hasbeen designedtoclassifyunknown signals[24].TheHCTC model employs a three-stageframework for featuresextractionfromin-phase/quadraturesignalsusingaconvolutionallayer,atransformerlayer,aswellasfeaturemapping. Whencomparing,theHCTCmodelgivessuperiorperformancebymaintaininghighaverageaccuraciesacrosstheSNRrange.This demonstrates complete robustness and reliability across various practical noise environments. This model shows good classificationaccuracyandworksathighSNRlevels.However,itslowsdownatverylowSNRlevelsfrom0dBandlower,and henceitislimitedinnoisycommunication.

ThelightweightradiotransformermethodforAMCwasusedin[25],leveragingbothlarge-scaleandsmaller-scaleRadioML 2018.01AdatasetstocomprehensivelyevaluatetheperformanceofMobileRaT,consideringvariousmodulationschemesand realisticcommunicationconditions.Althoughthemodelisuseful,itdoesnotdirectlyaddresstheadvantagesoftheCNNfrontends formaintainingthelocalwaveformstructure.Becauseitismostlytransformer-centricratherthanacompleteCNN–transformer hybrid,whichlimitsthefront-endadvantagesoftheCNNarchitecture.Theresearchersin[26]generatedasolutionforAMCnoise, NMformer,basedonavisiontransformer(ViT),whichissuitableforcomplexfeaturerelationshipsandimagereconstruction. Constellationdiagramshavebeengeneratedfrommodulatedsignals,convertingthesignalinformationintoa2-Drepresentation, whichachieveshighaccuracyacrossvariousSNRsanddemonstratesstrongresiliencetoout-of-distributiondata,outperforming baselineclassifiers,comparedwithtraditionalCNN-basedapproachesthatfocusonlocalfeatures.

TraditionalCNN-basedapproachesin[21]and[22]workwellincapturinglocalfeaturesbutoftenfailtogeneralizeinlowSNR environments owing to their limited capability to model long-range dependencies. Alternatively, transformer-based models [24][25][26], while are effective at global feature extraction, commonly overlook localized signal characteristics, which are essentialformodulationrecognition.RecenthybriddesignslikeHCTC[24]andNMformer[26]workwiththisgapbuteitherlack effective local feature encoding or exhibit instability at very low SNRs. The model proposed in this work addresses these shortcomingsbyjoiningconvolutionallayersforrobustfeatureextractionwithTransformerlayersforsequencemodeling.This strategyhelpsinachievinghighclassificationaccuracyandnoiseresilienceacrossabroadSNRspectrum.Thisdual-capability frameworkenhancesmodulationrecognitionperformancewherebothlocalandglobalsignalattributesareimportant

3. Proposed Methodology

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

3.1. Dataset

ThisworkutilizestheRML2016.10adataset,forverifyingtheresults.TheRML2016.10adatasetcomprises20SNRsthatranges from-20to18dBfor22,000samplesintotalacross11modulationschemes.TheseschemesincludeWBFM,4PAM,AM-SSB,BPSK, AM-DSB,16QAM,CPFSK,GFSK,QPSK,64QAM,and8PSK.ThesamplescountofeverymodulationmodeunderallSNRconditionis 1000.Also,thedataformatofallthesamplesis128×2,where128signifiesthelengthofthesignal.Theproposedarchitectureis showninfigure2.

3.1.1. Data Splitting

Datasplittingistheprocessofseparatingtheinputdatasetintotwodistinctsubsets,suchasthetraining,andtestsets.Thegoal istoletthemodelbetrainedonasubsetofthedata,checkedonanothersubsettochangehyperparameters,andestimatedona finalsubset(testset)toguesshowwellitwilldoondatathatithasn'tseenbefore.Itiscommonpracticetouse80%ofthedataset fortrainingpurposes.Themodelisfitted,anditsweightsareadjustedusingthisdata.Byadjustinghyperparametersincludingthe batchsizeandlearningrate,thevalidationsetfacilitatesmodelselection.Thisworkuses20%ofthedatafortesting.Thereisno biasinthetestfindingssincethemodeldoesnotviewthisdatawhentraining.

3.1.2. Normalization and One-Hot Encoding

Normalizationandone-hotencodingareimportantprocessestopreparethedataforDLtrainingandencoding.Scalingthe featuressuchthattheyhavecomparableranges,usuallybetween0and1or-1and1,iswhatnormalizationdoes.Thisisespecially importantforneuralnetworksbecauseitmakessurethatthemodeltreatsallfeaturesthesame.Thisway,problemsdon'thappen wheresomefeaturesdominatethelearningprocessbecausetheyarebiggerinequation(1). (1)

Where isthemeanand isthestandarddeviationofthefeature.Abinarymatrixmaybegeneratedfromcategoricallabels usingtheone-hotencodingprocess.EachoftheNclassesinthedatasetisrepresentedbyavectoroflengthN,exceptfortheclass's index,whichissetto1.Allothervaluesare0.

3.1.3. Shuffling and Augmentation

Ifthedataissortedbyclassortime,forexample,shufflingisconductedtopreventanypossiblebiasesinthetrainingprocess. To avoid the model from learning from data order effects, the dataset may be rearranged at random. Data augmentation approachesartificiallyenlargethedatasetbycreatingvariantsofthealready-existingdata.Thismightincludedatamodifications likeflipping,rotating,orintroducingnoise,whichenhancesthemodel'sgeneralizationinequation(2)

Tasksinvolvingsignalprocessingmayincludedoingthingslikeintroducingrandomnoisetotheinputsignalsoradjustingthe signalintensity.

3.2. Hybrid CNN-Transformer Model

Forsomesignalprocessingwork,thehybridCNN-transformermodelworkswithDLarchitectures.Thisworkmakesuseofthe bestfeaturesoftwodifferentnetworktypes,suchasthetransformermodelsandtheCNNs.

3.2.1. Convolutional Layers for Feature Extraction

Toextractlocalfeaturesfrominputdata,CNNsmakesuseoftheconvolutionallayers.CNNsareconsideredaviableoptionfor signal detectionsince they detectedgesandotherhigh-frequencypartsthatshow importantlocal patternsinthewaysignal representationisdoneintermsoftimeandfrequency.Inaconvolutionallayer,theinputdatapassesthroughafilters,orkernels

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

set.Eachfilterlooksforvariousfeatures,includingtextures,edges,orfrequencypatternsinsignalprocessing.A2Dconvolution operation’smathematicalformulationisinequation(3):

(3)

Where istheoutput, istheinputdata, isthekernel/filter,and and arethekerneldimensions.Forclassifying themodulations,itisnecessarytoremovethefrequencycomponentsandphaseshiftsfromthesignals.Additionally,theselayers alsoworkwellwithothersignalprocessingtasksrelatedtothiswork.

3.2.2. Transformer Layers for Sequence Modelling

Transformersarewell-suitedfortime-seriesdata,suchassignaldata,sincetheyareusedtodescribesequentialrelationships in data. Transformer layers in the hybrid CNN-transformer model can identify long-range relationships in signal data by concentrating on diverse input sequence parts. Mechanisms for self-attention are the basis of the transformer architecture. Prioritizing one part of a sequence above another is the main point. Here is the definition of the self-attention method as in equation(4) (4)

Where isthequerymatrix, isthekeymatrix, isthevaluematrix,and isthekeyvectordimension.Self-attention mechanismsmayhelpmodelsbyfocusingonsignalfeatures,suchascertainfrequencycomponents,thatmaybeimportantfor classifying

Figure2depictsthelayerarrangementoftheproposedhybridCNN-transformermodel.Thearchitecturecontainsnumerous vitalanddiverselayers,eachperformingauniquerole.Theinputtotheinputlayerisanissue-specific1Dor2Dsignal,suchas time-seriesorfrequency-domaindata.Theformofasignalconsistsofitslengthandthenumberofchannels(e.g.,1forunivariate dataandmoreformultivariatedata),whicharerepresentedasbatchsize,sequencelength,andNumchannels,respectively.Onthe otherhand,convolutionallayersextractlocalizedfeaturesandpatternsfromtheinputdata.

Theconvolutionlayerlearnstheimportantfeatures,likefrequencycomponents,edges,orshiftsfromtheinputsignal.Asingle convolutionalblockcomprisesvariouslayers,whichincludetheconvolutional,max-pooling,anddropoutlayers.

TheCNNnetworkutilizesfilterstogeneratefeaturemaps.Themaximumpoolingmethodhelpsintheidentificationofthe highestvalueintheareaofinterest.Thismethodkeepsthefeaturesthatareconsideredmostusefulalongwithdimensionality reduction.Bychangingaunit'sportionoftheinputto0,thedropoutlayerpreventsoverfittingduringthetrainingprocess

FollowingtheCNN layer, thiswork usesa flattenlayer, whichconverts the outputtoa 1Dvectorandpasses it onto the transformerlayers.ThischangeturnsthefeaturemapsfromtheCNNlayersintoasequencethatthetransformeraccepts.Thefour layersinthisnetworkincludeafeedforwardnetwork,multi-headself-attention,normalization,andadropoutlayer.Toidentify attentionscoresandhighlighttheimportantfeatures,multi-headself-attentionusesself-attentionontheinputsequence.Twofully connectedlayersinafeedforwardnetworkapplyaReLUactivationfunctiontotheattentionoutputs. Thedropoutlayerensures constanttrainingbystandardizingtheoutputofeverytransformerblock.Byrandomlydeactivatingneurons,thedropoutlayer preventsoverfitting.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Input Layer

[batch_size, sequence length, num channels]

Extract local features from signal

Flatten Layer

[batch_size, flattened length]

Block

(Attention, Feed-Forward, Norm, Drop)

Captures long-range dependencies

Global Average Pool

[batch_size, num features]

Dense Layer

[batch_size, num classes]

Output Layer

Predicted class probabilities

Aglobalaveragepoolinglayerhelpssummarizethelearningfeaturesinthesequence.Thisprocedureisdoneafterpassing throughthetransformerlayersbyreducingthesequencedimensionandoutputtingasinglevalueperfeature.Theoutputshapeof thistransformerlayeris(batch_size,num_features). Adenselayerisinsertedafterthetransformerblocktomapthelearned featurestotheoutputclasses.Thislayerisresponsibleforcarryingoutthelastchange.Inmostcases,thisworkintroducesnonlinearityusingReLUoranequivalentactivationfunction.Asalaststep,theoutputlayerusessoftmaxasanactivationfunction The finalresultforclassificationisgotfromthislayer.

3.3.

Model Training

TheMLmodellearnsfromthedatathroughthemodeltrainingphase.Thisworkusesmultiplemethodstoconfirmthatthe modeltrainsandgeneralizeswell.

Chart -3: HybridCNN-Transformermodel
CNN Block (Conv2D)
(Conv, Pool, Drop)
Transformer

3.3.1.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Early Stopping

Early stopping is meant to avoid the overfitting issue. By this method the training process ends once the validation loss improvesfora prearranged epochcount. After the model hasachieved excellentgeneralization,this preventsit from further learningnoisefromthetrainingdata.

3.3.2.

Learning Rate Scheduling

ThisworkusesatoolforlearningrateschedulingcalledReduceLROnPlateau,whichchangesthelearningrateduringtraining. Thismethodologyensuresthatthemodelisfine-tunedasitgrowsclosertoanoptimalstatebydecreasingthelearningratewhen thevalidationperformanceplateausinequation(5).

(5)

Thefactorisusuallylessthan1,andinthiswork,itis0.5.

3.3.3.

Model Checkpoints

Duringtraining,theproposedmodelkeepsanoteofitscheckpointsforsomeregularintervals,usuallywhenthevalidationloss islow.Thetrainingprocessmayresumefromtheoptimalconditionifoverfittingoccursorinterruptionisnecessary.

4. Results

Thissectiondiscussestheadvantagesoftheproposedapproachanditsimplementation.Thisstudyimplementationisdonein Pythonbyusingthewell-knownlibrarieslikescikit-learn,TensorFlow,andKeras.Theselibrariesserveasbuildingblocksfor creation,training,andassessmentoftheDLmodels.TheproposedhybridCNN-transformermodelcombinesthebestfeaturesof convolutionallayersforfeatureextractionalsotransformerlayersforsequencemodelling.Thismodelisalsosuitableforrealworldapplicationsincommunicationsystems,asitworksbetterthanexistingworksingeneralizationandaccuracy.Furthermore, theproposedworkmanagesnoisydataandlong-rangerelationshipsinsignals.

Table – 1: ModelArchitectureTable Layer

InputLayer (None,128,2,1)

Conv2D (None,128,2, 64)

MaxPooling2D (None,64,2,64)

Conv2D (None,64,2, 128)

MaxPooling2D (None,32,2, 128)

Flatten (None,8192)

Reshape (None,128,64)

MultiHeadAttention (None,128,64) 33,216

LayerNormalization (None,128,64)

Flatten (None,8192)

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Dense (None,128)

Dropout (None,128) 0

Dense (None,11) 1,419

Table1providestheoutputshapeandtheparametersusedinthiswork.Eachlayerdeterminesthefinalproduct'ssize.The firstparameter,None,intheinputlayerrepresentsthedynamicbatchsize.Theoutputtensor'swidth,height,andchanneldepth aretheremainingvaluesintheinputlayer.Therequiredparametersneededforlayeroperations,suchasweightsandbiases,add uptothetotalparameters.Thismodelhandlesfeatureextraction,sequencemodelling,andclassificationtasks.Theinputlayer analyzes the data first, then the CNN layers extract spatial features, and finally the transformer layers capture long-range relationships.

Fromthetransformerprocess,afterflatteningtheoutput,isamulti-headattentionlayerthatreshapesthedataandpassesit on.Thenthenormalizationlayerstabilizesthelearningprocess.Lastly,decision-makingandclassificationaredonebydenselayers, which include dropout layers to avoid overfitting. The final output layer predicts the class labels using a softmax activation function.

4.1. Testing and Evaluation

Oncetrained,itiscrucialtomeasurethemodel'sperformanceonunseentestdatatodetermineitsgeneralizationcapabilities. Theproportionofaccuratepredictionstoallpredictionsisknownasaccuracy.Theformulaisinequation(6): (6)

BERcalculatestherateofmisclassifiedbitsasfollowsinequation(7): (7)

Bycomparingthepredictedlabelswiththetruelabels,aconfusionmatrixdisplaystheclassificationmodelperformance.It revealsthekindsoferrorsthemodelcommits.

4.2. Comparison

TheCNNmodel,DNN,andtheexistingCLDNNmodelarethethreeadditionalmodels.Tables2,3,and4describethestructure oflayersintheseexistingmodels.

Table -2: ExistingCLDNNArchitecture

Layers

Convolution(Conv2d)

Input Parameters

Customactivation

MaxPooling2D poolsize=(1,2)

Dropout(rate) 0.3

Conv2D

Customactivation

MaxPooling2D poolsize=(1,2)

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Dropout(rate)

Dense

128hiddenlayer,Custom activation

Dropout(rate) 0.3

Reshape (2,4096) LSTM Customactivation

Dropout(rate)

Dense

Dropout(rate)

128hiddenlayer,Custom activation

Dense 11hiddenlayer,softmax

Table -3: ExistingCNNArchitecture Layers Input Parameters

Convolution(Conv2d)

Relu

BatchNormalization -

MaxPooling2D pool_size=(1,2)

Dropout(rate)

Conv2D

0.3

Relu

BatchNormalization -

MaxPooling2D poolsize=(1,2)

Dropout(rate) 0.3

Convolution(Conv2d)

Relu

BatchNormalization -

MaxPooling2D pool_size=(1,2)

Dropout(rate) 0.3

Convolution(Conv2d)

Relu

BatchNormalization -

MaxPooling2D poolsize=(1,2)

Dropout(rate) 0.3

2026, IRJET | Impact Factor value: 8.315 | ISO 9001:2008

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Flatten -

Dense 128hiddenlayer,Relu

BatchNormalization -

Dense 11hiddenlayers,SoftMax

Table – 4: ExistingDNNArchitecture

Layers

Dense

Input Parameters

256hiddenlayers,Relu

Dense 128hiddenlayers,Relu

Dense 250hiddenlayers,Relu

Dense 64hiddenlayer,Relu

Flatten -

Dense 128hiddenlayer,Relu

Dense 11hiddenlayer,softmax

Table - 5: Proposedaccuracy,loss,validationlossandvalidationaccuracy

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Chart – 4: ValidationLossComparison

Table – 6: Proposedotherperformancemeasures

Chart -3: ValidationAccuracyComparison

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

– 6: BERcomparison

– 7: PrecisionComparison

– 8: RecallComparison

Chart
Chart
Chart

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Chart – 9: F1scoreComparison

Theabovefiguresfrom3to9showtheevaluationresultsofseveralMLmodelsusingfiveimportantmetrics,whichinclude precision, recall, accuracy, F1 score, and BER. All the metric calculations are done for diverse SNR values in the dataset. The proposedmodelhasthebestaccuracyofallthemodelsatSNRsof18,16,and14dB.Theproposedmodelgets0.7830accuracyat SNR=18dBandkeepsthataccuracyprettyhighevenasthenoiselevelgoesup,whichismuchbetterthanCLDNN,CNN,andDNN. However,sinceithassuchachallengingtimedealingwithnoise,theCNNmodelconsistentlyshowsverypooraccuracy,hovering around0.1.Theresultshowsthattheproposedmodelisbetteratwithstandingnoiseandalsokeepsitsaccuracyhighevenin difficultenvironments.

Thegeneralizabilityoftheproposedmodelisshownbyusingthevalidationlossmetric.WhencomparingtheexistingCNNand DNNmodels,theproposedmodelshowsalowandconstantloss.WhileexistingmodelsshowsharplinesinlossathigherSNR levels,theproposedmodel'slossisconstant.At18dBSNR,thelossfortheproposedmodelis0.4861;however,CNN'slossis 4.7325.ThissuggeststhatCNNstruggleswithhighSNRdata,mostlikelybecauseofitslackofgeneralizability.Thus,theproposed modeloutperformsotherexistingworks,especiallyundersettingswithgreaterSNRs.Theproposedmodelidentifiessignalswitha probabilityof0.7645at18dBSNR.WithlowerSNRvalues,everymodelshowsareductionindetectionlikelihoodasthereisarise innoiselevels.Buttheproposedmodelmaintainshigherdetectionrates,whichindicateshighnoiseflexibility.

WithalowerBER,theproposedmodelperformsbettersinceitassessestherateoferrorbitsinitspredictionsincomparison totheexistingmodels.TheproposedmodelhasaBERof0.235forSNR=18dB,whileCNNhasaveryhighBERof5.8593.This showstheproposedmodel'sabilitytoconveymoretrustworthydata,withoutconsiderationofthenoiselevels.Theperformance metricsshowthetrade-offbetweencorrectlyidentifyingpositivesamples(precision)andgettingasmanytruepositives.Thisis shownbytheproposedmodel'shighF1scoresincomparisontotheexistingCNNalongwithDNN.Theproposedmodelismuch betterthanCNNandDNNatdealingwithunbalancedclassesornoisydata.ItsF1scoreisaround0.77atSNR=18dB,whichis muchhigherthantheirscoresofalmostzero.

4.3.

Discussion

Theproposedmodelshowsbetteroutcomesthantheexistingmodels(CLDNN,CNN,andDNN)onallimportantcriteria.This advantagemakesitthecleargoodinsituationswithdifferentamountsofnoise.EvenatlowerSNRlevels,theproposedmodel keeps validation accuracy high, loss low, and signal detection efficient. The proposed model maintains a more constant performanceatlowerSNRlevels,whileCNN'sperformancedeclinessignificantly.Thisstudyshowsthattheproposedmodelhas beenmodifiedtobemoreflexibletonoiseandisperfectforreal-worldapplicationswheresignaldeteriorationisprevalent.Asthe validationlossoftheproposedmodelislowandconstantoverarangeofSNRvalues,itappearstobeabletogeneralizeverywell withnewdata.Alternatively,CNNandDNNmodelsshowalotoflosswhenthere'salotofnoise,whichmeanstheyeitheroverfitor don'tgeneralizewell.Theproposedmodelisthepreferredalternativewhenadaptingtodifferentnoisesituations.

TheproposedmodeldoesanimprovedjobthantheCNNandDNNmodelsbecauseitcanbeusedrepeatedlyandisaccurateat findingpositivecases while reducingfalsepositivesand negatives.Foractivitieswheretheproperidentificationofsignalsis

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

critical,suchasincommunicationsystemsordetectiontasksinnoisysettings,thisisessential.Incommunicationsystemswhere theintegrityofthesentdataiscrucial,theproposedmodelconsistentlyshowsreducedbiterrorrates(BER).Whencomparedto CNNandDNN,whichhavemuchhigherBERvaluesandarethereforelessreliable,theproposedmodelhastheuniqueabilityto lowerBERevenwhenthereisalotofnoise.

Theproposedmodeldoesagoodjobwithdatasetsthataren'tbalancedornoisysituationswherebothfalsepositivesaswellas falsenegativesareimportant,asintheF1score,whichcomparesaccuracyandrecall.Becauseofitssuperioroverallperformance, theproposedmodelismoresuitedforreal-worldapplicationsthatneedprecisionanddependability,asshownbyitshighF1score. Duetoitstolerancetonoise,superiorgeneralization,increasedaccuracyandrecall,anddecreasederrorrates,theproposedmodel surpassestheexistingmodels.Itsabilitytohandledifferentamountsofnoisewithoutsacrificingperformancemakesitamore dependableoptionforreal-worldapplications.

Thehybridmodel’ssuperiorresultsqualifiedtoitsabilitytojointlyexploitlocalpatternsviaCNNsandcontextualdependenciesvia Transformerlayers.Thisallowsmorerobustmodulationclassificationevenunderseverenoise,unlikeclassicalmodelswhich strugglewithlong-rangedependencies.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

5. Conclusion

ThisstudyintroducedahybridCNN-TransformermodelforAMRinwirelesscommunicationsystemsoperatingundervarying noise conditions. The proposed model successfully overcomes the drawbacks of conventional CNNs and classical ML techniques, particularly in high SNR amplitude settings, by combining the long-range dependency modelling abilities of transformer architectures with the local feature extraction capabilities of convolutional layers. Its robustness is further demonstratedbytheconfusionmatrixanalysis,particularlyinaccuratelyclassifyingcomplexmodulationtypeseveninnoisy environments.Insummary,thisproposedhybridmodel,evaluatedonadatasetcomprisingelevenmodulationschemesacross awideSNRrange,demonstratesimprovedclassificationaccuracy,generalization,andnoiserobustness.Performancemetrics, includingconfusionmatricesandBER,confirmitssuperiorityoverbaselinemethods,particularlyindifferentiatingclosely spacedmodulationtypes.Makingitsuitableforreal-worldapplicationssuchasCR,software-definedradios,andadaptive communicationsystems.

REFERENCES

[1] A.S.JamwalandG.Kaur,“CognitiveRadio:AnEmergingTrendforBetterSpectrumUtilization,”IJCATR,vol.2,no.3,pp. 229–231,May2013.

[2] R.Zhang,Y.-C.Liang,andS.Cui,“DynamicResourceAllocationinCognitiveRadioNetworks:AConvexOptimization Perspective,”IEEESignalProcessingMagazine,vol.27,no.3,pp.102–114,May2010.

[3] M.U.MuzaffarandR.Sharqi,“AReviewofSpectrumSensinginModernCognitiveRadioNetworks,”Telecommunication Systems,vol.85,no.2,pp.347–363,Feb.2024.

[4] E.M.Ali,G.M.Salama,K.A.A.,andM.Ezz-Eldin,“AutomaticModulationClassificationforEnhancedCognitiveRadiofor IoTSystemsBasedonDeepLearning,”JournalofAdvancedEngineeringTrends,vol.44,no.1,pp.0–0,Jan.2025.

[5] B.Tang,Y.Tu,Z.Zhang,andY.Lin,“DigitalSignalModulationClassificationwithDataAugmentationUsingGenerative AdversarialNetsinCognitiveRadioNetworks,”IEEEAccess,vol.6,pp.15713–15722,2018.

[6] Q. Zheng, X. Tian, L. Yu, A. Elhanashi, and S. Saponara, “Recent Advances in Automatic Modulation Classification Technology:Methods,Results,andProspects,”InternationalJournalofIntelligentSystems,vol.2025,no.1,p.4067323, Jan.2025.

[7] B.Jdid,K.Hassan,I.Dayoub,W.H.Lim,andM.Mokayef,“MachineLearning-BasedAutomaticModulationRecognitionfor WirelessCommunications:AComprehensiveSurvey,”IEEEAccess,vol.9,pp.57851–57873,2021.

[8] X. Liu, C. J. Li, C. T. Jin, and P. H. W. Leong, “Wireless Signal Representation Techniques for Automatic Modulation Classification,”IEEEAccess,vol.10,pp.84166–84187,2022.

[9] O. A. Dobre, A. Abdi, Y. Bar-Ness, and W. Su, “A Survey of Automatic Modulation Classification Techniques: Classical ApproachesandNewTrends.”

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

[10] A.R.,L.C.,C.C.,J.C.W.A.,andA.B.R.,“ModulationClassificationinCognitiveRadio,”inFoundationsofCognitiveRadio Systems,S.Cheng,Ed.InTech,2012.

[11] K.Tekbıyık,A.R.Ekti,A.Görçin,G.K.Kurt,andC.Keçeci,“RobustandFastAutomaticModulationClassificationwithCNN underMultipathFadingChannels,”inProc.IEEE91stVehicularTechnologyConference(VTC-Spring),May2020,pp.1–6.

[12] S.Huangetal.,“GeneralizedAutomaticModulationClassificationforOFDMSystemsunderUnseenSyntheticChannels,” IEEETransactionsonWirelessCommunications,vol.23,no.9,pp.11931–11941,Sep.2024.

[13] H. Zhang, F. Zhou, H. Du, Q. Wu, and C. Yuen, “Revolution of Wireless Signal Recognition for 6G: Recent Advances, ChallengesandFutureDirections,”arXivpreprint,arXiv:2503.08091,Mar.2025.

[14] M. C. Park and D. S. Han, “Deep Learning-Based Automatic Modulation Classification with Blind OFDM Parameter Estimation,”IEEEAccess,vol.9,pp.108305–108317,2021.

[15] S.K.Jagatheesaperumal,I.Ahmad,M.Höyhtyä,S.Khan,andA.Gurtov,“DeepLearningFrameworksforCognitiveRadio Networks:ReviewandOpenResearchChallenges,”arXivpreprint,arXiv:2410.23949,Oct.2024.

[16] E.VijayandK.Aparna,“RNN-BIRNN-LSTMBasedSpectrumSensingforProficientDataTransmissioninCognitiveRadio,” e-Prime–AdvancesinElectricalEngineering,ElectronicsandEnergy,vol.6,p.100378,Dec.2023.

[17] T.XuandY.Ma,“SignalAutomaticModulationClassificationandRecognitioninViewofDeepLearning,”IEEEAccess,vol. 11,pp.114623–114637,2023.

[18] T. Zhang, C. Shuai, and Y. Zhou, “Deep Learning for Robust Automatic Modulation Recognition Method for IoT Applications,”IEEEAccess,vol.8,pp.117689–117697,2020.

[19] B.Xuetal.,“TowardsExplainabilityforAI-BasedEdgeWirelessSignalAutomaticModulationClassification,”Journalof CloudComputing,vol.13,no.1,p.10,Jan.2024.

[20] Y. Wang et al., “An Improved Modulation Recognition Algorithm Based on Fine-Tuning and Feature Re-Extraction,” Electronics,vol.12,no.9,p.2134,May2023,doi:10.3390/electronics12092134.

[21] S.Mohsen,A.M.Ali,andA.Emam,“AutomaticmodulationrecognitionusingCNNdeeplearningmodels,”MultimedTools Appl,vol.83,no.3,pp.7035–7056,Jan.2024,doi:10.1007/s11042-023-15814-y.

[22] Z.Elkhatib,F.Kamalov,S.Moussa,A.B.Mnaouer,M.C.E.Yagoub,andH.Yanikomeroglu,“RadioModulationClassification Optimization Using Combinatorial Deep Learning Technique,” IEEE Access, vol. 12, pp. 17552–17570, 2024, doi: 10.1109/ACCESS.2024.3357628.

[23] Y.Qu,Z.Lu,R.Zeng,J.Wang,andJ.Wang,“EnhancingAutomaticModulationRecognitionthroughRobustGlobalFeature Extraction,”arXivpreprint,arXiv:2401.01056,Jan.2024.

[24] J.D.Ruikar,D.-H.Park,S.-Y.Kwon,andH.-N.Kim,“HCTC:HybridConvolutionalTransformerClassifierforAutomatic ModulationRecognition,”Electronics,vol.13,no.19,p.3969,Oct.2024,doi:10.3390/electronics13193969.

[25] Q.Zhengetal.,“MobileRaT:ALightweightRadioTransformerMethodforAutomaticModulationClassificationinDrone CommunicationSystems,”Drones,vol.7,no.10,p.596,Sep.2023,doi:10.3390/drones7100596.

[26] A. Faysal, M. Rostami, R. G. Roshan, H. Wang, and N. Muralidhar, “NM former: A Transformer for Noisy Modulation ClassificationinWirelessCommunication,”Oct.30,2024,arXiv:arXiv:2411.02428.doi:10.48550/arXiv.2411.02428.

Turn static files into dynamic content formats.

Create a flipbook
Hybrid CNN-Transformer Architecture for Enhanced Signal Classification in Wireless Communications by IRJET Journal - Issuu