Skip to main content

Analysis of Framework for Robust Gender Recognition from Speech Signals

Page 1


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Analysis of Framework for Robust Gender Recognition from Speech Signals

Digambar B. Gote1 , Prof. Dr. T. B. Mohite-Patil2

ME (E & TC) Student, D. Y. Patil College of Engg. & Technolog, Kolhapur Prof. Dr.T.B.Mohite-patil D. Y. Patil College of Engg. & Technology Kolhapur,

Abstract: Speech-basedgenderrecognitionisafundamental paralinguistic task with wide applicability in speech-driven human–computer interaction, assistive technologies, and intelligentvoiceservices.Despitesignificantprogressachieved through deep learning, existing methods often suffer from limitedrobustnesstonoiseandchannelvariability,sensitivity to utterance duration, and poor interpretability of model decisions. This work proposes a compact and explainable framework for gender recognition from speech that emphasizes effective acoustic representation and attentiondriven feature refinement. Log-Mel and cepstral features are analyzed in conjunction with a lightweight convolutional neural network augmented by an attention mechanism to selectivelyemphasizeinformativespectro-temporalregions.A focused experimental analysis evaluates the impact of utterance duration, noise conditions, and channel mismatch on model behavior. In addition, attention-based visualization is employed to provide insights into the decision-making process, improving transparency and trustworthiness. The results demonstrate that the proposed framework achieves a balanced trade-off between robustness, efficiency, and interpretability, making it suitable for practical real-world deployment.

Keywords: Speech processing, Gender recognition, Attention-based learning, Acoustic feature analysis, Explainable AI

I. Introduction

Speech-based gender recognition has emerged as an important paralinguistic task in speech processing, with applicationsspanninghuman–computerinteraction,voicebased authentication, assistive technologies, and adaptive dialoguesystems.Humanspeechinherentlyencodesgenderrelatedcharacteristicsthroughphysiologicalandbehavioral factors such as vocal tract length, fundamental frequency distribution,formantstructure,andspeakingstyle.Advances indeeplearninghavesignificantlyimprovedtheabilityto modelthesecuesbylearningdiscriminativerepresentations directly from acoustic signals. However, recent studies reveal persistent challenges related to robustness under noisy and channel-mismatched conditions, sensitivity to utterance duration, and limited interpretability of model decisions. Moreover, the increasing reliance on complex architecturesoftenleadstotrade-offsbetweenperformance, computational efficiency, and transparency, which are criticalconsiderationsforreal-worlddeployment.

Motivated by these challenges, this work presents a compact and explainable speech-based gender recognition framework that synthesizes insights from recent deep learningandoptimization-drivenapproaches.Theproposed pipeline emphasizes effective acoustic representation, attention-basedfeaturerefinement,andsystematicanalysis ofrobustnessandinterpretability.Ratherthanintroducing excessive architectural complexity, the focus is placed on identifyingandvalidatingaminimalyeteffectivesetofdesign choices that contribute to reliable gender discrimination acrossvariedacousticconditions.

Contributionsofthisworkaresummarizedasfollows:

 Alightweightattention-enhancedCNNframework forspeech-basedgenderrecognition.

 A systematic analysis of key factors including feature representation, utterance duration, noise robustness,andchannelmismatch.

 Anexplainability-drivenevaluationusingattentionbasedvisualizationtoenhanceinterpretabilityand trust.

 Aconsolidatedexperimentalprotocolthatbalances performance, robustness, and deployment feasibility.

II. Literature Survey

Review of Recent Speech-Based Gender Recognition Studies (2022–2025)

SindhaandRana[1]developedanoptimizedartificialneural networkforvocalgenderrecognitionbyintegratingaselfattention mechanism to emphasize gender-discriminative acousticregionsintime–frequencyspace.Theworktypically beginsbyconvertingaspeechwaveform intoashorttimetime–frequencyrepresentationusingtheSTFT,

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

followed by log-compression on magnitude spectrograms

.Theproposedattentionmodule canbeabstractedasscaleddot-productattention, where arelearned projectionsof spectro-temporal embeddings. Their main achievement is improving robustness in capturing pitch- and formant-related cues while reducing reliance on handcrafted features; an importantremarkisthatattentiontendstofocusonvoiced segments,socarefulhandlingofsilence/non-speechframes iscrucialindeployment.

Yücesoy[2]studiedgenderrecognitionbystackinghybrid acoustic feature sets, where multiple descriptors (e.g., cepstral,prosodic,andspectralshapefeatures)arecombined in a hierarchical ensemble.Acommonfoundation isMFCC extraction,computedfromMel-filterbankenergies viathe DCT:

Stacking is realized by concatenating features and training meta-learners on base-model outputs. The key achievement is improved generalizationacrossrecordingconditionsthroughfeature diversity; an important remark is that feature stacking increases dimensionality and may require regularization (e.g., penalty)toavoidoverfitting.

Yücesoy[3]introducedanensemble-basedframeworkfor jointageandgenderrecognitionfromspeech,usingmultibranchlearningtoexploitsharedlow-levelrepresentations whilespecializingat task heads. Multi-task optimizationis commonlyexpressedas where and are task losses for gender and age, respectively, and control trade-offs. The main achievementisleveragingcorrelatedcues(e.g.,fundamental frequency and spectral tilt) while reducing redundant training;animportantremarkisthatmulti-tasklearningcan introducenegativetransferifonetaskdominatesgradients, sogradientbalancingoradaptiveweightsisbeneficial.

Yücesoy[4]investigated1Dand2DCNNsforspeakerage and gender recognition, comparing waveform/featuresequence convolutions against spectrogram-image convolutions. A 1D convolutional layer for frame-level embedding’scanbewrittenas

while 2D CNNs operate on . Their achievement is demonstratingthat2Dmodelscanbettercaptureformant trajectoriesandharmonicstructure,whereas1Dmodelssuit compact pipelines; an important remark is that 2D approachesdependstronglyonconsistentfront-endsettings (windowlength,hopsize,Melscale).

Mavaddati[5]developedaResNet-basedtransferlearning approach for voice-based age, gender, and language recognitionusingspectro-temporalrepresentations.Residual learningisexpressedas

whichstabilizesdeepoptimizationandhelpspreservelowlevel speech cues. The achievement is improved feature reuse across tasks and domains via transfer learning; an important remark is that transferring from large audio corpora to smaller demographic datasets requires careful fine-tuningtoavoiddatasetbiasamplification.

Younis et al. [6] introduced a comprehensive Arabic speechdataset(Hu-Int)designedforgenderdetectionand age estimation of Arab celebrities, enabling standardized evaluation in a language-specific setting. Their pipeline commonlyincludesfeaturenormalization and speaker-levelaggregation,suchastemporalpooling

torepresentvariable-lengthutterances.Theachievementis providing data diversity across speakers and recording conditionsforrobustmodeling;animportantremarkisthat celebritydatasetsmayincludestudio/editedaudio,which candifferfromreal-worldconversationalspeech.

Yue et al. [7] studied gender-aware speech emotion recognition using advanced differential evolution (DE) for featureselection,wheregenderinformationisusedtorefine feature subsets that remain discriminative under demographicvariation.DEevolvescandidatefeaturemasks viamutationandcrossover;acanonicalmutation is

followedbyselectionbasedonanobjectivetiedtoclassifier loss.Theachievementisdemonstratingthatoptimizationguided selection can reduce redundant descriptors while improvingstabilityacrossgenders;animportantremarkis thatfeature-selectionobjectivesshouldbevalidatedagainst leakage,especiallywhenspeakeroverlapexists.

Yueetal.[8]introducedagender-drivenspeechemotion recognition approach using a genetic algorithm (GA) and Fisher score ranking for selecting salient features. Fisher scoringforafeature canbeexpressedas

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

where areclass-wisestatistics.Theachievementis combiningfilter-basedrankingwithevolutionarysearchfor compact representations; an important remark is that demographic attributes (gender) can be used either as conditioningvariablesorasnuisancefactors,andthedesign choiceaffectsfairnessandgeneralization.

Garain et al. [9] developed GRaNN, a feature-selection framework using a golden ratio-aided neural network for emotion, gender, and speaker identification from voice signals.Themethodtypicallylearnsanembedding and applies a selection/weighting mechanism with constrained by sparsity. A common regularizeris

encouraging compactness. Their achievement is unifying multi-attribute recognition with feature economy; an importantremarkisthatmulti-labelsettingsrequirecareful loss design to avoid one attribute dominating shared embeddings.

Sánchez-Heviaetal.[10]studiedage-groupclassification andgenderrecognitionusingtemporalconvolutionalneural networks (TCNs), emphasizing long-range temporal modeling through dilated convolutions. A dilated 1D convolutionis where is dilation, enabling exponential receptive-field growth without deep recurrence. The achievement is capturingprosodicevolutionandphoneticdynamicsacross time;animportantremarkisthatpadding/causalitychoices (causalvs.non-causal)matterforreal-timeapplications.

RadhaandGowrisankari[11]developedadeeplearning approach combining spectral and prosodic features for speech gender classification, typically mixing short-term spectraldescriptorswithlonger-termstatisticssuchaspitch andenergycontours.Fundamentalfrequencyestimationis oftenusedasacue,wherepitch relatestoperiodicityin voiced speech and can be inferred from autocorrelation peaks. Prosodic sequences may be summarized via statistics

The achievement is demonstrating that combining complementary cues improves discriminability beyond single-family features; an important remark is that pitch trackingerrorsinnoisyconditionscandegradeperformance unlessrobustvoicingdetectionisincluded.

TrawickiandŻyła[12]studiedgenderclassificationusing emotionalspeech,comparingdeepfeaturelearningstrategies under affective variability. A typical approach learns

embeddingsfromlog-Melinputswithnormalizationandthen usesaclassifierhead .Thework highlightsthatemotionalstatesalterspectraltilt,speaking rate, and intensity, which can be modeled via learned representationsratherthanfixedheuristics.Theachievement is characterizing how emotion-conditioned speech shifts gendercues;animportantremarkisthatevaluationshould separate speaker identity from emotion to avoid confounding.

Guerrieri et al. [13] developed a two-level hierarchical system integrating gender identification within a speech emotionrecognitionpipeline,wheregenderisinferredfirst andthenusedtoadaptemotionclassification.Conditioning canbeexpressedbyfeaturemodulation

where and aregender-dependentscalingandbias functions.Theirachievementisshowingthatdemographic conditioningcanreduceintra-classvariancefordownstream tasks; an important remark is that such conditioning may encode bias if gender labels are noisy or non-binary categoriesareexcluded.

Zhang et al. [14] introduced a gender-specific deep learning method for speech emotion recognition by extractingandfusingfeaturestailoredtogender-dependent acousticpatterns.Featurefusioncanbeformulatedas

or via concatenation followed by projection. Their achievement is highlighting that gender-conditioned representationscanbettercapturedistinctpitchrangesand formantspacing;animportantremarkisthatgender-specific modelingimprovesseparabilitybutmayreduceportabilityif appliedtounseendemographicdistributions.

Vlaj and Zgank [15] studied acoustic gender and age classification as a tool for privacy-preserving speech processing, where sensitive demographic inference is acknowledged and controlled in downstream pipelines. A privacy-preserving objective may be expressed via adversariallearning,whereanencoder produces , andanadversary triestopredictasensitiveattribute :

Theachievementisframingdemographicinferencewithin privacy constraints; an important remark is that privacy goalsrequireexplicitthreatmodels(whatattackerknows andobserves).

Alkhammash[16]developedahybridensemblestacking modelforgendervoicerecognition,combiningmultiplebase learnersandameta-classifiertrainedontheiroutputs.Ifbase modelsyieldprobabilities ,stackingformsametafeaturevector andlearns .Regularized linearstackingcanbewrittenas

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

The achievement is increasing robustness to front-end variabilitythroughmodeldiversity;animportantremarkis thatstackingrequirescarefulcross-validationtoavoidoveroptimisticmeta-trainingduetoleakage.

Rizhinashvilietal.[17]studiedgenderneutralisationfor unbiasedspeechsynthesising,addressinghowgendercues canberemovedorcontrolledingeneratedspeechtoreduce bias.Acommonapproachuseslatent-variablemodelswhere anencodermapsspeechtoalatentvector andadecoder reconstructsspeech;genderneutralizationaimstoenforce invariance:

where denotesmutual information,often minimized implicitly via adversarial objectives. The achievement is demonstrating that controllable synthesis can decouple speaker attributes; an important remark is that neutralizationmayreducenaturalnessifconstraintsaretoo strong.

DeCarioetal.[18]introducededge-orientedmulti-task learningforgender,age,ethnicity,andemotionrecognition, emphasizingefficientinferenceforembeddeddevices.Model efficiency is typically achieved via depthwise separable convolution,

whichreducescomputationrelativetostandardconvolution. Their achievement is unifying multiple recognition tasks with shared computation under edge constraints; an importantremarkisthatdemographicinferenceon-device raisesethicalandconsentconsiderationsdespitetechnical feasibility.

Taranetal.[19]studiedspeakergenderidentificationfor voice assistants under device and channel variability, focusingonrobustnesstodomainshifts.Channeleffectscan be modeledas convolutionwithan impulse response andadditivenoise :

whichmotivatesaugmentationordomain-invariantlearning. Theachievementishighlightingtheneedforchannel-robust front-endsandtrainingstrategiesindeployedassistants;an important remark is that far-field microphones change spectral coloration and reverberation, requiring realistic augmentationbeyondsimplenoiseaddition.

ShagiandMoin[20]performedacomparativeanalysisfor gender recognition using acoustic features and machine learning under real-world speech conditions. Typical baselines include linear classifiers on MFCC/i-vectors or kernelmachineswheredecisionfunctionstaketheform

with such as the RBF kernel. Their achievement is establishingpracticaltrade-offsbetweenclassicalfeatures

and learned representations; an important remark is that real-worldevaluationshouldincludecross-corpustestingto reflectdeploymentreality.

Bhatetal.[21]developedadeeplearningapproachfor speech-based gender classification using spectral representations, typically employing CNN backbones over log-Mel or spectrogram images. Nonlinear activation in convolutionalstackscanberepresentedas

where is ReLU/GELU. The achievement is demonstrating that spectral images encode stable gender cuesrelatedtoharmonicspacingandformantpatterns;an importantremarkisthatconsistentamplitudenormalization isessentialtopreventthenetworkfromlearningloudness artifacts.

Lisetti et al. [22] studied identity, gender, age, and emotion recognition from speech using deep neural representations, emphasizing shared embeddings and disentanglementacrossattributes.Atypicaldisentanglement goalistofactorize intosubspaces and encourageconditionalindependencethroughauxiliaryheads andregularizers.Acommoncompactnessconstraintuses normonembeddings:

Their achievement is showcasing unified modeling for multiple paralinguistic tasks; an important remark is that multi-attribute systems must manage dataset annotation completeness,sincemissinglabelscanbiastraining.

Javid et al. [23] developed an attention-based deep learning method for multilingual voice-based gender recognition, handling variability across languages and phoneticinventories.Languagevariabilitycanbeapproached viasharedencoderswithlanguage-adaptivelayers;asimple adapterformulationis

where are small bottleneck parameters. The achievement is emphasizing that gender cues exist across languages but can shift due to phonotactics and speaking styles; an important remark is that multilingual training needs balanced sampling to prevent dominance of highresourcelanguages.

DeSimoneetal.[24]introducedamultimodalmulti-task approachintegratingvisualandaudiocuesforemotionand gender recognition, where audio embeddings and visual embeddingsarefusedforjointinference.Multimodalfusion iscommonlydoneby

or via cross-attention. The achievement is demonstrating that visual cues (lip motion, facial geometry) complement acousticcuesinchallengingnoiseconditions;animportant remark is that audio–visual synchronization errors can degrade fusion, so temporal alignment is a critical system component.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Markitantov and Verkholyak [25] studied occlusionrobust audiovisual gender recognition and age estimation using attention mechanisms, focusing on resilience when parts of the face are occluded. Attention can dynamically reweightmodalitiesviagating,

so the system relies more on audio when vision is compromised.Theachievementisstrengtheningrobustness under partial observability; an important remark is that robustness must be assessed across realistic occlusion patterns(masks,glasses)andreverberantaudio.

Anidjar et al. [26] introduced an objective evaluation methodology for gender classification from speech, emphasizingstandardizedprotocolsthatreduceconfounding effectssuchasspeakeroverlapandrecordingmismatch.A coreprincipleisensuringindependencebetweentrain/test speaker sets, i.e., , and consistent preprocessingpipelines.Theachievementisimprovingthe reproducibilityandinterpretabilityofgenderclassification claims; an important remark is that protocol design can significantlychangeconclusionsevenwithidenticalmodels.

Hu et al. [27] studied gender-sensitive speech emotion recognition via robust feature fusion, where genderdependent transformations are applied to stabilize representations under demographic variation. A typical approachusesconditionalnormalization, withparameters learnedpergendercondition.The achievementisdemonstratingthatconditioningcanreduce representation drift across demographic groups; an importantremarkisthatgender-sensitivedesignmustalso considernon-binaryidentitiesanddatasetlabellimitations.

Kirchhuebeletal.[28]identifiedlimitsofbinarygender recognition from speech by analyzing how voice carries genderdiversitybeyondbinarycategories.Fromamodeling view,theyemphasizethatobservedacousticfeatures are generatedfromoverlappingdistributions,e.g., wherelatentfactors maynotalignwithbinarylabels.The achievement is providing a critical perspective on classification assumptions and highlighting ambiguity regions; an important remark is that deploying binary genderclassifierscanbemisleadingandshouldbeframed withcaution,consent,andappropriateuncertaintyhandling.

Puri and Baghel [29] developed a voice-based gender recognition method that incorporates interpretability throughheatmap-styleanalysis,typicallybyattributingtime–frequency regions that drive predictions. A common attribution mechanism uses gradient-based saliency on spectrograms: where istheinputspectrogramand istheclassification loss. The achievement is improving transparency by localizing influential acoustic regions (often voiced harmonics);animportantremarkisthatattributionscanbe unstable,sosmoothingorintegratedgradientscanprovide moreconsistentexplanations.

Yıldırım and Bingöl [30] studied metaheuristic optimization to enhance voice-based gender classification andageestimation,usingsearchstrategiestotunefeature sets, model parameters, or classifier hyperparameters. A genericmetaheuristicformulationis where representstunableparametersand isanobjective tied totrainingloss or validationrisk.The achievement is demonstrating that systematic search can improve model stability across datasets without manual tuning; an importantremarkisthatoptimizationshouldbeconstrained topreventoverlycomplexsolutionsthatmaynotgeneralize acrossrecordingconditions.

Study Method Outcome

Sindhaetal. [1]

OptimizedANN withselfattentionon spectrotemporal features

Yücesoy[2] Stackedhybrid acoustic featureswith ensemble learning

Yücesoy[3] Multi-task ensemble learningforage andgender recognition

Yücesoy[4] 1Dand2D CNNson waveform-and spectrogrambasedinputs

Improved discrimination ofgenderspecificvocal cuesby focusingon informative voicedregions

Enhanced robustnessby combining cepstral, prosodic,and spectral descriptors

Jointlearning exploitsshared acoustic characteristics betweenage andgender

2DCNNs effectively capture formant structuresand

Remark / Limitation

Attentionmay overemphasize voicedframes; performancecan degradewith excessivesilence ornoisypitch tracking

High-dimensional featurestacking increases computational costandriskof overfitting

Negativetransfer mayoccuriftask importanceis imbalanced duringtraining

Strong dependencyon consistent spectrogram parameterization

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Study Method Outcome

Mavaddati [5]

Younisetal. [6]

Yueetal.[7]

ResNet-based transfer learningon spectrotemporal representations

Dataset-driven modelingwith speaker-level temporal pooling

Differential evolution-based feature selectionwith gender awareness

Yueetal.[8]

Garainetal. [9]

SánchezHeviaetal. [10]

Radhaand Gowrisankari [11]

Trawickiand Żyła[12]

Guerrieriet al.[13]

Zhangetal. [14]

Vlajand Zgank[15]

Genetic algorithmwith Fisherscore featureranking

Goldenratioaidedneural networkwith sparsefeature selection

TemporalCNN withdilated convolutions

Deeplearning withcombined spectraland prosodic features

Deepfeature learningon emotional speech

Hierarchical genderconditioned emotion recognition system

Gender-specific feature extractionand fusionnetwork

Adversarial learningfor privacy-aware gender inference

Remark / Limitation harmonic patterns and preprocessing

Deepresidual learning improves generalization acrossmultiple speaker attributes

Provides standardized Arabicspeech resourcesfor genderanalysis

Reduces redundant featureswhile preserving genderdiscriminative information

Compactand discriminative featuresubsets forgenderrelatedtasks

Unified frameworkfor emotion, speaker,and gender identification

Captureslongrangetemporal dependencies inspeech signals

Improved gender separation through complementary acousticcues

Maintains gender discrimination underaffective variability

Gender-aware conditioning reducesintraclassfeature variance

Captures genderdependent pitchand formant characteristics

Balances gender recognition withprivacy preservation objectives

Transferlearning maypropagate dataset-specific biasesifnot carefullyfinetuned

Celebrityspeech maynotfully represent conversationalor spontaneous speech

Evolutionary optimizationcan be computationally expensiveon largefeaturesets

Performance dependson stabilityofclass statisticsunder dataimbalance

Multi-attribute learningrequires carefulloss balancingto avoiddominance effects

Non-causal convolutions limitdirectrealtimedeployment

Pitchestimation errorsinnoisy speechcanaffect reliability

Emotionand speakeridentity mayactas confounding factors

Sensitiveto incorrectornoisy genderlabels

Limited generalizationto unseen demographic distributions

Requiresexplicit threatmodels andcareful adversarial tuning

Study Method Outcome

Remark / Limitation stackingof multiple classifiers throughmodel diversity leakagewithout strictvalidation

Rizhinashvili etal.[17]

DeCarioet al.[18]

Taranetal. [19]

Shagiand Moin[20]

Bhatetal. [21]

Lisettietal. [22]

Javidetal. [23]

DeSimoneet al.[24]

Markitantov and Verkholyak [25]

Anidjaretal. [26]

Huetal.[27]

Latent-variable modelingfor gender neutralization

Edge-oriented multi-task learningwith lightweight CNNs

Channel-robust gender recognitionfor voiceassistants

ClassicalML anddeep learning comparison underrealworldspeech

CNN-based gender classification usingspectral images

Shareddeep embeddingsfor identity, gender,and emotion

Attention-based multilingual gender recognition

Audio–visual multimodal fusionfor gender recognition

Attention-gated audiovisual fusionunder occlusion

Standardized evaluation protocolfor gender classification

Gendersensitive conditional normalization andfeature fusion

Decouples gender attributesfrom synthesized speech

Efficientondevicegender recognition alongsideother attributes

Improved resilienceto deviceand channel variability

Highlights trade-offs between handcrafted andlearned features

Stable extractionof harmonicand formant-based cues

Unified modelingof multiple paralinguistic attributes

Demonstrates cross-lingual consistencyof gendercues

Visualcues complement audiounder noisy conditions

Dynamic relianceon audioorvisual modality improves robustness

Improves reproducibility andfairnessof experimental results

Reduces demographicinducedfeature drift

Overneutralization canreduce naturalnessof generatedspeech

Ethicalconcerns regarding demographic inferenceonedge devices

Requiresrealistic augmentation beyondadditive noise

Cross-corpus generalization remains challenging

Sensitiveto amplitudescaling and normalization artifacts

Incomplete annotationscan biasshared representation learning

High-resource languagesmay dominate multilingual training

Performance dependson accurateaudio–visual synchronization

Limitedby availabilityof synchronized multimodaldata

Protocoldesign alonedoesnot addressinherent datasetbias

Binarygender assumptionslimit inclusivity

Kirchhuebel etal.[28]

Statistical analysisof gender diversityin speech

Reveals limitationsof binarygender classification models

Challenges applicabilityof conventional binarylabels

Alkhammash [16]

Hybrid ensemble Improves robustness

Stackingmay sufferfromdata

Puriand Baghel[29]

Interpretable gender recognition

Improves transparencyof modeldecisions

Attributionmaps canbeunstable without

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Study Method Outcome

Yıldırımand Bingöl[30]

Remark / Limitation withsaliency heatmaps smoothing techniques

Metaheuristic optimization forfeatureand parameter tuning

III. Proposed Work

Enhances stabilityacross datasets withoutmanual tuning

Riskofoverparameterization ifsearchspaceis unconstrained

IV.

Results and Analysis

Analysis of Selected Parameters

Figure 1 illustrates the overall workflow of the proposed speech-based gender recognition framework, designed by synthesizing insights from recent deep learning and optimization-driven studies. The process begins with raw speech acquisition, where continuous audio signals are collectedundervariedacousticconditions.Apreprocessing stage follows, incorporating voice activity detection and amplitude normalization to suppress silence, background noise, and channel-induced variability, thereby stabilizing theinputsignal.Thecleanedspeechisthentransformedinto compact acoustic representations using log-Mel spectrograms or MFCCs, which preserve gender-relevant cues such as pitch distribution, harmonic spacing, and spectraltilt.Theserepresentationsarefedintoalightweight convolutional neural network enhanced with an attention mechanism that selectively emphasizes informative time–frequency regions while suppressing redundant or noisy components.Themodelproducesagenderlabelalongwith anassociatedconfidencescore.Finally,ananalysismoduleis integratedtosupportablationstudies,robustnessevaluation under noise and channel mismatch, and explainability through attention or saliency visualization, ensuring both reliabilityandinterpretabilityoftheproposedframeworkin practicaldeploymentscenarios.

Figure 2: Effectofacousticfeaturetypeonspeech-based genderrecognitionperformance.

Figure2 presents the comparative analysis of acoustic feature representations. The log-Mel spectrogram consistently outperforms MFCC features, indicating its superiorabilitytopreserveharmonicspacingandspectral envelopevariationsthatarestronglycorrelatedwithgenderspecificvocaltraits.

Figure 3: Impact of utterance duration on gender recognition accuracy.

Figure 3 illustratestheimpactofutterancedurationon recognition performance. Short utterances of 1s show reduced reliability due to insufficient phonetic coverage, whereasperformancestabilizesbeyond2s,confirmingthata moderatetemporalcontextissufficientforeffectivegender discrimination.

Figure 1: Steps of Proposed Analysis

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net

4: Noise robustness analysis under varying signal-tonoise ratios.

Noise robustness is evaluated in Figure4 under clean, moderate,andseverenoiseconditions.Agradualdegradation isobservedasnoiseintensityincreases;however,themodel maintainsreasonablestabilityat10dBSNR,demonstrating robustness to moderate environmental noise commonly encounteredinreal-worldscenarios.

Figure 5: Effect of incorporating an attention mechanism in the proposed model.

Figure5 analyzes the contribution of the attention mechanism.Theinclusionofattentionleadstoanoticeable improvementbyselectivelyemphasizinginformativevoiced time–frequency regions while suppressing redundant or noisy components, validating its effectiveness for speechbasedgenderrecognition.

Figure 6: Channel mismatch analysis between studioquality and telephone-band speech.

p-ISSN: 2395-0072

ChannelmismatcheffectsareexaminedinFigure6 The performancereductionobservedfortelephone-bandspeech highlights the sensitivity of spectral representations to bandwidthlimitations,emphasizingthenecessityofchannelawaretrainingoraugmentationstrategiesfordeploymentorientedsystems.

Figure 7: Attention-based explainability analysis showing normalized attention responses.

Finally, Figure 7 demonstrates the explain ability behavior of the proposed framework. Strong attention localization on salient spectro-temporal regions confirms that model decisions are driven by acoustically meaningful cues, thereby enhancing interpretability and trust in the system for realworld applications.

The experimental analysis demonstrates that acoustic representation,temporalcontext,andarchitecturalchoices jointly influence the reliability of speech-based gender recognition. Log-Mel spectrograms consistently provide richergender-discriminativecuesthanMFCCsduetotheir superiorpreservationofharmonicandformantstructures. Performanceimproveswithincreasingutteranceduration and stabilizes beyond two seconds, indicating sufficient phoneticcoverageforrobustinference.Themodelexhibits gracefuldegradationundernoisyandchannel-mismatched conditions,confirmingitspracticalresilience.Incorporation ofanattentionmechanismfurtherenhancesdiscrimination byfocusingoninformativevoicedregionswhilesuppressing irrelevant components. Attention-based visualization validatesthatpredictionsareguidedbymeaningfulspectrotemporal patterns, supporting both interpretability and deploymentreadiness.

V. Conclusion

This study presented a compact and interpretable framework for speech-based gender recognition by systematically analyzing key factors that influence model reliability and robustness. Through focused experimentation,itwasobservedthatappropriateacoustic representations and sufficient temporal context play a critical role in preserving gender-specific vocal characteristics.Theintegrationofanattentionmechanism

Figure

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

proved effective in selectively emphasizing informative spectro-temporal regions, leading to improved discriminationwhilemaintainingalightweightarchitecture. Robustness analysis under noise and channel mismatch conditionshighlightedthemodel’ssuitabilityforreal-world deploymentscenarios,whererecordingenvironmentsand devices are often heterogeneous. Additionally, the incorporation of explain ability through attention visualization enhanced transparency, enabling a clearer understandingofthedecision-makingprocess.Overall,the proposed approach balances accuracy, robustness, and interpretability, offering a practical solution for speechbased gender recognition and establishing a strong foundationforfutureextensionsinvolvingcross-lingualdata, fairness-awaremodeling,andmulti-attributeparalinguistic analysis.

References:

[1] M. M. R. Sindha and D. K. Rana, “Optimized artificial neural network for vocal gender recognition using a self-attention mechanism,” ETRI Journal, 2024, doi: 10.4218/etrij.2024-0608.

[2] E.Yücesoy,“Genderrecognitionbasedonthestacking of different types of hybrid features created from speech,”AppliedSciences,vol.14,no.15,Art.no.6564, 2024,doi:10.3390/app14156564.

[3] E. Yücesoy, “Automatic age and gender recognition using ensemble models on speech datasets,” Applied Sciences, vol. 14, no. 16, Art. no. 6868, 2024, doi: 10.3390/app14166868.

[4] E.Yücesoy,“Speakerageandgenderrecognitionusing 1D and 2D convolutional neural networks,” Neural Computing and Applications, 2024, doi: 10.1007/s00521-023-09153-0.

[5] S.Mavaddati,“Voice-basedage,gender,andlanguage recognitionbasedonResNetdeepmodelandtransfer learning in spectro-temporal domain,” Neurocomputing,vol.580,Art.no.127429,2024,doi: 10.1016/j.neucom.2024.127429.

[6] H. A. Younis et al., “Creating the Hu-Int dataset: A comprehensive Arabic speech dataset for gender detection and age estimation of Arab celebrities,” BiomedicalSignalProcessingandControl,vol.96,Art. no.106511,2024,doi:10.1016/j.bspc.2024.106511.

[7] L. Yue et al., “Advanced differential evolution for gender-awareEnglishspeechemotionrecognitionwith optimal feature selection,” Scientific Reports, vol. 14, 2024,doi:10.1038/s41598-024-68864-z.

[8] L. Yue et al., “Gender-driven English speech emotion recognition with genetic algorithm optimization and Fisher score,” Biomimetics, vol. 9, no. 6, Art. no. 360, 2024,doi:10.3390/biomimetics9060360.

[9] A.Garainetal.,“GRaNN:Featureselectionwithgolden ratio-aided neural network for emotion, gender and speaker identification from voice signals,” Neural ComputingandApplications,vol.34,no.17,pp.14463–14486,2022,doi:10.1007/s00521-022-07261-x.

[10] H.A.Sánchez-Hevia,R.Gil-Pita,M.Utrilla-Manso,andM. Rosa-Zurera, “Age group classification and gender recognitionfromspeechwithtemporalconvolutional neuralnetworks,”MultimediaToolsandApplications, vol. 81, no. 3, pp. 3535–3552, 2022, doi: 10.1007/s11042-021-11614-4.

[11] J. Radha and N. Gowrisankari, “Speech gender classificationbasedonspectralandprosodicfeatures with deep learning,” International Journal of Speech Technology,2023,doi:10.1007/s10772-023-10039-8.

[12] J.TrawickiandM.Żyła,“Genderclassificationbasedon emotionalspeech:Deeplearningandfeaturelearning perspectives,” International Journal of Speech Technology,2024,doi:10.1007/s10772-024-10090-z.

[13] A.Guerrierietal.,“Genderidentificationinatwo-level hierarchical speech emotion recognition system,” Sensors, vol. 22, no. 5, Art. no. 1714, 2022, doi: 10.3390/s22051714.

[14] L. M. Zhang et al., “A deep learning method using gender-specific features for speech emotion recognition,”Sensors,vol.23,no.3,Art.no.1355,2023, doi:10.3390/s23031355.

[15] D. Vlaj and A. Zgank, “Acoustic gender and age classification as an aid to privacy-preserving speech processing,” Mathematics, vol. 11, no. 1, Art. no. 169, 2023,doi:10.3390/math11010169.

[16] E.H.Alkhammash,“Ahybridensemblestackingmodel forgendervoicerecognition,”Electronics,vol.11,no. 11, Art. no. 1750, 2022, doi: 10.3390/electronics11111750.

[17] D. Rizhinashvili et al., “Gender neutralisation for unbiasedspeechsynthesising,”Electronics,vol.11,no. 10, Art. no. 1594, 2022, doi: 10.3390/electronics11101594.

[18] A.DeCarioetal.,“Multi-tasklearningontheedgefor effective gender, age, ethnicity and emotion recognition,” Engineering Applications of Artificial Intelligence, vol. 117, Art. no. 105651, 2023, doi: 10.1016/j.engappai.2022.105651.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

[19] P. Taran, G. B. Kumar, and V. M. Suresh, “Speaker genderidentificationforvoiceassistantsunderdevice and channel variability,” Applied Acoustics, vol. 205, Art. no. 109271, 2023, doi: 10.1016/j.apacoust.2023.109271.

[20] S.ShagiandA.S.M.Moin,“Acomparativeanalysisfor genderrecognitionusingacousticfeaturesandmachine learninginreal-worldspeech,”AppliedAcoustics,vol. 190, Art. no. 108392, 2022, doi: 10.1016/j.apacoust.2021.108392.

[21] A. A. Bhat, R. K. Garg, and S. Jain, “A deep learning approachforspeech-basedgenderclassificationusing spectral representations,” Computer Systems Science and Engineering, 2024, doi: 10.32604/csse.2023.046730.

[22] C. Lisetti et al., “Identity, gender, age, and emotion recognition from speech using deep neural representations,” Cognitive Computation, 2024, doi: 10.1007/s12559-023-10241-5.

[23] Y. J. Javid, R. Ahmed, and A. Ali, “Voice-based gender recognitionusingattention-baseddeeplearningwith multilingualspeechsignals,”WirelessCommunications and Mobile Computing, vol. 2022, Art. no. 4444388, 2022,doi:10.1155/2022/4444388.

[24] G. De Simone, L. Greco, A. Saggese, and M. Vento, “Integrating visual and audio cues for emotion and gender recognition: A multi modal and multi task approach,” Information Fusion, 2025, doi: 10.1016/j.inffus.2025.104071.

[25] M.MarkitantovandO.Verkholyak,“Occlusion-robust audiovisual gender recognition and age estimation using attention mechanisms,” Expert Systems with Applications,2025,doi:10.1016/j.eswa.2025.127473.

[26] O.H.Anidjar,R.Marbel,andR.Yozevitch,“Anobjective gender classification evaluation methodology for speech,”ScientificReports,2025,doi:10.1038/s41598025-99011-x.

[27] Y. Hu, H. Zhang, and X. Li, “Gender-sensitive speech emotion recognition: A deep learning approach with robust feature fusion,” Scientific Reports, 2025, doi: 10.1038/s41598-025-14016-w.

[28] C.Kirchhuebel,H.Jones,andA.Simpson,“Voicecarries genderdiversity:Classificationlimitsofbinarygender recognitionfromspeech,”RoyalSocietyOpenScience, 2025,doi:10.1098/rsos.251193.

[29] S.PuriandV.Baghel,“Voice-basedgenderrecognition withHeatMapanalysisthroughmachinelearning,”SN Computer Science, 2025, doi: 10.1007/s42979-02503892-8.

[30] S.Yıldırımandİ.Bingöl,“Metaheuristicapproachesto enhance voice-based gender classification and age estimation,” Applied Sciences, vol. 15, no. 23, Art. no. 12815,2025,doi:10.3390/app152312815.

Turn static files into dynamic content formats.

Create a flipbook
Analysis of Framework for Robust Gender Recognition from Speech Signals by IRJET Journal - Issuu