
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Sanjeev Yadav1, Mrs. Arifa Khan2
1Master of Technology, Computer Science and Engineering, Lucknow Institute of Technology, Lucknow, India 2Assistant Professor, Department of Computer Science and Engineering, Lucknow Institute of Technology, Lucknow, India
Abstract - Large-scale language models (LLMs) have transformed natural language processing through selfsupervised pretraining on massive multilingual corpora. However,theexpansionofthesesystemsintomultilingualand low-resource linguistic settings has intensified concerns regarding bias transmission and amplification. This review synthesizescurrentempiricalevidenceonhowbiasesemerge, propagate, and persist in LLMs trained on multilingual lowresource data streams. Drawing from recent studies in computational linguistics, fairness-aware machine learning, and cross-lingual representation learning, the paper categorizesbiastransmissionmechanismsintothreeprincipal dimensions: data-driven bias, model-centric bias, and crosslingual transfer bias. Particular attention is given to dataset imbalance, tokenization artifacts, representational entanglement across languages, and the hierarchical dominance of high-resource languages during pretraining. Thereviewcriticallyexaminesexistingbiasdetectionmetrics, multilingualbenchmarklimitations,andmitigationstrategies, including data rebalancing, embedding debiasing, and fairness-constrained optimization. Despite notable progress, theliteraturerevealspersistentmethodologicalfragmentation and a lack of standardized evaluation protocols tailored to low-resourcecontexts.Thepaperconcludesbyidentifyingkey research gaps and proposing directions for developing linguistically inclusive, culturally sensitive, and empirically grounded fairness frameworks for next-generation multilingual language models.
Key Words: Bias transmission, Large-scale language models, Multilingual NLP, Low-resource languages, Fairness-aware machine learning, Cross-lingual transfer.
1.1 Background: Rise of Large-Scale Pretrained Language Models
1.1.1 Evolution of Transformer-Based Architectures
The emergence of Transformer architectures marked a paradigmshiftinnaturallanguageprocessingbyreplacing recurrent and convolutional structures with attention mechanismscapableofmodelinglong-rangedependencies (Vaswanietal.,2017).Thisarchitecturalinnovationenabled large-scale pretraining on massive corpora using self-
supervised objectives. Models such as BERT introduced masked language modeling for contextual representation learning (Devlin et al., 2019), while autoregressive frameworks such as GPT demonstrated the scalability of generative pretraining (Radford et al., 2019; Brown et al., 2020).Thesedevelopmentsestablishedthefoundationfor contemporarylargelanguagemodels(LLMs)characterized bybillionsofparametersandextensivedataexposure.
To extend performance beyond English-centric systems, multilingualmodelssuchasmBERTandXLM-Rweretrained on corpora spanning dozens to hundreds of languages (Conneauetal.,2020).Thesemodelsrelyonsharedsubword vocabularies and parameter sharing to facilitate crosslingual transfer. While multilingual pretraining improves performance in low-resource languages via transfer from high-resource counterparts, it also introduces representational entanglement that may transmit sociocultural biases across linguistic boundaries (Pires, SchlingerandGarrette,2019).
1.2.1
Bias in LLMs refers to systematic and unfair associations learnedfromtrainingdatathatdisproportionatelyfavoror disadvantage specific demographic, cultural, or linguistic groups.Empiricalstudieshavedemonstratedthatpretrained embeddingsencodegender,racial,andreligiousstereotypes reflective of societal distributions present in corpora (Bolukbasi et al., 2016; Caliskan, Bryson and Narayanan, 2017). In multilingual contexts, bias may manifest as linguistichierarchy,culturalmarginalization,ordifferential performanceacrosslanguages.
Fairness in NLP generally concerns equitable treatment across demographic groups and languages in both representationanddownstreamtaskperformance(Mehrabi et al., 2021). In multilingual settings, fairness extends to parity in accuracy, error distribution, and semantic representation across typologically diverse languages. However, defining fairness operationally remains

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
challenging due to differing sociolinguistic norms and contextualmeaningsacrosscultures.
Low-resource languages are characterized by limited digitized corpora, scarce annotated datasets, and minimal computationalinfrastructure.Multilingualmodelsattemptto mitigatetheseconstraintsthroughsharedparameterization andcross-lingualtransferlearning(WuandDredze,2019). Although beneficial, this approach often prioritizes highresource language structures, potentially reinforcing systemicimbalanceinrepresentationalquality.
1.3.1 Data Imbalance and Linguistic Dominance
Training corpora for multilingual LLMs are typically dominated by high-resource languages such as English, Mandarin, and Spanish. This disproportionate representationresultsinunevenparameterallocationand performancedisparities(Benderetal.,2021).Consequently, biasesembeddedindominant-languagedatacanpropagate into low-resource linguistic spaces through shared embeddingsandtransfermechanisms.
Inlow-resourcecontexts,languagetechnologiesoftenserve communities already marginalized in digital ecosystems. Biased outputs may therefore exacerbate exclusion, misrepresentation,orharmfulstereotyping.Empiricalaudits have shown that cross-lingual models may exhibit higher toxicity or stereotyping rates in certain non-English languagesduetoinsufficientcontextualgrounding(Nozza, BianchiandHovy,2022).Addressingsuchissuesisessential forequitableAIdeployment.
1.4.1 Scope
Thisreviewsystematicallysynthesizesempiricalresearchon bias transmission mechanisms in large-scale multilingual languagemodels,withparticularemphasisonlow-resource data streams. It encompasses data-centric, model-centric, andcross-lingualtransferperspectives,integratingfindings from computational linguistics, fairness research, and sociotechnicalAIstudies.
1.4.2 Objectives
Theprimaryobjectivesarethreefold:
To categorize and critically evaluate documented mechanisms through which bias is encoded and transmittedacrosslanguages.
To assess methodological approaches for measuring biasinmultilinguallow-resourcesettings.
Toidentifyunresolvedchallengesandproposeresearch directionsforconstructinglinguisticallyinclusiveand fairness-awarelargelanguagemodels.
2.1.1
BiasinNLPsystemsmanifestsinmultipleinterrelatedforms. Representationalbiasreferstotheunequalorstereotypical portrayalofsocialgroupswithintextualcorporaandlearned embeddings. Empirical evidence shows that word embeddingscapturegenderandracialstereotypesaligned with societal associations (Bolukbasi et al., 2016). Algorithmicbiasarisesfrommodeltrainingdynamicsand optimizationprocessesthatamplifyexistingdisparities,even whendatadistributionsappearneutral.Societalbiasreflects historical and cultural inequities embedded in source corpora, particularly web-scale data (Bender et al., 2021). Demographic bias is observed when systems yield systematically different performance outcomes across population groups, languages, or dialects, often disadvantagingmarginalizedcommunities(Blodgettet al., 2020). In multilingual settings, these categories overlap, especially when high-resource language norms dominate representationalspace.


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Quantitative bias assessment relies on embedding-based, probabilistic,andtask-levelmetrics.TheWordEmbedding AssociationTest(WEAT)measuresdifferentialassociations between target and attribute word sets (Caliskan, Bryson and Narayanan, 2017). Sentence-level extensions such as SEATandcontextualizedembeddingprobesevaluatebiasin transformer-based models. For generative models, bias is often quantified using log-likelihood differentials, toxicity classifiers, and stereotype scoring benchmarks (Nadeem, BethkeandziczandReddy,2021).Inmultilingualcontexts, cross-lingualalignmentmetricsandperformancedisparity indices are used to evaluate fairness across languages. However,thelackofstandardizedmultilingualbenchmarks complicatescross-studycomparability.
BiasinNLPsystemshassignificant ethical andregulatory implications. Biased outputs may reinforce harmful stereotypes, propagate misinformation, or marginalize linguistic minorities. From a governance perspective, fairness concerns intersect with emerging AI regulatory frameworksemphasizingaccountability,transparency,and non-discrimination.Sociotechnicalanalysesarguethatbias cannotbetreatedpurelyasatechnicalartifactbutmustbe contextualized within broader power structures shaping dataproductionandtechnologicaldeployment(Birhaneand Guest, 2021). In low-resource environments, biased NLP systems may exacerbate digital inequality by misrepresentingunder-documentedculturesandlanguages.
2.2.1
Large-scale language models are primarily built upon the Transformerarchitecture,whichleveragesmulti-headselfattentionmechanismstocapturecontextualdependencies across tokens (Vaswani et al., 2017). Unlike recurrent models, Transformers process sequences in parallel, enablingefficientscalingtobillionsofparameters.Positional encodingspreservesequenceorder,whilestackedencoder–decoderlayersenablehierarchicalrepresentationlearning. Thescalabilityofthisarchitecturehasfacilitatedthetraining ofincreasinglylargeanddata-intensivemodels.
2.2.2
Two dominant pretraining paradigms underpin modern LLMs. Masked Language Modeling (MLM), introduced in BERT, predicts randomly masked tokens to learn bidirectionalcontextualrepresentations(Devlinetal.,2019). Incontrast,autoregressivenext-tokenprediction,employed inGPT-stylemodels,generatestextsequentiallybymodeling conditionalprobabilitiesovertokens(Brownetal.,2020). Variantssuchassequence-to-sequencedenoisingobjectives
extend these paradigms for multilingual and generative tasks.Whiletheseobjectivesimprovegeneralization,they mayinadvertentlyencodecorrelationsreflectingbiasedcooccurrencepatternsinlarge-scalecorpora.
MultilingualmodelssuchasmBERTandXLM-Raretrained on concatenated corpora from multiple languages using shared subword vocabularies (Conneau et al., 2020). Parametersharingenablescross-lingualtransfer,allowing knowledgelearnedfromhigh-resourcelanguagestobenefit low-resourceones.Empiricalstudiesdemonstrateemergent cross-lingualalignmentwithoutexplicitsupervision(Pires, Schlinger and Garrette, 2019). However, this shared representation space may also facilitate the transfer of dominant-languagebiasesintounderrepresentedlinguistic contexts,leadingtorepresentationalimbalanceandfairness concerns.
Low-resource languages are those with limited digitized corpora, scarce annotated datasets, and minimal computational resources for NLP development. These languages often lack standardized orthography, domaindiversetextcollections,andlarge-scaleparallelcorpora.The disparityindigitalpresencebetweenhigh-andlow-resource languages creates structural inequities in model performanceandrepresentation(Joshietal.,2020).
Data scarcity introduces both quantitative and qualitative challenges.Limitedcorpussizerestrictsvocabularycoverage and contextual diversity, reducing model robustness. Additionally, multilingual training corpora are typically imbalanced,withhigh-resourcelanguagesdominatingtoken distribution. This imbalance leads to disproportionate parameter optimization toward majority languages, resulting in degraded performance and potential bias in minority language outputs (Wu and Dredze, 2019). Noise, code-switching, and orthographic variability further complicatemodeltraininginlow-resourceenvironments.
2.3.3SyntheticAugmentationandTransferLearning
To address scarcity, researchers employ synthetic data augmentation methods such as back-translation, machine translationbootstrapping,anddatasynthesisviagenerative models.Transferlearningapproachesleveragemultilingual pretraining to improve downstream performance in lowresource languages. While these techniques enhance accuracy, they may introduce translation artifacts or propagatestructuralbiasesfromsourcelanguages.Recent

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
studies advocate adaptive fine-tuning, balanced sampling strategies,andculturallygroundeddatacurationtomitigate these risks and improve equitable representation across linguisticcommunities.
This section synthesizes empirically identified pathways throughwhichbiasisencoded,amplified,andpropagatedin multilingual large-scale language models trained on lowresourcedatastreams.
3.1.1
Bias transmission often originates at the data acquisition stage. Large language models are typically pretrained on web-crawled corpora, encyclopedic sources, and usergenerated content, which reflect historical and societal asymmetries. Web-scale datasets such as Common Crawl disproportionately represent dominant linguistic and culturalnarratives,embeddinghegemonicperspectivesinto training corpora (Bender et al., 2021). For multilingual corpora, disparities in digital presence mean that highresource languages contribute substantially more textual volumeandtopicaldiversitythanlow-resourcelanguages, predisposingmodelstowardskewedrepresentations.
3.1.2
Token-level imbalance across languages constitutes a primary mechanism of bias transmission. During multilingualpretraining,samplingstrategiesoftenallocate training probability proportional to corpus size, favoring high-resourcelanguages.Empiricalstudiesdemonstratethat such imbalance leads to uneven parameter updates and performancegapsacrosslanguages(Conneauetal.,2020). Temperature-basedsamplingpartiallymitigatesdominance effects, yet residual disparities persist, particularly for morphologically rich or underrepresented languages (Arivazhagan et al., 2019). Consequently, low-resource languages may inherit structural biases embedded in majority-languagedistributions.
Beyond quantitative imbalance, qualitative cultural skew shapesrepresentationalbias.Linguisticcorporafrequently overrepresent urban, standardized, or majority dialects whilemarginalizingregionalorminorityspeechforms.This imbalanceresultsinsociolinguisticunderrepresentationand misclassificationindownstreamtasks(Blodgettetal.,2020). Inmultilingualcontexts,culturallyspecificmeaningsmaybe misaligned or flattened during joint training, leading to semantichomogenizationthatprivilegesdominantcultural frameworks.
Self-supervised objectives such as masked language modelingandautoregressivenext-tokenpredictionoptimize likelihood-based learning without fairness constraints. These objectives reinforce high-frequency associations present in training data, thereby amplifying stereotypical correlations (Zhao et al., 2019). In multilingual settings, biased co-occurrence patterns in high-resource languages may be statistically reinforced and propagated through shared representations, increasing the persistence of demographicstereotypes.
Cross-lingual transfer enables performance gains in lowresource languages through shared embedding spaces. However, empirical analyses reveal that representational alignment may transmit biases from dominant languages intostructurallydistinctones(Pires,SchlingerandGarrette, 2019). For example, gender associations embedded in Englishcorporacaninfluencerepresentationinlanguages withgrammaticalgendersystems,producingcompounded bias effects. Such transfer-induced bias is particularly pronounced when low-resource data lacks sufficient counterbalancingexamples.
Multilingual LLMs rely on shared parameters across languagestoachievescalability.Whileefficient,thisstrategy induces representational entanglement, whereby latent featuresencodeoverlappinglinguisticandculturalsignals. Research indicates that entangled representations can obscurelanguage-specificnuancesandpropagatemajoritylanguage dominance in embedding geometry (Wu and Dredze,2019).Thisstructuralcouplingservesasaconduit for biastransmission,particularlywhen model capacityis insufficienttodifferentiatediverselinguisticpatterns.
3.3.1
Attention mechanisms dynamically weight contextual tokens,influencinghowassociationsareformed.Analysesof transformerattentionpatternssuggestthatattentionheads may disproportionately focus on socially salient or stereotypically associated tokens (Vig et al., 2020). Such contextualweightingcanmagnifybiasedassociationsduring both pretraining and inference. In multilingual models, attentiondistributionsmayvaryacrosslanguages,leadingto inconsistent semantic emphasis and uneven bias manifestation.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Subwordtokenizationschemes,suchasByte-PairEncoding and SentencePiece, create shared vocabularies across languages.Whileenablingefficientparametersharing,these methodsmayfragmentlow-resourcelanguage wordsinto disproportionately many subword units, reducing representational fidelity (Rust et al., 2021). Tokenization artifacts can distort semantic coherence and introduce frequency-based distortions, effectively privileging morphologically simpler or high-frequency language segments.
Optimization strategies influence how gradients are allocated across languages and examples. Standard crossentropylossfunctionstreatalltokensequally,regardlessof language representation or social sensitivity. Without reweighting mechanisms, optimization tends to prioritize frequentpatterns,reinforcingmajority-languagedominance. Adaptive training methods, including fairness-aware regularization, have been proposed to counteract such biases, though empirical validation in multilingual lowresourcesettingsremainslimited(Mehrabietal.,2021).
3.4.1
Bias evaluation commonly employs benchmark datasets designedtoprobestereotypes,toxicity,orfairness-related disparities. Resources such as StereoSet and CrowS-Pairs measure stereotypical associations in generative and masked language models (Nadeem, Bethke and Reddy, 2021). However, these benchmarks are predominantly English-centric,limitingtheirgeneralizabilitytomultilingual contexts.Cross-lingualevaluationframeworksareemerging butremainuneveninlinguisticcoverage.
Existing bias metrics often assume culturally consistent semanticcategories,whichmaynotholdacrosslanguages. Direct translation of benchmarks can introduce semantic drift and fail to capture culturally specific stereotypes. Furthermore,low-resourcelanguagesfrequentlylackgoldstandardannotatedfairnessdatasets,constrainingreliable empiricalanalysis.Scholarsarguethatmultilingualfairness assessmentrequiresculturallygroundedevaluationdesign and community-informed annotation practices to ensure contextualvalidity(BirhaneandGuest,2021).Withoutsuch adaptations,benchmarkingprocessesriskunderestimating biasseverityinmarginalizedlinguisticsettings.
Thissectionorganizespriorscholarshipchronologicallyand thematically to trace the evolution of bias research from monolingual embeddings to multilingual large-scale languagemodels,withemphasisonlow-resourcecontexts.
Initial empirical investigations into bias in NLP focused primarily on English word embeddings. A seminal contribution was the Word Embedding Association Test (WEAT), which demonstrated that distributional embeddingsencodehuman-likeimplicitbiases,particularly alonggenderandracial dimensions(Caliskan,Brysonand Narayanan, 2017). Earlier work showed that analogical reasoning in embeddings reproduced stereotypical associations, such as linking professions with specific genders(Bolukbasietal.,2016).Thesestudiesestablished that statistical co-occurrence patterns in corpora are sufficient to encode socially salient stereotypes, even withoutexplicitlabeling.
Following detection, research turned toward mitigation strategies.Hardandsoftdebiasingmethodswereproposed to remove gender subspaces from static embeddings (Bolukbasi et al., 2016). Subsequent work introduced counterfactualdataaugmentationandadversarialtrainingto reduce demographic signal leakage in contextual models (Zhao et al., 2018). However, later analyses argued that many debiasing methods reduce measurable bias without fully eliminating latent associations, suggesting the persistence of deeper representational issues (Gonen and Goldberg, 2019). These early findings laid the methodological groundwork for bias evaluation in more complexarchitectures.
TheintroductionofmultilingualBERT(mBERT)markeda transitiontowardcross-lingualpretrainedrepresentations trainedonjointlyconcatenatedcorpora(Devlinetal.,2019). Cross-lingualLanguageModel(XLM)anditsrobustvariant XLM-Rfurtherexpandedcoverageandimprovedalignment across 100 languages (Conneau et al., 2020). Encoder–decoder architectures such as mT5 extended multilingual modelinginto generativeandtranslationtasks(Xue etal., 2021).Thesemodelsdemonstratedemergentcross-lingual transferwithoutexplicitalignmentobjectives.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Empiricalstudiesobservedthatmultilingualmodelsexhibit hierarchical language performance patterns, with highresource languages benefiting disproportionately from model capacity (Wu and Dredze, 2019). Cross-lingual probing experiments revealed that alignment quality depends on typological similarity and data volume, often privileging Indo-European languages. This “language hierarchyeffect”impliesthatmultilingualpretrainingmay structurally reinforce dominance relationships among languages,raisingconcernsaboutfairnessinlow-resource contexts.
4.3.1StudiesonAfrican,SouthAsian,andIndigenous
More recent research has shifted focus toward underrepresented linguistic communities. Analyses of multilingual embeddings have uncovered differential stereotypepatternsin Africanand SouthAsianlanguages, often shaped by translation artifacts and sociocultural variation (Nekoto et al., 2020). Investigations into Indigenouslanguagemodelinghighlightbothdatascarcity and cultural misrepresentation as sources of bias. These studies emphasize that bias manifestations are contextspecificandcannotbeassumedtomirrorEnglish-language patterns.
Quantitativebiasdetectionreliesontranslatedbenchmarks and embedding association tests adapted for multilingual
evaluation.However,scholarsarguethatpurelyquantitative metrics may overlook culturally embedded meanings and localized forms of discrimination (Blodgett et al., 2020). Qualitative analyses, including community-informed annotation and discourse-level examination, provide complementaryinsightsintonuancedbiasexpressions.The literatureincreasinglysupportsmixed-methodapproaches toensurecontextualvalidityinlow-resourcesettings.
Comparativeevaluationsacrosslanguagesrevealthatbias levels vary depending on linguistic structure, corpus composition,andculturalframing.Studiescomparinggender bias across multiple European and Asian languages demonstratebothsharedandlanguage-specificstereotype patterns(LauscherandGlavaš,2019).Suchfindingsindicate thatmultilingualmodelsdonotuniformlydistributebiasbut insteadreshapeitaccordingtorepresentationalalignment dynamics.
4.4.2EvidenceofBiasPropagationfromHigh-toLow-
Research on cross-lingual transfer suggests that biases presentindominant-languagecorpora canpropagateinto low-resource languages via shared embeddings and parameter coupling (Pires, Schlinger and Garrette, 2019). For instance, stereotypical occupational associations encoded in English may influence predictions in typologically distinct languages lacking equivalent data distributions.Thispropagationeffectunderscorestheroleof multilingualjointtrainingasastructuralmechanismforbias transmissionratherthanmerelyapassivereflectionoflocal corpora.
4.5.1
Data-centricmitigationstrategiesaimtoaddressbiasatthe source.Techniquesincludecorpusbalancing,oversampling underrepresented groups, and curating culturally diverse datasets.Temperature-basedresamplinghasbeenemployed to reduce dominance effects in multilingual pretraining (Arivazhaganetal.,2019).Counterfactualdataaugmentation further attempts to neutralize gendered or demographic associationsatscale.However,ensuringculturalauthenticity whilebalancingrepresentationremainsanopenchallenge.
Model-centricinterventionsoperatewithintrainingorfinetuningprocesses.Approachesincludeadversarialdebiasing, fairness-awareregularization,andconstrainedoptimization

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
techniques designed to minimize demographic signal retention (Mehrabi et al., 2021). Recent multilingual researchexploresdisentangledrepresentationlearning to separate language identity from sociocultural attributes. Despitepromisingresults,empiricalevidencesuggeststhat mitigation often reduces surface-level bias metrics while deeper structural associations persist, indicating the need formoretheoreticallygroundedfairnessframeworks.
This section integrates the reviewed literature to identify convergentpatterns,methodologicaltensions,andstructural insightsregarding bias transmissioninmultilingual largescalelanguagemodelstrainedonlow-resourcedatastreams.
5.1.1
Aconsistent pattern across empirical investigationsisthe structural dominance of high-resource languages in multilingual pretraining. Studies demonstrate that token imbalance and corpus heterogeneity lead to disproportionate parameter allocation, yielding higher representationalfidelityanddownstreamperformancefor majoritylanguages(Conneauetal.,2020).Thisdominance effect extends beyond accuracy disparities, influencing embeddinggeometryandsemanticalignment.Asa result, biasesembeddedinhigh-resourcecorporaoftenshapethe sharedlatentspaceofmultilingualmodels.
Research further indicates that large-scale pretraining amplifies pre-existing societal stereotypes rather than merely reflecting them. Likelihood-based optimization reinforces high-frequency co-occurrences, intensifying stereotypical associations related to gender, race, or profession(Zhaoetal.,2019).Inmultilingualcontexts,these amplifiedassociationsmaypermeatelinguisticallydistinct environments through shared parameters, producing translingual bias patterns even where local corpora lack equivalentdistributions.
The literature reports inconsistent bias magnitudes depending on evaluation methodology. Embedding associationtestsoftendetectsignificantstereotypeeffects, while task-basedevaluationssometimesreveal weaker or context-dependentdisparities(GonenandGoldberg,2019). Differences in template design, translation fidelity, and annotation quality contribute to divergent findings. In
multilinguallow-resourcesettings,translatedbenchmarks may distort semantic nuance, leading to measurement artifactsratherthangenuinebiasestimation.
5.2.2
Scholarlydisagreementalsoexists regardingtheextentto whichcross-lingualtransferpropagatesbias.Somestudies argue that multilingual alignment facilitates bias diffusion from dominant languages (Pires, Schlinger and Garrette, 2019),whereasothersreportlanguage-specificattenuation effectsdependingonmorphologicalorsyntacticdivergence. Such discrepancies often arise from differences in corpus composition, sampling temperature, and model capacity, indicating the absence of standardized experimental protocolsformultilingualfairnessresearch.
5.3.1Strengths:ScalabilityandDiagnosticInnovation
Current research demonstrates methodological sophistication in bias diagnostics. The development of benchmark datasets such as StereoSet and multilingual probing tasks has enabled systematic evaluation across architectures(Nadeem,BethkeandReddy,2021).Moreover, multilingual pretraining offers scalability advantages, enabling low-resource languages to benefit from shared representations.Dataaugmentationandadaptivesampling strategiesfurtherillustrateinnovativeattemptstoaddress imbalanceatscale.
5.3.2
Despite progress, existing approaches exhibit notable limitations.ManyfairnessbenchmarksareEnglish-centric, anddirecttranslationoftenfailstocaptureculturallyspecific stereotypes or linguistic subtleties (Blodgett et al., 2020). Additionally, mitigation techniques frequently target surface-level statistical parity without addressing deeper sociotechnical roots of bias. The absence of culturally grounded evaluation frameworks limits interpretability, particularlyinIndigenousandunder-documentedlanguage contexts.
5.4.1
Linguistic typology influences how bias manifests and propagates.Languagesdifferinmorphologicalcomplexity, grammatical gender systems, and word order structures. Research indicates that typological similarity facilitates cross-lingualalignment,therebyincreasingthelikelihoodof

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
bias transfer between structurally related languages (Wu and Dredze, 2019). Conversely, typologically distant languages may experience attenuated but not eliminated bias effects due to shared subword vocabularies and parameterentanglement.
5.4.2 Grammatical Gender and Structural Reinforcement
Languageswithgrammaticalgendersystemspresentunique biasdynamics.Whenpretrainedonmixedcorpora,models mayconflategrammaticalandsocialgendercues,reinforcing stereotypical occupational or role-based associations (Lauscher and Glavaš, 2019). In contrast, gender-neutral languages may exhibit bias primarily through semantic associations rather than morphological marking. This typologicalinterplayunderscoresthatbiastransmissionis not solely data-driven but also mediated by structural linguisticproperties.
Thisreviewsynthesizedempiricalandconceptualresearch on bias transmission mechanisms in large-scale language modelstrainedonmultilinguallow-resourcedatastreams. The analysis demonstrates that bias is not confined to isolated model components but emerges from the interactionbetweendataimbalance,pretrainingobjectives, architectural design, and cross-lingual parameter sharing. High-resource languages exert structural influence during multilingual training, shaping embedding geometry and facilitating the propagation of societal stereotypes into underrepresented linguistic contexts. While significant methodological advancements have been made in bias detection and mitigation, current evaluation frameworks remain predominantly English-centric and insufficiently adaptedtoculturallydiverseenvironments.Evidencefurther indicates that typological characteristics, tokenization strategies, and optimization dynamics modulate how bias manifests across languages. Despite progress in fairnessawaremodelinganddata-centricinterventions,mitigation efforts often address surface-level statistical disparities ratherthandeepersociotechnicaldeterminants.Overall,the literature underscores the need for standardized multilingual benchmarks, culturally grounded evaluation protocols, and typology-aware fairness frameworks. Advancing equitable multilingual NLP requires interdisciplinary collaboration integrating computational rigor,linguisticinsight,andethicalaccountabilitytoensure inclusiveandsociallyresponsiblelanguagetechnologies.
Thisreviewislimitedbytherapidlyevolvingnatureoflargescalelanguagemodelresearch,wherenewarchitecturesand evaluationbenchmarksemergefrequently.Althoughefforts were made to include diverse multilingual studies, the availableliteratureremainsskewedtowardwidelystudied
languages,potentiallyrestrictingcoverageofextremelylowresourceorIndigenouscontexts.Additionally,manyexisting studies rely on heterogeneous experimental settings, limiting direct comparability across findings. The review synthesizes published empirical evidence but does not include meta-analytic statistical aggregation of results. Finally,accesstoproprietarytrainingdataandlarge-scale model parameters constrains transparency in several referenced works, affecting comprehensive assessment of biastransmissionmechanisms.
1. Arivazhagan, N., Bapna, A., Firat, O., Lepikhin, D., Johnson, M., Krikun, M., Chen, M., Cao, Y., Foster, G., Cherry, C. and Macherey, W. (2019) ‘Massively multilingual neural machine translation in the wild: Findings and challenges’, arXiv preprint arXiv:1907.05019.
2. Bender, E.M., Gebru, T., McMillan-Major, A. and Shmitchell, S. (2021) ‘On the dangers of stochastic parrots:Canlanguagemodelsbetoobig?’,Proceedings of the 2021 ACM Conference on Fairness, Accountability,andTransparency(FAccT’21),pp.610–623.
3. Birhane,A.andGuest,O.(2021)‘Towardsdecolonising computationalsciences’,Patterns,2(10),100289.
4. Blodgett,S.L.,Barocas,S.,DauméIII,H.andWallach,H. (2020) ‘Language (technology) is power: A critical surveyof“bias”inNLP’,Proceedingsofthe58thAnnual Meeting of the Association for Computational Linguistics,pp.5454–5476.
5. Bolukbasi,T.,Chang,K.-W.,Zou,J.Y.,Saligrama,V.and Kalai,A.T.(2016)‘Manistocomputerprogrammeras womanistohomemaker?Debiasingwordembeddings’, Advances in Neural Information Processing Systems, 29,pp.4349–4357.
6. Brown,T.B.,Mann,B.,Ryder,N.,Subbiah,M.,Kaplan,J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell,A.etal.(2020)‘Languagemodelsarefew-shot learners’,AdvancesinNeuralInformationProcessing Systems,33,pp.1877–1901.
7. Caliskan, A., Bryson, J.J. and Narayanan, A. (2017) ‘Semantics derived automatically from language corpora contain human-like biases’, Science, 356(6334),pp.183–186.
8. Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek,G.,Guzmán,F.,Grave,E.,Ott,M.,Zettlemoyer, L.andStoyanov,V.(2020)‘Unsupervisedcross-lingual representation learning at scale’, Proceedings of the

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
58th Annual Meeting of the Association for ComputationalLinguistics,pp.8440–8451.
9. Devlin, J., Chang, M.-W., Lee, K. and Toutanova, K. (2019) ‘BERT: Pre-training of deep bidirectional transformersforlanguageunderstanding’,Proceedings ofNAACL-HLT2019,pp.4171–4186.
10. Gonen, H. and Goldberg, Y. (2019) ‘Lipstick on a pig: Debiasingmethodscoverupsystematicgenderbiases in word embeddings but do not remove them’, ProceedingsofNAACL-HLT2019,pp.609–614.
11. Joshi,P.,Santy,S.,Budhiraja,A.,Bali,K.andChoudhury, M.(2020)‘Thestateandfateoflinguisticdiversityand inclusion in the NLP world’, Proceedings of the 58th Annual Meetingof the AssociationforComputational Linguistics,pp.6282–6293.
12. Lauscher,A.andGlavaš,G.(2019)‘Areweconsistently biased? Multidimensional analysis of biases in distributionalwordvectors’,Proceedingsofthe2019 Workshop on Gender Bias in Natural Language Processing,pp.85–91.
13. Mehrabi,N.,Morstatter,F.,Saxena,N.,Lerman,K.and Galstyan, A. (2021) ‘A survey on bias and fairness in machinelearning’,ACMComputingSurveys,54(6),pp. 1–35.
14. Nadeem,M.,Bethke,A.andReddy,S.(2021)‘StereoSet: Measuring stereotypical bias in pretrained language models’,Proceedingsofthe59thAnnualMeetingofthe Association for Computational Linguistics, pp. 5356–5371.
15. Nekoto, W., Marivate, V., Matsila, T., Fasubaa, T., Fagbohungbe,T.,Akinola,S.O.,Muhammad,S.,Kabongo, S., Osei, S., Sackey, F. et al. (2020) ‘Participatory researchforlow-resourcedmachinetranslation:Acase studyinAfricanlanguages’,FindingsoftheAssociation for Computational Linguistics (EMNLP 2020), pp. 2144–2160.
16. Pires, T., Schlinger, E. and Garrette, D. (2019) ‘How multilingualismultilingualBERT?’,Proceedingsofthe 57th Annual Meeting of the Association for ComputationalLinguistics,pp.4996–5001.
17. Rust,P.,Pfeiffer,J.,Vulić,I.,Ruder,S.andGurevych,I. (2021) ‘How good is your tokenizer? On the monolingual performance of multilingual language models’,ProceedingsofACL2021,pp.3118–3135.
18. Vaswani,A.,Shazeer,N.,Parmar,N.,Uszkoreit,J.,Jones, L., Gomez, A.N., Kaiser, Ł. and Polosukhin, I. (2017) ‘Attention is all you need’, Advances in Neural InformationProcessingSystems,30,pp.5998–6008.
19. Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer,Y.andShieber,S.(2020)‘Investigatinggender bias in language models using causal mediation analysis’, Advances in Neural Information Processing Systems,33.
20. Wu,S.andDredze,M.(2019)‘Beto,Bentz,Becas:The surprising cross-lingual effectiveness of BERT’, ProceedingsofEMNLP-IJCNLP2019,pp.833–844.
21. Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A. and Raffel, C. (2021) ‘mT5: A massively multilingual pre-trained text-to-text transformer’, Proceedings of NAACL-HLT 2021, pp. 483–498.
22. Zhao,J.,Wang,T.,Yatskar,M.,Ordonez,V.andChang, K.-W. (2018) ‘Gender bias in coreference resolution: Evaluation and debiasing methods’, Proceedings of NAACL-HLT2018,pp.15–20.
23. Zhao, J., Zhou, Y., Li, Z., Wang, W. and Chang, K.-W. (2019) ‘Learning gender-neutral word embeddings’, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 4847–4853.
24. Adebayo,J.,Muelly,M.,Liccardi,I.andKim,B.(2020) ‘Debuggingtestsformodelexplanations’,Advancesin Neural Information Processing Systems, 33, pp. 700–712.
25. Barocas, S., Hardt, M. and Narayanan, A. (2019) Fairness and Machine Learning: Limitations and Opportunities.Cambridge,MA:fairmlbook.org.
26. Blodgett,S.L.andO’Connor,B.(2017)‘Racialdisparity innaturallanguageprocessing:Acasestudyofsocial media African-American English’, Proceedings of Fairness,Accountability,andTransparencyinMachine Learning(FAT/ML).
27. Costa-jussà, M.R., Cross, J., Çelebi, O. and Heafield, K. (2020) ‘Google’s multilingual neural machine translation system: Enabling zero-shot translation’, Transactions of the Association for Computational Linguistics,8,pp.339–351.
28. Dinan,E.,Fan,A.,Williams,A.,Urbanek,J.,Kiela,D.and Weston,J.(2020)‘Queensarepowerfultoo:Mitigating gender bias in dialogue generation’, Proceedings of EMNLP2020,pp.8173–8188.
29. Dodge, J., Gururangan, S., Card, D., Schwartz, R. and Smith, N.A. (2021) ‘Documenting large webtext corpora: A case study on the Colossal Clean Crawled Corpus’,ProceedingsofEMNLP2021,pp.1286–1305.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
30. Ethayarajh, K. (2019) ‘How contextual are contextualizedwordrepresentations?’,Proceedingsof ACL2019,pp.55–65.
31. Friedman, B. and Nissenbaum, H. (1996) ‘Bias in computersystems’,ACMTransactionsonInformation Systems,14(3),pp.330–347.
32. Garg,N.,Schiebinger,L.,Jurafsky,D.andZou,J.(2018) ‘Word embeddings quantify 100 years of gender and ethnic stereotypes’, Proceedings of the National AcademyofSciences,115(16),pp.E3635–E3644.
33. Hardt, M., Price, E. and Srebro, N. (2016) ‘Equality of opportunityinsupervisedlearning’,AdvancesinNeural InformationProcessingSystems,29,pp.3315–3323.
34. Hovy, D. and Spruit, S.L. (2016) ‘The social impact of naturallanguageprocessing’,ProceedingsofACL2016, pp.591–598.
35. Huang, P.-S., Stoyanov, V. and Zettlemoyer, L. (2019) ‘Addressing the data scarcity issue in multilingual languagemodeling’,ProceedingsofCoNLL2019,pp.1–10.
36. Jiang,Z.,Araki,J.,Ding,H.andNeubig,G.(2020)‘How can we know what language models know?’, Transactions of the Association for Computational Linguistics,8,pp.423–438.
37. Liang,P.P.,Wu,C.,Morency,L.-P.andSalakhutdinov,R. (2020)‘Towardsunderstandingandmitigatingsocial biasesinlanguagemodels’,ProceedingsofICML2020 WorkshoponResponsibleAI.
38. Nozza, D., Bianchi, F. and Hovy, D. (2021) ‘Honest: Measuring hurtful sentence completion in language models’,ProceedingsofNAACL2021,pp.2398–2406.
39. Raji,I.D.,Smart,A.,White,R.N.,Mitchell,M.,Gebru,T., Hutchinson,B.,Smith-Loud,J.,Theron,D.andBarnes,P. (2020)‘ClosingtheAIaccountabilitygap:Definingan end-to-end framework for internal algorithmic auditing’,ProceedingsofFAccT2020,pp.33–44.
40. Sheng, E., Chang, K.-W., Natarajan, P. and Peng, N. (2019)‘Thewomanworkedasababysitter:Onbiases inlanguagegeneration’,ProceedingsofEMNLP-IJCNLP 2019,pp.3407–3412.
41. Strubell,E.,Ganesh,A.andMcCallum,A.(2019)‘Energy and policy considerations for deep learning in NLP’, ProceedingsofACL2019,pp.3645–3650.
42. Talat,Z.,Rogers,A.,Schmaltz,A.,Suresh,H.,Blodgett,S., DauméIII,H.,Choi,Y.andSmith,N.A.(2022)‘Aholistic approach to documenting datasets and models for
responsibleNLP’,CommunicationsoftheACM,65(3), pp.90–99.
43. Touvron,H.,Lavril,T.,Izacard,G.,Martinet,X.,Lachaux, M.-A.,Lacroix,T.,Rozière,B.,Goyal,N.,Hambro,E.and Azhar,F.(2023)‘LLaMA:Openandefficientfoundation languagemodels’,arXivpreprintarXiv:2302.13971.
44. Ziems, N., Held, W., Shaikh, O., Chen, J., Zhang, D. and Yang,D.(2022)‘Canlargelanguagemodelstransform computational social science?’, Computational Linguistics,48(3),pp.1–29.