Skip to main content

A REVIEW OF DECENTRALIZED COLLABORATIVE MODEL TRAINING FOR MEDICAL DATA USING FEDERATED LEARNING WIT

Page 1


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

A REVIEW OF DECENTRALIZED COLLABORATIVE MODEL TRAINING FOR MEDICAL DATA USING FEDERATED LEARNING WITH ENHANCED PRIVACY CONTROLS

1Master of Technology, Computer Science and Engineering, Bansal Institute of Engineering & Technology, Lucknow, India

2Assistant Professor, Department of Computer Science and Engineering, Bansal Institute of Engineering & Technology, Lucknow, India ***

Abstract - Therapiddigitizationofhealthcaresystemshas generated vast volumes of sensitive medical data, creating significant opportunities for advanced machine learning applications while simultaneously raising critical privacy concerns.Traditionalcentralizedtrainingapproachesrequire dataaggregation,whichconflictswithregulatoryframeworks and institutional data governance policies. Federated Learning (FL) has emerged as a promising paradigm that enables collaborative model training across distributed medical institutions without transferring raw patient data. This review systematically examines decentralized collaborative model training frameworks for healthcare applications, with a particular focus on enhanced privacypreservingmechanismsintegratedintoFLsystems.Itanalyzes architectural models, secure aggregation protocols, differential privacy techniques, homomorphic encryption schemes, and block chain-assisted decentralized infrastructures. The paper further evaluates performance trade-offsbetweenmodelaccuracy,communicationefficiency, andprivacyguarantees,highlightingchallengessuchasnonIID data distribution, adversarial attacks, scalability limitations, and regulatory compliance. By synthesizing current research trends and identifying persistent technical and ethical gaps, this review outlines future research directions aimed at achieving secure, scalable, and trustworthy decentralized medical AI systems. The study provides a structured foundation for researchers and practitioners developing privacy-aware federated learning frameworksinhealthcareenvironments.

Key Words: Federated Learning; Decentralized Machine Learning; Medical Data Privacy; Differential Privacy; Secure Aggregation; Healthcare AI

1. INTRODUCTION

Theintegrationofartificialintelligence(AI)intohealthcare has transformed diagnostic systems, clinical decision support,medicalimaginganalytics,andpredictivemodeling. However,theeffectivenessofmachinelearning(ML)models largelydependsonaccesstolarge-scale,diverse,andhighquality datasets. In the medical domain, such data are inherently sensitive and governed by strict regulatory frameworks, making centralized data aggregation increasingly impractical. Federated Learning (FL) has

emerged as a distributed learning paradigm that enables collaborative model training without sharing raw data, therebyaddressingprivacyandcomplianceconcernswhile maintaining analytical utility (McMahan et al., 2017). This sectioncontextualizesthemotivation,challenges,andscope of decentralized collaborative training in medical environments.

1.1 Background and Motivation

Healthcareinstitutionsgenerateheterogeneousdatastreams including electronic health records (EHRs), radiological images,genomicsequences,andbiosignals.Theapplication of deep learning techniques to such datasets has demonstrated superior performance in disease detection andoutcomepredictioncomparedtoconventionalstatistical approaches (Esteva et al., 2017). Nevertheless, training robustmodelsrequiresmulti-institutionaldatacollaboration toovercomelocaldatabiasandlimitedsamplesizes.

Traditional centralized machine learning architectures requirepoolingdataintoasinglerepository,creatingrisks related to privacy breaches, data misuse, and regulatory non-compliance.LegislativeframeworkssuchasGDPRand HIPAAimposestrictlimitationsoncross-borderandinterinstitutional data sharing. Consequently, decentralized collaborative training has gained attention as a viable alternativethatenablesknowledgesharingwithoutdirect dataexchange.

1.2 Challenges in Medical Data Sharing and Model Training

Medical data sharing is constrained by legal, ethical, and technicalbarriers.Privacyconcernsariseduetothehighly sensitivenatureofpatientinformation,whereunauthorized access may lead to identity disclosure or discrimination. Even anonym zed datasets remain vulnerable to reidentification attacks when combined with auxiliary information(NarayananandShmatikov,2008).

From a technical standpoint, healthcare data are typically non-independent and identically distributed (non-IID), imbalanced,andinstitution-specific.Variationsinimaging protocols, demographic distributions, and disease

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

prevalence reduce model generalizability. Furthermore, centralizedinfrastructuresintroducesinglepointsoffailure andincreasesusceptibilitytocyber-attacks.

1.3 Role of Decentralized Learning and Federated Learning (FL)

1.3.1

Federated Learning Framework

FederatedLearningisadistributedoptimizationframework in which participating clients train local models using private data and transmit only model updates to a coordinating server for aggregation. The foundational FedAvg algorithm demonstrated that decentralized stochastic gradient descent can achieve competitive performance compared to centralized training while reducingcommunicationcosts(McMahanetal.,2017).

Inhealthcare,FLenableshospitalstocollaborativelydevelop diagnosticmodelsfortaskssuchastumorsegmentationand disease prediction without exposing patient-level records (Rieke et al., 2020). This paradigm enhances data sovereignty while leveraging collective intelligence across institutions.Recentextensionsincorporatepeer-to-peerand block chain-based architectures to eliminate reliance on a central aggregator, thereby improving trust and fault tolerance.

1.4 Need for Enhanced Privacy Controls

Although FL prevents direct data sharing, it does not inherentlyguaranteecompleteprivacy.Modelupdatescan leak sensitive information through gradient inversion or membershipinferenceattacks(Zhuetal.,2019).Therefore, enhancedprivacy-preservingmechanismsarenecessaryto strengthensecurityguarantees.

Differential privacy introduces calibrated noise to model updates to limit individual data contribution, providing formal privacy bounds (Dwork, 2006). Secure multi-party computationandhomomorphicencryptionenableencrypted aggregation of model parameters without revealing intermediateupdates.Secureaggregationprotocolsfurther ensure that the server cannot access individual client gradients.Thesetechniquescollectivelymitigateadversarial riskswhilemaintainingacceptableperformancetrade-offs.

2. FUNDAMENTALS

Thedeploymentofadvancedmachinelearningtechniquesin healthcare necessitates a clear understanding of computationalparadigms,architecturalmodels,andprivacypreserving mechanisms. This section outlines the foundational principles underpinning decentralized collaborativemodeltraining,particularlywithinfederated learning(FL)environmentsdesignedforsensitivemedical data.

2.1 Overview of Machine Learning in Healthcare

Machinelearning(ML)hassignificantlytransformedclinical diagnostics, disease prediction, medical imaging analysis, andpersonalizedtreatmentplanning.Supervisedlearning models,especiallydeepneuralnetworks,havedemonstrated highperformanceinradiology,dermatology,andpathology byleveraginglargeannotateddatasets(Litjensetal.,2017). Similarly, predictive analytics applied to electronic health records (EHRs) enable early detection of adverse clinical eventssuchassepsisandhospitalreadmission.

Despite these advancements, medical ML systems face inherent constraints. Healthcare datasets are typically fragmentedacrossinstitutions,heterogeneousinformat,and subject to strict regulatory controls. The traditional assumption of centralized data availability is often unrealistic in clinical settings. Consequently, alternative learning paradigms that preserve data locality while enabling collaborative model improvement have gained increasingattention.

2.2 Centralized vs Decentralized Learning Paradigms

Centralizedlearninginvolvesaggregatingall training data into a single repository, where model optimization is performed on unified datasets. This approach simplifies coordination and often yields strong performance due to completedatavisibility.However,itintroducessignificant privacy, security, and governance risks. Large centralized databases become attractive targets for cyber-attacks and mayviolatedataprotectionregulations.

Incontrast,decentralizedlearningdistributesthetraining processacrossmultiplenodes,eachretainingitslocaldata. Instead of transferring raw datasets, only intermediate model parameters or gradients are communicated. This paradigmenhancesdatasovereigntyandreducesexposure tobreaches.However,decentralizedsystemsmustaddress challenges such as communication overhead, synchronization complexity, and statistical heterogeneity (Kairouz et al., 2021). The trade-off between privacy preservationandsystemefficiencyremainsacentraldesign consideration.

2.3 Federated Learning: Definition and Key Properties

2.3.1

Core Concept and Operational Mechanism

FederatedLearningisadistributedoptimizationframework inwhichmultipleclientscollaborativelytrainasharedglobal model under the coordination of a central server or decentralizedprotocol.Eachclientperformslocaltraining usingprivatedataandperiodicallytransmitsmodelupdates for aggregation. The canonical Federated Averaging (FedAvg) algorithm reduces communication rounds by

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

combining locally computed gradients through weighted averaging(McMahanetal.,2017).

Key properties of FL include data locality,communication efficiency, and iterative global aggregation. In healthcare contexts, this allows hospitals to jointly train predictive modelswhilemaintainingpatientdatawithininstitutional boundaries.FLalsosupportscross-silocollaboration,where a limited number of reliable institutions participate, as opposed to cross-device settings common in mobile applications.

2.4 Privacy and Security Concepts in FL

AlthoughFLmitigatesdirectdatasharingrisks,itdoesnot inherently eliminate privacy vulnerabilities. Adversaries may reconstruct sensitive information from transmitted gradientsormanipulateupdatesthroughpoisoningattacks.

To address these threats, several cryptographic and statisticaltechniquesareintegratedintofederatedsystems.

2.4.1 Differential Privacy

DifferentialPrivacy(DP)providesamathematicallyrigorous framework for quantifying and limiting individual data contributionwithinadataset.Byinjectingcalibratednoise into gradients or model parameters, DP ensures that the inclusion or exclusion of a single record does not significantlyinfluencethemodeloutput(Dwork,2006).In FL, DP can be applied locally at client devices or globally duringaggregation.Whileenhancingprivacyguarantees,it introducesatrade-offbetweenmodelaccuracyandprivacy budget (ε), requiring careful calibration in clinical applications.

2.4.2 Secure Multi-Party Computation

Secure Multi-Party Computation (SMPC) enables multiple participantstojointlycomputeafunctionovertheirinputs without revealing the inputs themselves. In federated settings, SMPC protocols facilitate secure aggregation of model updates, ensuring that the server cannot access individualgradients.Techniquessuchassecretsharingand cryptographicmaskingarecommonlyemployedtoachieve confidentialityduringcollaborativeoptimization(Bonawitz etal.,2017).SMPCstrengthenstrustinmulti-institutional medicalcollaborations.

2.4.3 Homomorphic Encryption

HomomorphicEncryption(HE)allowscomputationstobe performed directly on encrypted data without requiring decryption. In FL architectures, clients encrypt model updatesbeforetransmission,andaggregationisconducted in encrypted form. Only the final aggregated result is decrypted, preventing intermediate exposure of sensitive parameters. Although HE provides strong confidentiality guarantees, it is computationally intensive and may

introducelatencyinlarge-scalehealthcaresystems(Gentry, 2009).

2.4.4 Trusted Execution Environments

Trusted Execution Environments (TEEs) are hardwarebased secure enclaves that isolate sensitive computations from the main operating system. Within FL frameworks, TEEs can securely perform model aggregation while protecting against malicious server-side interference. Technologies such as Intel SGX create encrypted memory regionsinaccessibletounauthorizedprocesses.WhileTEEs reduce cryptographic overhead compared to fully homomorphic approaches, they rely on hardware trust assumptionsandmaybevulnerabletoside-channelattacks (CostanandDevadas,2016).

3. ARCHITECTURAL FRAMEWORKS FOR FEDERATED LEARNING IN MEDICAL SETTINGS

Federatedlearning(FL)architecturesdeterminehowmodel updatesareexchanged,aggregated,andsynchronizedacross participating medical institutions. In healthcare environments, architectural design must balance privacy preservation,computationalefficiency,faulttolerance,and regulatorycompliance.Thissectionexaminestheprincipal architectural models adopted in federated medical AI systems.

3.1 Client–Server Federated Architecture

The client–server model represents the foundational architecture of federated learning. In this framework, a centralcoordinatingserverorchestratestrainingroundsby distributing a global model to participating hospitals or clinical centers (clients). Each client performs local optimization using its private dataset and returns model updatestotheserverforaggregation.Theservercomputesa weightedaverageofupdatesandredistributestheimproved modeliniterativerounds(McMahanetal.,2017).

In medical applications, this architecture is particularly suited to cross-silo federated learning, where a limited number of trusted institutions collaborate. It offers structuredcoordination,simplifiedconvergencecontrol,and relatively low system complexity. However, the central serverconstitutesasinglepointoffailureandmaybecomea bottleneckinlarge-scaledeployments.Additionally,despite secureaggregationmechanisms,trustassumptionsremain concentrated around the coordinating entity, raising governanceconsiderationsininter-hospitalcollaborations (Riekeetal.,2020).

3.2 Peer-to-Peer Decentralized Federated Models

Peer-to-peer (P2P) federated architectures eliminate the reliance on a central aggregator by enabling direct communication among participating nodes. In this decentralizedtopology,clientsexchangemodelparameters

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

withneighboringnodesanditerativelyupdatelocalmodels based on consensus algorithms. Distributed optimization techniques such as gossip protocols and decentralized stochasticgradientdescentaretypicallyemployed.

Such architectures enhance fault tolerance and reduce centralizedtrustdependencies.Inmedicalconsortiawhere institutions seek equal authority in collaborative model governance, P2P systems promote fairness and transparency.However,convergencecontrolbecomesmore complex,andcommunicationoverheadmayincreasedueto multi-hop exchanges. Ensuring robustness against adversarial participants also requires stronger validation mechanisms compared to centralized setups (Lian et al., 2017).

3.3 Hybrid and Blockchain-Based Federated Systems

Hybrid federated architectures combine centralized coordination with decentralized trust mechanisms. One prominent approach integrates blockchain technology to maintain a tamper-resistant ledger of model updates and training transactions. Smart contracts can automate validationrules,enforceparticipationincentives,andensure accountabilityacrossinstitutions.

Blockchain-assistedFLenhancestransparencyandmitigates risks associated with malicious parameter manipulation. Each update is cryptographically recorded, reducing the probability of undetected poisoning attacks. However, blockchain integration introduces computational latency, storage overhead, and scalability challenges, particularly whenhandlinghigh-dimensionalmedicalmodels(Lietal., 2020). Hybrid architectures therefore aim to balance coordination efficiency with decentralized trust reinforcement.

3.4 Communication ProtocolsandSynchronization

3.4.1 Federated Averaging (FedAvg)

FedAvgisthefoundationalaggregationprotocolinfederated learning. It reduces communication costs by allowing multiplelocaltrainingepochsbeforetransmittingupdatesto the server. The global model is computed as a weighted average of client parameters proportional to local dataset sizes(McMahanetal.,2017).InmedicalimagingandEHRbased prediction tasks, FedAvg demonstrates competitive accuracy while maintaining data locality.Nevertheless, its performance may degrade under non-independent and identicallydistributed(non-IID)dataconditions,whichare commoninhealthcare.

3.4.2 Federated Proximal (FedProx)

FedProx extends FedAvg by incorporating a proximal regularization term into the local objective function, constraining client updates to remain closer to the global model.Thismodificationimprovesstabilityandconvergence in heterogeneous environments with varying data distributions and computational capabilities (Li et al., 2020a).Inmulti-hospitalcollaborationswheredemographic and clinical variations are pronounced, FedProx enhances robustnessagainststatisticaldivergence.

3.4.3 Synchronization Strategies

Federated systems may operate under synchronous or asynchronousupdatemechanisms.Synchronousprotocols requireallselectedclientstocompletelocaltrainingbefore aggregation,ensuringconsistencybutpotentiallyincreasing latency. Asynchronous approaches allow incremental updates, improving scalability but introducing challenges relatedtostalegradientsandconvergencecontrol.Selecting an appropriate synchronization strategy depends on institutionalinfrastructure,networkbandwidth,andclinical workloadconstraints(Kairouzetal.,2021).

4. LITERATURE REVIEW

This section synthesizes prior research on federated learning(FL)inmedicalcontexts,structuredthematicallyto capture methodological evolution, privacy enhancements, architectural diversification, and domain-specific applications.Thereviewhighlightsdatasets,experimental settings, performance outcomes, and unresolved research challenges.

Figure-1: Decentralized Federated Models

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

4.1 Early Works on FL in Healthcare

4.1.1

Initial Implementations in Medical Imaging and EHRs

Theearliestapplicationsoffederatedlearninginhealthcare primarily focused on medical imaging tasks, where deep convolutionalneuralnetworkswerecollaborativelytrained acrossinstitutionswithoutcentralizingradiologicaldatasets. A landmark study demonstrated multi-institutional brain tumorsegmentationusingfederatedtrainingonMRIdata, achievingperformancecomparabletocentralizedbaselines whilepreservinginstitutionaldataboundaries(Shelleretal., 2018).Similareffortsextendedtohistopathologyandchest X-rayclassificationtasks,validatingthefeasibilityofcrosssilocollaboration.

Parallel research explored FL for electronic health record (EHR) analytics, particularly for predictive modeling of clinical outcomes such as mortality and readmission risk. These studies confirmed that decentralized gradient aggregation could maintain predictive accuracy despite heterogeneousdatadistributionsacrosshospitals.

4.1.2

Initial Privacy Challenges Identified

Although early implementations validated functional viability,theyalsoexposedprivacyvulnerabilitiesinherent ingradientsharing.Researchdemonstratedthatadversaries could reconstruct input data from model updates using gradient inversion techniques (Zhu et al., 2019). Additionally,membershipinferenceattacksrevealedrisksof identifyingwhetherspecificpatientrecordscontributedto model training. These findings underscored that FL alone does not guarantee privacy and necessitated additional protectivemechanisms.

4.2 Privacy Enhancements Integrated with FL

4.2.1

Differential Privacy in Federated Settings

Subsequentresearchincorporateddifferentialprivacy(DP) intofederatedoptimizationworkflows.Byaddingcalibrated noise to gradients before aggregation, DP-FL frameworks providedformalprivacyguaranteesquantifiedbyaprivacy budgetparameter.Empiricalstudiesshowedthatmoderate privacybudgetspreservedacceptableaccuracyinmedical image classification while limiting information leakage (Geyer et al., 2017). However, excessive noise injection degraded convergence stability, highlighting the need for adaptiveprivacycalibrationstrategies.

4.2.2

Cryptographic Secure Aggregation and SMPC

Tomitigaterisksfromuntrustedservers,secureaggregation protocolsbasedonsecuremulti-partycomputation(SMPC) were introduced. These methods ensured that the central aggregator could only access encrypted or masked parametersumsratherthanindividualupdates(Bonawitzet

al.,2017).Inhealthcarecollaborations,secureaggregation enhancedtrustamongparticipatinghospitalsbypreventing exposureofinstitution-specificgradientinformation.

4.2.3 Homomorphic Encryption and Trusted Execution

Homomorphic encryption (HE) and Trusted Execution Environments(TEEs)furtherstrengthenedconfidentiality guarantees. HE-enabled federated frameworks allowed encrypted model updates to be aggregated without decryption, albeit at higher computational cost (Gentry, 2009).TEEs,suchasIntelSGX,providedhardware-isolated aggregationenvironmentsthatprotectedmodelparameters duringcomputation.Comparativeevaluationsindicatedthat cryptographic methods offer stronger theoretical guarantees,whileTEEsprovideimprovedefficiencyunder controlled hardware trust assumptions (Costan and Devadas,2016).

4.3 Decentralized / Distributed Model Training Approaches

4.3.1

Peer-to-Peer Decentralized Medical FL

Moving beyond centralized orchestration, decentralized peer-to-peerFLmodelswereproposedtoeliminatereliance on a coordinating server. These systems employed consensus-based optimization and gossip protocols to propagate updates among hospitals. Studies reported improvedfaulttoleranceandresilienceagainstsingle-point failure, although convergence rates were sensitive to network topology and communication delays (Lian et al., 2017).

4.3.2 Block chain and Ledger-Assisted FL Systems

Block chain integration emerged as a trust-enhancing mechanism in federated healthcare systems. Distributed ledgers recorded model updates immutably, while smart contracts enforced aggregation policies and participation incentives. Empirical evaluations demonstrated improved transparencyandtamperresistance,particularlyinmultistakeholderenvironmentsinvolvingresearchhospitalsand diagnosticcenters(Lietal.,2020).Nevertheless,blockchain latencyandscalabilityremainlimitingfactorsinlarge-scale medicaldeployments.

4.3.3 Centralized vs Decentralized Paradigm Comparison

Comparative analyses reveal that centralized FL architectures generally achieve faster convergence due to structuredcoordination,whereasdecentralizedframeworks enhance trust distribution and robustness. Centralized models are computationally efficient but vulnerable to server compromise, while fully decentralized systems requiresophisticatedsynchronizationprotocols.Thechoice ofparadigmdependsoninstitutionalgovernancestructures andrisktolerancelevels.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

4.4 Applications and Case Studies

4.4.1

Medical Imaging

Federated learning has been widely applied to MRI-based tumor segmentation, CT-based COVID-19 diagnosis, and mammographic cancer detection. Multi-center imaging collaborations demonstrated that federated models can generalizebetteracrossdemographicvariationscompared to single-institution training (Rieke et al., 2020). These studiestypicallyemployedconvolutional neural networks andevaluatedperformanceusingDicesimilaritycoefficients andAUCmetrics.

Figure-2:

Federated learning in the Medical

4.4.2 Electronic Health Records

Analytics

FL-based predictive models for EHR data have addressed taskssuchasICUmortalitypredictionandsepsisdetection. Recurrent neural networks and transformer-based architecturesweretrainedacrossgeographicallydistributed hospitals. Results indicated improved robustness against localdatabiaswhilemaintainingcompliancewithregulatory constraints.

4.4.3 Genomic and Biosignal Frameworks

EmergingresearchextendsFLtogenomicsequenceanalysis and wearable biosignal monitoring. Distributed training across genomic repositories has enabled collaborative variant classification without exposing sensitive genetic markers. Similarly, federated frameworks applied to ECG andEEGdatademonstrateprivacy-preservingcardiacand neurologicalmonitoringsystems.

5. PRIVACY AND SECURITY ENHANCEMENTS

Federatedlearning(FL)mitigatesdirectdatasharingrisks by retaining patient data within institutional boundaries; however,theexchangeofmodelparametersintroducesnew securityandprivacyvulnerabilities.Adversarialparticipants orcompromisedaggregationserversmayexploitgradient information to infer sensitive data or manipulate model behaviour. Consequently, robust privacy-preserving algorithmsandthreat-awaresystemdesignsareessentialfor securedecentralizedmedicalAIdeployments.

5.1 Differential Privacy in FL: Algorithms and Trade-offs

5.1.1 Differentially Private Gradient Mechanisms

DifferentialPrivacy(DP)providesformalprivacyguarantees byensuringthattheinclusionorexclusionofasingledata record does not significantly influence model outputs. In federated settings, DP is commonly implemented through gradient clipping followed by calibrated noise injection beforetransmissiontotheaggregator.Thisapproachbounds sensitivity and limits potential information leakage from individual client updates (Dwork, 2006). Client-level DP furtherstrengthensprotectionbymaskingthecontribution ofentireinstitutionaldatasets,whichisparticularlyrelevant incross-silohealthcarecollaborations.

5.1.2

Privacy–Utility Trade-offs

While DP enhances confidentiality, it introduces a measurabletrade-offbetweenmodelaccuracyandprivacy budget(ε).Excessivenoisereducesconvergencespeedand predictive performance, especially in high-dimensional medical imaging tasks. Adaptive privacy accounting and dynamic noise scaling strategies have been proposed to balance diagnostic accuracy with regulatory compliance requirements (Geyer, Klein and Nabi, 2017). Selecting optimal privacy parameters remains a context-dependent challengeinsafety-criticalhealthcaresystems.

5.2 Secure AggregationandEncryptionTechniques

5.2.1 Secure Aggregation Protocols

Secure aggregation ensures that individual client updates remain confidential during the aggregation process. Protocols based on cryptographic masking and secret sharingallowtheservertocomputeonlytheaggregatedsum of gradients without accessing intermediate values (Bonawitzetal.,2017).This techniquereducestherisk of server-side inference attacks while maintaining computational efficiency suitable for multi-hospital federateddeployments.

5.2.2

Homomorphic Encryption and Multi-Party Computation

Homomorphic encryption (HE) enables mathematical operations to be performed directly on encrypted data, allowingaggregationwithoutexposingplaintextparameters. AlthoughHEprovidesstrongconfidentialityguarantees,it imposes significant computational overhead, which may limitreal-timemedicalapplications(Gentry,2009).Secure Multi-Party Computation (SMPC) offers an alternative cryptographic approach, distributing computation among participantssuchthatnosingleentitylearnsthecomplete input. SMPC-based federated systems enhance trust in collaborativeclinicalnetworkswheremutualdistrustmay exist.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

5.3 Block chain and Smart Contracts for Trust Management

Block chain technology introduces decentralized trust management into federated learning ecosystems. By recording model updates and aggregation events in an immutable distributed ledger, block chain enhances transparency and auditability. Smart contracts automate validationrules,enforceaccesscontrolpolicies,andmanage incentivemechanismsamongparticipatinginstitutions.

InhealthcareFLframeworks,blockchainmitigatesrisksof modeltamperingandunauthorizedmodificationbyensuring traceability of updates. However, consensus mechanisms introduce latency and scalability constraints, particularly whenhandlinglargeparameterexchanges.Hybriddesigns combining block chain logging with off-chain aggregation havebeenproposedtoaddressefficiencylimitations(Liet al.,2020).

5.4

Threat Models and Attack Vectors

Understandingadversarialstrategiesiscriticalfordesigning resilientfederatedmedicalsystems.Threatmodelstypically considermaliciousclients,compromisedservers,orexternal eavesdroppers.

5.4.1

Poisoning Attacks

Poisoning attacks occur when adversarial clients intentionallysubmitmanipulatedgradientstocorruptthe globalmodel.Inmedicaldiagnosissystems,suchattacksmay bias predictions or degrade accuracy. Data poisoning modifies local training data, whereas model poisoning directly alters gradient updates. Robust aggregation rules and anomaly detection mechanisms are necessary to mitigatesuchthreats(Kairouzetal.,2021).

5.4.2

Model Inversion and Reconstruction Attacks

Modelinversionattacksaimtoreconstructsensitiveinput datafromsharedgradientsormodelparameters.Empirical demonstrations show that gradient leakage can reveal patientimagesortrainingsamplesundercertainconditions (Zhu,LiuandHan,2019).Thesevulnerabilitieshighlightthe necessity of combining FL with differential privacy and secureaggregationtopreventunauthorizedreconstruction ofmedicalrecords.

5.4.3 Eavesdropping and Relay Attacks

Eavesdropping attacks exploit insecure communication channels to intercept model updates during transmission. Relayorman-in-the-middleattacksmayalterupdatesbefore theyreachtheaggregator.Securecommunicationprotocols employing end-to-end encryption and authenticated channels are essential to prevent interception and tampering. Network-level security must complement

algorithmic privacy measures to ensure comprehensive protectionindistributedhealthcareinfrastructures.

6. EVALUATION METRICS AND BENCHMARKING

Robust evaluation of federated learning (FL) systems in healthcare requires multidimensional assessment, encompassingpredictiveperformance,privacyguarantees, computational efficiency, and dataset representativeness. Unlike conventional centralized models, federated frameworks introduce additional variables such as communicationcost,clientheterogeneity,andcryptographic overhead. This section outlines the principal evaluation metricsandbenchmarkingstrategiesadoptedinmedicalFL research.

6.1 Performance Metrics for Federated Models

6.1.1

Accuracy

Accuracy represents the proportion of correctly classified instancesoverthetotalnumberofpredictions.Infederated medical classification tasks such as tumor detection or diseasediagnosis accuracyprovidesageneralperformance indicator. However, it may be misleading in imbalanced clinical datasets where negative cases significantly outnumber positive ones (Saito and Rehmsmeier, 2015). Consequently, additional metrics are required to ensure clinicallymeaningfulevaluation.

6.1.2

Area Under the Curve (AUC)

TheAreaUndertheReceiverOperatingCharacteristicCurve (AUC-ROC) evaluates a model’s ability to discriminate betweenclassesacrossdifferentthresholdsettings.Inmultiinstitutional FL studies, AUC is widely used for assessing diagnosticrobustnessacrossheterogeneousdatasources.It is particularly valuable in medical risk prediction models, where threshold-independent evaluation is necessary to balancesensitivityandfalse-positiverates(Bradley,1997).

6.1.3 Sensitivity and Specificity

Sensitivity (true positive rate) measures the ability to correctlyidentifypatientswithacondition,whilespecificity (truenegativerate)evaluatescorrectidentificationofnonaffected individuals. These metrics are critical in safetysensitivehealthcareapplications,suchascancerscreening orinfectiousdiseasedetection.Highsensitivity minimizes missed diagnoses, whereas high specificity reduces unnecessary interventions. Federated frameworks often report these metrics to demonstrate clinical reliability (Powers,2011).

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

6.2 Privacy and Security Metrics

6.2.1

Privacy Loss (ε)

In differentially private federated learning, privacy guarantees are quantified using the privacy budget parameter(ε),whichmeasuresthemaximuminformation leakageattributabletoanindividualrecord.Lowerεvalues correspond to stronger privacy but typically introduce greaternoiseandreducedmodelutility.Privacyaccounting techniques, including moment’s accountant methods, are usedtotrackcumulativeprivacylossacrosstrainingrounds (Abadietal.,2016).Selectinganappropriateεinhealthcare systems requires balancing regulatory compliance with diagnosticperformance.

6.2.2 Cryptographic Overhead

Cryptographic mechanisms such as secure aggregation, homomorphic encryption, and secure multi-party computation introduce additional computational and communication overhead. Evaluation commonly includes encryptiontime,decryptionlatency,andincreasedmessage size. These metrics are crucial for determining system feasibility in resource-constrained hospital networks. Empirical analyses indicate that fully homomorphic encryptionsignificantlyincreasesprocessingtimecompared to secure aggregation schemes, highlighting trade-offs between theoretical security strength and practical deployability(Gentry,2009).

6.3 Communication and Computation Overheads

Federated learning inherently requires repeated communication between clients and aggregators. Communicationoverheadistypicallymeasuredintermsof transmitted bytes per round and total communication rounds to convergence. Optimization strategies such as model compression, quantization, and sparse updates are frequently evaluated to reduce bandwidth consumption (Kairouzetal.,2021).

Computationoverheadincludeslocaltrainingtime,memory usage, and server-side aggregation latency. In medical environments with heterogeneous computational infrastructure,performancebenchmarkingmustaccountfor variability in hardware capabilities. Efficient federated protocols aim to minimize synchronization delays while maintainingconvergencestabilityundernon-independent andidenticallydistributed(non-IID)dataconditions.

7. CONCLUSION

Decentralizedcollaborativemodeltrainingusingfederated learning (FL) represents a transformative paradigm for privacy-preservingartificialintelligenceinhealthcare.This review has examined architectural frameworks, privacyenhancing mechanisms, decentralized coordination strategies,anddomain-specificapplicationsacrossmedical

imaging, electronic health records, and genomic analytics. The synthesis of existing literature indicates that FL effectivelymitigatesdirectdata-sharingriskswhileenabling multi-institutional knowledge integration. However, federatedsystemsarenotinherentlysecure;vulnerabilities suchasgradientleakage,poisoningattacks,andinference risks necessitate the integration of differential privacy, secure aggregation, cryptographic protocols, and trustenhancing mechanisms such as block chain. Evaluation metrics must extend beyond predictive accuracy to incorporate privacy budgets, communication costs, and computational feasibility, particularly in heterogeneous clinical environments. Although substantial progress has been achieved, unresolved challenges remain in handling non-IIDdatadistributions,scalabilityconstraints,regulatory harmonization, and real-world deployment readiness. Future research should prioritize adaptive privacy–utility optimization, robust aggregation under adversarial conditions, and standardized benchmarking frameworks. Overall, privacy-aware federated learning offers a viable pathwaytowardsecure,scalable,andethicallyresponsible medical AI systems capable of supporting collaborative healthcareinnovation.

8. LIMITATIONS OF REVIEW

Thisreviewissubjecttoseverallimitations.First,therapidly evolvingnatureoffederatedlearningresearchmeansthat emergingtechniquesandpreprintcontributionsmaynotbe comprehensively covered. Second, the analysis primarily focusesoncross-silohealthcaresettings,withcomparatively limited discussion of cross-device medical IoT scenarios. Third,manyreferencedstudiesrelyonsimulatedfederated environments rather than fully distributed real-world deployments, which may restrict external validity. Additionally,quantitativemeta-analysiswasnotperformed duetoheterogeneityinevaluationprotocols,datasets,and reporting standards across studies. Finally, regulatory, ethical, and socio-technical dimensions such as patient consentframeworksandgovernanceinteroperability were discussedconceptuallybutnotexaminedthroughempirical policyanalysis.

REFERENCES

1. Abadi, M., Chu, A., Goodfellow, I., McMahan, H.B., Mironov, I., Talwar, K. and Zhang, L. (2016) ‘Deep learningwithdifferentialprivacy’,Proceedingsofthe ACM Conference on Computer and Communications Security(CCS),pp.308–318.

2. Bonawitz, K., Ivanov, V., Kreuter, B., Marcedone, A., McMahan,H.B.,Patel,S.,Ramage,D.,Segal,A.andSeth, K. (2017) ‘Practical secure aggregation for privacypreservingmachinelearning’,ProceedingsoftheACM ConferenceonComputerandCommunicationsSecurity (CCS),pp.1175–1191.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

3. Bradley,A.P.(1997)‘TheuseoftheareaundertheROC curveintheevaluationofmachinelearningalgorithms’, PatternRecognition,30(7),pp.1145–1159.

4. Costan,V.andDevadas,S.(2016)‘IntelSGXexplained’, IACRCryptologyePrintArchive,2016(086),pp.1–118.

5. Dwork,C.(2006)‘Differentialprivacy’,Proceedingsof the 33rd International Colloquium on Automata, LanguagesandProgramming(ICALP),pp.1–12.

6. Esteva,A.,Kuprel,B.,Novoa,R.A.,Ko,J.,Swetter,S.M., Blau, H.M. and Thrun, S. (2017) ‘Dermatologist-level classificationofskincancerwithdeepneuralnetworks’, Nature,542(7639),pp.115–118.

7. Gentry,C.(2009)‘Fullyhomomorphicencryptionusing ideallattices’,Proceedingsofthe41stACMSymposium onTheoryofComputing(STOC),pp.169–178.

8. Geyer,R.C.,Klein,T.andNabi,M.(2017)‘Differentially privatefederatedlearning:Aclientlevelperspective’, NIPS Workshop on Private Multi-Party Machine Learning.

9. Kairouz,P.,McMahan,H.B.,Avent,B.,Bellet,A.,Bennis, M.,Bhagoji,A.N.,Bonawitz,K.,Charles,Z.,Cormode,G., Cummings, R., D’Oliveira, R.G.L., Eichner, H., El Rouayheb, S., Evans, D., Garcia-Saavedra, A., et al. (2021) ‘Advances and open problems in federated learning’,FoundationsandTrendsinMachineLearning, 14(1–2),pp.1–210.

10. Li, T., Sahu, A.K., Talwalkar, A. and Smith, V. (2020) ‘Federated optimization in heterogeneous networks’, ProceedingsofMachineLearningandSystems(MLSys), pp.429–450.

11. Li,X.,Jiang,M.,Zhang,X.,Kamp,M.andDou,Q.(2020) ‘Ablockchain-basedfederatedlearningframeworkfor data privacy preservation’, IEEE Transactions on IndustrialInformatics,16(6),pp.4287–4296.

12. Lian,X.,Zhang,C.,Zhang,H.,Hsieh,C.-J.,Zhang,W.and Liu,J.(2017)‘Candecentralizedalgorithmsoutperform centralizedalgorithms?Acasestudyfordecentralized parallel stochastic gradient descent’, Advances in NeuralInformationProcessingSystems,30,pp.5330–5340.

13. Litjens,G.,Kooi,T.,Bejnordi,B.E.,Setio,A.A.A.,Ciompi, F.,Ghafoorian,M.,VanderLaak,J.A.W.M.,VanGinneken, B.andSánchez,C.I.(2017)‘Asurveyondeeplearning inmedicalimageanalysis’,MedicalImageAnalysis,42, pp.60–88.

15. Narayanan, A. and Shmatikov, V. (2008) ‘Robust deanonymization of large sparse datasets’, IEEE SymposiumonSecurityandPrivacy,pp.111–125.

16. Powers, D.M.W. (2011) ‘Evaluation: From precision, recall and F-measure to ROC, informedness, markedness and correlation’, Journal of Machine LearningTechnologies,2(1),pp.37–63.

17. Rieke, N., Hancox, J., Li, W., Milletari, F., Roth, H.R., Albarqouni, S., Bakas, S., Galtier, M.N., Landman, B.A., Maier-Hein,K.,Ourselin,S.,Sheller,M.andCardoso,M.J. (2020) ‘The future of digital health with federated learning’,npjDigitalMedicine,3(119),pp.1–7.

18. Saito, T. and Rehmsmeier, M. (2015) ‘The precisionrecallplotismoreinformativethantheROCplotwhen evaluatingbinaryclassifiersonimbalanceddatasets’, PLoSONE,10(3),e0118432.

19. Sheller, M.J., Reina, G.A., Edwards, B., Martin, J. and Bakas, S. (2018) ‘Multi-institutional deep learning modeling without sharing patient data: A feasibility study on brain tumor segmentation’, Brainlesion: Glioma,MultipleSclerosis,StrokeandTraumaticBrain Injuries,pp.92–104.

20. Sheller,M.J.,Edwards,B.,Reina,G.A.,Martin,J.,Pati,S., Kotrotsou,A.,Milchenko,M.,Xu,W.,Marcus,D.,Colen, R.R. and Bakas, S. (2020) ‘Federated learning in medicine:Facilitatingmulti-institutionalcollaborations without sharing patient data’, Scientific Reports, 10, 12598.

21. Zhu,L.,Liu,Z.andHan, S.(2019)‘Deepleakage from gradients’,AdvancesinNeuralInformationProcessing Systems,32,pp.14774–14784.

22. Akhmetov, A., Latif, Z., Tyler, B. and Yazici, A. (2025) ‘Enhancinghealthcaredataprivacyandinteroperability with federated learning’, PeerJ Computer Science, 11:e2870.

23. Adnan, M., Kalra, S., Cresswell, J.C., Taylor, G.W. and Tizhoosh, H.R. (2022) ‘Federated learning and differential privacy for medical image analysis’, ScientificReports,12,1953.

24. Ali,M.S.,Ahsan,M.M.,Tasnim,L.etal.(2024)‘Federated learning in healthcare: Model misconducts, security, challenges, applications, and future research directions’,arXiv,2405.13832.

25. Daram, S. (2025) ‘Federated learning in medical AI: Advancing privacy-preserving data sharing for collaborative healthcare research’, International

14. McMahan,H.B.,Moore,E.,Ramage,D.,Hampson,S.and Arcas,B.A.Y.(2017)‘Communication-efficientlearning ofdeepnetworksfromdecentralizeddata’,Proceedings of the 20th International Conference on Artificial IntelligenceandStatistics(AISTATS),pp.1273–1282.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Journal of Artificial Intelligence, Data Science, and MachineLearning,6(2).

26. Dendukuri, S.V. (2025) ‘Federated learning in healthcare:Protectingpatientprivacywhileadvancing analytics’,JournalofComputerScienceandTechnology Studies,7(7),pp.840–845.

27. Gu,X.,Sabrina,F.,Fan,Z.andSohail,S.(2023)‘Areview ofprivacyenhancementmethodsforfederatedlearning in healthcare systems’, International Journal of Environmental Research and Public Health, 20(15), 6539.

28. Haripriya,R.,Khare,N.andPandey,M.(2025)‘Privacypreservingfederatedlearningforcollaborativemedical data mining in multi-institutional settings’, Scientific Reports,15,12482.

29. Haripriya, R., Khare, N., Pandey, M. and Biswas, S. (2025) ‘A privacy-enhanced framework for collaborative Big Data analysis in healthcare using adaptivefederatedlearningaggregation’,JournalofBig Data,12,113.

30. Manjula, N.J., Randhi, K. and Bandarapu, S.R. (2023) ‘Federated learning for healthcare: Balancing data privacy and model accuracy’, American Journal of ComputingandEngineering,6(1),pp.69–80.

31. Pati,S.,Kumar,S.etal.(2024)‘Privacypreservationfor federatedlearninginhealthcare’,Patterns(NY),5(7), 100974.

32. Sandhu, S.S., Taheri Gorji, H. et al. (2023) ‘Medical imagingapplicationsoffederatedlearning’,Diagnostics, 13(19),3140.

33. Shah, S.T., Ali, Z., Waqar, M. and Kim, A. (2025) ‘Federated learning in public health: Decentralized, equitable,andsecurediseasepreventionapproaches’, Healthcare,13(21),2760.

34. Verma,A.andGonzalez,M.(2023)‘Privacy-preserving federated learning for healthcare data sharing’, International Journal of Recent Advances in EngineeringandTechnology,12(2),pp.14–20.

35. Zafar, A., Saad, M. and Haque, S.B.U. (2025) ‘Efficient andprivacy-enhancedfederatedlearningformedical imaginginresource-limitedenvironments’,Journalof ElectricalSystems,21(01).

36. Koutsoubis, N., Waqas, A., Yilmaz, Y., Ramachandran, R.P., Schabath, M. and Rasool, G. (2024) ‘Futureproofing medical imaging with privacy-preserving federated learning and uncertainty quantification: A review’,arXiv,2409.16340.

37. Dendukuri, S.V. (2025) ‘Federated Learning in Healthcare: Protecting Patient Privacy …’, Journal of Computer Science and Technology Studies, 7(7), pp.840–845.

38. Leveraging federated learning for rare disease EHR analysis (2024) Journal of Electronic Clinical Data Science,DOI:10.1016/j.ject.2024.11.001.

39. Federated Learning in Smart Healthcare (2024) ‘Federated learning in smart healthcare: Privacy, security and IoT predictive analytics’, Healthcare, 12(24),2587.

40. Anonymizing Data for Privacy-Preserving FL (2020) ‘Anonymizing data for privacy-preserving federated learning’,arXiv,2002.09096.

41. Comprehensive surveys of FL methods in medical imaging(2023)‘Federatedlearningformedicalimage analysis:Asurvey’,PubMed.

42. Darzidehkalani, E., Ghasemi-Rad, M. et al. (2022) ‘Federated learning in medical imaging: Methods, challenges,andconsiderations’,JournaloftheAmerican CollegeofRadiology,19(8):975–982.

43. Federated Learning Approaches for Healthcare AI (2025) ‘Federated learning approaches for privacypreserving AI in healthcare data science’, Journal of InformaticsEducationandResearch,5(2).

44. Systematic review of FL in healthcare ethics (2026) ‘Federated learning in healthcare ethics: PrivacypreservingandequitablemedicalAI’,Healthcare,14(3), 306.

45. Systematic Pattern analysis of FL privacy (2024) ‘Privacypreservationforfederatedlearninginhealth care’,Patterns,5(7),100974.

46. Comprehensive review of FL challenges (2024) ‘Federatedlearninginhealthcare:Modelmisconducts, security,challenges’,arXiv:2405.13832.

47. IEEE-scalereviewofprivacyenhancementFL(2025) ‘Efficient and privacy-enhanced federated learning’, JournalofElectricalSystems.

48. MedicalAIreviewonFLandprivacy(2025)Daram,S. IJAIDSML.

49. FL in healthcare analytics framework (2025) Richardson, A. ‘Federated Learning Framework for Privacy-Preserving Healthcare Analytics’, Eureka JournalofComputingScience&DigitalInnovation,1(1), pp.8–14.

2026, IRJET | Impact Factor value: 8.226 | ISO 9001:2008 Certified Journal | Page

Turn static files into dynamic content formats.

Create a flipbook
A REVIEW OF DECENTRALIZED COLLABORATIVE MODEL TRAINING FOR MEDICAL DATA USING FEDERATED LEARNING WIT by IRJET Journal - Issuu