
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Annapurna Yadav1 , Mrs. Arifa Khan2
1Master of Technology, Computer Science and Engineering, Lucknow Institute of Technology, Lucknow, India
2Assistant Professor, Department of Computer Science and Engineering, Lucknow Institute of Technology, Lucknow, India
Abstract - The rapid proliferation of encrypted communication and data-driven network services has significantly increased the demand for accurate and privacypreserving traffic classification mechanisms. Traditional traffic analysis techniques, including deep packet inspection andcentralizedmachinelearningmodels,oftenrequireaccess to raw packet payloads, raising serious privacy, regulatory, and security concerns in privacy-sensitive environments such as healthcare, finance, and critical infrastructure networks. Recently, split learning has emerged as a promising distributed learning paradigm that partitions deep neural networks between clients and servers, enabling collaborative modeltrainingwithoutdirectsharingofrawdata.Thisreview systematically examines the state-of-the-art research onsplit learning–based secure traffic classification, focusing on architectural designs, privacy guarantees, communication overhead, and performance trade-offs. We analyze how split learning compares with related paradigms such as federated learninganddifferentialprivacy–basedmethodsinthecontext ofencryptedandlarge-scalenetworktraffic.Furthermore,this review synthesizes existing datasets, evaluation metrics, and threatmodelsadoptedinpriorstudies,identifying limitations in benchmarking practices and security analysis. Key challenges,includinggradientleakage,scalabilityconstraints, and adversarial robustness, are critically discussed. Finally, potential research directions are outlined to guide future developmentstowardpractical,secure,andhigh-performance deployment of split learning frameworks for traffic classification in privacy-sensitive networks.
Key Words: Split Learning, Secure Traffic Classification, Privacy-Preserving Machine Learning, Encrypted Network Traffic, Federated Learning, Network Security, Distributed Deep Learning
TheexponentialgrowthofInternet-enabledservices,cloud computing, IoT ecosystems, and mobile applications has intensified the need for efficient and intelligent network traffic classification. Traffic classification plays a fundamental role in network management, intrusion detection, quality-of-service (QoS) enforcement, and cybersecurity operations. However, the increasing deployment of encryption protocols and strict privacy regulationshassignificantlycomplicatedtraditionaltraffic analysis approaches. In this context, privacy-preserving
distributed learning paradigms particularly split learning haveemergedaspromisingsolutions.Thisreview examines the evolution, challenges, and emerging role of split learning in secure traffic classification for privacysensitivenetworks.
1.1.1
Network traffic classification has evolved through several methodologicalparadigms.Earlyapproachesreliedonportbased identification, where traffic flows were categorized basedonwell-knowntransportlayerportnumbers.While computationallyefficient,thistechniquebecameunreliable asapplicationsincreasinglyadopteddynamicportsandport obfuscationmechanisms(MooreandPapagiannaki,2005).
Subsequently,deeppacketinspection(DPI)techniqueswere introduced, enabling payload-level inspection to achieve fine-grained classification. DPI significantly improved accuracybutrequireddirectaccesstopacketcontent,raising scalability and privacy concerns (Nguyen and Armitage, 2008).
Withtheproliferationofencryptedandhigh-volumetraffic, statistical and machine learning-based methods gained prominence.Theseapproachesleveragedflow-levelfeatures such as packet inter-arrival time, flow duration, and byte distribution patterns. More recently, deep learning architectures including convolutional neural networks (CNNs) and recurrent neural networks (RNNs) have demonstrated superior performance in encrypted traffic classification tasks by automatically extracting discriminativefeatures(Lotfollahietal.,2020).
The widespread adoption of encryption protocols such as TLS has fundamentally altered the traffic classification landscape. Current reports indicate that a substantial majority of Internet traffic is encrypted, limiting visibility into packet payloads and rendering DPI ineffective (Andersonetal.,2017).
Simultaneously,regulatoryframeworkssuchastheGeneral Data Protection Regulation (GDPR) and sector-specific compliancemandateshaveimposedstrictrequirementson

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
how user data is processed and stored. These regulations emphasizedataminimization,consent,andsecurehandling, thereby restricting centralized collection of raw network data for machine learning purposes (Voigt and Von dem Bussche,2017).
Consequently, traffic classification systems must balance detection accuracy with privacy preservation, creating a complexdesigntrade-offbetweenvisibilityandcompliance.
Centralized machine learning frameworks require aggregationofrawtrafficdataatacentralserverformodel training.Althougheffectiveincontrolledenvironments,this approachintroducesmultiplelimitations.First,transferring raw network data across administrative boundaries increases exposure risk and attack surfaces. Second, centralized storage creates a single point of failure vulnerabletobreaches.Third,large-scaledatatransmission leads to significant communication overhead and latency (Kairouzetal.,2021).
Moreover, centralized deep learning models may inadvertentlymemorizesensitivetrafficpatterns,enabling model inversion or membership inference attacks. These vulnerabilities highlight the necessity for distributed privacy-awarelearningarchitecturesthatminimizerawdata exposure.
Rawpacketinspectioninherentlyinvolvesanalyzingheaders and payload contents, which may contain personally identifiable information (PII), authentication tokens, or proprietary communication details. Even flow-level metadatacanrevealbehavioralpatternswhenaggregatedat scale.Studieshaveshownthattrafficanalysistechniquescan infer user activities despite encryption, raising serious privacyimplications(Shbairetal.,2016).
Furthermore, in domains such as healthcare, financial systems, and critical infrastructure, traffic data may correspond to highly sensitive operational information. Unauthorized exposuremay resultin regulatorypenalties and reputational damage. Therefore, minimizing raw data visibilityduringmodeltraininghasbecomeafundamental requirement.
1.2.2
Modernnetworksareinherentlydistributed,spanningedge devices, IoT sensors, enterprise gateways, and cloud infrastructures. In such environments, data is naturally partitionedacrossmultiplenodes.Distributedcollaborative
learning enableslocal model training withoutcentralizing rawdata,therebyreducingexposurerisks.
Paradigms such as federated learning have demonstrated the feasibility of decentralized training by sharing model updates instead of data (McMahan et al., 2017). However, evengradientsharingmayleaksensitiveinformationunder certain threat models. These limitations motivate explorationofalternativearchitecturesthatfurtherrestrict information exchange, leading to the adoption of split learningforprivacy-sensitivetrafficclassificationscenarios.
1.3.1
Basic Architecture
Split learning is a distributed deep learning paradigm in which a neural network is partitioned between client and serversegments.Clientsperformforwardpropagationupto a predefined cut layer and transmit intermediate activations notrawdata toacentralserver.Theserver completesforwardandbackwardpropagationandreturns gradients to the client for local parameter updates (Vepakommaetal.,2018).
This architecture reduces direct data exposure while maintaining collaborative training efficiency. Because raw inputsremainattheclientside,splitlearningisparticularly suitedforprivacy-sensitiveapplicationswheredatacannot besharedacrossorganizationalboundaries.
1.3.2 Distinction from Other Distributed Paradigms
Unlikefederatedlearning,wherecompletelocalmodelsare trained independently and only model parameters are shared,splitlearningdividesasinglemodelacrossentities. This structural difference reduces computational requirements at the client side and may offer improved privacyunderspecificthreatassumptions.
Additionally,comparedtosecuremulti-partycomputationor homomorphic encryption which introduce significant computationaloverhead splitlearningprovidesapractical trade-offbetweenprivacyandefficiency(Singhetal.,2019). Nevertheless,recentresearchhasidentifiedpotentialrisks such as activation leakage and gradient reconstruction attacks, indicating that split learning is not inherently immunetoinferencethreats.
1.4 Contributions of the Review
Thisreviewprovidesacomprehensivesynthesisofresearch on split learning–based secure traffic classification in privacy-sensitivenetworks.First,itsystematicallyexamines the evolution of traffic classification methodologies and contextualizes the transition toward privacy-preserving learning frameworks. Second, it critically analyzes split learning architectures, their security properties, and comparative advantages over alternative distributed

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
paradigms.Third,itevaluatesexistingliteraturewithrespect to datasets, performance metrics, threat models, and deployment feasibility. Finally, it identifies research gaps and outlines future directions aimed at improving robustness, scalability, and regulatory compliance in realworldnetworkenvironments.
Secure traffic classification integrates traditional network traffic analysis with privacy-preserving computational paradigms. As modern networks increasingly handle encryptedandsensitivecommunications,thefoundationsof secure classification must encompass methodological accuracy, security guarantees, and regulatory compliance. This section reviews core traffic classification techniques, examines privacy and security concerns, and outlines distributed learning paradigms that enable privacy-aware modeltraining.
2.1.1
Port-basedclassificationrepresentstheearliesttechnique for identifying application-layer protocols in IP networks. This approach maps traffic flows to predefined transport layerportnumbers(e.g.,HTTPonport80,HTTPSonport 443).Itsprimaryadvantageliesincomputationalsimplicity and low processing overhead. However, the technique suffers from severe limitations due to dynamic port allocation, port masquerading, and tunneling mechanisms used by modern applications (Moore and Papagiannaki, 2005).
Furthermore, with the increasing adoption of encrypted protocolsandcontentdeliverynetworks,portnumbersno longerreliablycorrespondtospecificservices.Consequently, port-based methods have become largely obsolete in contemporaryhigh-speedandencryptedenvironments.
To address the limitations of port-based identification, statistical and flow-based classification techniques were introduced. These methods rely on traffic flow characteristics such as packet size distribution, flow duration,inter-arrivaltime,andbytecounts.Byextracting features from NetFlow or similar metadata records, classifiers can infer application types without accessing payloadcontents(NguyenandArmitage,2008).
Machine learning algorithms including Support Vector Machines (SVM), Random Forests, and k-Nearest Neighbors havebeenwidelyappliedtosuchfeaturesets. Althoughtheseapproachesimprovedclassificationaccuracy in encrypted environments, their performance depends heavily on feature engineering and dataset
representativeness. Additionally, flow-level metadata can still reveal sensitive behavioral information when aggregatedatscale.
Recent advancements in deep learning have significantly enhanced encrypted traffic classification. Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) enable automated feature extraction from raw packetsequencesorflowmatrices,reducingdependenceon manualfeaturedesign.Architecturessuchas1D-CNNsand Long Short-Term Memory (LSTM) networks have demonstrated superior performance in identifying encryptedapplicationsandintrusionpatterns(Lotfollahiet al.,2020).
Deep learning models offer scalability and adaptability to evolving traffic patterns. However, they typically require large volumes of labeled data and centralized training infrastructure, which introduces privacy and security concernsinsensitivedeploymentscenarios.
2.2.1
Evenwhentrafficpayloadsareencrypted,metadatasuchas timinginformation,packetlengths,andflowstatisticsmay enable inference of user behavior or application types. Traffic analysis attacks have shown that encrypted communicationscanstillleaksensitiveinformationthrough side-channelcharacteristics(Shbairetal.,2016).
Centralized storage of network traffic datasets further increasesexposurerisksbycreatinghigh-valuetargetsfor attackers.Inaddition,machinelearningmodelsthemselves may leak information through model inversion or membershipinferenceattacks,particularlywhentrainedon sensitivedatasets.Therefore,minimizingrawdatasharingis a critical design objective in secure traffic classification frameworks.
Theregulatorylandscapehasintensifiedtheimportanceof privacy-aware data processing. Frameworks such as the General Data Protection Regulation (GDPR) enforce strict requirements for lawful data collection, processing transparency, and data minimization (Voigt and Von dem Bussche,2017).
For network operators and service providers, centralized traffic monitoring systems may conflict with compliance requirements if personal or behavioral data is stored without explicit consent. Consequently, privacy-by-design principles must be incorporated into traffic classification systems,encouragingdecentralizedandprivacy-preserving learningapproaches.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Securetrafficclassificationmustbeevaluatedunderclearly defined threat models. Common adversarial assumptions include honest-but-curious servers, malicious clients, externaleavesdroppers,andinsiderthreats.Attackvectors mayinvolvegradientreconstruction,activationleakage,or trafficcorrelationattacks.
Threatmodelingisessentialforunderstandingtheprivacy guarantees of machine learning-based classification frameworks.Withoutrigorousadversarialanalysis,systems may provide a false sense of security while remaining vulnerabletoinference-basedattacks.
2.3.1
Centralizedlearningaggregatesrawdataatasingleserver for model training. While this approach simplifies optimization and coordination, it introduces significant communication overhead and data exposure risks. In contrast,distributedlearningretainsdatalocallyandshares model-relatedinformationinsteadofrawinputs.
Distributed paradigms enhance scalability and align with privacy requirements, particularly in multi-organization environments where direct data sharing is restricted (Kairouz et al., 2021). However, distributed training introduces challenges in synchronization, communication efficiency,androbustness.
2.3.2
Federated Learning (FL) is a decentralized learning framework in which clients locally train full models and share parameter updates with a coordinating server. The serveraggregatesupdates commonlyusingtheFederated Averaging(FedAvg)algorithm toconstructaglobalmodel (McMahanetal.,2017).
FLreducestheneedforrawdatatransmissionandhasbeen explored for network intrusion detection and traffic classificationtasks.Nevertheless,researchhasdemonstrated that shared gradients may still leak sensitive information under certain attack models, necessitating additional protectivemechanismssuchassecureaggregation.

2.3.3
DifferentialPrivacy(DP)providesamathematicallyrigorous frameworkforlimitinginformationleakagefromstatistical outputs.Byinjectingcalibratednoiseintomodelgradientsor queryresponses,DPensuresthattheinclusionorexclusion of a single data record does not significantly affect model outputs(Dworketal.,2014).
Intrafficclassification,DPmechanismscanprotectsensitive traffic patterns during collaborative learning. However, privacy guarantees come at the cost of reduced model accuracy, especially when noise levels are high. Balancing privacybudgetsandclassificationperformanceremainsan activeresearchchallengeinsecurenetworkanalytics.
Splitlearninghasemergedasapromisingdistributeddeep learning paradigm designed to address privacy, communicationefficiency,andcomputationalconstraintsin collaborative environments. Unlike traditional distributed learning frameworks that replicate full models across participants, split learning partitions a neural network between clients and a central server. This structural separation allows sensitive data to remain local while enablingjointmodeloptimization.Thefollowingsubsections discuss its architectural mechanisms, design variants, and associatedprivacyconsiderations.
3.1.1
Thedefiningfeatureofsplitlearningismodelpartitioning.A deepneuralnetworkisdividedatapredefined“cutlayer” into two segments: the client-side sub network and the server-sidesubnetwork.Theclientprocessesrawinputdata through its local layers and transmits intermediate activationstotheserver.Theservercompletestheforward passandcomputesthelossfunction.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
This partitioning significantly reduces the need for transmitting raw data across network boundaries. The concept was formally introduced to enable privacypreserving deep learning in distributed medical and enterprise environments, demonstrating that sensitive datasets can remain local without compromising collaborativetrainingefficiency(Vepakommaetal.,2018). Furthermore, partitioning allows lightweight clients to offloadcomputationallyintensivelayerstomorepowerful servers,improvingscalabilityinedge-baseddeployments.

3.1.2
In split learning, training follows a coordinated forward–backwardpropagationworkflow.Duringtheforwardpass, the client computes activations up to the cut layer and transmits them to the server. The server completes the forward propagation, evaluates the loss, and initiates backpropagation.Gradientscorrespondingtothecutlayer are sent back to the client, which then updates its local parameters.
This sequential interaction ensures that raw input data neverleavestheclientenvironment.Comparedtofederated learning where complete model gradients are shared splitlearningonlyexchangesintermediaterepresentations and partial gradients. This design reduces client-side computational load and may lower communication complexity in certain configurations (Gupta and Raskar, 2018).However,thesequentialdependencybetweenclient andserverintroduceslatencyconsiderationsthatmustbe managedinlarge-scaledeployments.
3.2.1
Vanillasplitlearningrepresentstheoriginalimplementation of the paradigm, involving a single client and a central server.Inthisconfiguration,onlyoneparticipantperforms localcomputationbeforetransmittingactivations.Although conceptuallystraightforward,vanillasplitlearningprimarily servesasabaselinearchitectureforevaluatingprivacyand efficiencycharacteristics.
Its applicability is limited in multi-client scenarios unless extended with scheduling mechanisms. Nevertheless, it providesfoundationalinsightsintocommunicationoverhead andinformationleakagedynamics.
To accommodate multiple data owners, multi-client split learning architectures have been proposed. In these configurations, multiple clients sequentially or asynchronously interact with a shared server-side model. Somedesignsincorporateclient-sideweightsynchronization orparallelizedcut-layerstrategiestoimprovescalability.
Advanced variants also introduce techniques such as activation compression and pipelining to mitigate communication bottlenecks. These enhancements aim to make split learning viable in bandwidth-constrained or latency-sensitiveenvironments,particularlyinIoTandedge computing scenarios (Singh et al., 2019). However, coordination complexity increases as the number of participatingclientsgrows.
Hybridarchitecturescombineelementsoffederatedlearning and split learning to leverage the advantages of both paradigms.Insuchsystems,clientsmaytrainpartialmodels locallyusingfederatedaveragingwhilestillsplittingdeeper layers with a server. This approach reduces sequential dependencyandcanenhanceparallelization.
Hybrid frameworks are particularly useful when computationalresourcesareheterogeneousacrossclients. By integrating model partitioning with federated aggregation, these architectures aim to balance privacy, scalability, and training efficiency. Recent studies have explored hybrid designs to strengthen resilience against inference attacks while maintaining acceptable model convergencerates(Thapaetal.,2022).
Althoughsplitlearningpreventsdirectsharingofrawdata, the transmission of intermediate activations introduces potential privacy risks. Research has demonstrated that under certain conditions, adversaries may reconstruct approximateinputdatafromsharedactivations,especially whenthecutlayerisshallow(Hitajetal.,2017).
The risk of activation leakage depends on network architecture, depth of partitioning, and adversarial capabilities. Deeper cut layers generally reduce reconstruction feasibility but may increase client computational burden. Therefore, selecting an optimal partitionpointiscriticalforbalancingprivacyandefficiency.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
In addition to activation leakage, gradient-based attacks posesignificantthreatstodistributedlearningframeworks. Gradient inversion techniques can potentially recover sensitive input information from shared gradients during backpropagation. Although such attacks were initially studiedinfederatedlearningcontexts,similarvulnerabilities may arise in split learning if adversaries gain access to intermediategradients(Zhuetal.,2019).
Mitigation strategies include adding noise, applying differentialprivacymechanisms,orencryptingintermediate representations.However,thesedefensesintroducetradeoffs in computational cost and model accuracy. Consequently, while split learning enhances privacy compared to centralize training, it does not inherently guarantee complete protection against sophisticated inferenceattacks.
Thissectioncriticallysynthesizespriorresearchontraffic classification and its evolution toward privacy-preserving distributedlearningparadigms,withparticularemphasison splitlearning–basedsecuretrafficanalytics.
4.1.1
Before the widespread adoption of deep learning, traffic classificationprimarilyreliedonclassicalmachinelearning algorithms trained on flow-level statistical features. Techniques such as Support Vector Machines (SVM), DecisionTrees,RandomForests,andk-NearestNeighbors werewidelyusedduetotheirinterpretabilityandmoderate computational requirements. Early empirical evaluations demonstratedthatstatisticalfeature-basedclassifierscould achievecompetitiveaccuracyinidentifyingapplication-layer protocolswithoutpayloadinspection(NguyenandArmitage, 2008).
However, these approaches were heavily dependent on manualfeatureengineeringandstruggledwithencryptedor obfuscated traffic. Additionally, model performance often degradedwhenexposedtoevolvingtrafficpatternsorzerodayapplications,highlightinggeneralizationlimitationsin dynamicnetworkenvironments.
The rise of encrypted traffic accelerated the transition toward deep learning architectures capable of automated featureextraction.ConvolutionalNeuralNetworks(CNNs) havebeenemployedtoanalyzerawpacketbytesequences, whileRecurrentNeuralNetworks(RNNs)andLongShort-
Term Memory (LSTM) networks capture temporal dependencies in traffic flows. Deep Packet, for example, demonstratedthefeasibilityofend-to-enddeeplearningfor encryptedtrafficclassificationwithouthandcraftedfeatures (Lotfollahietal.,2020).
Despiteimprovedclassificationaccuracy,centralizeddeep learningmodelsintroduceprivacyrisksduetotheneedfor large-scaledataaggregation.Moreover,highcomputational demandsandstoragerequirementslimittheirdeployment indistributedorresource-constrainedenvironments.
Federated Learning (FL) has been widely explored as a privacy-aware alternative to centralized training. In FLbasedtrafficclassificationsystems,eachclienttrainsalocal modelonitsownnetworkdataandshares modelupdates withacentralaggregatorusingalgorithmssuchasFederated Averaging(McMahanetal.,2017).
SeveralstudieshaveappliedFLtointrusiondetectionand encryptedtrafficclassification,demonstratingreduceddata exposure while maintaining competitive performance. However, subsequent analyses revealed that shared gradientsmayleaksensitiveinformationunderadversarial settings, raising concerns about gradient inversion and membershipinferencevulnerabilities(Zhuetal.,2019).
Differential Privacy (DP) introduces mathematically quantifiableprivacyguaranteesbyinjectingcalibratednoise intomodelparametersorgradients.Intrafficclassification, DP mechanisms have been integrated into distributed learning frameworks to limit information leakage from sharedupdates(Dworketal.,2014).
WhileDPenhancesprivacyprotection,theaddednoisemay degrade classification accuracy, particularly in complex encryptedtrafficscenarios. Achievinganoptimalprivacy–utility balance remains an ongoing research challenge, especiallywhenstrictprivacybudgetsareenforced.
Secure Multi-Party Computation (SMPC) enables collaborativecomputationwithoutrevealingprivateinputs amongparticipatingentities.Cryptographictechniquessuch assecretsharingandhomomorphicencryptionhavebeen exploredtotraintrafficclassificationmodelssecurelyacross multipleorganizations.
Although SMPC provides strong theoretical security guarantees,itscomputationaloverheadandcommunication complexity often limit scalability in high-throughput

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
networkenvironments.Consequently,practicaldeployment inreal-timetrafficmonitoringsystemsremainschallenging.
4.3.1
Architectural Setups
Initial studies applying split learning to network security tasks adopted client–server neural network partitioning models,whereearlylayerswereexecutedlocallyanddeeper layerswereprocessedcentrally.Thisconfigurationaimedto preventrawpacketexposurewhileleveragingcentralized computationalresources.Foundationalworkdemonstrated thefeasibilityofsplitlearninginprivacy-sensitivedomains, includinghealthcareanalytics,establishingthearchitectural principles later adapted for network security applications (Vepakommaetal.,2018).
Dataset Usage
Early implementations primarily relied on benchmark intrusiondetectionandtrafficclassificationdatasetssuchas NSL-KDD, UNSW-NB15, and CICIDS2017. These datasets providedlabeledflow-levelfeaturessuitableforevaluating distributedarchitectures.However,limitedexperimentation onreal-worldencryptedtrafficdatasetsconstrainedexternal validity.
Empiricalresultsindicatedthatsplitlearningcouldachieve classification accuracy comparable to centralized deep learning models while reducing raw data exposure. Nonetheless, communication overhead and sequential training dependencies were identified as potential bottlenecks.
4.3.2 Advanced Frameworks and Enhancements
Communication-Efficient Models
Recent advancements have focused on reducing communication latency through activation compression, model pruning, and pipelining techniques. These enhancements aim to optimize bandwidth utilization in distributedenvironments,particularlyforIoT-basedtraffic monitoringsystems(Singhetal.,2019).
Secure Aggregation Mechanisms
To mitigate intermediate information leakage, secure aggregation and encryption mechanisms have been integrated into split learning frameworks. These methods encrypt activations or gradients before transmission, strengthening resilience against eavesdropping and maliciousservers.
Advancedframeworksincorporateadversarialtrainingand noiseinjectionstrategiestodefendagainstinferenceattacks. Hybridarchitecturescombiningsplitlearningwithfederated learninghavealsobeenproposedtoimproverobustnessand parallelizationefficiency(Thapaetal.,2022).
Accuracy vs Privacy Trade-Off
Comparative analyses indicate that split learning offers a favorable balance between privacy preservation and classification accuracy relative to purely centralized approaches.However,privacygainsdependheavilyoncutlayerdepthandthreatassumptions.
Client-sidecomputationalloadinsplitlearningistypically lowerthaninfederatedlearning,asonlypartialmodelsare trained locally. Nevertheless, sequential client–server interactionsmayincreasetraininglatency.
Scalability remains a key concern in multi-client deployments. While model partitioning reduces data transmissionvolume,coordinationoverheadgrowswiththe number of participants. Efficient scheduling and asynchronous communication protocols are therefore criticalforlarge-scalenetworkapplications.
Mostreviewedstudiesrelyonpublicintrusiondetectionand trafficclassificationbenchmarkssuchasCICIDS2017,UNSWNB15, and ISCX VPN-nonVPN datasets. Although these datasetsfacilitatereproducibility,theymaynotfullycapture contemporaryencryptedtrafficcharacteristics.
Performanceistypicallyassessedusingaccuracy,precision, recall,F1-score,andAreaUndertheCurve(AUC).However, few studies report privacy leakage metrics or communication cost evaluations, limiting comprehensive comparison.
Reproducibility remains a challenge due to variations in preprocessing pipelines, feature extraction strategies, and split-layerconfigurations.Inconsistentexperimentalsetups hinderfairbenchmarkingandcross-studyvalidation.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Althoughsplitlearning–basedsecuretrafficclassificationhas demonstrated promising capabilities in privacy-sensitive environments,severalunresolvedchallengesremain.These challengesspantechnicalefficiency,privacyrobustness,and evaluation standardization. A critical examination of the literaturerevealsthatexistingframeworksoftenprioritize accuracy over holistic system design, leaving important research gaps in scalability, security guarantees, and benchmarkingconsistency.
5.1.1 Communication
Oneoftheprimarytechnicalconstraintsofsplitlearningis communication overhead resulting from iterative client–serverinteractions.Duringeachtrainingcycle,intermediate activationsmustbetransmittedfromtheclienttotheserver, followedbygradientupdatesreturnedtotheclient.Inhighfrequency traffic monitoring systems, this repeated bidirectionalcommunicationcangeneratesignificantlatency and bandwidth consumption, particularly in wide-area or edge-baseddeployments(Singhetal.,2019).
Unlikefederatedlearning,wheremodelupdatesaretypically exchangedperiodically,splitlearningrequiressynchronous communication at every batch iteration. This sequential dependency limits throughput and may hinder real-time trafficclassification.Althoughcompressionandquantization techniques have been proposed to reduce activation size, thesemethodsintroduceadditionalcomputationaloverhead andmayaffectnumericalstability.
5.1.2
Model convergence in split learning depends on stable coordination between distributed segments of the neural network. The partitioning of layers alters gradient propagation dynamics, potentially affecting optimization stability. Research in distributed optimization has shown thatasynchronousupdates,non-IIDdatadistributions,and heterogeneous client resources can slow convergence or causeoscillatorybehavior(Kairouzetal.,2021).
Intrafficclassificationscenarios,wherenetwork behavior varies across domains, non-identically distributed data further complicates convergence. Existing studies often evaluatemodelsundercontrolledexperimentalconditions, leaving open questions regarding robustness under realworldvariabilityandlarge-scalemulti-clientdeployments.
5.2.1
Althoughsplitlearningpreventsdirectsharingofrawtraffic data,exchangedgradientsandintermediaterepresentations
may still leak sensitive information. Gradient inversion attacks have demonstrated that input samples can be reconstructed from shared gradients under certain assumptions(Zhuetal.,2019).
In split learning, similar vulnerabilities arise when adversariesgainaccesstocut-layergradients.Theextentof leakage depends on model architecture, activation dimensionality,andadversarialknowledge.Whiledefensive strategies such as gradient clipping, noise injection, and encryptionhavebeenproposed,theirintegrationintotraffic classification systems remains limited. Consequently, the privacyguaranteesofmanysplitlearningimplementations areempiricalratherthanformallyproven.
Model inversion attacks attempt to reconstruct sensitive input features by exploiting access to trained model parametersoroutputs.Indistributedlearningframeworks, adversaries may leverage prediction confidence scores or intermediate activations to infer private traffic characteristics. Studies in collaborative learning contexts have demonstrated that deep models can inadvertently memorize training data patterns, enabling partial reconstructionofinputs(Hitajetal.,2017).
In privacy-sensitive network environments, such leakage may reveal behavioral signatures or communication metadata.Despitetheserisks,manyexistingsplitlearning studies evaluate privacy qualitatively rather than through rigorous adversarial testing. The absence of standardized attacksimulationsrepresentsasignificantresearchgap.
A critical limitation in current literature is the lack of standardizedbenchmarkingframeworksforevaluatingsplit learning–based traffic classification. Studies frequently employdifferentdatasets,preprocessingpipelines,feature extraction strategies, and split-layer configurations. As a result,directperformancecomparisonacrosspublications becomeschallenging.
Moreover, evaluation metrics predominantly focus on classificationaccuracy,precision,recall,andF1-score,while neglecting privacy leakage quantification and communicationcostanalysis.Fewworksreportmetricssuch as bandwidth consumption, latency, or privacy budgets, limitingcomprehensiveassessment.Theabsenceofunified experimental protocols undermines reproducibility and impedes fair comparison among distributed learning paradigms.
Despite notable progress in split learning–based secure traffic classification, several open challenges hinder its seamless deployment in privacy-sensitive network

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
environments.Thesechallengesextendbeyondarchitectural designandencompassoperationalfeasibility,generalization acrossheterogeneousdomains,adversarialrobustness,and regulatorycompliance.Addressingtheseissuesisessential for translating experimental prototypes into productiongradenetworksecuritysystems.
One of the most pressing challenges involves real-time deployment in high-throughput network environments. Traffic classification systems in enterprise backbones, 5G infrastructures,andclouddatacentersmustprocessmassive volumes of packets with minimal latency. Split learning introduces sequential client–server interactions during training, which can generate additional communication delayscomparedtofullylocalinferencesystems(Singhetal., 2019).
Althoughinferencecanbeoptimizedaftertraining,dynamic environments often require continuous model updates to adapttoevolvingtrafficpatterns.Thisiterativeretraining processmayconflictwithstrictlatencyrequirementsinrealtimeintrusiondetectionsystems.Furthermore,edgedevices participating in distributed training may have limited computational and energy resources, constraining the feasibilityofdeepmodelpartitioning.Efficientscheduling, lightweightarchitectures,andhardware-awareoptimization remainareasrequiringfurtherinvestigation.
Network traffic characteristics vary significantly across domains,suchasenterprisenetworks,IoTecosystems,and mobile carrier infrastructures. Models trained on one dataset often experience performance degradation when deployed in a different operational environment due to domainshiftandnon-identicallydistributed(non-IID)data (Kairouzetal.,2021).
In split learning scenarios, cross-domain heterogeneity furthercomplicatesconvergenceandgeneralization.Clients maypossessvastlydifferenttrafficdistributions,encryption protocols, or application usage patterns. Without domain adaptation mechanisms, global models risk bias toward dominantparticipants.Techniquessuchastransferlearning, domainadversarialtraining,andmeta-learninghavebeen proposed in broader machine learning contexts, but their integration with split learning for traffic classification remainsunderexplored.Developingadaptivearchitectures capableofhandlingheterogeneousnetworkenvironments constitutesacriticalresearchdirection.
Adversarialrobustnessrepresentsanothersignificantopen challenge.Machinelearningmodelsfortrafficclassification are vulnerable to evasion and poisoning attacks, where
adversariesmanipulatetrafficpatternsortrainingdata to degrade detection accuracy. In distributed learning frameworks,maliciousclientsmayinjectpoisonedupdates toinfluenceglobalmodelbehavior(Bhagojietal.,2019).
In split learning, threats may arise from compromised clients, curious servers, or external eavesdroppers intercepting intermediate activations. Additionally, adversarialexamplescraftedtomimiclegitimateencrypted traffic can bypass classifiers without altering encryption protocols. Defensive strategies including robust aggregation, anomaly detection for gradient updates, and adversarial training require further adaptation to split learning architectures. Formal security proofs and systematic red-team evaluations are largely absent in currenttrafficclassificationstudies.
Regulatory compliance poses both technical and organizational challenges. Data protection frameworks emphasizeaccountability,transparency,anduserconsentin data processing activities. While distributed learning reduces raw data transfer, it does not automatically guarantee compliance if intermediate representations or modeloutputsrevealsensitiveinformation(VoigtandVon demBussche,2017).
Moreover,practicalintegrationofsplitlearningintoexisting network management systems requires interoperability withlegacymonitoringtools,standardizedAPIs,andscalable orchestration frameworks. Enterprises must balance compliance requirements withoperational efficiency, cost considerations, and cybersecurity objectives. Auditable privacyguarantees,explainablemodelbehavior,andclear governancepoliciesareessentialforreal-worldadoption.
This review has systematically examined the evolution of securetrafficclassificationwitha particularfocusonsplit learning–based distributed architectures for privacysensitive networks. Traditional port-based and statistical approacheshavegraduallybeenreplacedbydeeplearning modelscapableofhandlingencryptedandhigh-dimensional traffic patterns. However, centralized training paradigms introduce substantial privacy risks and regulatory challenges. Distributed learning frameworks, including federated learning and differential privacy mechanisms, represent important milestones toward privacy-aware analytics. Within this landscape, split learning offers a distinctive architectural advantage by partitioning neural networksbetweenclientsandservers,therebyminimizing directexposureofrawtrafficdata.
The literature indicates that split learning can achieve classificationperformancecomparabletocentralizeddeep

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
learningwhilereducingdatasharingrisks.Nevertheless,its effectiveness depends heavily on model partitioning strategies, threat assumptions, communication efficiency, anddeploymentcontext.Comparativeanalysesrevealtradeoffs among accuracy, scalability, and privacy guarantees, particularly in heterogeneous network environments. Although promising, split learning does not inherently eliminate risks such as gradient leakage or activation reconstruction. Future advancements must therefore integrate formal privacy guarantees, communication optimization,andadversarialrobustnesstoenablepractical large-scale adoption. Overall, split learning represents a viableandevolvingparadigmforsecuretrafficclassification, butsustainedresearcheffortsarenecessarytoaddressits architecturalandoperationalconstraints.
This review is subject to several limitations. First, it primarilysynthesizespeer-reviewedacademicliteratureand may not fully capture proprietary industrial implementations of split learning in operational network environments. Second, variations in experimental setups, datasets, and evaluation metrics across studies limit the ability to provide strict quantitative comparisons. Third, rapid advancements in distributed learning and privacypreservingtechniquesmeanthatnewlyemergingmethods may not be comprehensively represented. Additionally, while privacy risks and adversarial threats are discussed conceptually, detailed empirical validation of attack resilience is beyond the scope of this review. Finally, the focus on split learning may underrepresent alternative cryptographic or hardware-based secure computation approachesthatcouldalsocontributetoprivacy-sensitive trafficclassification.
1) Anderson, B., Paul, S. and McGrew, D. (2017) ‘Deciphering malware’s use of TLS (without decryption)’,JournalofComputerVirologyandHacking Techniques,13(3),pp.195–211.
2) Bhagoji, A.N., Chakraborty, S., Mittal, P. and Calo, S. (2019) ‘Analyzing federated learning through an adversariallens’,Proceedingsofthe36thInternational ConferenceonMachineLearning(ICML),pp.634–643.
3) Dwork, C., Roth, A., et al. (2014) ‘The algorithmic foundations of differential privacy’, Foundations and Trends in Theoretical Computer Science, 9(3–4), pp. 211–407.
4) Gupta,O.andRaskar,R.(2018)‘Distributedlearningof deepneuralnetworkovermultipleagents’,Journalof NetworkandComputerApplications,116,pp.1–8.
5) Hitaj,B.,Ateniese,G.and Perez-Cruz,F.(2017)‘Deep models under the GAN: Information leakage from
collaborative deep learning’, Proceedings of the ACM SIGSACConferenceonComputerandCommunications Security(CCS),pp.603–618.
6) Kairouz, P., McMahan, H.B., Avent, B., et al. (2021) ‘Advances and open problems in federated learning’, FoundationsandTrendsinMachineLearning,14(1–2), pp.1–210.
7) Lotfollahi, M., Siavoshani, M.J., Zade, R.S.H. and Saberian,M.(2020)‘DeepPacket:Anovelapproachfor encryptedtrafficclassificationusingdeeplearning’,Soft Computing,24(3),pp.1999–2012.
8) McMahan,H.B.,Moore,E.,Ramage,D.,Hampson,S.and Arcas,B.A.Y.(2017)‘Communication-efficientlearning ofdeepnetworksfromdecentralizeddata’,Proceedings of the 20th International Conference on Artificial IntelligenceandStatistics(AISTATS),pp.1273–1282.
9) Moore,A.W.andPapagiannaki,K.(2005)‘Towardthe accurateidentificationofnetworkapplications’,Passive andActiveNetworkMeasurement(PAM),LNCS3431, pp.41–54.
10) Nguyen, T.T. and Armitage, G. (2008) ‘A survey of techniques for internet traffic classification using machine learning’, IEEE Communications Surveys & Tutorials,10(4),pp.56–76.
11) Shbair, W., Cholez, T., Francois, J. and Chrisment, I. (2016) ‘A multi-level framework to identify HTTPS services’, IEEE Conference on Communications and NetworkSecurity(CNS),pp.240–248.
12) Singh, A., Vepakomma, P., Gupta, O. and Raskar, R. (2019) ‘Detailed comparison of communication efficiencyofsplitlearningandfederatedlearning’,arXiv preprintarXiv:1909.09145.
13) Thapa, C., Arachchige, P.C.M., Camtepe, S. and Sun, L. (2022)‘SplitFed:Whenfederatedlearningmeetssplit learning’, Proceedings of the AAAI Conference on ArtificialIntelligence,36(8),pp.8485–8493.
14) Vepakomma, P., Gupta, O., Swedish, T. and Raskar, R. (2018) ‘Split learning for health: Distributed deep learning without sharing raw patient data’, arXiv preprintarXiv:1812.00564.
15) Voigt, P. and Von dem Bussche, A. (2017) The EU GeneralDataProtectionRegulation(GDPR):APractical Guide.Cham:Springer.
16) Zhu,L.,Liu,Z.andHan, S.(2019)‘Deepleakage from gradients’,AdvancesinNeuralInformationProcessing Systems(NeurIPS),32,pp.14774–14784.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
17) Li, P., Guo, C., Xing, Y., Shi, Y., Feng, L. and Zhou, F. (2024) ‘Core network traffic prediction based on verticalfederatedlearningandsplitlearning’,Scientific Reports, 14(1), p. 4663, doi:10.1038/s41598-02453193-y.
18) Jin, Z., Duan, K., Chen, C., He, M., Jiang, S. and Xue, H. (2024)‘FedETC:Encryptedtrafficclassificationbased onfederatedlearning’,Heliyon,10(16),e35962.
19) Trivedi, D., Boudguiga, A., Kaaniche, N. and Triandopoulos,N.(2026)‘SplitML:Aunifiedprivacypreservingarchitectureforfederatedsplit-learningin heterogeneous environments’, Electronics, 15(2), p. 267.
20) Alnasser, W., Beigi, G., Mosallanezhad, A. and Liu, H. (2022)‘PPSL:Privacy-preservingtextclassificationfor splitlearning’,in4thInt.Conf.onDataIntelligenceand Security(ICDIS),IEEE,pp.160–167.
21) Anonymous (2026) ‘PrivRobust-SL: A privacypreserving and adversarially robust split learning frameworkforIoTintrusiondetection’,JournalArticle, ScienceDirect.
22) Shalabi,E.,Khedr,W.,Rushdy,E.andSalah,A.(2025)‘A comparativestudyofprivacy-preservingtechniquesin federatedlearning:Performanceandsecurityanalysis’, Information,16(3),p.244.
23) Chaudhary,D.,Rajasegarar,S.andPokhrel,S.R.(2025) Towards Adapting Federated & Quantum Machine Learning for Network Intrusion Detection: A Survey, arXivpreprint.
24) Peng,Y.,He,M.andWang,Y.(2021)‘Afederatedsemisupervised learning approach for network traffic classification’,arXivpreprint,arXiv:2107.03933.
25) Jiang,L.,Wang,Y.,Zheng,W.,Jin,C.,Li,Z.andTeo,S.G. (2022)‘LSTMSPLIT:EffectivesplitlearningbasedLSTM on sequential time-series data’, arXiv preprint, arXiv:2203.04305.
26) Nguyen,P.T.,Dao,N.N.,Do,Q.T.,Nguyen,T.V.andCho,S. (Year) ‘Privacy-Preserving Traffic Flow Prediction: A Split Learning Approach’, Conference Proceeding, ElsevierPure.
27) (2025) ‘Efficient privacy-preserving ML for IoT: Cluster-basedsplitfederatedlearningschemefornonIID data’, Journal of Network and Computer Applications,doi:10.1016/j.jnca.2025.104105.
28) (2024)‘Encryptednetworktrafficclassificationbased onmachinelearning’,AinShamsEngineeringJournal, 15(2),102361.
29) (2025) ‘Enhanced IoT security: Privacy-preserving federated learning model for accurate, real-time intrusion detection across devices’, ScienceDirect Article.
30) (2024) ‘Federated distributed network traffic classification based on deep mutual learning’, MDPI Electronics,14(24),p.4928.