
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Divya Vishwakarma1 , Mrs. Arifa Khan2
1Master of Technology, Computer Science and Engineering, Lucknow Institute of Technology, Lucknow, India
2Assistant Professor, Department of Computer Science and Engineering, Lucknow Institute of Technology, Lucknow, India ***
Abstract - The effectiveness of machine learning models is oftenconstrainedbytheavailabilityoflarge,labeleddatasets, which are difficult and costly to obtain in many real-world scenarios such as healthcare, scientific research, and specialized industrial applications. This challenge has intensified interest in learning paradigms that can perform effectively in low-data environments. This review paper presents a systematic comparative study of supervised learning and self-supervised learning approaches with a specific focus on data-scarce settings. Supervised learning methods traditionally rely on annotated data and employ techniquessuchastransferlearning,dataaugmentation,and meta-learning to mitigate label scarcity; however, their performance often degrades significantly as labeled data decreases. In contrast, self-supervised learning leverages unlabeled data through pretext tasks and representation learning, enabling models to learn robust and transferable featurerepresentationspriortodownstreamfine-tuning.The paper critically examines recent literature across computer vision,naturallanguageprocessing,andtime-seriesdomains to analyze performance trends, data efficiency, and computational trade-offs between these paradigms. The review highlights that self-supervised learning generally exhibits superior generalization and label efficiency in lowdata regimes, while supervised methods remain competitive when high-quality labeled samples are available. Finally, key research gaps and future directions for hybrid and scalable learning frameworks are identified.
Key Words: Supervised learning, Self-supervised learning, Low-data environments, Representation learning, Data efficiency, Machine learning
Machine learning (ML) has become a foundational technology across diverse application domains, enabling automated decision-making, pattern recognition, and predictive analytics. However, the success of most ML modelshashistoricallydependedontheavailabilityoflargescale, high-quality labeled datasets. In many practical scenarios,suchextensivelabeleddataisunavailable,giving rise to the problem of learning in low-data environments. Thislimitationhasmotivatedgrowingresearchinterestin alternativelearningparadigms,particularlyself-supervised learning,whichaimstoreducerelianceonlabeleddatawhile maintainingrobustperformance(LeCunetal.,2015).
1.1.1.
Data-scarce domains such as healthcare, cyber security, remotesensing,andscientificdiscoveryincreasinglyrelyon machinelearningtoextractmeaningfulinsightsfromlimited observations. In medical imaging, for example, expert annotations are expensive and time-consuming, while privacyconstraintsfurtherrestrictdataavailability.Despite these challenges, accurate ML models are critical for diagnosis,monitoring,anddecisionsupport,makingdataefficientlearningapproachesessential(Estevaetal.,2019). Similar constraints exist in domains such as autonomous systems and industrial fault detection, where rare events limitlabeleddataavailability.
Low-data environments introduce several technical challenges for machine learning systems. Models trained withinsufficientlabeleddataarepronetooverfitting,poor generalization, and high variance in predictions. Additionally, deep learning architectures with millions of parameters require substantial supervision to converge effectively,whichexacerbatesperformancedegradationin data-limitedsettings(Zhangetal.,2017).Thesechallenges necessitatealternativestrategiesthatcanlearnmeaningful representationswithminimalornoexplicitlabeling.
1.2.1.
Supervised learning is a traditional machine learning paradigm in which models are trained using explicitly labeled input–output pairs. The objective is to learn a mappingfunctionthatminimizespredictionerroronunseen data,typicallyusinglossfunctionssuchascross-entropyor mean squared error (Bishop, 2006). While supervised learninghasachievedremarkablesuccessacrosstaskssuch asimageclassificationandnaturallanguageprocessing,its dependenceonlargelabeleddatasetsmakesitlesssuitable forlow-datascenarios.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
Self-supervised learning is an emerging paradigm that enablesmodelstolearnusefulfeaturerepresentationsfrom unlabeled data by formulating surrogate or pretext tasks. Thesetasksgeneratesupervisorysignalsdirectlyfromthe dataitself,suchaspredictingmaskedtokens,imagepatches, ortransformations(Chenetal.,2020).Bylearninggeneralpurpose representations prior to downstream task finetuning, self-supervised learning significantly reduces the reliance on labeled data and has demonstrated strong performanceinlow-dataregimes.
1.2.3.
Low-data environments refer to learning settings where labeleddataisscarce,incomplete,orcostlytoobtain.This includes few-shot learning scenarios, where only a small numberoflabeledexamplesperclassareavailable,aswell as settings with severely imbalanced or partially labeled datasets (Wang et al., 2020). In this review, low-data environments are considered across multiple domains, focusingonhowdifferentlearningparadigmsadapttosuch constraints.
1.3.
1.3.1. Comparative Analysis of Learning Paradigms
The primary objective of this review is to systematically comparesupervisedandself-supervisedlearningparadigms in the context of low-data environments. The comparison focuses on learning efficiency, representation quality, and generalizationcapabilityacrossdifferenttasksanddomains.
1.3.2.
Another key objective is to identify performance trends reported in existing literature, particularly how model accuracy and robustness scale as labeled data availability decreases.Understandingthesetrendsprovidesinsightsinto thepracticaltrade-offsbetweenthetwoparadigms(Heetal., 2022).
1.3.3.
This review also aims to survey commonly used datasets, benchmarks,andevaluationprotocolsemployedinlow-data learning research. Emphasis is placed on identifying inconsistenciesinexperimentalsetupsandhighlightingthe needforstandardizedevaluationframeworks.
Thisreviewadoptsasystematicandstructuredmethodology to ensure comprehensive coverage, transparency, and reproducibilityinsurveyingexistingliteratureonsupervised andself-supervisedlearninginlow-dataenvironments.The methodological framework follows established guidelines
for systematic reviews in computer science research, focusingonrigorousarticleselection,well-definedinclusion andexclusioncriteria,andaclearcategorizationstrategyto enable meaningful comparative analysis (Kitchenham and Charters,2007).
2.1.1.
Toensuretheinclusionofhigh-qualityandpeer-reviewed studies, relevant articles were retrieved from wellrecognized academic databases, including IEEE Xplore, Scopus,andWebofScience.Thesedatabaseswereselected due to their extensive coverage of leading journals and conferencesinmachinelearning,artificialintelligence,and data science. The use of multiple databases reduces publicationbiasandenhancesthebreadthoftheliterature considered.
2.1.2.
Asystematickeyword-basedsearchstrategywasemployed toidentifyrelevantstudies.Searchstringswereconstructed usingcombinationsoftermssuchas“supervisedlearning,” “self-supervisedlearning,”“representationlearning,”“lowdata,” “few-shot learning,” and “data-efficient machine learning.” Boolean operators (AND, OR) were applied to refine search results and capture studies addressing both learning paradigms within data-scarce settings. This approach enabled the identification of both foundational worksandrecentmethodologicaladvances.
2.2.
2.2.1.
Thereviewfocusesonstudiespublishedbetween2015and 2025, a period that captures the rapid evolution of deep learning and the emergence of modern self-supervised learning techniques. Earlier works were excluded unless theyprovidedessentialtheoreticalfoundationsrelevantto thecomparativeanalysis.
2.2.2.
Onlypeer-reviewedjournalarticlesandtop-tierconference papers were included to maintain academic rigor. Studies were excludediftheylackedempirical evaluation,did not explicitly address low-data settings, or focused solely on fully supervised learning with large labeled datasets. This filtering process ensured that the reviewed literature directly aligned with the objectives of the study and met acceptablequalitystandards(Petersenetal.,2015).

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
2.3.1. Supervised versus Self-Supervised Learning
Selectedstudieswerefirstcategorizedbasedontheprimary learningparadigmemployed eithersupervisedlearningor self-supervised learning. Supervised learning studies typicallyreliedonlabeleddatasetsandincludedtechniques suchastransferlearningandmeta-learning,whereasselfsupervised learning studies emphasized representation learning through pretext tasks and unlabeled data utilization.Thisdistinctionenabledastructuredcomparison ofmethodologicalassumptionsandlearningefficiency.
2.3.2.
Within each learning paradigm, studies were further classified according to the type of low-data environment addressed, such as few-shot learning, limited-label finetuning, or partially labeled datasets. This secondary categorizationfacilitatedtheanalysisofperformancetrends andgeneralizationbehavioracrossvaryingdegreesofdata scarcity,allowingforamorenuancedcomparisonofthetwo paradigms(Wangetal.,2020).
Thissectionoutlinesthetheoreticalprinciplesunderpinning supervisedandself-supervisedlearningparadigms,witha particular emphasis on their behavior in low-data environments.Understandingthesefoundationsisessential forinterpretingempiricalresultsreportedintheliterature and for explaining observed differences in data efficiency and generalization performance between the two approaches.
3.1.
3.1.1.
Supervised learning is grounded in the optimization of a predefinedobjectivefunctionthatquantifiesthediscrepancy between predicted outputs and ground-truth labels. Commonly used loss functions include cross-entropy for classificationtasksandmeansquarederrorforregression problems.Modelparametersareoptimizedthroughiterative algorithms such as stochastic gradient descent and its variants, which aim to minimize empirical risk over the labeledtrainingdataset(Bishop,2006).Theeffectivenessof thisoptimizationprocessiscloselytiedtothequantityand qualityoflabeleddataavailable.
3.1.2.
Adefiningcharacteristicofsupervisedlearningisitsstrong relianceonlabeleddatatoguidemodeltraining.Inlow-data environments,thisdependencyoftenleadstooverfitting,as modelsmaymemorizelimitedtrainingsamplesratherthan learn generalizable patterns. Although techniques such as
regularization,dataaugmentation,andtransferlearningcan partially mitigate this issue, supervised models typically exhibitreducedrobustnessanddegradedperformancewhen labeledsamplesarescarce(Zhangetal.,2017).
3.2.1.
Self-supervised learning addresses label scarcity by constructing auxiliary tasks, known as pretext tasks, that generate supervisory signals directly from the data itself. Examplesincludepredictingmaskedportionsofinputdata, temporalordering,ortransformationsappliedtosamples. Through these tasks, models learn meaningful latent representations that capture underlying data structure without requiring manual annotation. These learned representations can later be transferred to downstream supervised tasks with minimal labeled data (LeCun et al., 2015).
Amongself-supervisedapproaches,contrastivelearningand masked prediction have emerged as particularly effective. Contrastivelearningmethodsaimtomaximizeagreement betweendifferentaugmentedviewsofthesamedatawhile minimizing similarity to other samples, leading to discriminativefeaturerepresentations.Maskedprediction techniques, widely used in vision and natural language processing,train modelsto reconstruct missing orhidden input components, thereby encouraging contextual understanding.Suchapproacheshavedemonstratedstrong performanceacrossavarietyoflow-databenchmarks(Chen etal.,2020;Heetal.,2022).
Data efficiency and generalization are central theoretical considerations when comparing supervised and selfsupervised learning. From a representation learning perspective,self-supervisedpre-trainingcanbeviewedasa formofinductivebiasthatconstrainsthehypothesisspace, enablingbettergeneralizationwithfewerlabeledsamples. Theoretical and empirical studies suggest that models initializedwithrich,task-agnosticrepresentationsrequire fewerlabeledexamplestoachievecompetitiveperformance, thereby improving sample efficiency. In contrast, purely supervisedmodelsoftenrequiresignificantlylargerlabeled datasets to achieve similar generalization capabilities (Belkinetal.,2019).
Thissectionpresentsacriticalsynthesisofpriorresearchon supervised and self-supervised learning in low-data environments.Ratherthanprovidinganexhaustivelistingof studies,thefocusisonidentifyingdominantmethodological trends,comparativefindings,andunresolvedchallenges.The

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
reviewedliteraturespanscomputervision,naturallanguage processing, and time-series analysis, reflecting the crossdomainrelevanceofdata-efficientlearningparadigms.
Supervised learning has traditionally dominated machine learning research; however, its effectiveness in low-data environments has been a persistent concern. To address data scarcity, researchers have proposed several enhancementstostandardsupervisedtrainingpipelines.
4.1.1.
Data augmentation is one of the most widely adopted strategiesforimprovingsupervised learning performance underlimiteddataconditions.Byapplyingtransformations such as rotation, cropping, noise injection, or synonym replacement, augmentation artificially increases dataset diversityandreducesoverfitting.Whileeffectiveinvision andspeechtasks,augmentationstrategiesareoftendomainspecific and may introduce unrealistic samples if not carefullydesigned(ShortenandKhoshgoftaar,2019).
Transfer learning has emerged as a more robust solution, whereinmodelspre-trainedonlarge-scaledatasetsarefinetunedonsmallertargetdatasets.Thisapproachenablesthe reuse of learned feature representations, significantly improvingconvergencespeedandperformanceinlow-data settings.Despiteitssuccess,transferlearningperformance dependsheavilyonthesimilaritybetweensourceandtarget domains,limitingitsgeneralizabilityacrossheterogeneous tasks(PanandYang,2010).
4.1.2.
Meta-learning,oftenreferredtoas“learningtolearn,”aims to train models that can rapidly adapt to new tasks using onlyafewlabeledexamples.Model-agnosticmeta-learning (MAML)isaprominentexamplethatlearnsaninitialization optimizedforfastadaptationacrosstasks.Empiricalstudies haveshownthatmeta-learningcanoutperformconventional supervisedlearninginfew-shotscenarios,particularlywhen task distributions are well defined (Finn et al., 2017). However,meta-learningframeworksofteninvolvecomplex trainingproceduresandhighcomputationalcosts.
Across the literature, supervised learning approaches generallyexhibitsteepperformancedegradationaslabeled data decreases. Although augmentation, transfer learning, andmeta-learningprovidemeasurableimprovements,these techniquesdonotfullyovercomethefundamentalreliance onlabeledsupervision.Moreover,supervisedmethodstend tooverfitinextremelylow-dataregimesandstrugglewith robustnessunderdistributionshifts,highlightinginherent limitationsoflabel-dependentlearningparadigms.
Self-supervisedlearninghasgainedsignificantattentionasa scalable alternative to supervised learning, particularly in scenarioswhereunlabeleddataisabundantbutannotations are scarce. Recent advances demonstrate that selfsupervised pre-training can produce high-quality representationstransferabletodiversedownstreamtasks.
4.2.1.
Contrastive learning methods such as SimCLR and Momentum Contrast (MoCo) learn representations by bringing semantically similar samples closer in the embedding space while pushing dissimilar samples apart. Thesemethodshaveachievedperformancecomparableto fully supervised pre-training on image classification benchmarks, even with limited labeled fine-tuning data (Chen et al., 2020; He et al., 2020). However, contrastive methodsoftenrequirelargebatchsizesormemorybanks, increasingcomputationaldemands.

4.2.2.
Generative self-supervised approaches rely on reconstruction-basedobjectives,suchasautoencodingand maskedprediction.Innaturallanguageprocessing,models like BERT employ masked language modeling to learn contextual representations, while masked auto encoders (MAE) have demonstrated strong performance in vision tasks.Theseapproachesemphasizesemanticunderstanding rather than instance discrimination, offering improved robustnessinlow-datafine-tuningscenarios(Devlinetal., 2019;Heetal.,2022).
4.2.3.
Beyondcomputervision,self-supervisedlearninghasbeen successfullyappliedtotext,speech,andtime-seriesdata.In NLP,large-scaleself-supervisedpertaininghasbecomethe defactostandard.Intime-seriesanalysis,techniquesbased on temporal prediction and contrastive objectives have shown promise for tasks such as anomaly detection and forecasting,indicatingthemodality-agnosticnatureofselfsupervisedrepresentations.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

4.2.4. Role in Low-Data Settings
The literature consistently shows that self-supervised pretraining substantially improves downstream task performancewhenlabeleddataisscarce.Modelsinitialized withself-supervisedrepresentationstypicallyrequirefewer labeled samples to reach competitive accuracy, demonstrating superior sample efficiency compared to purelysupervisedcounterparts(LeCunetal.,2015).
4.3. Comparative Analyses in Prior Work
4.3.1. Direct Comparisons Between Paradigms
Several studies have conducted direct empirical comparisons between supervised and self-supervised learningundercontrolledlow-dataconditions.Thesestudies generallyreportthatself-supervisedpretrainingfollowedby limited supervised fine-tuning outperforms end-to-end supervised training, particularly in extreme few-shot regimes(Chenetal.,2020).
4.3.2. Representation
Comparative analyses reveal that self-supervised models learnmoregeneralizableandtransferablerepresentations, capturing semantic structure rather than task-specific decisionboundaries.Thisdistinctionexplainstheirsuperior generalizationwhenlabeled data islimitedandhighlights representation learning as a key factor in data efficiency (Belkinetal.,2019).
4.3.3.
Despitepromisingresults,manycomparativestudiessuffer from inconsistencies in experimental setups, including varying datasets, architectures, and fine-tuning protocols. Theseinconsistencieslimitthegeneralizabilityofreported findingsandcomplicatefaircomparisonsbetweenlearning paradigms.
Low-datalearningresearchcommonlyreliesonbenchmark datasets such as CIFAR-10 and CIFAR-100 subsets, MiniImageNet,andreduced-labelversionsofGLUEforNLPtasks. Evaluation protocols typically involve training with progressivelyfewerlabeledsamplesperclassandreporting averagedperformanceacrossmultipleruns.However,the lack of standardized low-data benchmarks and evaluation metrics remains a significant challenge for comparative analysis(Wangetal.,2020).
Thereviewedliteraturerevealsseveralopenresearchgaps. First, most studies focus on vision and language tasks, leavingdomainssuchashealthcaretime-series,graphdata, and multimodal learning under-explored. Second, there is limited theoretical analysis explaining why certain selfsupervisedobjectivesyieldbetterdataefficiency.Finally,the absence of standardized benchmarks and reproducible evaluationprotocolshindersfaircomparisonandpractical adoption.Addressingthesegapsisessentialforadvancing researchondata-efficientlearningparadigms.
Thissectionprovidesastructuredcomparisonofsupervised and self-supervised learning paradigms in low-data environments. The analysis focuses on empirical performance,dataefficiency,computationalrequirements, anddomain-specificapplicability,synthesizingfindingsfrom prior studies to highlight practical trade-offs and methodologicalstrengths.
Performanceevaluationinlow-datalearningtypicallyrelies on metrics such as accuracy and F1-score, which capture predictivecorrectnessandclass-wisebalance,respectively. Whilesupervisedmodelsmayachievehighaccuracywhen sufficientlabeleddataisavailable,theirperformanceoften deterioratessharplyinlabel-scarcesettings.Incontrast,selfsupervised models fine-tuned with limited labels tend to maintainmorestableF1-scores,particularlyinimbalanced datasets.Robustness,measuredthroughsensitivitytonoise or domain shifts, is increasingly recognized as a critical metric, with self-supervised representations generally exhibiting greater resilience to perturbations due to their relianceonbroaderdatastructureratherthanlabel-specific cues(Belkinetal.,2019).

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
5.2.1.
Dataefficiencyreferstoamodel’sabilitytoachievestrong performance withminimal labeleddata.Empirical studies consistently demonstrate that supervised learning performancescalespoorlyasthenumberoflabeledsamples decreases,oftenrequiringexponentialincreasesindatato achievemarginalgains.Self-supervisedlearning,bycontrast, exhibits more graceful degradation, with pertained representationsenablingcompetitiveperformanceevenin few-shotscenarios.Thisdifferenceisparticularlyevidentin imageandlanguagetasks,whereself-supervisedpertaining significantlyreducesthenumberoflabelsrequiredtoreach performanceparitywithfullysupervisedmodels(Chenetal., 2020).
5.3.1.
Computational cost is a critical factor when evaluating learningparadigmsforreal-worlddeployment.Supervised learning pipelines are typically less computationally intensive during training but incur high annotation costs. Self-supervisedlearningshiftsthiscosttowardcomputation, often requiring large-scale pertaining on unlabeled data usingsubstantialhardwareresources.Contrastivemethods, inparticular,demandextensivememoryandlongtraining times. However, once pretrained, self-supervised models often require minimal fine-tuning, resulting in reduced overalltrainingtimefordownstreamlow-datatasks(Heet al.,2020).
5.4.1. Vision, Natural Language Processing, Healthcare, and Robotics
At the application level, the relative advantages of supervisedandself-supervisedlearningvarybydomain.In computer vision and natural language processing, selfsupervised pretraining has become a standard practice, consistentlyoutperformingsupervisedapproachesinlowdataregimes.Inhealthcare,wherelabeleddataisscarceand expensive,self-supervisedmethodshaveshownpromisein medical imaging and physiological signal analysis. In robotics,however,supervisedlearningremainsrelevantfor task-specific control, although self-supervised representation learning is increasingly used to improve generalization and adaptability. These domain-specific trends underscore the contextual nature of paradigm selection(Estevaetal.,2019).
This review paper has presented a comprehensive comparative analysis of supervised and self-supervised learningparadigmsinlow-dataenvironments,synthesizing theoretical foundations, methodological trends, and empiricalfindingsacrossmultipleapplicationdomains.The analysis indicates that while supervised learning remains effective when high-quality labeled data is available, its performance degrades significantly under label-scarce conditions due to overfitting and limited generalization. Techniques such as data augmentation, transfer learning, andmeta-learningpartiallyalleviatethesechallengesbutdo not eliminate the fundamental dependency on labeled supervision.
In contrast, self-supervised learning has emerged as a powerfulalternativebyleveragingunlabeleddatatolearn transferable and robust representations. Advances in contrastive learning and masked modeling have demonstrated superior data efficiency, robustness, and scalability, particularly in vision and natural language processing tasks. Comparative studies consistently show that self-supervised pretraining followed by limited supervised fine-tuning outperforms purely supervised approachesinfew-shotandlow-labelregimes.However,the review also highlights trade-offs related to computational costandpretrainingcomplexity.
Overall, the findings suggest that self-supervised learning representsapromisingdirectionfordata-efficientmachine learning,especiallyindomainswherelabeleddataisscarce orexpensive.Futureprogressislikelytobedrivenbyhybrid frameworks that integrate supervised and self-supervised objectives, supported by standardized benchmarks and strongertheoreticalguarantees(LeCunetal.,2015;Chenet al.,2020).
Despite its comprehensive scope, this review has several limitations. First, the analysis is primarily based on published peer-reviewed literature, which may introduce publication bias by underrepresenting negative or inconclusiveresults.Second,althoughmultipleapplication domains are discussed, the majority of reviewed studies focusoncomputervisionandnaturallanguageprocessing, limitingthegeneralizabilityofconclusionstodomainssuch asgraphlearning,reinforcementlearning,andmultimodal systems.Third,directcomparisonsbetweensupervisedand self-supervisedmethodsareoftenaffectedbyinconsistent experimentalsetups,architectures,andevaluationprotocols, which complicates fair assessment. Finally, this review emphasizes empirical findings, while theoretical analyses explaining the fundamental advantages of self-supervised learninginlow-dataregimesremainrelativelylimitedand warrantdeeperinvestigation(Belkinetal.,2019;Wangetal., 2020).

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
1. Belkin, M., Hsu, D., Ma, S. and Mandal, S. (2019) ‘Reconcilingmodernmachine-learningpracticeandthe classical bias–variance trade-off’, Proceedings of the National Academy of Sciences, 116(32), pp. 15849–15854.
https://doi.org/10.1073/pnas.1903070116
2. Bishop,C.M.(2006)PatternRecognitionandMachine Learning.NewYork:Springer
3. Chen,T.,Kornblith,S.,Norouzi,M.andHinton,G.(2020) ‘Asimpleframeworkforcontrastivelearningofvisual representations’,Proceedingsofthe37thInternational Conference on Machine Learning (ICML), pp. 1597–1607.
4. Devlin, J., Chang, M.-W., Lee, K. and Toutanova, K. (2019) ‘BERT: Pre-training of deep bidirectional transformersforlanguageunderstanding’,Proceedings ofthe2019ConferenceoftheNorthAmericanChapter of the Association for Computational Linguistics (NAACL-HLT),pp.4171–4186.
5. Esteva,A.,Robicquet,A.,Ramsundar,B.,Kuleshov,V., DePristo,M.,Chou,K.,Cui,C.,Corrado,G.,Thrun,S.and Dean,J.(2019)‘Aguidetodeeplearninginhealthcare’, Nature Medicine, 25(1), pp. 24–29. https://doi.org/10.1038/s41591-018-0316-z
6. Finn, C., Abbeel, P. and Levine, S. (2017) ‘Modelagnostic meta-learning for fast adaptation of deep networks’, Proceedings of the 34th International Conference on Machine Learning (ICML), pp. 1126–1135.
7. He, K., Fan, H., Wu, Y., Xie, S. and Girshick, R. (2020) ‘Momentum contrast for unsupervised visual representationlearning’,ProceedingsoftheIEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR),pp.9729–9738.
8. He,K.,Chen,X.,Xie,S.,Li,Y.,Dollár,P.andGirshick,R. (2022) ‘Masked autoencoders are scalable vision learners’,ProceedingsoftheIEEE/CVFConferenceon ComputerVisionandPatternRecognition(CVPR),pp. 16000–16009.
9. Kitchenham,B.andCharters,S.(2007)Guidelinesfor performing systematic literaturereviews in software engineering.EBSETechnicalReport,KeeleUniversity andDurhamUniversity.
10. LeCun, Y., Bengio, Y. and Hinton, G. (2015) ‘Deep learning’, Nature, 521, pp. 436–444. https://doi.org/10.1038/nature14539
11. Pan, S.J. and Yang, Q. (2010) ‘A survey on transfer learning’, IEEE Transactions on Knowledge and Data Engineering, 22(10), pp. 1345–1359. https://doi.org/10.1109/TKDE.2009.191
12. Petersen, K., Vakkalanka, S. and Kuzniarz, L. (2015) ‘Guidelinesforconductingsystematicmappingstudies insoftwareengineering:Anupdate’,Informationand Software Technology, 64, pp. 1–18. https://doi.org/10.1016/j.infsof.2015.03.007
13. Shorten,C.andKhoshgoftaar,T.M.(2019)‘Asurveyon imagedata augmentation’,Journal ofBigData,6(60). https://doi.org/10.1186/s40537-019-0197-0
14. Wang, Y., Yao, Q., Kwok, J.T. and Ni, L.M. (2020) ‘Generalizing from a few examples: A survey on fewshotlearning’,ACMComputingSurveys,53(3),Article 63. https://doi.org/10.1145/3386252
15. Zhang,C.,Bengio,S.,Hardt,M.,Recht,B.andVinyals,O. (2017) ‘Understanding deep learning requires rethinking generalization’, Proceedings of the InternationalConferenceonLearningRepresentations (ICLR).
16. Arora, S., Du, S.S., Hu, W., Li, Z. and Wang, R. (2019) ‘Fine-grained analysis of optimization and generalizationforoverparameterizedtwo-layerneural networks’, Proceedings of the 36th International ConferenceonMachineLearning(ICML),pp.322–332.
17. Balestriero,R.,Cosentino,R.,Glotin,H.andBaraniuk,R. (2023) ‘A theory of self-supervised learning: From alignment and uniformity to spectral methods’, IEEE Signal Processing Magazine, 40(2), pp. 33–44. https://doi.org/10.1109/MSP.2022.3214709
18. Dosovitskiy,A.etal.(2021)‘Animageisworth16×16 words: Transformers for image recognition at scale’, InternationalConferenceonLearningRepresentations (ICLR).
19. Grill, J.-B. et al. (2020) ‘Bootstrap your own latent: A newapproachtoself-supervisedlearning’,Advancesin NeuralInformationProcessingSystems(NeurIPS),33, pp.21271–21284.
20. Hendrycks, D. and Gimpel, K. (2017) ‘A baseline for detecting misclassified and out-of-distribution examplesinneuralnetworks’,InternationalConference onLearningRepresentations(ICLR).
21. Hinton,G.,Vinyals,O.andDean,J.(2015)‘Distillingthe knowledge in a neural network’, arXiv preprint arXiv:1503.02531.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072
22. Jing, L. and Tian, Y. (2020) ‘Self-supervised visual featurelearningwithdeepneuralnetworks:Asurvey’, IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(11), pp. 4037–4058. https://doi.org/10.1109/TPAMI.2020.2992393
23. Kolesnikov,A.etal.(2020)‘Bigtransfer(BiT):General visualrepresentationlearning’,EuropeanConference onComputerVision(ECCV),pp.491–507.
24. Noroozi, M. and Favaro, P. (2016) ‘Unsupervised learning of visual representations by solving jigsaw puzzles’, European Conference on Computer Vision (ECCV),pp.69–84.
25. Raghu,A.etal.(2021)‘Visiontransformersoutperform ResNets in low-data regimes’, Advances in Neural Information Processing Systems (NeurIPS), 34, pp. 2798–2810.
26. Schroff, F., Kalenichenko, D. and Philbin, J. (2015) ‘FaceNet:Aunifiedembeddingforfacerecognitionand clustering’, Proceedings of the IEEE Conference on ComputerVisionandPatternRecognition(CVPR),pp. 815–823.
27. Tian,Y.,Krishnan,D.andIsola,P.(2020)‘Contrastive multiviewcoding’,EuropeanConferenceonComputer Vision(ECCV),pp.776–794.
28. Xie, Q. et al. (2020) ‘Self-training with noisy student improvesImageNetclassification’,Proceedingsofthe IEEE/CVFConferenceonComputerVisionandPattern Recognition(CVPR),pp.10687–10698.
29. Yang,L.etal.(2022)‘Self-supervisedlearningfortimeseries analysis: Taxonomy, progress, and prospects’, IEEETransactionsonKnowledgeandDataEngineering, 34(11), pp. 5426–5448. https://doi.org/10.1109/TKDE.2021.3092587
30. Zoph,B.etal.(2020)‘Rethinkingpre-trainingandselftraining’, Advances in Neural Information Processing Systems(NeurIPS),33,pp.3833–3845.