
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Ravindra. B. Murumkar¹, Pradyumna Wagholikar², Kaustubh Pawar³, Ved Pingle⁴, Pallavi Ranamale5
1 Professor, Dept. of IT, Pune Institute of Computer Technology, Pune, Maharashtra, India. 2 , 3 , 4 ,5 Student, Dept. of IT, Pune Institute of Computer Technology, Pune, Maharashtra, India.
Abstract - IndianSignLanguage(ISL)istheprimarymeans of communication for millions of individuals within India's Deaf and Hard of Hearing (DHH) community. However, a significant communication gap persists due to the scarcity of certified human interpreters, motivating the development of automated ISL recognition and translation systems. This survey presents a comprehensive review of the technical evolutionofISLprocessing,tracingtheprogressionfromearly computer vision methods based on handcrafted features to modern deep learning paradigms. Convolutional Neural Networks (CNNs) have advanced spatial feature extraction, while Recurrent Neural Networks (LSTMs, GRUs) enhanced temporal modeling of sequential gestures. More recently, Transformer architectures have enabled end-to-end translation with improved contextual understanding. A prominenttwo-stageparadigmhasemerged combiningpose estimation frameworks such as MediaPipe for skeletal keypoint extraction with deep learning models for dynamic sequence interpretation. The survey also emphasizes the growing importance of datasets, highlighting the transition from small, custom-collected corpora to large-scale public benchmarks like CISLR, ISLTranslate, and iSign. Despite notable progress, challenges remain, including limited data standardization, occlusions in two-handed signs, and insufficient modeling of non-manual markers. The paper concludesbyoutliningfutureresearchdirections,emphasizing large-scale unified datasets, multimodal integration, and advanced Transformer-based architectures for robust, realtime ISL understanding.
Key Words: Indian Sign Language, Deep Learning, Computer Vision, Sign Language Recognition, Sign LanguageTranslation,PoseEstimation,MediaPipe,CNN, LSTM, GRU, Transformers.
Language is the cornerstone of human connection and societal participation. In India alone, approximately 63 millionpeoplerelyonIndianSignLanguage(ISL)astheir primary meansofcommunication[1], [2], [5], [8]. ISLisa completeandcomplexvisual-gesturallanguage farmore thanacollectionofgestures withitsowngrammar,syntax, andlinguisticstructuredistinctfromthespokenlanguages oftheregion[3].AccordingtoEthnologue,ISLranksamong
themostwidelyusedsignlanguagesglobally,underscoring itssocialandculturalimportance[4].
Despiteitswidespreaduse,aprofoundcommunicationgap persists between the Deaf and Hard-of-Hearing (DHH) communityandthehearingpopulation.Thisgapislargely due to a critical shortage of qualified human interpreters. Official estimates suggest that there are only around 300 certified ISL interpreters in India, a number vastly insufficienttoserveacommunityofmillions[2],[4],[5].This deficithasfar-reachingconsequences,creatingbarriersto essential services such as healthcare, education, and legal support [1], [5]. In these critical contexts, the absence of effective communication can result in misdiagnoses, educationaldisadvantages,andinequitablelegaloutcomes.
The urgent need to bridge this gap provides strong motivation for developing automated, technology-driven solutions for ISL recognition and translation. Advances in computer vision and deep learning have opened new possibilities for such systems, which hold the potential to democratizeaccesstoinformationandfosterinclusionfor theDHHcommunity[1],[3].
However, building robust automated systems for ISL is a formidable task due to several intrinsic linguistic and computationalchallenges.UnlikeAmericanSignLanguage (ASL), which is predominantly one-handed, ISL is fundamentally a two-handed language, with many signs requiring coordinated hand movements [3], [11]. This introducesthepersistentproblemofocclusion,whereone handcanobscuretheother,leadingtoalossofcriticalvisual informationforrecognitionsystems[3],[10].
Furthermore,ISLcommunicationisinherentlymulti-modal. Non-ManualMarkers(NMMs) suchasfacialexpressions, head movements, and body posture serve as essential grammaticalelementsthatmodifyordefinethemeaningof signs [5], [7], [10]. The absence of a single standardized versionofISLhasalsoledtosignificantregionalvariations, creating a major generalization challenge for data-driven models[3].
To effectively address these challenges, it is essential to clearlydefinethecorecomputationaltasksinautomatedISL processing. The two principal tasks are Sign Language Recognition(SLR)andSignLanguageTranslation(SLT).SLR

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
focuses on identifying and classifying individual signs or sequencesofsigns,typicallyproducingasequenceofglosses (written labels representing signs) [3], [11]. SLT, on the otherhand,seekstotranslatecontinuoussequencesofsigns into grammatically correct sentences in a target spoken language such as English. This represents a true machine translationproblem,asthesyntacticstructuresofISLand spokenlanguagesdifferfundamentally[3],[11].
The evolution of ISL recognition research reflects a significantmethodologicalprogression.Earlysystemsrelied on traditional computer vision and image processing techniques. For instance, a 2011 study by Rajam and Balakrishnanproposedarule-basedsystemforrecognizing 32signsusingbinaryfingerpositionanalysiscombinedwith Canny Edge Detection and handcrafted feature extraction [6].Whileinnovativefortheirtime,suchapproacheswere brittle,computationallyintensive,andhighlydependenton controlledconditions
Theadventofdeeplearningintroducedaparadigmshift moving from handcrafted features to hierarchical feature learning directly from raw data. Convolutional Neural Networks (CNNs) demonstrated strong capabilities for spatialfeatureextraction,whilelatermethodsincorporated poseestimationwithRecurrentNeuralNetworks(RNNs)to modeltemporaldependencies[1].Thistransitionhasgreatly enhanced the robustness and scalability of modern ISL recognitionsystems.
Thispaperpresentsacomprehensivesurveyoftheprogress in automated ISL processing. We review the critical evolution of datasets that have driven the field’s growth, tracethemethodologicaltrajectoryfromtraditionalvisionbased techniques to state-of-the-art deep learning architectures, and analyze the transformative impact of enablingtechnologiessuchasposeestimation.Furthermore, weidentifykeychallenges,highlightperformancegaps,and outline future research directions aimed at developing linguisticallygrounded,inclusive,andreal-world-deployable ISLrecognitionandtranslationsystems.
The advancement of automated ISL processing is inextricably linked to the availability and quality of data resources. For years, the field was hampered by a “data desert,” which limited research to small-scale, often incomparablestudies.Therecentemergenceoflarge,public benchmarkshasbeenthesinglemostimportantcatalystfor progress, enabling the application of data-hungry deep learningmodels.
Early research was characterized by the use of small, bespokedatasetscreatedbyindividualresearchgroups[3], [6].Thesedatasetstypicallyfeaturedalimitedvocabulary,a small number of signers, and were recorded in highly controlled laboratory environments to simplify the
computervision task [3], [6].Whilea necessaryfirst step, this practice led to models that lacked real-world generalizability and made it difficult to compare different approachesmeaningfully[4].
Aturningpointforthefieldwasthereleaseofthefirstlargescale,publiclyavailabledatasets.TheINCLUDEdataset,for instance,providedover4,000videosfor263commonword signs, offering a more substantial resource for word-level recognition[1],[3].ThiswasfollowedbytheCISLR(Corpus for Indian Sign Language Recognition) dataset, which represented a major leap forward with a vocabulary of approximately 4,700 words across 7,050 videos [4]. Its uniquestructure,withalowaverageof1.5videosperword, spurredresearchintoone-shotlearningparadigms,where modelsmustlearntorecognizeasignfromasingleexample [4].
The focus of the research community shifted from recognition to translation with the publication of ISLTranslate.Asthefirstlarge-scaletranslationdatasetfor ISL,itcontainedapproximately31,000parallelISL-English sentenceandphrasepairs,enablingthetrainingofmodern end-to-end translation architectures for the first time [2], [10].
The current state of the art in ISL data resources is representedbytheiSignbenchmark,releasedin2024[10]. By consolidating previous major datasets like CISLR and ISLTranslate and augmenting them with new data from authentic sources such as ISLRTC and ISH News, iSign providesamassiveresourceofover118,000video-sentence pairs[10]
Crucially, iSign is a multi-task framework that proposes a standardized set of challenges, including SignVideo2Text Translation, Text2Pose Generation, and Sign Semantic Similarity,therebyguidingtheresearchcommunitytoward more holistic and capable systems [10]. The “benchmark effect”ofthesedatasetshasbeenprofound,activelyshaping theresearchagendafromisolatedsignclassificationtoward the more ambitious goal of continuous, multi-modal sign languageunderstandingandgeneration.
TheevolutionofmethodologiesforISLprocessingmirrors broader trends in computer vision and natural language processing, marked by a progression from handcrafted features to sophisticated, end-to-end deep learning architectures.
InitialforaysintoSLRwererootedintraditionalcomputer visionparadigms.Thesesystemstypicallyinvolvedatwostage process: first, extracting handcrafted visual features fromimagesusingdescriptorslikeScale-InvariantFeature

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Transform(SIFT),HistogramofOrientedGradients(HOG), orcontour analysis[3],[6], [7];and second,feedingthese feature vectors into classical machine learning classifiers such as Support Vector Machines (SVM), K-Nearest Neighbors (KNN), or Hidden Markov Models (HMM) for recognition[3],[5],[7].
While foundational, these methods were often brittle and highly sensitive to variations in lighting, background, and signerappearance[4],[5].
B. The Rise of Deep Learning: CNNs, LSTMs, and Transformers.
Theadventofdeeplearning triggereda paradigmshiftby enabling the automatic learning of feature hierarchies directlyfromdata.
•ConvolutionalNeuralNetworks(CNNs):CNNsbecamethe standardforspatialfeatureextraction,firstappliedtostatic sign recognition by treating it as an image classification problem[3],[7].ArchitectureslikeInceptionV3andcustom CNNs achieved high accuracy on constrained tasks like alphabetrecognition[3].Fordynamicsigns,3DCNNswere introducedtolearnspatio-temporalfeaturesdirectlyfrom videovolumes,thoughatahighcomputationalcost[7].
•RecurrentArchitectures:Toexplicitlymodelthetemporal dynamics of sign language, Recurrent Neural Networks (RNNs),particularlyLongShort-TermMemory(LSTM)and GatedRecurrentUnit(GRU)networks,wereadopted[1],[5], [8],[9].Thisledtoadominanthybridarchitecturewherea CNNextractsspatialfeaturesfromeachvideoframe,andan LSTMorGRUprocessestheresultingsequenceoffeaturesto model temporal dependencies [7], [9]. Attention mechanismswereoftenaddedtoallowthemodeltofocuson themostsalientframesinasequence[7].
•Transformers:Thecurrentstateoftheartinsequence-tosequencetasksistheTransformerarchitecture,whoseselfattention mechanism allows it to model long-range dependenciesmoreeffectivelythanRNNs[10],[11].Akey breakthroughwasthedevelopmentofanarchitecturethat jointlylearnsrecognitionandtranslationinasingleend-toendmodel[11].Inthisframework,aConnectionistTemporal Classification(CTC)lossisappliedtotheencoder’soutputto supervisetherecognitionofsignglosses.
This intermediate supervision forces the model to learn a meaningfulsignrepresentation,whichisthenusedbythe decoderforthefinaltranslationtask,significantlyimproving performanceoverdirectvideo-to-textmodels[11].
C. The Role of Pose Estimation Frameworks
Acriticalenablingtechnologythathasreshapedthefieldis real-time pose estimation. Frameworks like Google’s MediaPipeHolisticcanprocessavideoframeandoutputa structured, skeletal representation comprising 3D
coordinatesforkeypointsonthehands(21perhand),body (33),andface(468)[5],[8],[9].Thistransformstheinput fromahigh-dimensionalpixelspacetoalow-dimensional, semanticallyrichfeaturespace.
This approach offers several profound advantages. It is computationallyefficient,makingreal-timeapplicationson consumer hardware feasible [1], [5]. The skeletal data is inherentlymorerobusttovariationsinlighting,background, andclothing[1].Italsohelpsmitigateprivacyconcernsby anonymizingthesigner’sappearance[9].
This pipeline effectively decouples the computer vision problem from the sequence modeling task, allowing researchers to leverage a powerful, off-the-shelf feature extractor and focus on designing sophisticated sequential models [5], [8], [9]. Consequently, many state-of-the-art systems now follow this pipeline: video is fed into MediaPipe,andtheresultingsequenceoflandmarkvectors is processed by an LSTM, GRU, or Transformer network, oftenachievingrecognitionaccuraciesexceeding95%[1], [5],[8].
This section presents a comparative evaluation of major approaches developed for Indian Sign Language (ISL) recognition and translation. The analysis highlights performancetrends,methodologicaladvancements,andthe datasetsthathavedrivenprogressinthisdomain.
TheprogressionofresearchinISLrecognitionismarkedby increasingmethodologicalsophisticationandperformance, largelydrivenbytheavailabilityofbetterdatasetsandmore powerfulcomputationalmodels.Acomparativeanalysisof keystudies,assummarizedinTable??,revealscleartrends inthefield.
Early works, such as that by Rajam et al. [6], relied on traditional image processing techniques like feature point extraction on small, custom datasets of static signs. While achievinghigh accuracy(98.1%) withintheirconstrained, single-userenvironment,thesemethodslackedtheabilityto generalizetonewsignersordynamicgestures.
Theadoptionofdeeplearning,particularlyCNNs,markeda significant step forward. Studies like Sharma et al. [3] focused on CNN-based recognition of static alphabets, demonstrating the power of learned features but also highlightingthelimitationsofexistingdatasetsintermsof diversityandthechallengeoftwo-handedsigns.
A major breakthrough came with the integration of pose estimation frameworks and recurrent neural networks to handle dynamic signs. The work of Shetty et al. [1] and Rawat et al. [5], both leveraging MediaPipe for feature extraction and LSTMs for sequence modeling, achieved

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
impressiveaccuraciesof98%and96.97%,respectively,on isolated dynamic signs. These studies underscore the effectiveness of the pose-estimation pipeline for creating robust,real-timesystemsthatareinvarianttobackground andlightingconditions.
Subramanian et al. further refined this approach by proposing an optimized GRU (MOPGRU) model, which demonstrated faster convergence and higher learning efficiency[9].
The focus has recently expanded from recognition to translation, a shift enabled by new datasets and architectures.TheworkbyCamgozetal.ontheTransformer
architecture, while not on ISL, introduced a seminal joint recognition-translationframeworkthathasbecomehighly influential[11].
ThecreationoftheISLTranslate[2]andiSign[10]datasets has now provided the necessary resources to apply such advancedtranslationmodelstoISL,settingthestageforthe nextgenerationofresearchinthefield.
Paper Core Methodology
Rajametal. (2011)[6]
Feature-point extractionand image processing
Sharmaetal. (2025)[3]
Dataset (reported)
Custom(32 staticsigns)
Shettyetal. (2024)[1]
Convolutional NeuralNetwork (CNN)
Custom (2,600+static images)
Task / Protocol
Isolatedstatic-sign recognition;singleuser
98.1% Accuracy (isolated, controlled)
Letter/number/staticsignrecognition
MediaPipepose estimationwith LSTM
INCLUDE(as perpaper)
Isolatedword/sign recognition;real-time setting
Highaccuracy onlettersand numbers
Highaccuracy under controlled conditions; earlyISLwork
Addressestwohandedsigns; optimizedfor affordable hardware
98%Accuracy (isolatedsigns)
Singlesigner, static-only signs;limited generalization tocontinuous signing
Primarilystatic signs;limited cross-paper comparability duetocustom dataset
Real-time performance; robustto clothingand background variations
Evaluatedon isolatedsigns; not demonstrated forcontinuous translation
Subramanian etal.(2022) [9]
MediaPipewith optimizedGRU (MOPGRU)
Custom+ WLASL,LSA64
Small-gesture recognition;signerdependentandsignerindependentsplits
95%Accuracy (taskdependent)
Faster convergence andimproved learning efficiency
Mostlysmall datasets; limited vocabularyfor large-scale generalization
Rawatetal. (2025)[5]
MediaPipe Holisticwith sequential LSTM
Joshietal. (2022)[4]
One-shot learningusing ASL-pretrained I3Dfeatures
Camgozetal. SignLanguage
Custom(11 dynamicand staticgestures)
Isolateddynamic/static gesturerecognition; signer-independent evaluation
96.97% Accuracy (isolated gestures)
CISLR(4,700 words,7,050 videos)
Large-vocabulary word-levelrecognition; one-shotevaluation
Robustacross lighting conditions; signerindependent evaluation
16.8%Top-1 Accuracy Introduced largevocabularyISL dataset; exploredcrosslingualtransfer
Strongend-to-
Verysmall dataset;not evaluatedon sentence-level translation
Lowabsolute accuracy; highlights challengeof one-shot learninginISL
Evaluatedon

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net
(2020)[11] Transformer (joint recognitionand translation usingCTC+ seq2seq)
Joshietal. (2024)[10]
Multi-task benchmark creation(iSign): translation, generation, similarity
(GermanSL) languagetranslation (video-to-text) improvement overprevious baselines end architecture; introduced intermediate CTC supervision
iSign(118k pairs)
Multi-taskbenchmark: SignVideo2Text, Text2Pose,semanticsimilarity
In summary, the comparative analysis highlights a clear paradigmshiftfromhandcraftedandpixel-basedrecognition techniquestowardpose-driven,deeplearningarchitectures, culminating in end-to-end Transformer-based translation frameworks empowered by large-scale datasets like ISLTranslateandiSign.
Despite significant progress, the field of automated ISL processingfacesseveralpersistentchallengesthatdefinethe frontierofcurrentresearch.Theselimitationsarefrequently citedacrosstheliteratureandmustbeaddressedtodevelop trulypracticalsystems.
A primary challenge remains the scarcity and diversity of data.WhilebenchmarkslikeiSignrepresentamonumental stepforward,thevolumeofdataisstillordersofmagnitude smaller than that available for spoken languages [10]. Furthermore,existingdatasetsoftenlacksufficientdiversity in terms of signers, regional dialects, and recording environments, which can lead to models that do not generalizewelltoreal-world“in-the-wild”conditions[3].
The inherent linguistic complexity of ISL poses another majorhurdle.
•Two-HandedSignsandOcclusion:ThefrequentuseoftwohandedgesturesinISLleadstopersistentocclusion,where onehandblockstheotherfromview.Thisremainsadifficult computervisionproblemthatcancausethelossofcrucial handshapeinformation[3],[10],[11].
• Non-Manual Markers (NMMs): The grammatical informationconveyedthroughfacialexpressions,headtilts, and body posture is critical for accurate interpretation. While pose estimation frameworks can extract facial landmarks,currentmodelsstruggletoeffectivelyintegrate theseNMMsinalinguisticallymeaningfulway[5],[7],[10].
• Fingerspelling and Role Shifts: Signers often use fingerspelling for names or technical terms, requiring modelstoswitchfromword-leveltocharacter-level
p-ISSN: 2395-0072
GermanSL,not ISL;crosslingual generalization uncertain
BaselineBLEU4scoreof1.47 Largest consolidated ISLresource; establishes standardtasks andevaluation splits
Baseline performance low,indicating scopefor improved modeling
recognition. Additionally, role shifts, where a signer embodiesacharacterinanarrative,arecomplexphenomena thatarenotwell-handledbycurrentarchitectures[10].
There is also a significant “lab-to-life” performance gap. Modelsthatachievehighaccuracyoncleanbenchmarkdata often experience a sharp performance degradation when deployed in real-world scenarios with cluttered backgrounds,variablelighting,andnovelsigners[1].
Finally, the community faces a challenge in evaluation. Standard NLP metrics like BLEU, designed for linear text, maybeinadequateforassessingthequalityoftranslationfor avisual-spatiallanguagelikeISL,astheyfailtocapturethe preservation of spatial grammar or other visual linguistic features[10].
Addressing the prevailing challenges in ISL processing requires a multi-pronged research effort focused on data, models,andevaluation.Basedonthelimitationsidentifiedin theexistingliterature,severalkeydirectionsforfuturework emerge.
First, there is a continued need for dataset expansion and enrichment.Futuredatacollectioneffortsshouldprioritize capturingawiderdiversityofsignersfromdifferentregions tobetterrepresentISL’sdialectalvariations.Datasetsshould also include more continuous, conversational signing recorded in naturalistic “in-the-wild” environments. Crucially,annotationsshouldbeexpandedtoincludeexplicit labels for non-manual markers and other linguistic phenomenatofacilitatethedevelopmentofmorenuanced models.
Second,thefieldmustmovetowardlinguistically-informed model architectures. Rather than relying on generic sequence-to-sequence models, future research should explorearchitecturesthatexplicitlyincorporatetheunique properties of ISL. This could involve using graph neural networks to model the dynamic spatial relationships between body parts or designing specialized attention

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
mechanisms that are aware of the signing space and the distinctrolesofthedominantandnon-dominanthands. Developing more sophisticated multi-modal fusion techniquestoeffectivelyintegratemanualandnon-manual channelsisacriticalpriority.
Third,giventhepersistentdatascarcity,self-supervisedand few-shotlearningtechniquesarevital.Thesuccessofusinga model pre-trained on ASL for one-shot recognition in ISL suggeststhatcross-lingualtransferlearningisapromising avenue [4]. Pre-training models on large, unlabeled sign language video corpora could enable the learning of powerful,generalizablerepresentations.
Finally, as the technology matures, research must increasingly focus on human-centered evaluation and community engagement. This involves developing new evaluation metrics that better correlate with human judgments of translation quality and are sensitive to the linguisticnuancesofISL[10].
Moreimportantly,itrequiresashiftfromdesigningsystems for the Deaf community to co-designing them with the community, ensuring that the resulting technology is not onlyaccuratebutalsoculturallyappropriate,respectful,and genuinelyusefultoitsintendedusers.
This survey has charted the rapid evolution of automated Indian Sign Language processing, a field that has transitionedfromfoundationalcomputervisionexercisesto theforefrontofmulti-modaldeeplearningresearch.
The journey has been characterized by a methodical progression from recognizing isolated, static signs to the ambitious goal of translating continuous, dynamic sign language.Thisadvancementhasbeenfundamentallydriven by two parallel forces: the development of increasingly sophisticatedneuralarchitectures fromCNNsandLSTMs toTransformers andthecriticalemergenceoflarge-scale, public datasets like CISLR, ISLTranslate, and the comprehensiveiSignbenchmark.
Thecurrentstateoftheartisdefinedbyend-to-endsystems thatleverageposeestimationframeworkslikeMediaPipefor efficient and robust feature extraction, coupled with Transformer-based models that can jointly learn to recognizeandtranslatesigns.
Despite achieving impressive performance on benchmark tasks, significant challenges remain. The linguistic complexitiesofISL,includingtheprevalenceoftwo-handed signs,thegrammaticalimportanceofnon-manualmarkers, andtheuseofspatialgrammar,continuetotestthelimitsof currentmodels.
Overcoming these hurdles and closing the gap between benchmark performance and real-world reliability will definethenextphaseofresearch.Futureworkmustfocuson creating larger and more diverse datasets, designing linguistically-informedarchitectures,and engagingdeeply with the Deaf community to ensure the development of technologythatisbothpowerfulandpurposeful.
The continued progress in this domain holds immense potentialtodismantlecommunicationbarriersandfostera moreinclusiveandaccessiblesociety.
[1]S.Shetty,E.Hirani,A.Singh,andR.Koshy,“Gesture-toText: A Real-Time Indian Sign Language Translator with Pose Estimation and LSTMs,” Procedia Computer Science, vol.235,pp.2684-2692,2024.
[2]A.Joshi,S.Agrawal,andA.Modi,“ISLTranslate:Dataset for Translating Indian Sign Language,” in Findings of the AssociationforComputationalLinguistics:ACL2023,2023.
[3] N. Sharma, J. Mandal, A. Chaudhury, and S. V, “PerformanceAnalysisofCNN-BasedIndianSignLanguage Recognition,”inProceedingsoftheInternationalConference on Advanced Research in Electronics and Communication Systems(ICARECS2025),AtlantisHighlightsinEngineering, vol.38,2025,pp.15-24.
[4]A.Joshietal.,“CISLR:CorpusforIndianSignLanguage Recognition,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP),2022,pp.10357-10366.
[5] P. Rawat, P. Kumar, V. K. Tamta, and A. Kumar, “A Comprehensive Approach to Indian Sign Language Recognition: Leveraging LSTM and MediaPipe Holistic for Dynamic and Static Hand Gesture Recognition,” EAI EndorsedTransactionsonAIandRobotics,vol.4,2025.
[6]P.S.RajamandG.Balakrishnan,“RealtimeIndianSign LanguageRecognitionSystemtoaiddeaf-dumbpeople,”in 2011IEEE13thInternationalConferenceonCommunication Technology,2011,pp.737-742.
[7]A.Singh,A.Wadhawan,M.Rakhra,U.Mittal,A.AlAhdal, andS.K.Jha,“IndianSignLanguageRecognitionSystemfor DynamicSigns,”in202210thInternationalConferenceon Reliability,InfocomTechnologiesandOptimization(Trends andFutureDirections)(ICRITO),2022.
[8]R.Ba,R.Holla,A.Azam,andP.S.Kiran,“Signlanguage recognitionusingLSTM,”SSRNElectronicJournal,2024.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
[9] B. Subramanian et al., “An integrated mediapipeoptimizedGRUmodelforIndiansignlanguagerecognition,” ScientificReports,vol.12,no.1,p.11964,2022.
[10] A. Joshi et al., “iSign: A Benchmark for Indian Sign Language Processing,” in Findings of the Association for ComputationalLinguistics:ACL2024,2024.
[11]N.C.Camgoz,O.Koller,S.Hadfield,andR.Bowden,“Sign Language Transformers: Joint End-to-end Sign Language Recognition and Translation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR),2020,pp.10023-10033.