
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 01 | Jan 2026 www.irjet.net
p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 01 | Jan 2026 www.irjet.net
p-ISSN: 2395-0072
Nakibinge Gavin¹
¹Department
of Information and Communication Technology,Marwadi University, India
Abstract -The rapid growth of digital scholarly repositories has resulted in an overwhelming volume of academic publications, making it increasingly difficult for researchers to efficiently discover relevant research papers. Conventional keyword-based search mechanisms often fail to capture semantic relationships between documents and require significant manual effort. To address this challenge, this paper presents an intelligent research paper recommendation system based on Natural Language Processing (NLP), similarity modeling, clustering, and topic modeling techniques. The proposed system utilizes the arXiv dataset comprising over 250,000 research papers across multiple scientific domains. Titles and abstracts are preprocessed and represented using Term Frequency–Inverse Document Frequency (TF-IDF) vectors. Cosine similarity is employed to compute semantic similarity between research papers. Dimensionality reduction techniques such as Principal Component Analysis (PCA) and t-Distributed Stochastic Neighbor Embedding (t-SNE) are applied for visualization and analysis of high-dimensional data. Furthermore, K-Means clustering is used to group similar papers, and Latent Dirichlet Allocation (LDA) is applied to extract dominant topics from each cluster. A Flask-based web application deployed on Hugging Face Spaces provides an interactive interface for searching and recommending research papers. Experimental results demonstrate that the system effectively identifies semantically related research articles and organizes them into meaningful clusters, thereby improving the efficiency and quality of academic literature discovery.
Keywords: Research Paper Recommendation, Natural Language Processing, TF-IDF, Cosine Similarity, Clustering, Topic Modeling, Information Retrieval, NLP
1. INTRODUCTION
TheexponentialgrowthofdigitalresearchrepositoriessuchasarXivandGoogleScholarhasresultedinamassiveincreasein scholarlypublicationsacross allscientificdomains.Whilethisgrowthpromotesknowledgedissemination,italsocreates significantchallengesforresearchersattemptingtoidentifyrelevantliteratureefficiently.Traditionalkeyword-basedsearch mechanisms relyprimarilyonlexical matchingandoftenfailtocapturesemanticrelationshipsbetweenresearchpapers, leadingtoincompleteorirrelevantresults[1].
Recommendersystemshaveemergedasaneffectivesolutionformanaginginformationoverloadbyprovidingpersonalized contentsuggestions.Inacademicresearch,recommendationsystemscanassistscholarsbyidentifyingpaperscloselyaligned withtheirresearchinterestsandreducingthetimerequiredformanualliteraturesurveys[2].Content-basedrecommendation approachesanalyzetextualinformationsuchastitlesandabstractstocomputesimilaritybetweendocumentswithoutrelying onuserinteractiondata,therebyavoidingthecold-startproblem[3].
RecentadvancesinNaturalLanguageProcessing(NLP)haveenabledtheuseofstatisticalandsemantictechniquessuchasTFIDF, cosine similarity, clustering, and topic modeling to enhance document similarity analysis [4]. Motivated by these developments,thispaperproposesanNLP-basedresearchpaperrecommendationsystemthatintegratessimilaritymodeling, clustering,andtopicmodelingtechniques.Aweb-basedimplementationdeployedonHuggingFaceSpacesdemonstratesthe practicalapplicabilityoftheproposedsystem.
Several studies have explored academic paper recommendation using collaborative filtering, content-based filtering, and hybrid approaches. Collaborative filtering methods rely on user behaviour data such as citations, downloads, or reading history,buttheyoftensufferfromcold-startandsparsityissueswhenuserdataislimited[5].
Content-basedapproachesutilizetextualfeaturesofresearchpaperstorecommendsimilardocuments.TF-IDFcombinedwith cosinesimilarityhasbeenwidelyadoptedfordocumentsimilarityanalysisduetoitssimplicityandinterpretability[1].Topic modelling techniques such as Latent Dirichlet Allocation (LDA) have been used to uncover latent thematic structures in academiccorpora[2].MorerecentstudiesemploydeeplearningmodelssuchasWord2Vec,Doc2Vec,andBERTforsemantic representation;however,thesemethodsrequirehighercomputationalresourcesandcomplextrainingpipelines[6].

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 01 | Jan 2026 www.irjet.net p-ISSN: 2395-0072
Incontrast,thisworkfocusesonascalableandinterpretableNLP-basedrecommendationsystemusingclassicalmachine learningtechniquesthatprovideeffectiveresultswhilemaintainingcomputationalefficiency.
3. Dataset Description
TheproposedsystemutilizesthearXivdatasetobtainedfromKaggle,whichcontainsmetadataforapproximately250,000 research papers across multiple scientific domains [7]. Each record includes the paper title, abstract, authors, subject categories,anduniqueidentifiers.Titlesandabstractsareusedastheprimarytextualfeaturesinthisstudy,astheyprovide concisesummariesofresearchcontributionsandarecommonlyemployedincontent-basedrecommendationsystems[3].
Parameter Description
DatasetSource arXiv(Kaggle)
NumberofPapers 250,000
TextFieldsUsed Title,Abstract
DomainsCovered ComputerScience,Physics,Mathematics,Engineering
Language English
Table1summarizesthekeycharacteristicsofthedatasetusedforexperimentation.
4. Proposed Methodology
4.1 Text Preprocessing
Textpreprocessingisappliedtotitlesandabstractstoreducenoiseandimprovefeaturequality.Thepreprocessingsteps includetokenization,conversiontolowercase,removalofstopwords,andeliminationofpunctuationandspecialcharacters.
4.2 Feature Extraction Using TF-IDF
ThecleanedtextualdataistransformedintonumericalfeaturevectorsusingTermFrequency–InverseDocumentFrequency (TF-IDF).TF-IDFeffectivelycapturestheimportanceoftermswithindocumentsrelativetotheentirecorpus[1].
4.3 Similarity Computation
CosinesimilarityisusedtocomputesemanticsimilaritybetweenTF-IDFvectors.Paperswithhighercosinesimilarityscores areconsideredmoresemanticallyrelatedandarerankedhigherintherecommendationlist[4]
4.4 Dimensionality Reduction
PrincipalComponentAnalysis(PCA)isappliedtoreducethedimensionalityoftheTF-IDFmatrixwhilepreserving95%ofthe variance.t-DistributedStochasticNeighborEmbedding(t-SNE)isfurtherusedtoprojecthigh-dimensionalvectorsintoatwodimensionalspaceforvisualizationandqualitativeanalysis[8].
4.5 Clustering and Topic Modelling
K-Meansclusteringisemployedtogroupresearchpapersintoclustersbasedonsemanticsimilarity.LatentDirichletAllocation (LDA) is applied to extract dominant topics and keywords from each cluster, improving interpretability and thematic understanding[2].

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 01 | Jan 2026 www.irjet.net p-ISSN: 2395-0072

Fig.1illustratestheoverallarchitectureoftheproposedsystem.TheprocessbeginswiththearXivdataset,followedbytext preprocessingandTF-IDFvectorization.Cosinesimilarityisusedtocomputedocumentsimilarity.Dimensionalityreduction techniques support visualization, while clustering and topic modeling organize papers into thematic groups. The recommendationresultsaredeliveredthroughaFlask-basedwebapplicationdeployedonHuggingFaceSpaces.
6. Recommendation Process Flow


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 01 | Jan 2026 www.irjet.net p-ISSN: 2395-0072
Fig. 2 presents the flowchart of the recommendation process. The user inputs a research paper title, which undergoes preprocessingandvectorgeneration.Similaritycomputationidentifiesthetop-Nrelatedpapers,whicharemappedtoclusters andtopicsbeforebeingdisplayedtotheuser.
7. Performance Evaluation
7.1 Precision@K
Precision@Kmeasurestherelevanceofthetop-Krecommendedpapersandisdefinedas:
Relevance is determined based on subject category overlap and semantic similarity. Experimental results indicate high Precision@5andstablePrecision@10values,demonstratingthatrelevantpapersarerankedprominently.

7.2 Similarity Score Analysis
Cosinesimilarityscoresareanalyzedtoassesssemanticclosenessbetweenrecommendedpapersandtheinputpaper.The systemconsistentlyproduceshighsimilarityscoresforrecommendedpapers,confirmingtheeffectivenessofTF-IDF-based similaritymodeling.

-2: Interactivephysicpapervisualization

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 01 | Jan 2026 www.irjet.net p-ISSN: 2395-0072
Clustercoherenceandtopicconsistencyareexaminedqualitatively.Paperswithinthesameclustersharesimilarresearch themes,andLDA-generatedkeywordsaccuratelyrepresentclustertopics.t-SNEvisualizationfurtherconfirmsclearseparation betweenclusters.

Chart -3: physicspaperclusteringvisualization
Table – 2: Performance Evaluation Summary
Metric
Observation
Precision@5 Highrelevance
Precision@10 Consistentperformance
SimilarityScores Strongsemanticalignment
ClusterCoherence Well-definedtopics
SystemResponseTime Fast
8. Web Application Implementation
A web-based application is developed using the Flask framework and deployed on Hugging Face Spaces. The application enables users to search for research papers by title and receive recommendations based on semantic similarity. Paper metadatasuchastitles,authors,andidentifiersaredisplayedtoenhanceusability.
9. Conclusion and Future Work
Thispaperpresenteda professional NLP-basedresearchpaperrecommendationsystemintegratingsimilarity modelling, clustering,andtopicmodellingtechniques.Thesystemeffectivelyaddressesthechallengeofacademicliteraturediscoveryby

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 01 | Jan 2026 www.irjet.net p-ISSN: 2395-0072
identifying semantically related research papers and organizing them into meaningful clusters. Experimental evaluation demonstratesstrongrecommendationrelevanceandclusteringquality.
Futureworkincludesintegratingdeeplearning-basedembeddingssuchasBERT,incorporatinguserprofilingandcollaborative filtering,andextendingthesystemtosupportreal-timeandcross-domainrecommendations.
References
[1]J.Ramos,“UsingTF-IDFtoDetermineWordRelevanceinDocumentQueries,”Proc.FirstInstructionalConf.onMachine Learning,2003.
[2]D.Blei,A.Ng,andM.Jordan,“LatentDirichletAllocation,”JournalofMachineLearningResearch,2003.
[3]C.Manning,P.Raghavan,andH.Schütze,IntroductiontoInformationRetrieval,CambridgeUniversityPress,2008.
[4] G. Salton and C. Buckley, “Term-Weighting Approaches in Automatic Text Retrieval,” Information Processing & Management,1988.
[5]J.Bobadillaetal.,“RecommenderSystemsSurvey,”Knowledge-BasedSystems,2013.
6]J.Devlinetal.,“BERT:Pre-trainingofDeepBidirectionalTransformersforLanguageUnderstanding,”NAACL,2019.
[7]CornellUniversity,“arXivDataset,”Kaggle,2023.
[8]L.vanderMaatenandG.Hinton,“VisualizingDatausingt-SNE,”JournalofMachineLearningResearch,2008
BIOGRAPHIES

afinal-yearundergraduatestudent pursuing a Bachelor’s degree in Information and Communication Technology.Thisresearchpaperis based on his final year project, which focuses on designing and implementing an NLP-based research paper recommendation system.