
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Mrs.K.Sandhya1 , Sujana Haritha.K2 , Sai kiriti.K3 , Sai sandeep.K4
1 Assistant Professor, Department of Computer Science and Engineering 2,3,4 B.Tech Students, Department of Computer Science and Engineering
Teegala Krishna Reddy Engineering College , Telangana, India
Abstract - Identifying potential side effects of drug molecules is a critical and time-consuming step in the drug discovery and development process. Failure to detect adverse effects early can lead to increased research costs, regulatory delays, and potential risks to patient safety. Traditional machine learning approaches such as Decision Tree (DT) and Random Forest (RF) have been widely used to predict drug side effects based on molecular descriptors and fingerprints. However, these models rely heavily on handcrafted features and often fail to capture the sequential relationships present in molecular representations. Toaddresstheselimitations, this project proposes a deep learning–based approach using Recurrent Neural Networks (RNN), specificallyLSTMandGRU architectures, for predicting the side effects of drug molecules. In the proposed system, drug molecules are represented as SMILES sequences, which are tokenized and converted into numerical vectors suitable for neural network processing. The RNN model processes these sequences step-by-step, learning complex dependencies between atoms and bonds within the molecular structure. The final hidden representationispassed through dense layers to classify and predict the probability of various drug side effects. Molecular data from databases such as DrugBank, SIDER, and PubChem are used for training and evaluation. Experimental results demonstrate that the proposed RNN-based model achieves improved predictive performance while significantly reducing parameter complexity compared to large-scale models. This approach enables efficient and accurate prediction of drug side effects, supporting safer and faster drug development.
Keywords DrugSideEffectPrediction,RecurrentNeural Network(RNN),GRU,LSTM,SMILESRepresentation,Deep Learning,MolecularDataAnalysis,Pharmacovigilance.
1.INTRODUCTION
Thediscoveryanddevelopmentofnewdrugmoleculesisa complex and costly process that requires extensive evaluationofmolecularpropertiessuchastoxicity,efficacy, andpotentialsideeffects.Identifyingadversedrugreactions (ADRs)atanearlystageisessentialtoensurepatientsafety and to reduce the financial burden associated with failed drugtrials.Traditionalexperimentalmethodsfordetecting drugsideeffectsinvolvelaboratorytestingandclinicaltrials, which are time-consuming and expensive. As a result, computational approaches have become increasingly importantinassistingresearcherstopredictpotentialside effects before drugs reach clinical testing stages. Public
biomedical databases such as DrugBank, SIDER, and PubChem provide large volumes of molecular and pharmacologicaldatathatenablethedevelopmentofdatadrivenpredictivemodelsfordrugsafetyanalysis[6],[7],[8]. Machinelearningtechniqueshavebeenwidelyappliedinthe pharmaceutical domain to analyze molecular data and predict drug properties. Traditional algorithms such as DecisionTree(DT)andRandomForest(RF)arecommonly usedbecausetheyprovideinterpretablepredictionsandcan workeffectivelywithmolecularfingerprintsordescriptors. RandomForestmodelscombinemultipledecisiontreesto improvepredictionaccuracyandreduceoverfitting,while Decision Trees provide rule-based structures that help understandtherelationshipbetweenmolecularfeaturesand predictedoutcomes. These modelshave beensuccessfully used to identify toxicity patterns and predict biological activities of drug compounds. However, these approaches depend heavily on handcrafted features and may fail to capture complex structural relationships present within molecularsequences[1].Inrecentyears,deeplearninghas shownsignificantpotentialinsolvingcomplexproblemsin bioinformatics and drug discovery. Recurrent Neural Networks (RNNs) are particularly effective for handling sequentialdatabecausetheymaintaininternalmemorythat captures dependencies between elements in a sequence. Advanced RNN architectures such as Long Short-Term Memory (LSTM) and Gated Recurrent Units (GRU) were introduced to address the limitations of traditional RNNs, especially the vanishing gradient problem, enabling the modeltolearnlong-termdependenciesmoreeffectively[3], [4]. These networks have been widely applied in natural language processing, healthcare data analysis, and pharmacovigilancestudiestoidentifyadversedrugevents frommedicalrecordsandclinicaltexts[2].Drugmolecules areoftenrepresentedusingSMILES(SimplifiedMolecular InputLineEntrySystem)strings,whichencodemolecular structuresassequencesofcharacters.SinceSMILESstrings containsequentialrelationshipsbetweenatomsandbonds, they are well suited for sequence-based deep learning models.RNNscanprocessSMILESstringsstep-by-stepand automaticallylearnstructuralpatternsthatinfluencedrug side effects. Compared with traditional machine learning models, RNN-based approaches can extract meaningful representationsdirectlyfrommolecularsequenceswithout relying solely on manually engineered features. Recent research has shown that RNN models, particularly GRUbased architectures, can achieve comparable prediction

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
performance while using significantly fewer parameters thanlarge-scalemodels[1].
The proposed system introduces a deep learning–based approach for predicting potential side effects of drug moleculesusingRecurrentNeuralNetworks(RNN).Unlike traditionalmachinelearningmodelsthatrelyonhandcrafted molecular descriptors, the proposed method processes molecular structures directly as sequences using SMILES representations. This allows the model to automatically learn patterns and dependencies between atoms and chemical bonds within a molecule. By leveraging RNN architectures such as LSTM and GRU, the system can effectively analyze sequential molecular information and predict the probability of possible drug side effects. The proposedframeworkimprovespredictioncapabilitywhile maintaininglowerparametercomplexitycomparedtolargescaledeeplearningmodels.
Theoverallarchitectureoftheproposedsystemconsistsof three major components: the input layer, the RNN processingmodule,andtheoutputlayer.Thesystembegins byreceivingdrugmoleculesrepresentedinSMILESformat. These sequences are then converted into numerical representations through tokenization and encoding. The encoded sequences are fed into the Recurrent Neural Network module, which contains LSTM or GRU cells that process the sequence step-by-step while maintaining memoryofpreviouslyprocessedelements.Thissequential processing enables the network to learn structural dependenciespresentinmoleculardata.Duringthetraining phase, the RNN model learns the relationship between molecularsequencesandtheirassociatedsideeffectsusing labeled datasets obtainedfrom biomedical databases. The model parameters are optimized to minimize prediction error and improve classification performance. In the predictionphase,newdrugmoleculesareprovidedasinput, andthetrainedmodelanalyzestheirstructuralpatternsto estimatetheprobabilityof potential sideeffects.Thefinal outputlayerproducespredictedsideeffectcategoriesalong withconfidencescores.Thesystemarchitectureusedinthe proposed model is illustrated in Figure 1, where the molecularinputlayer,theRNNprocessingmodule,andthe output prediction layer collectively form the side-effect predictionframework.

In the proposed system, drug molecules are represented using the Simplified Molecular Input Line Entry System (SMILES).SMILESencodeschemicalstructuresassequences of characters that describe atoms, bonds, and molecular branching.Thissequentialrepresentationmakesitsuitable for processing by RNN-based deep learning models. The SMILESstringsarefirsttokenizedintoindividualcharacters ortokensandthenconvertedintonumericalvectorsthrough encodingtechniques.Theseencodedsequencesserveasthe input data for theRNN model. Thecorecomponentof the proposedsystemistheRecurrentNeuralNetwork.RNNsare designedtohandlesequentialdatabymaintainingahidden statethatcapturesinformationfromprevioustimesteps.In thisproject,advancedRNNarchitecturessuchasLongShortTermMemory(LSTM)andGatedRecurrentUnit(GRU)are used to improve the model’s ability to learn long-term dependencieswithinmolecularsequences.Thesenetworks analyze the SMILES sequence step-by-step and generate meaningful representations of molecular structures. The final hidden state produced by the RNN is passed to fully connected layers for classification. The training process involvesfeedinglabeledmoleculardataintotheRNNmodel. Themodellearnstoassociatemolecularsequencepatterns with known side effects using supervised learning techniques.Duringtraining,optimizationalgorithmssuchas Adam are used to update the model parameters and minimize the loss function. Performance metrics such as accuracy,precision,recall,andF1-scoreareusedtoevaluate theeffectivenessofthemodel.Oncethemodelistrained,it can be used in the prediction phase to analyze new drug molecules. The SMILES sequence of the new molecule is processedthroughthetrainedRNNmodel,andthesystem outputs the predicted probability of possible side effects. This helps researchers and pharmaceutical developers identify potential risks associated with drug molecules beforeconductingexpensivelaboratoryexperiments.
The implementation of the proposed system focuses on building an end-to-end pipeline for predicting drug side effects using Recurrent Neural Networks. The system integrates multiple stages including molecular data collection,preprocessing,featureextraction,deeplearning model development, training, evaluation, and prediction. Initially, drug-related data is collected from publicly available biomedical databases such as DrugBank, SIDER, and PubChem, which provide detailed information about molecularstructures,properties,andknownsideeffects.The collected datasets, which include SMILES representations andcorrespondinglabels,arecleanedtoremoveduplicate andincompleterecords,andmissingvaluesarehandledto ensureconsistency.Theprocesseddatasetisthendivided intotrainingandtestingsetstoevaluatemodelperformance onunseendata.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
DrugmoleculesarerepresentedusingSMILESstrings,which encode chemical structures as sequences of characters. Thesesequencesaretokenizedandconvertedintonumerical representationsusingencodingtechniquessuchasinteger encoding or embedding layers. Additionally, molecular descriptors and fingerprints are extracted using tools like RDKittocaptureimportantstructuralfeatures.Thecoreof the system is a Recurrent Neural Network model implemented using deep learning frameworks such as TensorFloworPyTorch.AdvancedarchitectureslikeLSTM or GRU are used to effectively capture long-term dependencies within molecular sequences. The model typicallyconsistsofanembeddinglayerfollowedbyoneor moreRNNlayers,whoseoutputsarepassedtodenselayers forclassificationofsideeffects.
Duringtraining,themodellearnstherelationshipbetween molecularsequencesandtheirassociatedsideeffectsusing supervised learning. Loss functions such as binary crossentropy or categorical cross-entropy are used to measure predictionerror,whileoptimizationalgorithmslikeAdam are applied to update model weights. Hyperparameters including learning rate, batch size, number of epochs, and hidden units are tuned to improve performance, and regularization techniques such as dropout are used to prevent overfitting. After training, the model is evaluated usingmetricssuchas accuracy,precision,recall,F1-score, andROC-AUCtoassessitseffectiveness.Theperformanceof theRNNmodelisalsocomparedwith traditionalmachine learningapproacheslikeDecisionTreeandRandomForest tohighlighttheadvantagesofsequence-baseddeeplearning. Overall,theimplementedsystemprovidesanefficientand accurate framework for predicting drug side effects from moleculardata.
The performance of the proposed system is evaluated to determine its ability to accurately predict potential side effects of drug molecules. The Recurrent Neural Network (RNN)modelistrainedusingmoleculardatasetscontaining SMILESrepresentationsandcorrespondingsideeffectlabels.
After thetraining phase,the model is testedusingunseen datatoevaluateitsgeneralizationcapabilityandprediction accuracy.
Several evaluation metrics such as Accuracy, Precision, Recall, F1-score, and ROC-AUC are used to assess the effectivenessoftheproposedmodel.Thesemetricshelpin understanding how well the model identifies drug side effectswhileminimizingincorrectpredictions.Theresults demonstrate that the RNN-based model performs better than traditional machine learning methods in capturing sequentialpatternswithinmolecularstructures.
During the training phase, the model gradually learns the relationship between molecular structures and their associatedsideeffects.Thelossvaluedecreaseswitheach
trainingepochwhiletheaccuracyimproves,indicatingthat themodelsuccessfullylearnstheunderlyingpatternsinthe dataset. Proper hyperparameter tuning and optimization techniques help improve the stability and performance of themodel.
Toevaluatetheeffectivenessoftheproposedapproach,the performanceoftheRNNmodeliscomparedwithtraditional machinelearningalgorithmssuchasDecisionTree(DT)and Random Forest (RF). These models rely on handcrafted molecular descriptors, whereas the RNN model learns featuresdirectlyfromSMILESsequences.
Table 1: Performance Comparison of Different Models
From Table 1, it can be observed that the proposed RNNbased model achieves higher accuracy and improved performance across all evaluation metrics compared to traditionalmodels.TheabilityofRNNstocapturesequential dependencies within molecular structures significantly contributestoimprovedpredictionresults
Thepredictionresultscanalsobevisualizedusinggraphical methods such as accuracy curves, loss curves, and ROC curves. These visualizations help in analyzing the model’s learning behavior and classification capability. The ROC curveshowstherelationshipbetweenthetruepositiverate and false positive rate, providing insights into the classificationperformanceofthemodel. Theexperimental resultsdemonstratethattheproposedRNN-basedsystem can effectively analyze molecular sequences and predict drug side effects with high accuracy. The improved performanceandreducedparametercomplexitymakethe proposedapproachsuitableforpracticalapplicationsindrug discoveryandpharmacovigilance.
Inthisproject,adeeplearning–basedapproachfordetecting the side effects of drug molecules using Recurrent Neural Networks(RNN)hasbeenpresented.Identifyingpotential adversedrugreactionsisacriticalstepinthedrugdiscovery process,asithelpsensurepatientsafetyandreducestherisk ofcostlyfailuresduringclinicaltrials.Traditionalmachine learningmodelssuchasDecisionTreeandRandomForest relyheavilyonhandcraftedmoleculardescriptorsandmay noteffectivelycapturethesequentialrelationshipspresent inmolecularstructures.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Toaddresstheselimitations,theproposedsystemutilizes SMILES representations of drug molecules and processes themusingRNNarchitecturessuchasLSTMandGRU.These networks are capable of learning complex dependencies between atoms and bonds within molecular sequences, allowing the model to automatically extract meaningful structural patterns. The implementation involves data collection from biomedical databases, molecular preprocessing,sequenceencoding,RNNmodeltraining,and evaluationusingmultipleperformancemetrics.
Experimental results demonstrate that the RNN-based modelachievesimprovedpredictionaccuracycomparedto traditionalmachinelearningmethods.Themodeleffectively learnsfromsequentialmoleculardataandpredictspotential sideeffectswithhigherprecisionandrecall.Thisapproach reduces the need for manual feature engineering while maintainingefficientcomputationalperformance.
Although the proposed system demonstrates promising resultsinpredictingthesideeffectsofdrugmoleculesusing RecurrentNeuralNetworks,thereareseveralopportunities forfurtherimprovementandexpansion.Futureresearchcan focus on integrating more advanced deep learning architecturessuchastransformer-basedmodelsandgraph neuralnetworks(GNNs),whichcanbettercapturecomplex molecular structures and relationships between atoms. Thesemodelsmayfurtherimprovepredictionaccuracyand provide deeper insights into molecular behavior.Another potential direction is the use of larger and more diverse datasets collected from multiple biomedical databases. Incorporating additional pharmacological and biological information,suchasproteintargets,drug–druginteractions, andpatient-specificdata,couldenhancetherobustnessand reliability of the prediction system.Future work can also exploreexplainableartificialintelligence(XAI)techniquesto improvetheinterpretabilityofthemodel.Methodssuchas SHAP and LIME can help researchers understand how specific molecular features contribute to predicted side effects,increasingtrustinthemodel'sdecisions.Additionally, thesystemcanbeextendedtosupportreal-timeprediction throughaweb-basedorcloud-basedplatform.Thiswould allow pharmaceutical researchers and healthcare professionalstoinputdrugmoleculesandinstantlyobtain predicted side effects. Integrating visualization tools to displaymolecularstructuresandpredictedrisklevelswould furtherenhanceusability.Overall,expandingthemodelwith advanced algorithms, larger datasets, and improved interpretability techniques can significantly enhance the effectiveness of drug side effect prediction systems and contributetosaferdrugdevelopmentinthefuture.
[1]A.Authoretal.,“PredictingSideEffectsofDrugMolecules Using Recurrent Neural Networks,” arXiv preprint arXiv:2305.09421,May2023.
[2]Y.JagannathaandH.Yu,“BidirectionalRNNforMedical Event Detection in Electronic Health Records,” in Proceedings of the Conference on Empirical Methods in NaturalLanguageProcessing(EMNLP),2016,pp.473–482.
[3] S. Hochreiter and J. Schmidhuber, “Long Short-Term Memory,”NeuralComputation,vol.9,no.8,pp.1735–1780, 1997.
[4] K. Cho et al., “Learning Phrase Representations using RNNEncoder–DecoderforStatisticalMachineTranslation,” inProceedingsoftheConferenceonEmpiricalMethodsin NaturalLanguageProcessing(EMNLP),2014.
[5] G. Landrum, “RDKit: Open-source cheminformatics software,”2016.[Online].Available:http://www.rdkit.org
[6] M. Kuhn et al., “The SIDER database of drugs and side effects,”NucleicAcidsResearch,vol.44,no.D1,pp.D1075–D1079,2016.
[7]D.Wishartetal.,“DrugBank:Acomprehensiveresource forinsilicodrugdiscoveryandexploration,”NucleicAcids Research,vol.46,no.D1,pp.D1074–D1082,2018.
[8]S.Kimetal.,“PubChemin2021:Newdatacontentand improvedwebinterfaces,”NucleicAcidsResearch,vol.49, no.D1,pp.D1388–D1395,2021.