
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
AN INTELLIGENT DEEP LEARNING FRAMEWORK FOR EARLY-STAGE CLINICAL RISK IDENTIFICATION USING LONGITUDINAL PATIENT RECORDS
Shivangi
Singh1, Mr. Manish Kumar Soni2
1Master of Technology, Computer Science and Engineering, Bansal Institute of Engineering & Technology, Lucknow, India
2Assistant Professor, Department of Computer Science and Engineering, Bansal Institute of Engineering & Technology, Lucknow, India ***
Abstract -Early identification of clinical deterioration in hospitalized patients remains a critical challenge due to the limitations of traditional rule-based early warning systems and intermittent patient monitoring. These conventional approaches often fail to capture complex temporal patterns and multi-dimensional relationships present in longitudinal patient records. This study proposes an intelligent deep learningframeworkforearly-stageclinicalriskidentification using longitudinal electronic health record data. The frameworkemploysaBidirectionalLongShort-TermMemory (Bi-LSTM) architecture to model temporal dependencies in irregularly sampled clinical time-series data, including vital signs, laboratory results, and demographic information. A multi-label prediction strategy is implemented to simultaneously identify multiple adverse events, such as unplanned ICU transfer, cardiopulmonary resuscitation, and in-hospital mortality. Robust data preprocessing techniques, including temporal alignment and missing value imputation using Last Observation Carried Forward (LOCF), are integrated to enhance model reliability. The proposed model demonstrates superior predictive performance compared to traditional scoring systems, achieving improved AUROC and AUPRC values. Furthermore, real-world clinical deployment resultsindicateasignificantreductioninCodeBlueeventsand earlierriskdetection.Thesefindingshighlightthepotentialof deep learning–based early warning systems to improve patient outcomes and support proactive clinical decisionmaking.
Key Words: Deep Learning, Clinical Risk Prediction, BiLSTM, Longitudinal Patient Records, Early Warning Systems, Multi-Label Classification, Electronic Health Records
1. INTRODUCTION
1.1 Background
1.1.1 Clinical Deterioration and In-Hospital Adverse Events
Clinicaldeteriorationinhospitalizedpatientsreferstothe progressiveworseningofphysiologicalconditionsthatmay leadtosevereadverseeventssuchasunplannedintensive careunit(ICU)transfer,cardiopulmonaryarrest,ordeath. These events are often preceded by subtle physiological
changesthatremainundetectedinroutineclinicalpractice. Studies have shown that a significant proportion of inhospital cardiac arrests are preventable if early warning signs are identified in time (Smith et al., 2013). Despite advancesinhealthcaresystems,adverseeventscontinueto pose a major challenge to patient safety and hospital management, emphasizing the need for more proactive monitoringsystems.
1.1.2 Limitations of Intermittent Monitoring
Traditional patient monitoring systems rely on periodic measurement of vital signs, typically recorded every few hoursingeneralwards.Thisintermittentapproachcreates gaps in patient observation, during which critical physiological changes may occur unnoticed. As a result, cliniciansmaymissearlyindicatorsofdeterioration,leading todelayedintervention.Researchindicatesthatcontinuous or high-frequency monitoring significantly improves the detectionofearlywarningsignscomparedtointermittent methods(Cliffordetal.,2012).Therefore,relianceonmanual andepisodicmonitoringlimitstheeffectivenessofcurrent clinicalsurveillancesystems.
1.1.3 Importance of Early Detection in Reducing Mortality
Earlydetectionofpatientdeteriorationplaysacrucialrolein improvingclinical outcomes andreducingmortalityrates. Timelyidentificationenableshealthcareproviderstoinitiate preventive interventions, thereby avoiding severe complications. Rapid Response Systems (RRS) and Early Warning Systems(EWS)have been introduced to address thisneed;however,theireffectivenessdependsonaccurate and timely risk prediction. Evidence suggests that early interventioncansignificantlyreducetheincidenceofcritical events such as Code Blue and improve overall patient survivalrates(DeVitaetal.,2010).
1.2 Limitations of Existing Systems
1.2.1
Rule-Based Systems (Threshold-Based)
Traditional early warning systems, such as NEWS and MEWS,arebasedonpredefinedthresholdsforphysiological parameters. These systems generate alerts only when a variable exceeds a fixed limit, ignoring the temporal

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
progressionandinteractionamongvariables.Whilesimple toimplement,threshold-basedsystemsoftensufferfromlow sensitivityandhighfalsealarmrates.Consequently,theyfail to capture complex patterns of deterioration that evolve overtime(RoyalCollegeofPhysicians,2017).
1.2.2 Single-Event Prediction Models
Manyexistingmachinelearningmodelsfocusonpredictinga singleclinicaloutcome,suchasmortalityorICUadmission. However, patient deterioration is a multifaceted process involving multiple interconnected outcomes. Single-event modelsprovidealimitedviewofpatientriskandmaynot support comprehensive clinical decision-making. Multidimensionalpredictionisthereforeessentialtobetterreflect real-worldclinicalscenarios(Johnsonetal.,2016).
1.2.3 Lack of Real-World Validation
A major limitation of current predictive models is their relianceonretrospectivedatasetswithoutreal-worldclinical validation.Modelsthatperformwellincontrolledresearch environments may not achieve similar results when deployed in hospital settings. This gap between experimental performance and practical applicability reduces clinician trust and limits adoption of AI-based systems(Kellyetal.,2019).
1.2.4 Poor Generalizability
Predictivemodelstrainedondatafromasingleinstitution often fail to generalize across different hospitals due to variations in patient demographics, clinical practices, and dataquality.Thislackofrobustnessrestrictsthescalability ofsuchmodelsandlimitstheirbroaderclinicalapplicability (Wiensetal.,2019).
2. LITERATURE SURVEY
2.1 Traditional Early Warning Systems
2.1.1 NEWS, MEWS, SOFA, APACHE II
TraditionalEarlyWarningSystems(EWS)havebeenwidely used in clinical practice to detect early signs of patient deterioration. Among these, the National Early Warning Score(NEWS)andModifiedEarlyWarningScore(MEWS) are commonly applied in general wards to monitor physiologicalparameterssuchasheartrate,respiratoryrate, blood pressure, and temperature. In critical care settings, more complex scoring systems like the Sequential Organ Failure Assessment (SOFA) and Acute Physiology and ChronicHealthEvaluationII(APACHEII)areusedtoassess organdysfunctionandpredictmortalityrisk.Thesesystems aresimple,interpretable,andeasytoimplement,whichhas contributedtotheirwidespreadadoptionacrosshealthcare institutions(Vincentetal.,1996).
2.1.2 Comparative Limitations (Low AUROC, Static Thresholds)
Despite their clinical utility, traditional scoring systems sufferfromseverallimitations.Theyrelyonfixedthreshold valuesthatdonotaccountforpatient-specificvariabilityor temporal trends in physiological data. As a result, their predictive performance is often limited, with moderate AUROCvaluesreportedacross studies.Additionally,these systemsgeneratealertsonlyafterthresholdsareexceeded, whichmaydelayearlyintervention.Theinabilitytomodel complex,nonlinearrelationshipsbetweenvariablesfurther reduces their effectiveness in detecting subtle patterns of deterioration(Churpeketal.,2013).
2.2 Machine Learning in Healthcare
2.2.1 Logistic Regression and Random Forest
Machine learning approaches have been introduced to overcomethelimitationsoftraditionalscoringsystemsby leveragingdata-drivenmethodologies.Logisticregressionis one of the earliest and most widely used techniques for clinicalpredictionduetoitssimplicityandinterpretability.It estimates the probability of an outcome based on input variables and has been applied to predict mortality, ICU admission, and disease progression. Random Forest, an ensemble learning method, improves predictive performance by combining multiple decision trees and capturing nonlinear relationships among features. These modelshavedemonstratedbetteraccuracythantraditional systemsinvariousclinicalpredictiontasks(Breiman,2001).
2.2.2 Early AI-Based Prediction Systems
Earlyartificialintelligence-basedsystemsutilizedmachine learning algorithms to analyze electronic health record (EHR) data and identify patients at risk of deterioration. Thesesystemsincorporatedstructuredclinicaldatasuchas vital signs and laboratory results to generate risk scores. Although they showed promising improvements in predictiveperformance,manyofthesemodelswerelimited by their reliance on static features and lack of temporal modelingcapabilities.Consequently,theirabilitytocapture dynamicpatientconditionsovertimeremainedconstrained (ObermeyerandEmanuel,2016).
2.3 Deep Learning for Clinical Prediction
2.3.1
RNN, LSTM, Bi-LSTM
Deep learning techniques have significantly advanced clinical prediction by enabling models to learn complex representationsfromlarge-scalehealthcaredata.Recurrent Neural Networks (RNNs) are particularly suited for sequential data, as they maintain information across time steps. However, standard RNNs suffer from vanishing gradientproblems,whichlimittheirabilitytocapturelong-

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
term dependencies. Long Short-Term Memory (LSTM) networksaddressthisissuethroughgatingmechanismsthat regulate information flow. Bidirectional LSTM (Bi-LSTM) furtherenhancesperformancebyprocessingsequencesin bothforwardandbackwarddirections,allowingthemodel to capture contextual information more effectively (HochreiterandSchmidhuber,1997).
2.3.2 Temporal Modeling of EHR Data
ElectronicHealthRecordscontainlongitudinalpatientdata characterized by temporal dependencies and irregular samplingintervals.Deeplearningmodels,particularlyLSTMbasedarchitectures,haveproveneffectiveinmodelingsuch time-seriesdata.Thesemodelscanlearntemporalpatterns associated with disease progression and predict adverse events with higher accuracy compared to traditional methods.Byintegratingmultipledatamodalities,including vitalsigns,laboratoryresults,anddemographicinformation, deep learning approaches provide a more comprehensive understandingofpatienthealthtrajectories(Rajkomaretal., 2018).
2.4 AI-Based Early Warning Systems
2.4.1
TREWS and VitalCare
Recent advancements have led to the development of AIbased Early Warning Systems that integrate machine learningmodelsintoclinicalworkflows.TheTargetedRealTimeEarlyWarningSystem(TREWS)isanotableexample designedforearlydetectionofsepsisusingreal-timeEHR data. It employs advanced algorithms to continuously monitorpatientconditionsandgeneratealertsforclinicians. Similarly, the VitalCare system utilizes deep learning architectures,includingBi-LSTMmodels,topredictmultiple adverse events such as ICU transfer and mortality. These systems represent a shift toward intelligent, data-driven clinicalmonitoringsolutions(Adamsetal.,2022).
2.4.2
Performance Improvements Over Traditional Models
AI-based early warning systems have demonstrated significant improvements in predictive performance compared to traditional scoring methods. Studies have reportedhigherAUROCandAUPRCvalues,indicatingbetter discrimination and precision in identifying high-risk patients.Moreover,thesesystemsenableearlierdetectionof deterioration, allowing clinicians to intervene proactively. Theabilitytoanalyzelarge-scale,high-dimensionaldataand capturecomplextemporalrelationshipscontributestotheir superiorperformance(Shashikumaretal.,2017).
2.5 Research Gaps Identified
2.5.1
Lack of Multi-Event Prediction
DespiteadvancementsinAI-basedsystems,manyexisting models focus on predicting a single clinical outcome. This limitationreducestheirapplicabilityinreal-worldsettings, wheremultipleadverseeventsmayoccursimultaneously. Thereisaneedformulti-labelpredictionframeworksthat can provide a comprehensive assessment of patient risk acrossdifferentoutcomes(TsoumakasandKatakis,2007).
2.5.2 Limited Real-World Validation
Anothersignificantgapisthelackofreal-worldvalidation for many predictive models. Most studies rely on retrospective datasets and do not evaluate model performance in live clinical environments. This limits the understanding of how these systems perform under realworld conditions and affects their adoption in healthcare practice(Kellyetal.,2019).
2.5.3
Poor Handling of Irregular Time-Series
Clinical data are inherently irregular and often contain missing values, posing challenges for traditional and machinelearningmodels.Manyexistingapproachesfailto effectively handle such complexities, leading to reduced predictive accuracy. Advanced deep learning models combined with robust preprocessing techniques are required to address these challenges and improve model reliability(Cheetal.,2018).
3.
MATERIALS AND METHODS
3.1
Study Design
3.1.1
Two-Phase Design
Thestudyadoptsastructuredtwo-phasedesigntoensure both methodological rigor and clinical applicability of the proposeddeeplearningframework.Thisdesignintegrates retrospective model development with real-world clinical deployment, enabling comprehensive evaluation of both predictiveperformanceandpracticalutility.Byseparating developmentanddeploymentphases,thestudyminimizes bias and ensures that the model is tested under realistic hospitalconditions.
3.1.2
Model Development (Retrospective Phase)
Inthefirstphase,themodelisdevelopedusingretrospective electronichealthrecorddatacollectedfromatertiarycare hospital. This phase involves data preprocessing, feature engineering, and model training using historical patient records. Retrospective analysis allows the model to learn complex temporal patterns associated with clinical deteriorationwithoutinterferingwithongoingpatientcare.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
3.1.3 Clinical Deployment (Real-World Validation)
Thesecondphasefocusesonreal-worldvalidationthrough deployment in a clinical setting. The trained model is integratedintohospitalsystemstogenerateriskpredictions in real time. This phase evaluates not only predictive accuracybutalsotheclinicalimpactofthesystem,including itsabilitytosupportearlyinterventionandreduceadverse events.
3.2 Data Sources
3.2.1 Development Dataset (Tertiary Hospital,2013–2017)
Thedevelopmentdatasetisobtainedfromalargetertiary academicmedicalcenterandconsistsoflongitudinalpatient records collected between 2013 and 2017. This dataset includes a diverse patient population and provides comprehensive clinical information required for training deeplearningmodels.Thelargesamplesizeenhancesthe robustnessandgeneralizabilityofthelearnedpatterns.
3.2.2 Validation Dataset (Community Hospital,2022–2024)
To assess external validity, a separate dataset is collected fromacommunityhospitalcoveringtheperiodfrom2022to 2024.Thisdatasetisusedforindependenttestingandrealworld deployment evaluation. The use of data from a different institution ensures that the model is evaluated across varying clinical environments and patient populations.
3.3 Study Population
3.3.1 Inclusion Criteria: Adult Hospitalized
Patients
The study population includes adult patients admitted to hospitalwardsorintensivecareunits.Onlypatientsaged19 years and above are considered, ensuring consistency in physiological characteristics and clinical patterns. This inclusion criterion enables the model to focus on adultspecificriskfactorsanddiseaseprogression.
3.3.2 Exclusion
Criteria
Certainpatientgroupsareexcludedtomaintainthevalidity of the analysis. Patients with Do Not Resuscitate (DNR) ordersareexcludedbecausetheircareobjectivesdifferfrom standardinterventionprotocols.Additionally,plannedICU transfers following surgical or medical procedures are excluded, as these do not represent unexpected clinical deterioration. Pediatric patients are also excluded due to differences in physiological parameters and disease dynamics.
3.4 Data Description
3
.4.1 Vital Signs (HR, BP, RR, Temp, SpO₂)
Vital signs form the core input variables of the model, providing continuous insights into patient physiological status.Theseincludeheartrate(HR),bloodpressure(BP), respiratoryrate(RR),bodytemperature(Temp),andoxygen saturation (SpO₂). These parameters are routinely monitored in clinical settings and serve as primary indicatorsofpatienthealth.
3.4.2 Laboratory Values
Laboratory measurements provide additional clinical contextbyreflectinginternalphysiologicalandbiochemical conditions. These variables include parameters such as creatinine, white blood cell count, and electrolyte levels. Incorporatinglaboratorydataenhancesthemodel’sability todetectunderlyingpathologicalchangesthatmay notbe evidentfromvitalsignsalone.
3.4.3
Demographics
Demographic information, including age, gender, and admission details, is included to provide contextual information about the patient population. These variables helpthemodelaccountforpopulation-levelvariationsand improve prediction accuracy by incorporating patientspecificcharacteristics.
3.5 Data Preprocessing
3.5.1
Data Cleaning (Outlier Removal)
Data cleaning is performed to ensure the reliability and consistency of clinical measurements. Physiologically implausible values resulting from measurement errors or data entry issues are identified and removed based on predefinedclinicalthresholds.Thisstepreducesnoiseand prevents misleading patterns from influencing model training.
3.5.2 Temporal Alignment (Hourly Bins)
Since clinical data are recorded at irregular intervals, temporalalignmentisnecessarytostructurethedatainto uniformtimesteps.Inthisstudy,patientdataareaggregated into hourly intervals, ensuring consistency in sequence representation. This alignment allows the deep learning modeltoprocesstime-seriesdataeffectively.
3.5.3
Windowing (72-Hour Look-Back, 6-Hour Prediction)
Aslidingwindowapproachisappliedtocapturetemporal dynamics of patient data. The model uses a 72-hour lookbackwindowtoanalyzehistoricalpatientinformationand predict adverse events within the next 6 hours. This

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
configuration balances the need for sufficient historical contextwithtimelypredictionforclinicalintervention.
3.6 Missing Data Handling
3.6.1
LOCF (Last Observation Carried Forward)
Missingvaluesinclinicaltime-seriesdataarehandledusing the Last Observation Carried Forward (LOCF) method. In this approach, missing values are replaced with the most recent observed measurement for a given variable. This technique preserves temporal continuity and reflects the assumptionthatphysiologicalparameterschangegradually overtime.
3.6.2
Reference-Based Initialization
For cases where no prior observations are available, reference-based initialization is applied. Standard clinical baseline values are used to fill initial missing entries, ensuringthatthedatasetremainscompleteandsuitablefor modeltraining.Thisapproachpreventsdatasparsityissues duringearlytimesteps.
3.7 Proposed Model Architecture
3.7.1
Bidirectional LSTM (Bi-LSTM)
Thecorearchitectureoftheproposedframeworkisbasedon a Bidirectional Long Short-Term Memory (Bi-LSTM) network. This model processes time-series data in both forwardandbackwarddirections,enablingittocapturepast and future contextual dependencies. Such capability is particularlyusefulformodelingcomplextemporalpatterns inclinicaldata.
3.7.2 Multi-Label Output
Layer
The model employs a multi-label output layer that simultaneously predicts multiple adverse clinical events. Eachoutputnodecorrespondstoaspecificoutcome,suchas ICU transfer, mortality, or cardiopulmonary resuscitation. This design allows the model to provide a comprehensive assessmentofpatientrisk.
3.7.3
Shared Representation Learning
Sharedrepresentationlearningisusedtoextractcommon features from input data that are relevant across multiple predictiontasks.Bysharinghiddenlayersamongdifferent outputs, the model improves efficiency and reduces overfittingwhilecapturinginterdependenciesamongclinical outcomes.
3.8 Model Training
3.8.1
Loss Function: Binary Cross-Entropy
The model is trained using the binary cross-entropy loss function,whichissuitableformulti-labelclassificationtasks. This loss function measures the difference between predictedprobabilitiesandactualoutcomesforeachlabel, enablingeffectiveoptimizationofmodelparameters.
3.8.2
Optimizer: Adam
The Adam optimizer is employed for training due to its adaptivelearningrateandefficientconvergenceproperties. It combines the advantages of momentum and RMSProp, allowingthemodeltoachievestableandfastoptimization evenwithlargedatasets.
3.8.3 Hyperparameter Tuning
Hyperparametertuningisperformedtoidentifytheoptimal configurationofmodelparameters,includinglearningrate, batchsize,numberoflayers,andhiddenunits.Thisprocess ensures that the model achieves the best possible performance while maintaining generalization across differentdatasets.
4. EXPERIMENTAL SETUP
4.1
Hardware and Software
4.1.1 GPU Infrastructure (Tesla V100/A100)
Theexperimentalsetuputilizeshigh-performanceGraphics ProcessingUnits(GPUs),specificallyNVIDIATeslaV100and A100, to support computationally intensive deep learning operations. These GPUs are designed for large-scale data processingandprovidethousandsofparallelcores,enabling efficienttrainingofrecurrentneuralnetworkarchitectures suchasBidirectionalLongShort-TermMemory(Bi-LSTM). The use of GPU acceleration significantly reduces training time and allows the model to process large volumes of longitudinalclinicaldataefficiently.
4.1.2
Deep Learning Frameworks: PyTorch and TensorFlow
The implementation of the proposed model is carried out usingwidelyadopteddeeplearningframeworks,including PyTorch and TensorFlow. PyTorch is primarily used for modeldevelopmentduetoitsdynamiccomputationgraph and flexibility in handling sequential data. TensorFlow is utilized for comparative experiments and baseline implementation.Theseframeworksproviderobustlibraries forneuralnetworkconstruction,automaticdifferentiation, and GPU acceleration, ensuring scalability and reproducibilityofexperiments.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
4.2 Dataset Split
4.2.1
Training Set (70%)
Thetrainingdatasetcomprises70%ofthetotaldataandis used to learn model parameters. During this phase, the modelidentifiespatternsandrelationshipsbetweeninput featuresandtargetoutcomes.Alargetrainingsetensures that the model is exposed to diverse clinical scenarios, improvingitsabilitytogeneralize.
4.2.2 Validation Set (15%)
Thevalidationdatasetaccountsfor15%ofthedataandis usedtomonitormodelperformanceduringtraining.Itplays acrucialroleinhyperparametertuningandhelpsprevent overfitting by evaluating the model on unseen data. Performancemetricscomputedonthevalidationsetguide theselectionofoptimalmodelconfigurations.
4.2.3
Testing Set (15%)
Thetestingdataset,alsocomprising15%ofthetotaldata,is reservedforfinalevaluationofthemodel.Thisdatasetisnot used during training or tuning, ensuring an unbiased assessment of model performance. The testing phase providesarealisticestimateofhowthemodelwillperform inreal-worldclinicalsettings.
4.3 Baseline Models
4.3.1
NEWS and MEWS
The National Early Warning Score (NEWS) and Modified EarlyWarningScore(MEWS)areusedasbaselinemodels for comparison in general ward settings. These scoring systemsarebasedonpredefinedthresholdsforvitalsigns andgenerateriskscoresaccordingly.Theyserveasstandard clinical benchmarks for evaluating early warning systems andprovideareferencepointforassessingimprovements achievedbytheproposeddeeplearningframework.
4.3.2 SOFA and APACHE II
For critically ill patients, the Sequential Organ Failure Assessment (SOFA) and Acute Physiology and Chronic HealthEvaluationII(APACHEII)scoresareusedasbaseline comparators. These models incorporate multiple physiological andclinical parameters to assessseverity of illnessandpredictmortalityrisk.Includingtheseestablished scoringsystemsenablesacomprehensivecomparisonacross bothgeneralandintensivecaresettings.
4.4 Evaluation Metrics
4.4.1
Area Under the Receiver Operating Characteristic Curve (AUROC)
AUROCisusedtoevaluatetheoveralldiscriminativeability ofthemodel.Itmeasureshowwellthemodeldistinguishes betweenpatientswhoexperienceadverseeventsandthose who do not, across different classification thresholds. A higherAUROCvalueindicatesbettermodelperformancein separatingpositiveandnegativecases.
4.4.2 Area Under the Precision–RecallCurve (AUPRC)
AUPRC is particularly important for imbalanced datasets, where adverse events occur less frequently than normal cases.Thismetricfocusesonthetrade-offbetweenprecision andrecall,providingamoreinformativeevaluationofmodel performanceinidentifyingrarebutcriticalclinicalevents.
4.4.3
Precision, Recall, and F1-Score
Precision measures the proportion of correctly predicted positivecasesamongallpredictedpositives,reflectingthe reliability of alerts generated by the model. Recall, also known as sensitivity, measures the proportion of actual positivecasescorrectlyidentifiedbythemodel,indicatingits abilitytodetecttrueadverseevents.TheF1-score,whichis the harmonic mean of precision and recall, provides a balanced measure of model performance. Together, these metricsofferacomprehensiveevaluationofbothpredictive accuracyandclinicalusefulness.
5. RESULTS
5.1 Predictive Performance
5.1.1
Comparison with Baseline Models
Thepredictiveperformanceoftheproposeddeeplearning framework was evaluated against established clinical scoringsystems,includingNEWS,MEWS,SOFA,andAPACHE II.Thecomparisondemonstratesthattheproposedmodel significantly outperforms traditional rule-based systems acrossmultipleevaluationmetrics.Unlikebaselinemodels thatrelyonstaticthresholds,thedeeplearningframework captures temporal dependencies and complex feature interactions,resultinginimproveddiscriminationbetween high-riskandlow-riskpatients.Thisenhancementhighlights the advantage of data-driven approaches in clinical risk prediction.
5.1.2
AUROC and AUPRC Improvements
ThemodelachievedsuperiorperformanceintermsofArea UndertheReceiverOperatingCharacteristicCurve(AUROC) andAreaUnderthePrecision–RecallCurve(AUPRC).Higher AUROCvaluesindicatebetterclassificationcapabilityacross thresholds, while improved AUPRC values demonstrate

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
enhanced detection of rare adverse events. The gains in AUPRC are particularly important in healthcare settings, wherepositivecasessuchascardiacarrestorICUtransfer arerelativelyinfrequent.

Graph-1: Predictive Performance Comparison
5.2 Multi-Label Prediction Performance
5.2.1 ICU Transfer, Mortality, and CPR Prediction
Theproposedframeworkemploysamulti-labelprediction approachtosimultaneouslyidentifymultipleadverseclinical outcomes, including unplanned ICU transfer, in-hospital mortality,andcardiopulmonaryresuscitation(CPR)events. This approach enables the model to capture interdependencies among different clinical events, improvingoverallpredictiveaccuracy.Resultsindicatethat the model performs consistently well across all target outcomes, demonstrating its capability to provide comprehensiveriskassessment.
Table -1: Multi-Label Prediction Performance
5.3.2 Earlier Risk Detection (Up to 6 Hours)
The model demonstrated the ability to predict adverse events up to 6 hours in advance. This early detection window provides clinicians with sufficient time to initiate preventive measures, thereby reducing the likelihood of criticalevents.ThetemporalmodelingcapabilityoftheBiLSTM architecture plays a key role in identifying early warningpatterns.
5.3.3 Increased Early Interventions
Followingtheimplementationofthepredictivesystem,an increase in early clinical interventions was observed. Healthcare providers were able to respond proactively to risk alerts, leading to improved patient monitoring and management. This shift from reactive to proactive care significantly enhances patient safety and treatment outcomes.
Table -2: Clinical Outcome Improvements
5.3 Clinical Outcome Analysis
5.3.1 Code Blue Reduction (~24.97%)
The deployment of the proposed early warning system resulted in a substantial reduction in Code Blue events. A comparativeanalysisbetweenpre-implementationandpostimplementation periods shows an approximate 24.97% decrease in emergency events. This reduction reflects the effectiveness of the model in enabling timely clinical interventionsandpreventingseverepatientdeterioration.
6. DISCUSSION
6.1 Interpretation of Results
6.1.1
Why Bi-LSTM Performs Better
The superior performance of the proposed Bidirectional Long Short-Term Memory (Bi-LSTM) model can be attributed to its ability to capture complex temporal dependenciesinlongitudinalpatientdata.Unliketraditional machine learning models that treat input variables independently,Bi-LSTMprocessessequential data in both forward and backward directions, enabling it to learn contextual relationshipsacross time.Thisdual-directional processing allows the model to identify subtle changes in physiologicaltrendsthatmayindicateearlystagesofclinical deterioration.Furthermore,thegatingmechanismsinLSTM effectivelyhandlelong-termdependencies,preventingissues such as vanishing gradients and ensuring stable learning acrossextendedtimesequences.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
6.1.2 Importance of Temporal Learning
Temporal learning plays a critical role in clinical risk prediction, as patient conditions evolve dynamically over time. The incorporation of longitudinal data enables the modeltoanalyzenotonlycurrentphysiologicalvaluesbut also their progression and variability. This temporal perspectiveallowsforearlierandmoreaccuratedetectionof adverse events compared to static models. By leveraging time-seriesinformation,theproposedframeworkprovidesa more comprehensive understanding of patient health trajectories, which is essential for proactive clinical intervention.
6.2 Clinical Implications
6.2.1
Improved Patient Monitoring
The implementation of the proposed deep learning framework significantly enhances patient monitoring by enablingcontinuousandautomatedassessment ofclinical risk. Unlike traditional monitoring systems that rely on periodicmeasurements,themodelprocessesreal-timedata and provides ongoing risk evaluations. This continuous monitoring approach helps clinicians identify high-risk patientsearlierandprioritizemedicalattentionaccordingly.
6.2.2
Reduced Emergency Events
Oneofthemostimportantclinicaloutcomesobservedinthis study is the reduction in emergency events, such as Code Blue incidents. By detecting early warning signs of deterioration,thesystemallowsfortimelyinterventionsthat preventtheescalationofpatientconditions.Thisreduction notonlyimprovespatientsurvivalratesbutalsodecreases theburdenonemergencyresponseteamsandcriticalcare resources.
6.2.3
Support for Decision-Making
Theproposedsystemservesasadecision-supporttoolfor healthcareprovidersbydeliveringactionableinsightsbased on patient data. Risk predictions generated by the model assist clinicians in making informed decisions regarding patientmanagement,treatmentadjustments,andresource allocation.ThisintegrationofAI-driveninsightsintoclinical workflowsenhancestheoverallefficiencyandeffectiveness ofhealthcaredelivery.
6.3 Comparison with Existing Systems
6.3.1
Advantages Over Traditional Scores
ComparedtotraditionalearlywarningscoressuchasNEWS and MEWS, the proposed model demonstrates clear advantagesinpredictiveaccuracyandflexibility.Traditional systems rely on fixed thresholds and do not account for temporal patterns or interactions between variables. In contrast, the deep learning framework adapts to complex
datastructuresandprovidespersonalizedriskassessments based on individual patient trajectories. This results in improved sensitivity and specificity in detecting adverse events.
6.3.2 Benefits Over Single-Event AI Models
WhilemanyexistingAI-basedmodelsfocusonpredictinga single outcome, the proposed framework adopts a multilabelapproach,enablingsimultaneouspredictionofmultiple adverse events. This capability reflects the multifactorial natureofclinicaldeteriorationandprovidesamoreholistic viewofpatientrisk.Bycapturinginterdependenciesamong different outcomes, the model enhances predictive performanceandoffersgreaterclinicalutility.
6.4 Challenges
6.4.1 Alert Fatigue
Amajorchallengeassociatedwithautomatedearlywarning systems is alert fatigue, which occurs when clinicians are exposed to excessive or unnecessary alerts. High falsepositiveratescanleadtodesensitization,causingimportant warnings to be overlooked. Therefore, optimizing alert thresholds and ensuring high precision in predictions are essentialtomaintaincliniciantrustandsystemeffectiveness.
7.CONCLUSION
Thisstudypresentsanintelligentdeeplearningframework forearly-stageclinicalriskidentificationusinglongitudinal patientrecords,addressingcriticallimitationsoftraditional earlywarningsystems.ByleveragingaBidirectionalLong Short-TermMemory(Bi-LSTM)architecture,theproposed model effectively captures temporal dependencies and complexrelationshipswithinelectronichealthrecorddata. The integration of multi-label prediction enables simultaneous identification of multiple adverse clinical outcomes,includingICUtransfer,in-hospitalmortality,and cardiopulmonary resuscitation events, providing a comprehensiveassessmentofpatientrisk.
The experimental results demonstrate that the proposed frameworksignificantlyoutperformsconventionalscoring systems such as NEWS, MEWS, SOFA, and APACHE II in terms of AUROC and AUPRC. Furthermore, real-world deployment highlights its practical applicability, with notable improvements in early detection of patient deterioration and a substantial reduction in emergency events,includingCodeBlueincidents.Theabilitytopredict adverse outcomes up to several hours in advance allows clinicians to implement timely interventions, thereby improvingpatientsafetyandclinicaloutcomes.
Overall, this research underscores the potential of deep learning–based early warning systems in transforming healthcare delivery from reactive to proactive care. The proposedframeworknotonlyenhancespredictiveaccuracy

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
but also supports clinical decision-making, making it a valuable tool for modern healthcare systems aiming to improveefficiencyandpatientoutcomes.
7.1.Future Recommendations
Future research should focus on enhancing the scalability andinterpretabilityoftheproposedframeworktofacilitate broaderclinicaladoption.Incorporatingexplainableartificial intelligence techniques can improve transparency and cliniciantrustinmodelpredictions.Additionally,integrating multi-modal data sources, such as medical imaging and clinicalnotes,mayfurtherenhancepredictiveperformance. Expandingvalidationacrossdiversehealthcaresettingsand geographicregionsisessentialtoensuregeneralizabilityand robustness.
Real-timedeploymentandcontinuouslearningmechanisms shouldalsobeexploredtoenableadaptivemodelupdates based on new patient data. Furthermore, addressing challenges related to data quality, interoperability, and system integration will be crucial for successful implementation. These advancements will strengthen the role of intelligent systems in improving patient care and healthcareoutcomes.
REFERENCES
1. Adams, R., Henry, K.E., Sridharan, A., Soleimani, H., Zhan,A.,Rawat,N.,Johnson,L.,Hager,D.N.,Cosgrove, S.E. and Markowski, A., 2022. Prospective, multi-site study of a real-time electronic health record-based early warning system for sepsis. Nature Medicine, 28(2),pp.275–283.
2. Breiman,L.,2001.Randomforests.MachineLearning, 45(1),pp.5–32.
3. Che,Z.,Purushotham,S.,Cho,K.,Sontag,D.andLiu,Y., 2018.Recurrentneuralnetworksformultivariatetime series with missing values. Scientific Reports, 8(1), p.6085.
4. Churpek,M.M.,Yuen,T.C.andEdelson,D.P.,2013.Risk stratification of hospitalized patients on the wards. Chest,143(6),pp.1758–1765.
5. Clifford,G.D.,Clifton,D.A.,Reisner,A.andMoody,G.B., 2012. Advanced methods and tools for ECG data analysis.ArtechHouse.
6. DeVita, M.A., Bellomo, R., Hillman, K., Kellum, J., Rotondi,A.,Teres,D.,Auerbach,A.,Chen,W.J.,Duncan, R. and Kenward, G., 2010. Findings of the first consensus conference on medical emergency teams. CriticalCareMedicine,34(9),pp.2463–2478.
7. Hochreiter, S. and Schmidhuber, J., 1997. Long shortterm memory. Neural Computation, 9(8), pp.1735–1780.
8. Johnson,A.E.W.,Pollard,T.J.,Shen,L.,Lehman,L.W.H., Feng,M.,Ghassemi,M.,Moody,B.,Szolovits,P.,Celi,L.A. and Mark, R.G., 2016. MIMIC-III, a freely accessible criticalcaredatabase.ScientificData,3,p.160035.
9. Kelly,C.J.,Karthikesalingam,A.,Suleyman,M.,Corrado, G. and King, D., 2019. Key challenges for delivering clinical impact with artificial intelligence. BMC Medicine,17(1),p.195.
10. Lipton,Z.C.,Kale,D.C.andWetzel,R.,2016.Learningto diagnosewithLSTMrecurrentneuralnetworks.arXiv preprintarXiv:1511.03677.
11. Obermeyer,Z.andEmanuel,E.J.,2016.Predictingthe future big data, machine learning, and clinical medicine.NewEnglandJournalofMedicine,375(13), pp.1216–1219.
12. Rajkomar, A., Oren, E., Chen, K., Dai, A.M., Hajaj, N., Hardt, M., Liu, P.J., Liu, X., Marcus, J., Sun, M. and Sundberg,P.,2018.Scalableandaccuratedeeplearning with electronic health records. npj Digital Medicine, 1(1),p.18.
13. Royal College of Physicians, 2017. National Early Warning Score (NEWS) 2: Standardising the assessmentofacuteillnessseverityintheNHS.London: RCP.
14. Shashikumar,S.P.,Stanley,M.D.,Sadiq,I.,Li,Q.,Holder, A., Clifford, G.D. and Nemati, S., 2017. Early sepsis detectionincriticalcarepatientsusingmultiscaleblood pressure and heart rate dynamics. Journal of Electrocardiology,50(6),pp.739–743.
15. Smith,G.B.,Prytherch,D.R.,Meredith,P.,Schmidt,P.E. andFeatherstone,P.I.,2013.TheabilityoftheNational EarlyWarningScore(NEWS)todiscriminatepatients at risk of early cardiac arrest. Resuscitation, 84(4), pp.465–470.
16. Topol, E.J., 2019. High-performance medicine: the convergence of human and artificial intelligence. NatureMedicine,25(1),pp.44–56.
17. Tsoumakas, G. and Katakis, I., 2007. Multi-label classification: an overview. International Journal of DataWarehousingandMining,3(3),pp.1–13.
18. Vincent, J.L., Moreno, R., Takala, J., Willatts, S., De Mendonça, A., Bruining, H., Reinhart, C.K., Suter, P.M. andThijs,L.G.,1996.TheSOFA(Sepsis-relatedOrgan Failure Assessment) score to describe organ

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
19. Wiens, J., Saria, S., Sendak, M., Ghassemi, M., Liu, V.X., Doshi-Velez, F., Jung, K., Heller, K., Kale, D., Saeed, M. and Ossorio, P.N., 2019. Do no harm: a roadmap for responsible machine learning for health care. Nature Medicine,25(9),pp.1337–1340.
20. Zhang, M.L. and Zhou, Z.H., 2014. A review on multilabel learning algorithms. IEEE Transactions on Knowledge and Data Engineering, 26(8), pp.1819–1837.
21. Nasarudin, N.A., Al Jasmi, F., Sinnott, R.O., Zaki, N., Al Ashwal,H.andMohamed,E.A.,2024.Areviewofdeep learning models and online healthcare databases for electronic health records and their use for health prediction.ArtificialIntelligenceReview,57,p.249.
22. Solares, J.R.A., Raimondi, F.E., Zhu, Y., Rahimian, F., Canoy, D., Tran, J., Nguyen, H., Hassaine, A., NoriegaCampero,A.,Marlin,B.M.andClifton,D.A.,2020.Deep learningforelectronichealthrecords:Acomparative review. Artificial Intelligence in Medicine, 101, p.101840.
23. Si,Y.,Du,J.,Li,Z.,Jiang,X., Miller,T.,Wang,F.,Zheng, W.J. and Roberts, K., 2021. Deep representation learningofpatientdatafromelectronichealthrecords: Asystematicreview.JournalofBiomedicalInformatics, 115,p.103671.
24. Xie,F.,Yuan,H.,Ning,Y.,Ong,M.E.H.,Feng,M.,Hsu,W., Chakraborty, B. and Liu, N., 2021. Deep learning for temporal data representation in electronic health records. Journal of Biomedical Informatics, 119, p.103839.
25. Shickel,B.,Tighe,P.J.,Bihorac,A.andRashidi,P.,2018. DeepEHR:Asurveyofrecentadvancesindeeplearning techniquesforelectronichealthrecordanalysis.IEEE Journal of Biomedical and Health Informatics, 22(5), pp.1589–1604.
26. Choi,E.,Bahadori,M.T.,Song,L.,Stewart,W.F.andSun, J., 2016. Doctor AI: Predicting clinical events via recurrent neural networks. Machine Learning for HealthcareConference,pp.301–318.
27. Pham, T., Tran, T., Phung, D. and Venkatesh, S., 2017. DeepCare: A deep dynamic memory model for predictive medicine. Pacific-Asia Conference on KnowledgeDiscoveryandDataMining,pp.30–41.
28. Aczon,M.,Ledbetter,D.,Ho,L.,Gunny,A.,Flynn,A.and Wetzel, R., 2017. Dynamic mortality risk predictions using recurrent neural networks. arXiv preprint arXiv:1701.06675.
29. Baytas,I.M.,Xiao,C.,Zhang,X.,Wang,F.,Jain,A.K.and Zhou,J.,2017.Patientsubtypingviatime-awareLSTM networks.KDDConference,pp.65–74.
30. Che,Z.,Purushotham,S.,Cho,K.,Sontag,D.andLiu,Y., 2018.Recurrentneuralnetworksformultivariatetime serieswithmissingvalues.ScientificReports,8,p.6085.
31. Rajkomar, A., Dean, J. and Kohane, I., 2019. Machine learninginmedicine.NewEnglandJournalofMedicine, 380(14),pp.1347–1358.
32. Esteva,A.,Robicquet,A.,Ramsundar,B.,Kuleshov,V., DePristo,M.,Chou,K.,Cui,C.,Corrado,G.,Thrun,S.and Dean,J.,2019.Aguidetodeeplearninginhealthcare. NatureMedicine,25(1),pp.24–29.
33. Miotto,R.,Wang,F.,Wang,S.,Jiang,X.andDudley,J.T., 2018. Deep learning for healthcare: review and opportunities. Briefings in Bioinformatics, 19(6), pp.1236–1246.
34. Harutyunyan,H.,Khachatrian,H.,Kale,D.C.,VerSteeg, G. and Galstyan, A., 2019. Multitask learning and benchmarkingwithclinicaltimeseriesdata.Scientific Data,6(1),p.96.
35. Purushotham, S., Meng, C., Che, Z. and Liu, Y., 2018. Benchmarking deep learning models on large healthcaredatasets.JournalofBiomedicalInformatics, 83,pp.112–134.
36. Futoma,J.,Morris,J.andLucas,J.,2020.Acomparison of models for predicting early hospital readmissions. JournalofBiomedicalInformatics,56,pp.229–238.
37. Rasmy,L.,Wu,Y.,Wang,N.,Wang,J.,Zhi,D.,Wu,Y.and Zhi, D., 2018. Med-BERT: pretrained contextualized embeddings on large-scale structured EHR data. npj DigitalMedicine,4(1),p.86.
38. Li,Y.,Rao,S.,Solares,J.R.A.,Hassaine,A.,Ramakrishnan, R.,Canoy,D.,Zhu,Y.,Rahimi,K.andSalimi-Khorshidi, G., 2020. BEHRT: Transformer for electronic health records.ScientificReports,10(1),p.7155.
39. Zhang,D.,Yin,C.,Zeng,J.,Yuan,X.andZhang,P.,2020. Combining structured and unstructured data for predictive models. BMC Medical Informatics and DecisionMaking,20(1),p.280.
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072 © 2026, IRJET | Impact Factor value: 8.226 | ISO 9001:2008 Certified Journal | Page160 dysfunction/failure. Intensive Care Medicine, 22(7), pp.707–710.