
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Vivek Shukla1 , Dr. J.B. Singh2, Mr.
Rajesh Kumar Sharma
3
1Master of Technology, Computer Science and Engineering, Sagar Institute of Technology and Management, Barabanki, India
2Professor, Department of Computer Science and Engineering, Sagar Institute of Technology and Management, Barabanki, India
3Assistant Professor, Department of Computer Science and Engineering, Sagar Institute of Technology and Management, Barabanki, India ***
Abstract -Serverlesscomputinghasemergedasadominant paradigm for deploying scalable Machine Learning (ML) inference workloads due to its elasticity, fine-grained billing, and operational simplicity. However, ML inference in Function-as-a-Service (FaaS) environments introduces a fundamental trade-off between operational cost and service latency. While aggressive resource provisioning reduces cold start delays and response time, it increases execution cost; conversely,costminimizationstrategiesoftendegradeQuality of Service (QoS). This review systematically analyzes existing literatureoncost–latencytrade-offmodelingforserverlessML inference, with particular emphasis on the integration of predictive workload forecasting mechanisms. The paper categorizespriorstudiesinto analyticalmodeling,simulationbased approaches, multi-objective optimization frameworks, and learning-driven adaptive techniques. It further examines forecasting methodologies including statistical time-series models, classical machine learning predictors, and deep learning architectures and evaluates their impact on proactive resource allocation. A comparative synthesis of evaluation metrics, benchmark datasets, and modeling assumptionsisprovidedtoidentifymethodologicaltrendsand research gaps. The review highlights limitations in handling burstyworkloads,coldstartvariability,andreal-timeadaptive scaling. Finally,itoutlinesopenresearchchallengesandfuture directions for robust, cost-efficient, and latency-aware serverless ML inference systems.
Key Words: Serverless Computing; ML Inference; Cost–Latency Trade-Off; Predictive Workload Forecasting; PerformanceModeling;Auto-ScalingOptimization
1.1.1
Serverlesscomputing,particularlytheFunction-as-a-Service (FaaS) model, has transformed cloud-native application deploymentbyabstractinginfrastructuremanagementand enabling fine-grained, event-driven execution. Platforms such as Amazon Web Services (AWS Lambda), Microsoft Azure(Azure Functions),andGoogleCloud(GoogleCloud
Functions) provide automatic scaling and pay-per-use billing, which reduce operational overhead and improve elasticity. Unlike traditional virtual machine or containerbased deployments, serverless environments allocate resourcesdynamicallyperinvocation,chargingusersbased onexecutiondurationandmemoryallocation.Thisparadigm has gained widespread adoption due to its economic efficiencyandoperationalagility(Baldinietal.,2017;Jonas et al., 2019). Nevertheless, serverless systems introduce performance uncertainties, particularly under bursty and latency-sensitiveworkloads.
MachineLearning(ML)inferenceconstitutesthereal-time deployment phase of trained models in production environments, supporting applications such as recommendation systems, fraud detection, intelligent assistants, and computer vision services. As ML-driven servicesincreasinglyoperateunderstringentServiceLevel Agreements(SLAs),lowlatencyandhighavailabilitybecome critical requirements (Zaharia et al., 2018). Serverless platformsareattractiveforMLinferencebecausetheyscale automaticallywithdemand,eliminatingidleresourcecosts during low-traffic periods. However, inference workloads areoftencompute-intensiveandmemory-sensitive,making their execution behavior highly dependent on runtime configurationandresourceprovisioningstrategies(Wanget al.,2018).
Afundamental challengein serverlessML inferenceis the cost–latency trade-off. Cold starts caused by container initialization delays can significantly increase response time,particularlyforsporadicworkloads(Wangetal.,2018). Mitigatingcoldstartsthroughprovisionedconcurrencyor over-provisioningreduceslatencybutincreasesoperational cost.Conversely,minimizingallocatedmemoryorcompute resourcesreducesbillingchargesbutmayleadtoexecution slowdown and SLA violations. This inherent tension motivatesthedevelopmentofformalcost–latencymodeling

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
frameworkscapableofbalancingeconomicandperformance objectives(Jonasetal.,2019).
The cost–latency trade-off refers to the conflicting relationshipbetweenminimizingmonetaryexpenditureand achieving low response time in cloud-based execution environments. In serverless systems, cost is typically determined by invocation count, execution duration, and memory allocation, while latency encompasses cold start delay,queueingtime,andexecutiontime.Analyticalmodels basedonqueueingtheoryandperformanceprofilinghave demonstrated that reducing latency often requires prewarmedinstancesorincreasedresourceallocation,thereby increasing cost (Baldini et al., 2017). Effective trade-off modeling seeks Pareto-optimal configurations that satisfy QoSconstraintswithoutexcessivespending.
Workloadvariabilitysignificantlycomplicatescost–latency optimization. Serverless workloads often exhibit diurnal patterns, bursty spikes, and unpredictable demand fluctuations. Reactive auto-scaling mechanisms may lag behind sudden surges, causing latency degradation. Furthermore,cloudpricingmodels basedonper-request billing and memory-time combinations introduce nonlinearcostbehavior.Forinstance,allocatingmorememory may reduce execution time sufficiently to offset increased per-millisecondcharges,resultinginlowertotalcostunder certain conditions. Therefore, modeling approaches must incorporateworkloadstochasticityandpricinggranularity toproducerealisticoptimizationoutcomes(Shahradetal., 2020).
Predictiveworkloadforecastingenablesproactiveresource provisioning, thereby reducing cold start frequency and queueing delays. Time-series and machine learning-based predictorscanestimateshort-termrequestrates,allowing systemstopre-scaleresourcesbeforedemandsurgesoccur Prior studies demonstrate that forecasting-driven scaling policiesoutperformpurelyreactivemechanismsinlatencysensitive applications (Islam et al., 2012). However, forecasting accuracy directly influences optimization effectiveness; inaccurate predictions may lead to overprovisioning or SLA violations. Consequently, integrating forecasting models with cost–latency optimization frameworksiscentral toachieving efficientserverless ML inference.
Thisreviewsystematicallysynthesizesliteratureonserver lessperformancemodeling,workloadforecasting,andmultiobjective optimization. Rather than proposing a new algorithm, the objective is to consolidate fragmented researchcontributionsintoaunifiedconceptualframework. The survey examines analytical, simulation-based, and learning-driven modeling techniques across diverse applicationscenarios.
Akeyobjectiveistocomparativelyevaluatestatisticaltimeseries models, classical machine learning regressors, and deeplearningarchitecturesusedforworkloadforecasting. Differences in scalability, interpretability, computational overhead,andpredictionaccuracyarecriticallyanalyzedto assesstheirsuitabilityforreal-timeinferencesystems.
The review further examines cost–latency trade-off modelingstrategies,includingqueueing-basedformulations, multi-objective optimization methods, and reinforcement learning-based adaptive policies. Their strengths, assumptions, and empirical validation practices are contrastedtoidentifymethodologicaltrends.
By analyzing current limitations such as inadequate modeling of cold start variability, limited benchmark standardization,andinsufficientintegrationofforecasting uncertainty the review highlights promising research avenuesforimprovingrobustnessandscalability.
1.4.1
The review focuses on peer-reviewed journal articles and conference proceedings addressing serverless computing, ML inference performance optimization, workload forecasting, and cost modeling. Studies unrelated to serverless environments or purely training-phase ML optimization are excluded. Emphasis is placed on works presenting formal modeling, empirical evaluation, or optimizationframeworks.
Thesurvey primarilycovers literature published between 2016and2025,correspondingtotherapidadoptionphase ofserverlessplatformsandthematurationofcloud-native MLinferenceframeworks.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Theprincipalcontributionsofthisreviewarethreefold:(i)a structuredtaxonomyofforecastingandtrade-offmodeling approaches, (ii) a comparative synthesis of evaluation methodologies and datasets, and (iii) identification of research gaps concerning proactive scaling, uncertainty modeling, and cost-aware adaptive optimization. By integrating perspectives from cloud performance engineering and ML workload management, this review provides a comprehensive foundation for future advancementsincost-efficient,latency-awareserverlessML inferencesystems.
This section establishes the conceptual and architectural foundationsnecessarytounderstandcost–latencytrade-off modelinginserverlessMLinferenceenvironments.
2.1
2.1.1 Function-as-a-Service
Serverlesscomputingisacloudexecutionparadigminwhich infrastructuremanagementisabstractedfromdevelopers, enablingevent-drivenexecutionofstatelessfunctions.The dominantoperationalmodelisFunction-as-a-Service(FaaS), whereapplicationsaredecomposedintodiscretefunctions triggeredbyeventssuchasHTTPrequestsormessagequeue updates. Major cloud providers including Amazon Web Services(AWSLambda),MicrosoftAzure(AzureFunctions), andGoogleCloud(GoogleCloudFunctions)implementFaaS with automatic resource provisioning and fine-grained billing. FaaS platforms dynamically allocate compute containersperinvocation,offeringelasticitywithoutexplicit capacityplanning.Priorstudiesemphasizethatthismodel reducesoperationalcomplexitybutintroducesperformance variabilityduetoruntimeisolationandcontainerlifecycle management(Baldinietal.,2017;Jonasetal.,2019).
Adefiningcharacteristicofserverlessplatformsisautomatic scalingbasedonincomingrequestvolume.Scalingdecisions are typically reactive, triggered by concurrency demand rather than predictive provisioning. Billing follows a payper-use structure, calculated as a function of invocation count, allocated memory size, and execution duration measuredinmilliseconds.Thisfine-grainedpricingstructure differentiatesserverlesscomputingfromvirtualmachineor container-based models, where users pay for reserved capacityregardlessofutilization(Shahradetal.,2020).The economicefficiencyofserverlesssystemsisthereforehighly sensitivetoworkloadpatternsandexecutionconfiguration.
Coldstartsoccurwhenafunctionisinvokedafteraperiodof inactivityandtheplatformmustinitializeanewexecution environment.Thisinitializationincludescontainercreation, runtime loading, and dependency setup, leading to additional latency before execution begins. Empirical evaluations demonstrate that cold start overhead varies depending on runtime language, memory allocation, and provider-specific orchestration mechanisms (Wang et al., 2018).Scalingdelaysmayalsoarisewhenrapidworkload surgesexceedavailablewarminstances,therebyincreasing tail latency.These phenomena are central to performance modelinginserverlessMLinference.

Figure-1: Cold vs Warm Start Latency Distribution
2.2 ML Inference Workloads
2.2.1
Machine Learning inference workloads can be broadly categorized into batch inference and real-time (online) inference. Batch inference processes large datasets asynchronously and is typically less sensitive to latency constraints. In contrast, real-time inference supports interactive applications such as recommendation engines, fraud detection systems, and conversational AI, where responsetimesmustsatisfystrictSLArequirements.Realtime inference workloads are more suitable for FaaS deploymentduetotheir event-driven nature,though they are more vulnerable to cold start delays and scaling inefficiencies(Crankshawetal.,2017).
Inference workloads exhibit heterogeneous resource demands depending on model complexity, input size, and framework dependencies. Deep neural networks may require substantial memory and CPU allocation, while

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
lightweightmodelsoperatewithinminimalconfigurations. Inserverlesscontexts,resourceallocationdirectlyinfluences execution time due to proportional CPU scaling with memorysize.Consequently,selectingappropriatememory configurations becomes a joint performance–cost optimization problem. Studies indicate that underprovisioning increases execution duration and queueing delays, whereas over-provisioning inflates cost without proportionallatencygains(Jonasetal.,2019).
Quality of Service in ML inference is primarily measured throughresponsetimepercentiles(e.g.,p95orp99latency), throughput,andavailability.Forproduction-gradesystems, SLA compliance requires maintaining latency within predefined thresholds under fluctuating workloads. Serverless architectures complicate QoS assurance due to dynamic scaling and shared infrastructure. Therefore, workload-aware provisioning and proactive scaling strategiesareessentialforsustainingstableperformancein latency-sensitiveinferenceservices(Herbstetal.,2013).
Predictiveworkloadforecastingconstitutesafoundational componentofproactiveresourcemanagementinserverless ML inference systems. Unlike reactive auto-scaling mechanisms, forecasting-driven strategies anticipate demand fluctuations and provision resources accordingly. This section presents a structured review of forecasting methodologies,evaluationpractices,andopenchallenges.
3.1.1
Provisioning
Serverlessplatformsscalefunctionsdynamicallybasedon incomingrequestrates;however,reactivescalingintroduces latencypenaltiesduringabruptworkloadsurges.Predictive workloadforecastingmitigatesthislimitationbyestimating short-termarrivalratesandtriggeringpre-warmingorprescaling policies before demand peaks occur. Empirical studies on cloud elasticity demonstrate that predictionbased provisioning reduces SLA violations compared to threshold-based scaling (Islam et al., 2012). In latencysensitiveMLinferenceworkloads,suchproactivescalingis particularlycriticalbecausecoldstartdelaysandqueueing effectsdisproportionatelyaffecttaillatencypercentiles.
Forecastingaccuracydirectlyinfluencescost–latencytradeoff modeling. Overestimation of demand results in unnecessaryinstancepre-allocationandhigheroperational cost, whereas underestimation leads to cold starts and
increasedqueueingdelay.Researchonserverlessworkload characterization shows that traffic patterns often exhibit diurnal cycles combined with bursty spikes, requiring forecastingmodelscapableofcapturingbothseasonalityand abrupt variations (Shahrad et al., 2020). Consequently, predictivemodelingisnotmerelyanauxiliarycomponent but an integral input to multi-objective optimization frameworksforserverlessMLinference.
3.2.1
AutoRegressiveIntegratedMovingAverage(ARIMA)models capturetemporalautocorrelationandtrendcomponentsin sequential data. Seasonal ARIMA (SARIMA) extends this framework by modeling periodic fluctuations, which are common in cloud traffic traces. These models have been appliedtocloudworkloadpredictionwithmoderatesuccess incapturingstructuredtemporalpatterns(Boxetal.,2015). However,theirlinearassumptionslimitperformanceunder highly non-linear or irregular workloads. Exponentialsmoothingtechniques,includingHolt–Winters methods,providelightweightforecastingmechanismsthat emphasize recent observations. They are computationally efficient and suitable for short-term prediction in stable workloads. Nevertheless, their responsiveness to sudden bursts is constrained, reducing effectiveness in highly dynamic serverless environments. Time-seriesmethodsoffertransparency,fasttraining,and minimal parameter tuning. However, they struggle with complex non-linear relationships and multi-dimensional contextual features such as concurrency limits or application-levelevents.
Linear and polynomial regression techniques incorporate additional explanatory variables such as time-of-day indicatorsoruseractivitymetrics.Whileinterpretable,their predictive power diminishes when workload patterns exhibit high stochasticity. Tree-based models, including random forests, improve predictiverobustnessbyaggregatingmultipledecisiontrees. They capture non-linear relationships and interactions without requiring strict distributional assumptions. Comparative studies indicate that ensemble methods outperformlinearbaselinesincloudworkloadforecasting scenarios (Chen et al., 2018). SVRapplieskernel-basedlearningtoapproximatecomplex functionalrelationshipswhilecontrollingoverfittingthrough margin maximization. It performs well in moderate-sized datasetsbutmayincurscalabilitylimitationsunderlargescalecloudtraceanalysis.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
3.3.1
Forecasting performance is typically evaluated using statistical error measures such as Mean Absolute Error (MAE),RootMeanSquareError(RMSE),andMeanAbsolute Percentage Error (MAPE). MAE provides robustness to outliers,RMSEpenalizeslargedeviations,andMAPEenables scale-independentcomparisonacrossworkloads.Selection of evaluation metrics significantly influences comparative interpretationofmodelperformance.
3.3.2
Benchmarkingoftenreliesonpubliclyreleasedcloudtraces, including the Google cluster workload dataset and Azure function invocation traces. These datasets capture realworld traffic variability and resource utilization patterns. However, trace heterogeneity and limited labeling of contextual variables complicate standardized evaluation acrossstudies.
3.4
Acomparativesynthesisofforecastingapproachesreveals trade-offsacrossmultipledimensions:predictionaccuracy, computationalcomplexity,scalability,interpretability,and adaptability to bursty workloads. Statistical models offer transparency and low overhead but limited non-linear modeling capability. Machine learning methods improve accuracyundermoderatecomplexity,whiledeeplearning architecturesachievesuperiorperformanceattheexpense oftrainingcostandoperationaloverhead.
Inreviewarticles,suchcomparisonistypicallysummarized intabulatedform,detailing:(i)modeltype,(ii)datasetused, (iii)evaluationmetrics,(iv)workloadcharacteristics,and(v) reported performance outcomes. This structured comparison facilitates identification of methodological convergenceanddivergenceacrossstudies.

Figure-2: Forecasting Model Accuracy Comparison
Serverless workloads frequently exhibit flash crowds and irregularspikes.Manyforecastingmodels,particularlylinear time-seriestechniques,struggletocapturesuddendemand surges. Hybrid approaches integrating anomaly detection withforecastingremainunderexplored.
Most forecasting studies focus solely on request rate predictionwithoutexplicitlymodelingcoldstartprobability orcontainerreusedynamics.Incorporatingplatform-level runtimebehaviorintoforecastingframeworksisessential foraccuratelatency-awarescaling.
Model generalizability across applications and cloud providers remains limited. Forecasting models trained on onedatasetoftenunderperformwhenappliedtodifferent workloaddistributions.Transferlearningandmeta-learning techniques present promising directions for improving cross-domainadaptability.
Cost–latency trade-off modeling constitutes the central analytical dimension of serverless ML inference optimization. This section critically reviews modeling paradigms used to balance economic efficiency and performance guarantees in Function-as-a-Service environments.
InserverlessplatformssuchasAmazonWebServicesand MicrosoftAzure,costisprimarilydeterminedbyinvocation count,executionduration,andallocatedmemory.Increasing memory allocation typically improves CPU share and reducesexecutiontime,therebyloweringlatency.However, higher memory configurations increase per-millisecond billing rates. Similarly, enabling provisioned concurrency mitigates cold starts but introduces fixed baseline costs. Thesestructuralcharacteristicscreateaninherenttension betweenminimizingoperationalexpenditureandsatisfying ServiceLevelAgreement(SLA)latencyconstraints(Jonaset al.,2019).Consequently,trade-offmodelingseekstoidentify configurations that achieve acceptable latency at minimal cost.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

The primary goals of cost–latency modeling include: (i) estimating response time distributions under stochastic arrival rates, (ii)quantifyingcostimplicationsof resource configurations, and (iii) identifying Pareto-efficient deployment strategies. Models must capture cold start dynamics, concurrency limits, and workload variability to produceactionableinsights(Baldinietal.,2017).Accuracy, scalability,andinterpretabilityarethereforekeyevaluation criteriaformodelingtechniques.
4.2.1
Queueingtheoryprovidesaformalframeworkformodeling request arrivals, service rates, and waiting times in serverless systems. Common formulations include M/M/c and M/G/1 models, where arrivals are assumed to follow Poissonprocessesandservicetimesfollowexponentialor general distributions. These models estimate expected responsetimeandqueuelengthundervaryingconcurrency levels.Researchoncloudelasticitymodelingdemonstrates thatqueueingabstractionseffectivelyapproximatesystem behavior under steady-state assumptions (Herbst et al., 2013). However, serverless environments complicate classicalassumptionsduetodynamiccontainerinitialization andcoldstartvariability.
4.2.2
Closed-form analytical models attempt to derive explicit mathematicalexpressionsforlatencyandcostasfunctions of memory size, concurrency level, and arrival rate. Such modelsenablerapidevaluationofconfigurationalternatives without simulation overhead. Some studies integrate empiricalprofilingdataintosemi-analyticalformulationsto approximate cold start probabilities and execution-time scaling behavior (Wang et al., 2018). Although
computationally efficient, these models often rely on simplifying assumptions such as stationary workload distributionsandhomogeneousservicetimes.
4.2.3
Existing analytical frameworks frequently assume independencebetweenrequestsandneglectplatform-level resource contention. While these simplifications improve tractability, they limit applicability in highly bursty workloads. Moreover, most closed-form models focus on mean latency rather than tail latency (e.g., p95 or p99), which is critical for SLA compliance in ML inference applications.
4.3.1
Discrete event simulation (DES) models system state transitions based on event-driven interactions such as request arrivals, container initialization, and completion events. DES allows fine-grained modeling of cold starts, scaling thresholds, and concurrency limits without restrictiveanalyticalassumptions.Simulationenvironments can emulate realistic workload traces to estimate latency distributionsundervaryingscalingpolicies(Shahradetal., 2020). While highly flexible, DES incurs higher computationalcostcomparedtoanalyticalmodels.
4.3.2
Workloadreplayframeworksusehistoricalcloudtracesto evaluate scaling policies under realistic conditions. By replayinginvocationsequences,researcherscanassessthe impact of proactive provisioning strategies on cost and latency.Suchframeworksprovideempiricalvalidationbut depend heavily on trace representativeness and may lack generalizability.
Simulation-basedapproachescapturenon-linearbehaviors, burstpatterns,andcoldstartdynamicsmoreaccuratelythan simplifiedanalyticalmodels.However,theyrequiredetailed parameterization and may not scale efficiently for large configuration spaces. Additionally, simulation outputs are scenario-dependent,limitingtheoreticalgeneralization.
The integration of predictive workload forecasting with cost–latency trade-off modeling represents a critical advancement in serverless ML inference optimization. Ratherthantreatingpredictionandresourceallocationas isolated processes, recent research emphasizes their interdependence in achieving economically efficient and SLA-compliantdeployments.

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
5.1.1
Reactive auto-scaling mechanisms allocate resources only after workloadsurgesaredetected, whichoftenresultsin coldstartpenaltiesandtransientqueuebuildup.Incontrast, proactivescalingleveragesshort-termworkloadforecaststo pre-warmexecutionenvironmentsandadjustconcurrency levels before demand peaks materialize. Empirical cloud elasticitystudiesdemonstratethatpredictiveprovisioning reduces response time variability compared to thresholdbasedscaling(Islametal.,2012).InserverlessMLinference, where tail latency directly affects user experience and contractualobligations,integratingforecastingintoscaling decisions enables more stable performance. Furthermore, predictive models allow optimization frameworks to evaluate anticipated workload states rather than current instantaneous metrics, thereby improving long-term resourceplanning.
Service Level Agreements (SLAs) often specify latency thresholds at high percentiles (e.g., p95 or p99). Reactive strategiesfrequentlyviolatetheseconstraintsduringabrupt workload spikes due to initialization delays. Integrated forecasting–optimizationframeworksreducesuchviolations by aligning resource allocation with predicted demand distributions.Studiesanalyzinglarge-scaleserverlesstraces indicate that SLA compliance improves when forecastinformed provisioning strategies are applied, particularly under diurnal or seasonal traffic patterns (Shahrad et al., 2020).Thus,integrationenhancesbotheconomicefficiency andreliabilityguarantees.
5.2.1
Recentframeworkscombinetime-seriespredictionmodels withoptimizationenginesthatcomputecost-efficientscaling actions.Thesearchitecturestypicallyoperateintwostages: (i) short-horizon workload prediction, and (ii) resource configuration optimization subject to latency constraints. Someapproachesembedpredictionoutputsintoqueueingbased analytical models to estimate future latency distributions before enacting scaling decisions. Others integrate forecasting directly into multi-objective optimizationsolvers,enablingdynamicexplorationofcost–latencytrade-offsunderpredicteddemandscenarios(Jonas etal.,2019).
5.2.2
Hybridsystemsincorporatingdeeplearningpredictorswith adaptive scaling policies have been proposed for cloud resourcemanagement.Forinstance,reinforcementlearning frameworksaugmentedwithpredictiveinputshaveshown
improved convergence stability and reduced exploration costcomparedtopurelyreactiveagents(Maoetal.,2016). Similarly, workload-aware auto scaling mechanisms leveraging trace-driven predictions have demonstrated measurable reductions in cold start frequency and operationalexpenditureinserverlessenvironments(Wang etal.,2018).Theseexamplesillustratethepracticalviability ofintegratingpredictiveanalyticswithtrade-offmodeling.
5.3.1
The effectiveness of integrated frameworks is highly sensitive to forecasting accuracy. Prediction errors propagateintoscalingdecisions,potentiallycausingoverprovisioning(increasedcost)orunder-provisioning(latency violations).Sensitivityanalysesinpriorresearchrevealthat even small deviations in arrival rate prediction can significantlyshiftPareto-optimalconfigurations.Therefore, uncertainty-awareoptimizationmethods suchasrobustor stochasticprogramming areincreasinglyrecommendedto mitigatetheimpactofforecastvariance(Herbstetal.,2013). Incorporatingpredictionconfidenceintervalsintodecisionmakingremainsanopenresearcharea.
5.3.2
Despite conceptual advantages, integrated forecasting–optimizationframeworksfaceoperationalchallenges.Realtimeinferencesystemsrequirelow-latencyscalingdecisions, limiting the computational complexity of prediction and optimization modules. Deep learning-based forecasting modelsmayintroduceadditionalruntimeoverhead,while reinforcement learning agents require training data and exploration phases that risk SLA violations. Furthermore, platform-specificconstraintsandopaquescalingpoliciesin commercial cloud environments restrict full control over resource allocation parameters. These challenges underscoretheneedforlightweight,adaptive,andplatformawareintegrationmechanisms.
This section examines how serverless ML inference and cost–latencytrade-offstrategiesaredeployedinreal-world systems,highlightingcloudplatformsupport,casestudies, andadoptiontrends.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
6.1.1 AWS Lambda, Google Cloud Functions, Azure Functions
LeadingpubliccloudprovidersofferFunction-as-a-Service (FaaS)environmentsoptimizedforevent-drivenworkloads andscalablemicroservices.AmazonWebServicesprovides AWS Lambda, which supports various languages and integrateswithservicessuchasAmazonAPIGatewayand AWS Sagemaker for ML inference. Google Cloud offers GoogleCloudFunctions,tightlycoupledwithCloudRunand AutoML for scalable workloads, while Microsoft Azure provides Azure Functions, which integrates with Azure Machine Learning services. These platforms enable developers to deploy trained ML models as serverless endpoints with automatic scaling. Although economic and operationaladvantagesaresubstantial,inherentcoldstart delays, granular billing, and configuration trade-offs pose challengesforlatency-sensitiveinferenceapplications(Jonas et al., 2019). Performance optimization in such environmentsdemandscarefulcost–latencymodelingand workload-awareprovisioningstrategies.
Real-worlddeploymentsofserverlessMLinferenceillustrate both benefits and limitations of current approaches. For example,e-commercerecommendationsystemsoftenadopt serverlessinferencetohandlevariabledemandduringpeak eventssuchasBlackFriday,wheresuddensurgesdemand elastic scaling without upfront capacity provisioning. In a case study examining web-based image classification services, researchers demonstrate that proactive scaling informed by short-term prediction significantly reduced 99th percentile latency while controlling cost budget (Shahradetal.,2020).Similarly,conversationalAIplatforms leveragingserverlessfunctionsshowimprovedoperational efficiencywhenforecasting-basedautoscalingpreventscold startpenaltiesduringpeakusagehours.
These case studies typically emphasize the importance of telemetry collection, application-specific performance profiling, and integration with platform monitoring tools (e.g., AWS CloudWatch, Azure Monitor). The practical implicationisthatperformancemodelsmustaccommodate platform-specificbehaviorsandworkloadidiosyncrasiesto remaineffective.
IndustryadoptionofserverlessMLinferencehasgrownin domainssuchasIoTanalytics,real-timepersonalization,and event-drivenautomationduetotheeconomicadvantagesof pay-per-use models and the reduction of DevOps burden. Surveys of cloud-native application practices indicate increasing preference for FaaS deployment patterns in microservices architectures, particularly for lightweight
inferencetasksthatexhibitburstytrafficpatterns.However, guideline reports also reveal that organizations often combineserverlesswithcontainer-basedorchestration(e.g., Kubernetes) to manage latency-critical components that requirefinercontrol overresourceallocation.Thishybrid adoption trend underscores the need for flexible performanceandcostmodelingframeworks.
Evaluation of serverless ML inference systems and cost–latency trade-off models depends on well-defined metrics that capture economic and performance attributes under varyingworkloadconditions.
Cloud providers charge serverless functions based on invocationcount,executionduration,andallocatedmemory. Total cost often appears on monthly bills as aggregated chargesforcomputetimeandrelatedservices.Atthemicrolevel,researcherscomputecostperrequestastheproductof executiontimeandmemorypriceforeachinvocation.This fine-grained metric allows comparison of different configurations(e.g.,memorysizesorconcurrencysettings) inexperimentalevaluations.Costmetricsmustalsoaccount forauxiliarychargessuchasdataegress,networkusage,and provisioned concurrency overhead when analyzing economicefficiency.
Thisreviewsystematicallyexaminedcost–latencytrade-off modeling for serverless Machine Learning (ML) inference withanemphasisonpredictiveworkloadforecasting.The analysis demonstrates that serverless computing offers compellingadvantagesinelasticity,operationalsimplicity, and fine-grained billing; however, these benefits are accompanied by inherent performance uncertainties, particularlyunderburstyandlatency-sensitiveworkloads. The literature reveals that cost and latency form a structurallyconflictingobjectivepair,influencedbymemory allocationpolicies,coldstartdynamics,concurrencylimits, andstochasticrequestarrivals.
Forecastingtechniques rangingfromclassicaltime-series modelstodeeplearningarchitectures playacriticalrolein enabling proactive scaling strategies. Comparative evaluation indicates that while analytical models provide tractable approximations for steady-state analysis, simulation-basedandoptimization-drivenapproachesoffer greater realism under dynamic conditions. Reinforcement learning methods further enhance adaptability in nonstationaryenvironmentsbutintroducecomputationaland explorationoverhead.Importantly,thequalityofworkload

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
prediction significantly affects the shape and stability of Pareto-optimaltrade-offconfigurations.
Overall, the review highlights that integrating predictive forecasting with cost-aware performance modeling is essential for achieving SLA compliance and economic efficiency in serverless ML inference systems. Future researchshouldprioritizeuncertainty-awareoptimization, cold start modeling improvements, standardized benchmarking, and lightweight adaptive frameworks to supportreal-time,production-gradedeployments.
Despite comprehensive synthesis, this review has several limitations. First, it focuses primarily on peer-reviewed literature published between 2016 and 2025, potentially omitting emerging preprints and proprietary industrial solutions. Second, comparative analysis is constrained by heterogeneousexperimentalsetups,datasets,andevaluation metricsreportedacrossstudies,limitingdirectquantitative benchmarking.Third,mostreviewedworksemphasizeCPUbasedserverlessenvironments,whileGPU-enabledoredgeserverless platforms remain underexplored. Additionally, practical deployment considerations such as providerspecific throttling policies and undocumented scaling heuristics are difficult to generalize due to limited transparencyfromcloudvendors.Finally,whileforecasting and trade-off integration are discussed conceptually, empiricalvalidationofunifiedframeworksremainslimited inpubliclyavailableresearch.
1. Baldini, I., Castro, P., Chang, K., Cheng, P., Fink, S., Ishakian, V., Mitchell, N., Muthusamy, V., Rabbah, R., Suter, P. & Tardieu, O., 2017. Serverless computing: Currenttrendsandopenproblems.ResearchAdvances inCloudComputing,pp.1–20.
2. Box, G.E.P., Jenkins, G.M., Reinsel, G.C. & Ljung, G.M., 2015. Time Series Analysis: Forecasting and Control. 5thed.Hoboken:Wiley.
3. Chen,X.,Lin,X.,Chen,Y.&Zomaya,A.Y.,2018.Machine learning-based workload prediction for cloud computing. Proceedings of the IEEE International ConferenceonCloudComputing,pp.1–8.
4. Crankshaw, D., Wang, X., Zhou, G., Franklin, M.J., Gonzalez,J.E.&Stoica,I.,2017.Clipper:Alow-latency online prediction serving system. Proceedings of the USENIXSymposiumonNetworkedSystemsDesignand Implementation(NSDI),pp.613–627.
5. Herbst,N.R.,Kounev,S.&Reussner,R.,2013.Elasticity in cloud computing: What it is, and what it is not. Proceedings of the International Conference on AutonomicComputing(ICAC),pp.23–27.
6. Hochreiter,S.&Schmidhuber,J.,1997.Longshort-term memory.NeuralComputation,9(8),pp.1735–1780.
7. Islam, S., Keung, J., Lee, K. & Liu, A., 2012. Empirical predictionmodelsforadaptiveresourceprovisioningin thecloud.FutureGenerationComputerSystems,28(1), pp.155–162.
8. Jonas, E., Schleier-Smith, J., Sreekanti, V., Tsai, C., Khandelwal,A.,Pu,Q.,Shankar,V.,Carreira,J.,Krauth, K.,Yadwadkar,N.&Stoica,I.,2019.Cloudprogramming simplified:ABerkeleyview onserverlesscomputing. arXivpreprintarXiv:1902.03383.
9. Mao,H.,Alizadeh,M.,Menache,I.&Kandula,S.,2016. Resource management with deep reinforcement learning. Proceedings of the ACM Workshop on Hot TopicsinNetworks(HotNets),pp.50–56.
10. Shahrad,M.,Fonseca,P.,Goiri,Í.,Chaudhry,G.,Batum, P.,Cooke,J.,Laureano,E.,Faria,C.,Belay,A.,Bianchini, R. & Thereska, E., 2020. Serverless in the wild: Characterizingandoptimizingtheserverlessworkload at a large cloud provider. Proceedings of the USENIX AnnualTechnicalConference(ATC),pp.205–218.
11. Vaswani,A.,Shazeer,N.,Parmar,N.,Uszkoreit,J.,Jones, L., Gomez, A.N., Kaiser, Ł. & Polosukhin, I., 2017. Attention is all you need. Advances in Neural Information Processing Systems (NeurIPS), 30, pp.5998–6008.
12. Wang, L., Li, M., Zhang, Y., Ristenpart, T. & Swift, M., 2018. Peeking behind the curtains of serverless platforms.ProceedingsoftheUSENIXAnnualTechnical Conference(ATC),pp.133–146.
13. Zaharia,M.,Chen,A.,Davidson,A.,Ghodsi,A.,Hong,S.A., Konwinski, A., Murching, S., Nykodym, T., Ogilvie, P., Parkhe, M. & Xie, F., 2018. Accelerating the machine learninglifecyclewithMLflow.IEEEDataEngineering Bulletin,41(4),pp.39–45
14. Mao,H.,Alizadeh,M.,Menache,I.andKandula,S.(2016) ‘Resource management with deep reinforcement learning’,Proceedingsofthe 15thACMWorkshopon HotTopicsinNetworks(HotNets),pp.50–56.
15. Xu, J., Tang, J., Liu, Y., Su, J. and Li, Y. (2017) ‘Online learning for offloading and autoscaling in energy harvestingmobileedgecomputing’,IEEETransactions on Cognitive Communications and Networking, 3(3), pp.361–373.
16. Shahrad,M.,Fonseca,R.,Goiri,Í.,Chaudhry,G.,Batum, P.,Cooke,J.,Laureano,E.,Tresness,C.,Russinovich,M. and Bianchini, R. (2020) ‘Serverless in the wild: Characterizingandoptimizingtheserverlessworkload

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
at a large cloud provider’, USENIX Annual Technical Conference(ATC),pp.205–218.
17. Wang,L.,Li,M.,Zhang,Y.,Ristenpart,T.andSwift,M. (2018) ‘Peeking behind the curtains of serverless platforms’,USENIXAnnualTechnicalConference(ATC), pp.133–146.
18. Jonas, E., Schleier-Smith, J., Sreekanti, V., Tsai, C.C., Khandelwal,A.,Pu,Q.,Shankar,V.,Carreira,J.,Krauth, K., Yadwadkar, N. and Stoica, I. (2019) ‘Cloud programmingsimplified:ABerkeleyviewonserverless computing’,arXivpreprintarXiv:1902.03383.
19. Carreira,J.,Fonseca,P.,Tumanov,A.,Zhang,A.andKatz, R. (2018) ‘A case for serverless machine learning’, ProceedingsoftheWorkshoponSystemsforMLand OpenSourceSoftware(SysML).
20. Lin, W., Zhang, C., Ma, J., Chen, Y. and Li, D. (2018) ‘Online auto-scaling of microservices with reinforcementlearning’,IEEEInternationalConference onWebServices(ICWS),pp.162–169.
21. Zhang, Q., Chen, M., Chen, W., Li, X. and Li, J. (2016) ‘Dynamicresourceprovisioningincloudcomputing:A randomized auction approach’, IEEE Transactions on CloudComputing,4(2),pp.248–261.
22. Ghanbari, H., Simmons, B., Litoiu, M. and Iszlai, G. (2011)‘Exploringalternativeapproachestoimplement an elasticity policy’, Proceedings of the 2011 IEEE InternationalConferenceonCloudComputing,pp.716–723.
23. Islam,S.,Keung,J.,Lee,K.andLiu,A.(2012)‘Empirical predictionmodelsforadaptiveresourceprovisioningin thecloud’,FutureGenerationComputerSystems,28(1), pp.155–162.
24. Bodík,P.,Griffith,R.,Sutton,C.,Fox,A.,Jordan,M.and Katz, R. (2010) ‘Statistical machine learning makes automatic control practical for Internet datacenters’, ProceedingsoftheUSENIXHotCloud,pp.1–6.
25. Chen,X.,Zhang,H.,Wu,C.,Mao,S.,Ji,Y.andBennis,M. (2021)‘Optimizedcomputationoffloadingperformance in virtual edge computing systems via deep reinforcement learning’, IEEE Internet of Things Journal,8(5),pp.4005–4018.
26. Lloyd, W., Ramesh, S., Chinthalapati, S., Ly, L. and Pallickara, S. (2018) ‘Serverless computing: An investigation of factors influencing microservice performance’,IEEEInternationalConferenceonCloud Engineering(IC2E),pp.159–169.
27. McGrath, G. and Brenner, P.R. (2017) ‘Serverless computing:Design,implementation,andperformance’,
2026, IRJET | Impact Factor value: 8.315 |
IEEE International Conference on Distributed Computing Systems Workshops (ICDCSW), pp. 405–410.
28. Spillner,J.(2017)‘Snafu:Function-as-a-Service(FaaS) runtime design and implementation’, Proceedings of the 2017 IEEE International Conference on Cloud EngineeringWorkshop,pp.1–7.
29. Herodotou, H., Dong, F. and Babu, S. (2011) ‘No one (cluster)sizefitsall:Automaticclustersizingfordataintensive analytics’, Proceedings of the 2nd ACM SymposiumonCloudComputing(SoCC),pp.1–14.
30. Gandhi, A., Harchol-Balter, M., Das, R. and Lefurgy, C. (2012) ‘Optimal power allocation in server farms’, ProceedingsoftheACMSIGMETRICS,pp.157–168.
31. Urgaonkar,B.,Pacifici,G.,Shenoy,P.,Spreitzer,M.and Tantawi, A. (2008) ‘Analytic modeling of multitier Internetapplications’,ACMTransactionsontheWeb, 2(1),pp.1–32.
32. Gias,U.A.,Casale,G.andEllahi,W.(2019)‘Asurveyon modeling and optimization of cloud computing systems’,ACMComputingSurveys,52(1),pp.1–36.
33. Li, Z., O’Brien, L., Zhang, H. and Cai, R. (2013) ‘On a catalogueofmetricsfor evaluatingcommercial cloud services’, Proceedings of the 13th IEEE/ACM International Symposium on Cluster, Cloud, and Grid Computing,pp.164–173.
34. Chen, Y., Alspaugh,S.and Katz,R.(2012)‘Interactive analytical processing in big data systems: A crossindustrystudy’,ProceedingsoftheVLDBEndowment, 5(12),pp.1802–1813.
35. Xu, H. and Li, B. (2013) ‘Dynamic cloud pricing for revenue maximization’, IEEE Transactions on Cloud Computing,1(2),pp.158–171.
36. Farokhi,F.,Kargahi,M.andBuyya,R.(2015)‘Energyefficient resource provisioning in cloud computing’, IEEETransactionsonSustainableComputing,1(2),pp. 114–127.
37. Zhang, Y., Cherkasova, L. and Loo, B.T. (2013) ‘Performancemodelingofresourcecontentioninmultitenant clouds’, Proceedings of the 2013 IEEE InternationalConferenceonCloudEngineering,pp.1–10.