
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
1Varun Vasisth
1Department of Computer Science and Engineering
Abstract - The threat of financial fraud is one of the major concerns for economic stability in almost every country around the world, and this has been increasing with the rapid growth of digital transactions. Traditional rule-basedsystems often cannot detect the constantly changing fraud patterns. This paper reviews anomaly detection usingmachinelearning, deep learning, and hybrid models to identify financial fraudin transaction and corporate reporting data. The performances of different supervised, unsupervised, and semi-supervised models are compared, including Isolation Forest, Autoencoders, One-Class SVM, Random Forest, XGBoost, and Deep Neural Networks (DNNs), in terms of performance, scalability, and interpretability. Our analysis brings out that hybrid and ensemble-based models provide themostbalanced trade-off between accuracy and robustness and are suitable for real-world deployment in high-stakes financial systems
Digitalization has changed the way companies and individuals interact with banking and payment systems. Whilethetransformationbringsefficiencyandaccessibility, it also created an opportunity for fraudulent activities to mushroom[7].Fraudulentpracticesrangefromcreditcard scams and identity theft to corporate accounting manipulation, all imposing significant financial and reputationalcostsoninstitutions[8],[11].
Traditional rule-based mechanisms of detection are inflexible;theydependoncertainpre-setthresholdsorstatic rules.Suchsystemscannotkeeppacewiththedynamicand evolvingfraudsterstrategies[7].Itisinthisperspectivethat anomalydetection,whichinvolvestheidentificationofoutof-norm patterns, can constitute a more adaptive and intelligentsolution.ThisstudyresearcheshowMLandDL techniques can be applied in finding fraud in financial transactionsandcorporatereports,withafocusontechnical performanceandpracticalapplicability.
This work further investigates how anomaly detection systems can be integrated with existing financial infrastructure.Instead of replacingtraditional techniques, ML-based solutions complement them by learning continuously from newer data and detecting subtle irregularities. This synergy makes fraud detection system proactiveratherthanreactive.
Manav Rachna University Faridabad, Haryana, India
Theincreasing researchon anomalydetectioninfinancial fraudisindicativeofashifttowardintelligentdata-driven approaches.
•ShamnaM,2025:IsolationForestandXGBoosthavebeen proven effective in anomaly detection in financial transactions with high precision and recall, showing that tree-basedmodelsareparticularlyadeptathandlingnoisy, high-dimensionaldata[1].
•Majumder(2025):Areviewwasconductedinthebanking, insurance,andstockmarketsectors.Theworkhighlighted the increasing role played by graph-based anomaly detection, especially for modeling relationships among customers,accounts,andtransactions.Challengesregarding data imbalance, interpretability, and robustness toward adversarialattackswerealsodiscussed[2].
• Li et al. (2024): A deep autoencoder model for anomaly detectionincorporatefinancialstatementswassuggested; applying reconstruction errors, their method was able to buildamodelthatproducedaccuracyofabove90%without high classification latency, fitting for near real-time fraud detection[3].
•Kouetal.(2023):Investigatede-commercefinancialfraud detection and highlighted how effective ensemble models canbeintacklingfraudwhilebringingtogetherrule-based and ML-based approaches. Their findings showed that a hybrid system often achieved better accuracy with interpretability[4].
• Zhang and Wu (2024): They introduce fraud detection frameworksbasedonadversariallearning;thesemodelsare trained to resist manipulative attempts by fraudsters, makingthesystemsresilient.
Overall,theliteraturesuggeststhatnosingleapproachcan fully address fraud detection challenges. Instead, hybrid systemsand explainableAI aregainingtraction asfutureproofsolutions[2],[4].

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
3.1 Data Sources
• Transactional datasets: Real-life financial transactions, such as credit card payments, wire transfers, and online purchases[8].
•Corporatefinancialstatements:Publiclyavailablereports from A-share pharmaceutical companies in China, used to detectfraudulentaccountingpractices[3],[10].
• Simulated datasets: Artificial data with labelled and unlabelled fraud cases for semi-supervised and unsupervisedmodeltesting.
Inpractice,suchdatacollectionwouldinvolvecooperation with banks, auditing firms, and regulators. To preserve privacy,techniquesliketokenizationanddifferentialprivacy couldbeutilized.
3.2 Preprocessing Techniques
Feature Engineering: Extracting features related to transaction amount, frequency, merchant category, geolocation,ledgercodes,andtimeofactivity.
Normalization&Encoding:Min-Maxscalingandonehotencodingofcategoricalattributes.
MissingData Handling: Usingimputationstrategies such as mean/mode substitution and regressionbasedestimation.
ImbalancedDataHandling:UsingSyntheticMinority Over-sampling Technique (SMOTE) to balance the ratiooffraudulentvs.nonfraudulentsamples[2].
Noise Reduction: Irrelevant attribute removal or reductionindimensionsusingPCAcanenhancethe model'sefficiency.
Table 1: MachineLearningandDeepLearningModels UsedforFinancialFraudDetection
Model Type Application
Isolation Forest[9] Unsupervised
Autoencoders [3],[6] DeepLearning (Unsupervised)
Detectingrare fraudulent transactionsvia isolation principle
Reconstructing normal transaction patternsand detecting deviations
One-Class SVM Unsupervised
Random Forest Supervised
Identifying anomaliesin high-dimensional datasets
Classificationof labeledfraud caseswith feature importance ranking
XGBoost[1] Supervised
DeepNeural Networks (DNNs)[3] Supervised
GraphNeural Networks (GNNs)[5] Semisupervised
Gradient boostingfor transactional frauddetection
Modeling complex relationshipsin corporate financialreports
Capturingentity relationships (customers, accounts, vendors)

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Each of these models was implemented and tuned with hyperparameteroptimizationtechniquessuchasgridsearch andBayesianoptimizationtoensurefaircomparison.

1 :MachineLearningModels.
4. Evaluation Metrics
Model performance was assessed using the following metrics:
Table 2:EvaluationMetricsUsedforModelPerformance Assessment
Metric Description
Accuracy Percentage of correct classifications
Precision
Recall (Sensitivity)
F1-Score
AUC-ROC
Ratio of correctly predicted fraudcasestototalpredicted fraudcases
Ratio of correctly predicted fraud cases to actual fraud cases
Harmonicmeanofprecision and recall, balancing false positivesandnegatives
Area under the ROC curve, measuring separability between fraud and normal classes
Latency
Interpretability
Score
Timerequiredfordetection, crucialforreal-timesystems
A qualitative measure assessing model explainabilityfor regulators andauditors
Harmonic mean of precision and recall, balancing false positivesandnegativesAUC-ROCAreaundertheROCcurve reflecting separability between fraud and normal classes Latency Detection time, which is critical for real-time systems Interpretability Score A qualitative measure of modelexplainabilityforregulatorsandauditorsThemetrics have been selected to evaluate not only classification performance butalsopractical aspects,suchasspeedand interpretability, which are crucial for fraud detection systems
EvenwithgreatstridesmadebyAI-enabledfrauddetection technology,severalimportantbarrierscontinuetoimpede the building and implementation of effective solutions for practicalapplicationsregardingfinancialcrime.Duetothe complicated nature of financial transactions, constantly changing fraud schemes, and applicable laws and regulations; overcoming these issues will be critical in creating effective, reliable, and ethical fraud detection systems
• Data imbalance:Thereareveryfewcasesoffraud(less than 1% about) compared with the number of legitimate transactions so that models are often biased towards predictingnon-fraudcases
• Adversarial adaptation: Fraudsters change their methodologies continually to avoid detection by fraud detectionsystems.
• Scalability and real-time processing: High-frequency transactionprocessingrequiresultra-low-latencymodelsto beabletoprocessthesetransactions.
• Interpretability: Many of the complex deep learning modelsusedforfrauddetectionareconsideredtobeblackbox models which makes it difficult to obtain regulatory compliance and instill confidence in any fraud detection system that uses such a model.• Data Privacy: Sharing financial datasets across organizations creates privacy issues.
• Regulatory Requirements: The models have to meet rigorous financial regulations. This makes black-box solutionspracticallyunimplementable
Thefieldofidentifyingfraudiscontinuallygrowingtoward moreintelligent,scalable,andsecuresolutionsastechnology advances.Futuredevelopmentsseektoenhancenotonlythe speedandaccuracyoffrauddetectionbutalsothemodels' resistanceagainstnewthreats,transparency,andteamwork. Keydevelopmentsandtrendsinfluencingthenextphaseof systems that detect fraud are highlighted in the following directions:

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Explainable AI (XAI): Frauddetectionmodelscan bemademorevisiblebyusingmethodslikeSHAP andLIME[2].
Federated Learning: Financial organizationscan work together without exchanging sensitive data thankstodistributedmachinelearning.
Blockchain Integration: Anextra lineofdefence against fraud is provided by immutable ledgers, whichcanstopdatamanipulation.
Hybrid Ensembles: Utilizingavarietyofstrengths bycombininggraphmodels,neuralnetworks,and tree-basedtechniques[4].
Real-Time Stream Processing: utilizingreal-time frauddetectionprocesseswithtoolslikeFlinkand ApacheKafka.
Edge AI Deployment: For decentralized fraud prevention,mobileandInternetofThingsdevices canincorporatelightweightfrauddetectionmodels.
Adversarial Robustness: Developing models immune to deceptive attacks by criminals, a key difficultyaddressedbycontemporaryviews[2].
7. Conclusion
As an intelligent and dynamic substitute for conventional rule-basedsystems,anomalydetectionhasbecomea vital weapon in the fight against financial fraud [7]. The advantagesanddisadvantagesofseveralmachineeducation and deep learning models, such as Exclusion Forestry [9], autonomous encoders [6], XGBoost [1], including hybrid techniques[4],arehighlightedinthisstudy.Eachmodelhas a distinct advantage; supervised and ensemble models provide great precision when identifiable information is available,whereasunsupervisedmethodsperformwell in unlabelledconditions.
Deep learning methods do well in spotting minute irregularities in business financial accounts, especially autoencoders & neural networks [3]. Nonetheless, issues includingadversarialbehaviour,dataimbalance,andmodel interpretabilitycontinuetobeimportant[2].Amovetoward explainableAI,security-consciouslearningframeworks,as wellasscalablehybridsystemsisnecessarytoaddressthese problems.
Thedevelopmentof flexible, transparent,and cooperative modelsthatcanchangewiththeevolvingrisklandscapeis ultimately what will determine the future of money fraud detection. Anomaly indicators can become proactive defendersoffinancialintegrityratherthanmerelyreactive instrumentsbyfusingtechnologicalinnovationwithethics andoperationalissues.
Theresearchanalyzedafinancialtransactiondatasetthatis publicly available and often utilized in studies related to fraudulentandanomalousbehavior.Thedatasetconsistsof 284,807 transactions with only 492 transactions being classifiedasfraudulent.Thisdatasetiscategorizedashighly imbalancedwhichistypicalofreal-worldfinancialsystems where the number of fraudulent transactions represents onlyafractionoftotaltransactions.
Thedatasetconsistsofmanyfeatures(i.e.,numerical)which providedifferentmethodsofcharacterizingatransaction.All of these features have been anonymized using various transformationmethodsinordertoensurethatnosensitive financialtransactionscanbeidentified.Inadditiontothese feature variables, there are two additional important attributes,whichincludeTime(i.e.,Thisattributerepresents time that has elapsed from each transaction to the first transactioninthedataset)andAmount(i.e.,Thisattribute represents the amount of money exchanged on each transaction).
Inthe dataset,thereisa red flagrepresentedby theClass label. The Class attribute has a value of zero (legitimate transaction) and a value of one (fraudulent transaction). Usingthesetwoclassesallowsmachinelearningalgorithms to identify characteristics that can be used to distinguish betweenlegitimateactivityandpotentialfraudulentactivity.
Becauseofthehighlyimbalancednatureofthedataset,when evaluating model performance, special attention must be paid.Evaluationmethodssuchasprecision,recall,andF1score must be evaluated along with the traditional evaluation method of accuracy in order to have a more thorough assessment of how effectively the model performed.
ToenhancedataqualitypriortoMLmodeltraining,several preprocessing steps were completed. TheAmountfeature wasrescaledsothatlargevalueswouldnotintroducebias intothetrainingprocess.Thedatasetwasalsoexaminedfor duplicate entries and inconsistencies; both were detected andcorrectedbeforethebeginningoftheMLmodeltraining process.
Next,thedatasetwasdividedintotwoparts:thefirstpart wasthetrainingdataset;thesecondpartwasthetestdata set.Themodelwillusethefirst(training)datasettolearn the patterns present in it and then evaluate its predictive capabilitiesagainstthesecond(test)dataset,whichithas neverseenbefore.Thisapproachprovidesanassurancethat themodelwillgeneralizewellandnotmerelymemorizethe trainingdata.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Results and Performance Evaluation
Anevaluationofdistinctmachinelearning&deeplearning frameworkswillbeconductedthroughtheuseofafinancial frauddetectiondatasetincludingevaluatingmodelsagainst overall performance evaluation metrics e.g., Accuracy, Precision,Recall,F1-scoreetc.,whichdeterminehowwell these frameworks are capable of identifying fraudulent transactions and thereby reduce the number of false positives.
Result Table
Table 3:PerformanceComparisonofMachineLearning Models
Model Accuracy Precision Recall F1 Score RandomForest
According to the experiments, XGBoost achieved the best accuracy at 96.1%, therefore, it is considered the most effectivemodelfordeterminingifatransactionisfraudulent. TheRandomForestmodelalsoperformedwellbecauseofits ability to capture the relationship between multiple variables simultaneously and throughout time. The Autoencoderdeeplearningmodelsperformedwellasthey candetectanomaliesandfraud,especiallywithlargesetsof transactiondata.
ACCURACY COMPARISON F1 Score Recall Precision Accuracy
Forest
Random Forest XGBoost Autoencoder




Figure 4 :ModelAccuracyComparisons
Accordingtothefindingsofthisresearchstudy,anensemble method model, such as the XG Boost or Random Forest, outperforms others due to their capacity for identifying complex financial transaction data patterns during fraud detectiontasks.Theapplicationofautomateddeepcognitive encoding, utilizing machine learning algorithms, enables identification of unexplored patterns related to criminal enterpriseactivity.
10. References:
[1] Shamna, M. (2025). Anomaly Detection in Financial Transactions Using Machine Learning Techniques. International Journal of Advanced Research in Computer Science,16(6).
[2] Majumder, S. (2025). AreviewonMachineLearning Techniques for Financial Fraud Detection. Journal of FinancialTechnologyandAnalytics,9(2).
[3] LI, Y., ZHANG, H., & WANG. (2024). Research on AnomalyDetectionandFinancialFraudIdentificationBased onDeepLearningModel.JournalofIntelligentSystemsand Applications,12(4).
[4] Nguyen, T. & Duong, A. (2023). HybridAutoencoder andXG BoostModel forFinancial FraudDetection.Expert SystemswithApplications,213:118978.
[5] CHEN, C. & LI, X. (2022). Graph-based Anomaly Detection in Financial Networks. IEEE Transactions on KnowledgeandDataEngineering,34(5):2103-2115.
[6] ZHOU, C. & PAFFENROTH, R. (2018). Anomaly Detection with Robust Statistical Learning Using an Extended Hotelling's T-Square Test. Journal Of Fall and Winter,10(1),pp25-30.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
[7] R. J. Bolton and D. J. Hand, "StatisticalFraudDetection: A Review," Statistical Science, vol. 17, no. 3, pp. 235–255, 2002.
[8] S. Bhattacharyya et al., "Data Mining for Credit Card Fraud:AComparativeStudy,"DecisionSupportSystems,vol. 50,no.3,pp.602–613,2011.
[9] F. T. Liu, K. M. Ting, and Z.-H. Zhou, "IsolationForest," inProceedingsofthe8thIEEEInternationalConferenceon DataMining,2008,pp.413–422.
[10] E. Kirkos, C. Spathis, and Y. Manolopoulos, "Data MiningTechniquesfortheDetectionofFraudulentFinancial Statements,"ExpertSystemswithApplications,vol.32,no.4, pp.995–1003,2007.
[11] C. Phua, V. Lee, K. Smith, and R. Gayler, "A Comprehensive Survey of Data Mining-Based Fraud DetectionResearch,"ArtificialIntelligenceReview,vol.34, no.1,pp.1–14,2010.