
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
C Bhargava1 , Dr K. Venkataramana2
1Student, MCA 2nd Year, KMMIPS, Tirupati, Affiliated to S.V. University, Tirupati, A.P, India
2Professor, Dept. of MCA, KMMIPS, Tirupati, Affiliated to S.V. University, Tirupati, A.P, India
Abstract - Online advertising has become a fundamental revenue model for digital businesses, primarily operating through the Pay-Per-Click (PPC) mechanism. However, click fraud - the act of artificially inflating ad clicks using bots, automated scripts, or human click farms - poses a severe threat to this ecosystem, costing advertisers approximately $42billionin2021 alone.Thisprojectproposesaneffective systemfordetectingandpreventingadclickfraud usingtwo supervised machine learning algorithms: Decision Tree (DT) and Extreme Gradient Boosting (XGBoost). The system is trained and evaluated on the publicly available TalkingDataAdTrackingdataset, whichcontainsover184 million real-time mobile ad click records. Key features includingtemporalpatterns,IPbehavior,deviceinformation, and click frequency are extracted and engineered to train both models. Performance is evaluated using accuracy, precision, recall, F1-score and AUC-ROC metrics. Experimental results demonstrate that XGBoost significantly outperforms Decision Tree in detecting fraudulentclickswhilehandlingclassimbalanceeffectively
Key Words: Click Fraud, XGBoost, Decision Tree, PPC, TalkingData, Imbalanced Data, Feature Engineering
Advertising campaigns on websites and smartphone applicationshavebecomeanintegralpartofpeople'sdaily lives. The digital advertising industry has witnessed unprecedented growth over the past two decades, transforming the way businesses promote their products and services to potential customers. Among the various models of online advertising, the Pay-Per-Click (PPC) modelhasemergedasoneofthemostdominantandwidely adoptedapproaches.
InPPCadvertising,campaignproviderschargeadvertisers a fee for each click made on an advertisement link, under the assumption that every click represents a genuinely interested potential customer. According to Google AdWords statistics, the average cost of a click for Google Adsis$0.89,andGooglealoneearned$209.49billionfrom advertising in 2021. Despite its success, the PPC model is highlyvulnerabletoclickfraud-theintentionalgeneration offakeclicksusingbots,scripts,orhumanclickfarms.Click fraud cost advertisers approximately $42 billion in 2021 andcontinuestogrowinscaleandsophistication.
The online advertising ecosystem involves multiple stakeholdersworking together-advertiserswhocreate and fund ad campaigns, publishers who host advertisements,adnetworkssuchasGoogleAdSenseand Meta Ads that act as intermediaries, and end users who interactwithadvertisements.Advertisingcampaignsare designed to target specific groups of users based on interests, demographics, and online behavior, with the ultimategoalofdrivingsalesandincreasingbrand.Social mediaplatformslikeFacebook,Instagram,andYouTube havefurtherexpandedthereachofonlineadvertising.
2.1 Sources and Nature of Click Fraud
Clickfraudiscarriedoutthroughadvancedanddiverse methods,includingbotnets,VPN-basedIPmasking,and geographically distributed systems. Modern bots can mimic human behavior such as scrolling and mouse movement, making detection increasingly difficult. In some cases, even publishers engage in fraudulent clicking to boost their own ad revenue, further complicatingtheissue.
2.2 Need for AI-Based Detection
Traditional rule-basedsystemsare nolongersufficient to detect such sophisticated fraud patterns. Artificial Intelligence (AI), particularly Machine Learning (ML) and Deep Learning (DL), has gained importance in identifying fraudulent activities. These techniques can analyze large-scale data and detect complex patterns, making them highly effective in fraud detection applications.
This project applies two supervised machine learning algorithms - Decision Tree (DT) and Extreme Gradient Boosting(XGBoost)-fordetectingfraudulentadclicks. The models are trained using the TalkingData AdTrackingdataset.Keyfeaturessuchastimepatterns, IPbehavior,devicedetails,andappdataareused.Model performance is evaluated using accuracy, precision, recall, F1-score, and AUC-ROC to determine the most effective approach. Fig. 1 illustrates the overall system concept.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

Fig -1: Ad Click Fraud Detection System Overview - Showing Advertiser to Publisher flow, Real User vs Fraudster/Bot paths, the $42 billion loss in 2021, and the AI-based detection layer using Decision Tree and XGBoost.
3. RELATED WORK
3.1 Overview
Adclickfrauddetectionhasbeenwidelystudied,with manyresearchersproposingMachineLearning(ML)and DeepLearning(DL)approaches.Mostworkfocuseson tree-basedandgradientboostingmodels,suchasDecision Tree(DT)andXGBoost,duetotheireffectivenessin handlinglargeandimbalanceddatasets.
3.2 Tree-Based Approaches
Several studies highlight the effectiveness of tree-based models. Li et al. introduced the MadTracer system using DecisionTreestodetectmaliciousadbehaviors,including a new type of fake click redirection. Berrar applied Random Forests with time-based click features, showing that temporal patterns are strong indicators of fraud.Yan and Jiang demonstrated that tree-based models outperform Bayesian methods on imbalanced datasets [1][3][4].
3.3 Gradient Boosting Approaches
Gradient boosting methods have shown superior performance in many studies. LightGBM achieved high accuracy and efficiency in detecting suspicious click patterns. Multiple works using XGBoost reported accuraciesexceeding90%,duetoitsrobustness,ability to handle missing data, and resistance to overfitting. Hybrid models combining XGBoost with other ensemble techniquesfurtherimprovedperformance[6][9].
3.4 Deep Learning Approaches
Deep learning methods have also been explored for detecting complex fraud patterns. Approaches such as
Convolutional Neural Networks (CNNs) using mobile sensor data, Fully Connected Neural Networks (FCNNs), and hybrid models combining ANN, Autoencoders, and GANs have achieved very high accuracy, particularly effective in capturing subtle behavioral patterns that traditionalMLmodelsmaymiss[13].
3.5
Overall, tree-based and gradient boosting modelsespecially XGBoost and LightGBM - consistently deliver strong performance in click fraud detection. Feature engineering, particularly temporal features, plays a crucial role in model effectiveness. Class imbalance remains a major challenge and must be carefully addressed. This project builds on these insights by implementingandcomparingDecisionTreeandXGBoost on the TalkingData dataset, focusing on feature engineering and imbalance handling to achieve reliable frauddetection.
4.1 System Overview
Theproposedsystemisdesignedtodetectandpreventad click fraud using two supervised machine learning algorithms- Decision Tree (DT) and Extreme Gradient Boosting (XGBoost) The system follows a structured pipeline including data collection, preprocessing, feature engineering, model training, classification, andperformanceevaluation.Thecoreobjectiveisto classify each ad click as either legitimate (0) or fraudulent (1). The overall system architecture is illustratedinFig.2.

Fig -2: Proposed System Architecture - End-to-end pipeline from TalkingData collection (184M+ records) through preprocessing, feature engineering, model training (Decision Tree and XGBoost), and evaluation to classify clicks as Legitimate or Fraudulent.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
4.2 Dataset Description
The TalkingData AdTracking dataset is publicly availableonKaggle.TalkingDataisaChinesemobiledata servicecompanythatprocesses approximately 3 billion clicks per day, of which nearly 90% are potentially fraudulent. ThedatasetwasreleasedaspartofaKaggle competition in 2017. It contains 184,903,890 training records captured over four days, with an overall size of approximately 7 GB.
Table -1: TalkingDataAdTracking Dataset - Feature Descriptions
Feature Description ip IPaddressof
app AppIDformarketing device DevicetypeIDofmobilephone
os OSversionIDofmobilephone channel ChannelIDofadpublisher click_time TimestampofclickinUTC
attributed_time Timeofappdownload(ifany) is_attributed Target:1=fraud,0=legit
4.3 Data Preprocessing
Raw click data requires several preprocessing steps beforemodeltraining:
1) Handling Missing Values - attributed_time contains manymissingvaluessinceitisonlyrecordedwhenanapp isdownloaded;theseareremovedtopreventdataleakage.
2) Handling Class Imbalance - the dataset is highly imbalanced (99.75% legitimate vs 0.25% fraudulent); randomundersamplingandstratifiedsamplingareapplied.
3) Data Type Conversion - click_time is converted to datetimeformat.
4) Dropping Irrelevant Features - attributed_time is removedbeforemodeltraining.
4.4 Feature Engineering
Featureengineeringiscrucialforcapturingfraudpatterns:
4.4.1 Temporal Features -click_timeisdecomposedinto hour,minute,day,andweek.
4.4.2 Click Frequency Features -aggregatedclickcounts for IP, IP-App, IP-App-OS, and IP-App-Channel combinations.
4.4.3 Unique Count Features -uniqueapps,devices,and channelsperIPaddress
4.4.4 Time-to-Next-Click -timegapbetweenconsecutive clicksfromthesameIP,asbotsgeneraterapidandregular clicks.
4.5 Model Implementation
Decision Tree (DT): Uses Gini impurity criterion to split data and generate classification rules. Key parameters: max_depth=10, min_samples_split tuned to prevent
overfitting. Simple and interpretable but prone to overfittingonlargedatasets.
XGBoost(Extreme Gradient Boosting): Builds an ensemble of trees using gradient boosting and regularization. Trained with 200 estimators, learning rate=0.1,max_depth=6.Usesscale_pos_weighttohandle classimbalance.Well-suitedforlarge,high-dimensional, imbalanceddatasets.
4.6 Evaluation Metrics
Both models are evaluated using the standard classificationmetricsshowninTable2.
Table -2: Evaluation Metrics - Formulas and Descriptions
TP/(TP+FP) Ofpredictedfrauds,how manyreal Recall
TP/(TP+FN) Ofactualfrauds,how manydetected
F1-Score 2x(PxR)/(P+R) HarmonicmeanofPand R
AUC-ROC AreaunderROC Overalldiscrimination ability
Since the dataset is highly imbalanced, F1-Score and AUC-ROC are the primary comparison metrics, as accuracyalonecanbemisleadinginsuchscenarios
5.1 Experimental Setup
All experiments were conducted using the TalkingData AdTracking dataset on a standard computing environment using Python 3.x with libraries: pandas, numpy, scikit-learn, xgboost, and matplotlib. Due to the dataset size (184M+ records), a stratified sample of 1 million records was used - a common practice in click fraud detection literature. The dataset was split 80% training and 20% testing using stratified sampling to preserve class distribution. Class imbalance was handledusingrandomundersamplingduringtraining.
5.2 Decision Tree Performance
The Decision Tree classifier was trained using Gini impurity criterion with max_depth=10. The results are showninTable3.
Table -3: Decision Tree Model - Performance Results

Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
TheDecisionTreemodeldemonstratedreasonablystrong performance. However, it showed signs of overfitting when tree depth was increased. The recall score of 84.76% indicates approximately 15% of actual fraudulent clicks were missed -asignificantconcernin real-worldscenarioswheremissingfraudiscostly.
5.3
The XGBoost classifier was trained with 200 estimators, learningrate=0.1,max_depth=6,andscale_pos_weightto handleclassimbalance.ResultsareshowninTable4.
-4: XGBoost Model - Performance Results

Fig -3: Performance Comparison - Bar chart showing Decision Tree vs XGBoost across Accuracy, Precision, Recall, F1-Score, and AUC-ROC. XGBoost consistently outperforms Decision Tree with the largest gain in Recall (+11.55%).
5.5 FeatureImportance Analysis
XGBoost significantly outperformed the Decision Tree across all evaluation metrics. The high recall of 96.31% indicates the model successfully identified the vast majority of fraudulent clicks, while precision of 95.84% confirms very few legitimate clicks were incorrectly flagged. The AUC-ROC score of 98.45% demonstrates excellent overall discrimination ability between legitimateandfraudulentclicks,makingXGBoostahighly reliablemodelforreal-worldadclickfrauddetection.
Table5presentsaside-by-sidecomparisonofbothmodels.
Table -5: Comparative Analysis - Decision Tree vs XGBoost
Both Decision Tree and XGBoost provide feature importance scores indicating which features contributed most to fraud classification. Table 6 summarizes the top features ranked by the XGBoost model.
Table -6: Feature Importance Scores – Top Features (XGBoost)
The results clearly demonstrate that XGBoost outperforms Decision Tree across every metric, with the most significant improvement seen in Recall (+11.55%) and F1-Score (+9.62%). This aligns with findings from the literature review, where XGBoost and gradient boosting modelsconsistentlyachievedsuperior resultscomparedto standalone tree-based classifiers in clickfrauddetectiontasks.
Fig.3providesthevisualbarchartcomparison,hereis thevisualcomparisonofbothmodels:
The results confirm thattemporal features (hour,day) and click frequency features (clicks per IP combinations) are the most powerful indicators of fraudulent activity. Engineered features such as clicks grouped by IP-App-OS contributed the highest importance score (0.312), validating that feature engineeringisessentialforeffectivefrauddetection.
5.6 Discussion ofResults
Keyobservationsfromtheexperimentalresults: XGBoost is clearly superior - outperforms Decision Tree across all five metrics, most notably in Recall (+11.55%) and F1-Score (+9.62%), confirming ensemble-basedboostingmethodsarebettersuitedfor

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
complex,imbalancedclickfrauddata.
Decision Tree remains useful - achieves 91.34% accuracy and offers interpretability and training speed advantagesasafastbaselinemodel.
Feature engineering is critical - IP-App-OS click frequency features contributed the highest importance scoresinbothmodels.
Class imbalance handling is essential - without stratified sampling and undersampling, both models defaulted to predicting the majority class, yielding nearzerorecallonthefraudclass.Thestratifiedsamplingand undersampling approach applied in this project successfullyresolvedthisissue.
6.1 Summary:
This project proposed an effective system for Ad Click Fraud Detection and Prevention using Decision Tree (DT) and Extreme Gradient Boosting (XGBoost)trained on the TalkingData AdTracking dataset containing over 184 million real-time mobile ad click records. Click fraud remains one of the most financially damaging threats in digital advertising, costing approximately $42 billion in 2021 andaffecting90%of all PPC campaigns. The proposed system addressed this problem through a structured pipeline of data preprocessing, feature engineering, model training, and comparativeperformanceevaluation.
6.2 Key Findings:
Experimental results demonstrated that XGBoost outperformed Decision Tree across all metricsachievingaccuracyof 97.20%, F1-Scoreof 96.07%, and AUC-ROC of 98.45%, compared to Decision Tree's accuracyof 91.34% andF1-Scoreof 86.45%. Engineered featurescombiningIPaddresswithAppIDandOSproved to be the strongest fraud indicators. Proper handling of the severe class imbalance (99.75% vs 0.25%) through stratifiedsamplingwascriticalforproducingmeaningful results.
6.3Limitations:
1.only a 1 million record sample was used due to computationalconstraints
2.post-clickbehavioralfeaturessuchasmouse movementswereunavailableinthedataset
3.modelsweretrainedofflineanddonotsupportrealtimeretrainingasfraudpatternsevolve.
6.4 Future work:
Future improvements includes applying deep learning models (LSTM, Transformers), enabling real-time detection systems , incorporating behavioral data, and exploring for privacy-preserving techniques such as federatedlearning.
1. Alzahrani,R.A.;Aljabri,M.AI-BasedTechniquesforAd Click Fraud Detection and Prevention: Review and ResearchDirections. J.Sens.ActuatorNetw. 2023, 12,4.https://doi.org/10.3390/jsan12010004
2. Clickcease. The State of Click Fraud in SME Advertising. 2022. Available online:https://www.clickcease.com/blog/wpcontent/uploads/2020/09/SME-Click-Fraud-2020.pdf
3. Li,Z.;Zhang,K.;Xie,Y.;Yu,F.;Wang,X.F.Knowingyour enemy: Understanding and detecting malicious Web advertising.In Proceedingsofthe2012ACMConference on Computer and Communications Security, Raleigh, NC,USA,2012.
4. Berrar, D. Random forests for the detection of click fraudinonlinemobileadvertising.In Proceedingsofthe 1st International Workshop on Fraud Detection in MobileAdvertising(FDMA),Singapore,2012.
5. Oentaryo, R.; Lim, E.P.; Finegold, M. et al. Detecting click fraud in online advertising: A data mining approach. J.Mach.Learn.Res. 2014,15,99–140.
6. Minastireanu,E.A.;Mesnita,G.LightGBMMachine LearningAlgorithmtoOnlineClickFraudDetection. J.Inf.Assur.Cybersecur. 2019,1–12.
7. Viruthika, B.; Das, S.S.; Manishkumar, E.; Prabhu, D. Detectionofadvertisementclickfraudusingmachine learning. Int.J.Adv.Sci.Technol. 2020,29,3238–3245.
8. Thejas,G.S.;Dheeshjith,S.;Iyengar,S.S.;Sunitha,N.R.; Badrinath,P.Ahybridandeffectivelearningapproach forClick Frauddetection. Mach. Learn. Appl. 2021,3, 100016.
9. Gohil, N.P.; Meniya, A.D. Click Ad Fraud Detection UsingXGBoostGradientBoostingAlgorithm.Springer: Cham,Switzerland, 2021
10. Dash, A.; Pal, S. Auto-Detection of Click-Frauds using Machine Learning. Int. J. Eng. Sci. Comput. 2020, 10, 27227–27235.
11. Mikkili, B.; Sodagudi, S. Advertisement Click Fraud DetectionUsingMachineLearningAlgorithms. Smart Innov.Syst.Technol. 2022,282,353–362.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
12. Chari, H.; Aswale, S.; Pawar, V.N. Advertisement Click Fraud Detection Using Machine Learning Techniques. In Proceedings of the 2021 InternationalConference on Technological Advancements and Innovations (ICTAI), Tashkent,Uzbekistan,2021.
13. Shi, C.; Song, R.; Qi, X.; Song, Y.; Xiao, B.; Lu, S. ClickGuard: Exposing Hidden Click Fraud via Mobile SensorSide-channelAnalysis.In ProceedingsoftheICC 2020 -IEEE International Conference on Communications,Dublin,Ireland,2020.
14. Gabryel, M.; Scherer, M.M.; Sułkowski, L.; Damaševičius,R.DecisionMakingSupportSystemfor Managing Advertisers by Ad Fraud Detection. J. Artif. Intell.SoftComput.Res. 2021,11,331–339.
15. Sadeghpour, S.; Vlajic, N. Click fraud in digital advertising: A comprehensive survey. Computers 2021,10,164.
16. Statista. Advertising Revenue of Google from 2001 to 2021. Available online: https://www.statista.com/statistics/266249/adve rtising-revenue-of-google/