Skip to main content

IDENTIFICATION OF SPAMBOTS AND FAKE FOLLOWERS ON SOCIAL NETWORK VIA INTERPRETABLE AI-BASED MACHINE L

Page 1


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 05 | May 2026 www.irjet.net p-ISSN: 2395-0072

IDENTIFICATION OF SPAMBOTS AND FAKE FOLLOWERS ON SOCIAL NETWORK VIA INTERPRETABLE

AI-BASED MACHINE LEARNING

Sk.Nabeel 1 Dr. K.Venkataramana 2

1student, Mca 2nd Year Kmmips, Tirupati, Affiliated To S.V. University, Tirupati, A.P, India

2professor, Dept Of Mca, Kmmips, Tirupati, Affiliated To S.V. University, Tirupati, A.P, India

ABSTRACT - SocialnetworkingplatformslikeX(Twitter) serve as hubs for open human interaction, but they are also increasingly infiltrated by automated accounts masquerading as human users. These bots often engage in activities such as spreading fake news and manipulating publicopinionduringpoliticallysensitivetimeslikeelections. Most of the current bot detection methods rely on black-box algorithms, raising concerns about their transparency and practical usability. This study aims to address these limitations by developing a novel methodology for the detection of spambots and fake followers using annotated data. To this end, we propose an interpretable machine learning (ML) framework, leveraging multiple Algorithms with hyper parameters optimized through cross-validation, to enhance the detection process. Fur the more results showcase the model’s ability to identify key distinguishing attributes between bots and legitimate users which offers a transparent and effective solution for social network bot detection

Key words : Spambots , Identification, Social Network, Ai-Based Machine Learning Fake Followers Via Interpretable

1.INTRODUCTION

Socialnetworkshavebecomethekeysourceofinformation inthenewageofmankind.Xformerly knownasTwitterm is presently among the most prevalent and widely used nature and expanding user base. These bots can be useful aslegitimatebotsproducealotofeducationaltweets,such as blogs and news updates. Malicious bots, however, disseminate ante spam or harmful material. The characteristics used by current Twitter bot identification algorithms are often derived from user data, including timestamps, friendship, behaviour, and network connection. Nevertheless, feature Engi nearing requires a lot of work and effort. Social bots have the potential to facilitate the dissemination of misinformation, including fakenews,rumours,andhatespeech,byrapidlyamplifying low-credibility content on X through interactions with high-profile users and strategic mentions. Most of the aforementioned issues are controlled through the use of bots. A botnet is a collection of bots designed to execute specific tasks, while a Sybil account represents a fabric catted identity that does not correspond to or originate from a real human user. These botnets and Sybil accounts

are frequently employed to amplify disinformation and disruptgenuinediscourse,contributingtothechallengesof maintainIngintegrityinonlineplatformsMachinelearning (ML)hasbeen

successfullyutilizedinavastrangeofareassuchassports analytics, sentiment analysis, fake news detection and social bot detection. Our study focuses on interpretable machine learning (XAI) as it has been used in different areas to improve performance and to gain better comprehension of the model. Figure 1 provides the most commonlyusedInterprintableAItechniquesamongwhich SHAP and LIME are the most popular. Interpretable ML provides insight into how a particular data point or data point affects the prediction model using a variety of methods such as factor analysis, local interpretation model-agnostic interpretation (LIME), and Shapley additive interpretation (SHAP). The added transparency helps users understand and trust AI systems while it also allowsstakeholderstoidentifybiasesinthesesystemsthus promoting accountability and fairness in AI applications. Overall, descriptive ML plays an important part in closing the disparity between AI algorithms and human comprehension which supports informed decision-making andincreasingtrustinAItechnology.Thus,utilizingXAIfor socialnetwork

bot detection (SNBD) is an important step to gain a better understanding of its detection process. social media sites andthusitplaysanimportantroleinonline conversations and helps connect millions of active users. However, its substantial social and economic influx Ence has also made it an attractive target for malicious actors seeking to manipulate and influence public opinion and decisionmaking. X has for some time been a prime target for automated programs, or ‘‘bots,’’ due to its open Existing research utilizes various characteristics of the social network to differentiate between human and auto mated accounts.Thesefeaturesincludeuseractivitypatterns(e.g., tweet frequency, timestamps), account metadata (e.g., follower/followingratios,accountage),andsocialnetwork structures (e.g., retweet and mention networks) etc. Supervised ML models and deep neural networks have been widely employed for this purpose. Traditional bot detection systems such as heuristic methods fail against evolving spambots, network-based approaches depend on narrow social networks, and earlier ML models employ

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 05 | May 2026 www.irjet.net p-ISSN: 2395-0072

lime tied characteristics, disregarding linguistic, temporal, and sentiment trends. Furthermore, the majority are not explainingable,whichmakesitchallengingtoevaluatethe data. Our Interpretable AI-based model addresses these gaps by intel grating diverse feature sets. We enhance transparency with XAI which ensures improved accuracy, robustness,andinterprobability.

FIGURE 1. InterpretableAItechniques

Furthermore, clustering and anomaly detection methods have been explored for unsupervised detection of anomalous behaviours linked to bots. While these meth odds have shown promising results, they often lack scalability and adaptability due to their dependency on handcrafted feature engineering and static datasets. Moreover, the heavy reliance on black-box ML models limits their interpretability and creates barriers to understanding how decisions remade. Several challenges reduce the effectiveness of current bot detection methodologies. One of these challenges is feature engineering which is a labour-intensive process that requires domain expertise and manual effort to adapt the models to newer datasets and bots. Furthermore, Bot exhibit dynamic and adaptive behaviour through the evolution of their strategies to mimic human users more effectively and evade detection algorithms. As a result, black-box detection models struggle to adapt to the constantly evolving nature of bot activities. Additionally, the lack of model interpretability in these methods undermines trust and transparency. Evaluation without interpretability is a challenge as we don’t know if the model is identifying bots based on meaningful patterns or merely overfitting to noise in the data. Additionally, most methods are designed to optimize detection accuracy without considering the broader goals of generalizability and adapt ability which are critical for real-world deployment on social networks. These gaps highlight the need for more transparent and interpretable detection framework

2. LITERATURE REVIEW

The burgeoning interest in bot-detection challenges has precapitatedaproliferationofacademicinquiry,yieldinga plethora of articles that proffer diverse methodologies. That notwithstanding, a gap exists in the extant literature, as the overwhelming majority of these approaches fail to provide transparent and interpretable results. In the subsequent sec tons, a concise review of prevailing bot detection strategies will be presented, accompanied by an examination of the challenges that necessitate further investigation. The Prepon durance of bot detection methodologies relies on supervised ML paradigms, which necessitate the utilization of one or multiple annotated datasets to train ML classifiers and develop an efficacious framework. These annotated datasets are frequently generatedthroughhumanannotation,althoughalternative approaches such as leveraging pre-existing established models, crowdsourcing, or automated annotation tech inquest have also been employed to construct datasets for both detection purposes. Table 1 highlights the key literature for interpretable AI-based bot detection. This literature discusses the purpose of each study and our findingsoneachresearch.

3. SNBD METHODOLOGIES

In, the authors devised a novel approach by create Ing a corpus of honeypot accounts, specifically designed to attractspammerinteractions,andsubsequentlyloggedthe corresponding profile information. This dataset was then augmented with a collection of regular user profiles, thereby enabling the development of a comprehensive classificationalgorithmthatincorporatesbothuser-centric and content-centric features. In another research, authors took a similar method, attempting to detect botnets that were run by the same person. Reference employs crowd sourcingtechniquesforbothrecognitiononFacebook,and while it appeared to provide decent results; however, the inherentlimitationsofthismethodbecameapparentwhen the perpetual evolution and proliferation of bots rendered theapproachincreasinglyunsalable,therebyunderscoring the need for more adaptive and dynamic bot detection strategies.

4. CHALLENGES OF SNBD

Despite the plethora of scientific endeavours that have yieldedvariousmethodsfordetectingonlinesocialbots,as indicated in the aforementioned studies, there are still many outstanding difficulties. Even though many SNBD approaches employ more than 1,000 attributes to train their method, it remains unclear whether increasing the numberoffeetruesnecessarilyenhancesmodelefficiency. Moreover, the authors of the highlight the significant impactofutilizinganextensivefeaturesetonthescalability of bot detect toing systems. Interestingly, they also note that employing various subsets of publicly available

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 05 | May 2026 www.irjet.net p-ISSN: 2395-0072

labelled datasets can enhance model generalizability, as observed in the same study. Notably, the performance of machinelearning-basedbotdetectionmodelsvariesacross different datasets. Cones quaintly, the accumulation of additional datasets is essential to ensure that our training data encompasses a comprehensive range of bot behavioural features. The same conclusion is parties’ tweetsfrombeforeandthroughoutthe2017electioncycle, demonstrating an increase in the use of social bots. It is clear that Twitter bot identification is a difficult process that frequently needs thorough and robust treatment. Several ML-based methods, such as the Butternut have been offered with a total of 1200 distinct characteristics combined with an ML classifier. An enhanced version of thissystem,Bathometer,is detailedin,which needsX API keys to obtain user data during real-time calculations, making it inefficient to utilize real-time labelling tools in thecaseoflargedatasets.Thereisanincreasingnumberof Twitter bot identification programs that use machine learning and data (statistical) analysis such as the Steeler, theDebut,andtheRetweet-Buster(Robust)obtainedfrom thead,whichgeostationary-grainedcategorizationofbots, givingdistinctdatasetsforeachsortofbot.Asaresult,one major difficulty in online social bot identification is determining whatqualities genuinelyconstitutesocial bot. X bots are often used for malevolent objectives ranging from distributing fake news to propaganda and astroturfing.Thewritersofexamined245,000profilesonX between the 2016 US presidential election and the 2018 midterm elections, detecting around 31,000 bots. The authors of conducted an exhaustive analysis of 43 million election-related tweets pertinent to the U.S. Congress investigation into Russian interference during the 2016 U.S.electioncampaigns.

5.CONCLUSION

This research presents a unique way to differentiate betweenbotsandrealusersonXbyusinganinterpretable MLframeworkthatextractsandanalysesattributesforthe theextractionofa diversesetoffeaturesderivedfromthe datasets discussed in Section III-A. The model was trained onvariousfeaturesthatwerefinalizedthroughexplainable AI techniques to improve the detection of social and spam botsaswellasfakefollowers.Thisapproachincreasedthe accuracy and reliability of our model and gave important insights into potential patterns which enhanced transparency for social security. This is done through the incorporation of the XAI techniques SHAP and LIME into the model which allows the researchers to understand the impact of the features on the model. This information allowed us to reduce the size of the feature set to include the most important features which reduced the workload for the ML model. The significance of this study lies in its ability to bridge the gap between model accuracy and transparency thus addressing the key challenges in bot detectionbyofferinganinterpretablemethodology.

6. REFERENCES

[1] E. Canoe-Marin, M. Mora-Cantal lops, and S. SánchezAlonso, ‘‘Twitter as a predictive system: A systematic literature review,’’ J.Bus.Res.,vol.157,Mar.2023,Art. no. 113561,doe:10.1016/j.jbusres.2022.113561.

[2] F. Tabassum, S. Mubarak, L. Liu, and J. T. Du, ‘‘How many features do we need to identify Bots on Twitter?’’ in Information for a Better World: Normality, Virtuality, Physicality, Inclusivity, I. Ssemwanga, A. Goulding, H. Modulation-Sandy,J.T.Du,A.L.Soares,V.Hisami,andR.D. Frank, Eds., Cham, Switzerland: Springer, 2023, pp. 312–327.

[3]R.Al-AzawiandS.O.AL-Memory,‘‘Feature extractions and selection of bot detection on Twitter a systematic literature review,’’ Intelligences Arif., vol. 25, no. 69, pp. 57–86, Apr. 2022, due: 10.4114/intertie. vol25iss69pp5786.

[4] Zhigang A.A. Ghorbani, ‘‘An overview of online fake news: Chirac erization, detection, and discussion,’’ Inf. Process.Manage.,vol.57,no.2,Mar.2020,Art.no.102025, due:10.1016/j.ipm.2019.03.004.

[5] Y. Bosham, I. Malakhov, K. Beznosov, and M. Ripeanu, ‘‘Designandanalysisofasocialbotnet,’’Compu.Newt.,vol. 57, no. 2, pp. 556–578, Feb. 2013, Doi: 10.1016/j.comnet.2012.06.006.

[6]Z.Yang ,C.Wilson ,X.Want.Gao, B.Y.Zhao, andY.Dai,‘ ‘Uncoveringsocialnetwork Sybilsinthewild,’’ACMTrans. Know. Discovery from Data, vol. 8, no. 1, pp. 1–29, Feb. 2014,dui:10.1145/2556609.

[7]D.Javed,N.Z.Janghi,andN.A.Khan,‘‘Footballanalytics forgoalpredictiontoassessplayerperformance,’’inProc. Int. Conf. Innov. Technol. Sports (Reveal DNA ICITS), Apr. 2023,pp.245–257,Doi:10.1007/978981-99-0297-2_20.

[8]M.Humayun, DavedJhanjh i, M.F.Almufareh, andS. N.Almuayqil, ‘‘Deep learning based sentiment analysis of COVID-19 tweets via resam pling and label analysis,’’ Comput.Syst.Sci.Eng.,vol.47,no.1,pp.575–591,2023.

[9] S. N. Almuayqil, M. Humayun, N. Z. Jhanjhi, M. F. Almufareh, and D. Javed, ‘‘Framework for improvedsentiment analysis via random minor ity oversampling for user tweet review classification,’’Electronics,vol.11,no.19,p.3058,Sep. 2022, doi:10.3390/electronics11193058.

[10]F.Al-Quayed,D.Javed,N.Z.Jhanjhi,M.Humayun,and T. S. Alnusairi, ‘‘A hybrid transformer-based model for optimizing fake news detection,’’ IEEE Access, vol. 12, pp. 160822160834,2024,doi:10.1109/ACCESS.2024.3476432.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 05 | May 2026 www.irjet.net p-ISSN: 2395-0072

[11]D.Javed,N.Jhanjhi,N.A.Khan,S.K.Ray,A.A.Mazroa, F. Ashfaq, and S. R. Das, ‘‘Towards the future of bot detection: A comprehensive tax onomical review and challenges on Twitter/X,’’ Comput. Netw., vol. 254, Dec. 2024,Art.no.110808,doi:10.1016/j.comnet.2024.110808.

[12] S.Lundberg and S. Lee, ‘‘A unified approach to interpreting model pre dictions,’’ in Proc. 31st Int. Conf. NeuralInf.Process.Syst.(NIPS).RedHook,NY,USA:Curran AssociatesInc.,Jan.2017,pp.4768–4777.

[13] M. Aljabri, R. Zagrouba, A. Shaahid, F. Alnasser, A. Saleh, and D. M. Alomari, ‘‘Machine learning-based social media bot detection: A comprehensive literature review,’’ Social Netw. Anal. Mining,vol.13,no.1, pp. 1–40, Jan. 2023, doi:10.1007/s13278-022-01020-5.

[14]K.Hayawi,S.Saha,M.M.Masud,S.S.Mathew,andM. Kausar, ‘‘Social media bot detection with deep learning methods:Asystematicreview,’’NeuralCompute.Appl.,vol. 35, no. 12, pp. 8903–8918, Mar. 2023, doi: 10.1007/s00521-023-08352-z.

[15] S. Kaguta and E. Ferrara, ‘‘Deep neural networks for bot detect tion,’’ Inf. Sci., vol. 467, pp. 312–322, Oct. 2018, doi:10.1016/j.ins.2018.08.

Turn static files into dynamic content formats.

Create a flipbook
IDENTIFICATION OF SPAMBOTS AND FAKE FOLLOWERS ON SOCIAL NETWORK VIA INTERPRETABLE AI-BASED MACHINE L by IRJET Journal - Issuu