
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
Shubhi Katiyar
1 ,
Vartika Choudhary
1 ,
Bhargavee Singh
1, Aditi
Srivastava
1, Daood Saleem2 1, 2School of Computing Science and Engineering, VIT Bhopal University, Sehore, India
Abstract - Online harassment has become a critical issue on social media. Women repeatedly face the risk of various abusive messages. The existing system generally atplatformslikeInstagram dependsonuserreporting or basic keyword filtering which make them incapable of detecting repetitive forms of harassment. To bridge this gap, Guardian is developed as an AI powered, privacypreservingChromeextensionwhichpreemptivelydetects harmfulmessagesduringliveInstagramchat. Thesystem uses a fine-tuned DistilBERT transformer model for multiclass text classification, categorizing messages as safe, creepy, or different levels of harassment. Total privacy is guaranteed as all the processing occurs locally ontheuser’sdevice,makingsurethatnomessagecontent isstoredexternally.Astrike-basedtrackingsystemkeeps track of repeat offenders. It is activated after three times ofcontinuousharassingbehaviour.Afterthatanalertpop up comes automatically bifurcating the type of harassment being done and also highlighting the text at which it’s been done. Also, it auto-hides the chats after multiple offenses. It is hooked into using content scripts and a local FastAPI server. Evaluation shows strong modelperformance,achieving78%accuracyandformost critical categories of stalker behaviour and severe harassment, both precision and recall are very high. Further, real-time tests confirm this system indeed hides harmful messages while keeping normal conversations untouched.
Keywords - Online harassment detection, privacypreserving classification, strike-based offender tracking, real-time message moderation, transformerbasedtextanalysis.
Thefastexpansionofdigitalcommunicationtools andsocialmediaplatformshaschangedhowpeoplein society today build relationships and exchange information and show their identity. People can use Instagram and various online communities to create extensive opportunities for communication and networking and learning and self-expression. The platforms enable users to create connections with others who live in different locations while they engage in social conversations and display their artistic work. The widespread ability to connect with othershasledtoanincreaseinonlineharassment[1] which now stands as a major societal problem. The digital platforms which people use nowadays have
resultedinincreasedexposuretoharmfulinteractions among users from various age groups and gender identities and social backgrounds. Online users from different demographic groups experience cyber harassment as a widespread problem which damages their mental health and social relationships and their abilitytouseonlineservices.
Cyber harassment includes a wide range of dangerous activities which involve sending abusive messages and trolling and cyberstalking and using fake identities and spreading disinformation and making threats and sending unwanted messages. The target of these actions experiences emotional distress becausetheycreateapersonalattackwhichgenerates fear and insecurity. The studies on gender-based online abuse demonstrate that women and marginalized groups experience [2,3] greater online abuse than other demographics. The psychological effects of harassment increase because of repetitive abuse which combines with escalating patterns of abuse to create more severe long-term effects on victims. Cyber harassment occurs as a sequence of abusive events whichrepeatthemselves because they represent social and cultural and technological forces thatoperateinsociety.
1.1.1 Nature and Forms of Cyber Harassment
Cyber harassment exists on various digital platforms which include social networking sites and messaging apps and online gaming environments and forums and professional networks. Users commonly encounter identity-based abuse and cybercrime activities [4,5] which include offensive messages and threats and unauthorized sharing of personal information and online impersonation. Harassment can occur through two types of behavior which include minor exclusion and disrespectful comments and major threats and forced disclosure of private details.
Recent studies highlight the increasing role of technology-facilitated violence [6] which demonstrates how digital tools and bots and automated systems are used to track and attack specific people. The internet permits users to view objectifyingcontentandharmfulmediacontentwhich teachesthemtoadoptabusivebehaviorwhileitbuilds social prejudice against victims and reduces their

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
empathy. Research on cyberbullying and online aggression [7,8] shows that these two behaviors link with each other while they create community standards which result in ongoing harassment patternsthatimpactmultipleusersatonce.
1.1.2
Online harassment usually starts with minor harmful acts which eventually develop into severe forms of online abuse. Minor negative interactions, such as unsolicited comments or subtle insults, can escalate into persistent abuse, which develops throughtheprocessofcontinuoustargeting[9,10]and increased harassment activity by offenders. The severity and impact of cyber harassment are influenced by multiple factors, including the type of abuse, duration, frequency, and the victim’s online presence. Studies on abuse intensity and impact variation [11-13] show that repetitive harassment exposure causes people to experience higher psychologicalstresswhichresultsinincreasedanxiety andsocialwithdrawal.Theidentificationofdangerous situations depends on understanding these patterns, which also helps to create preventive measures and educationalprogramsandpolicyframeworks.
1.1.3
The process of identifying and managing cyber harassment offenses requires complete knowledge about actual online communication [14] settings which includes all elements of prior interactions together with their associated emotional expressions and their intended meanings and the distinctive communication methods that each online platform permits. The ability to identify contextual elements enables people to recognize between normal conflicts and dangerous conduct which assists authorities and moderatorsinmakingcorrectdecisions.
The online environment depends on user engagement because it establishes the framework for digital spaces. Research on bystander behavior and intervention patterns [15,16] shows that online communities influence the spread and mitigation of harassment. Bystanders who recognize abusive behavior can intervene by reporting incidents or supporting victims, whereas inaction can reinforce harassment. The studies on user interaction and engagement behavior [11,12] demonstrate that active participation together with community moderation efforts results in decreased harmful interactions whichhappeninonlinespaces.
1.1.4
Cyber harassment shows a connection to the cultural and societal standards which exist in society. In areas which experience increasing cybercrime attacks against women users face greater risk of
online harassment especially women and vulnerable populations. Thecombinationofsocial status systems and gender roles and male-dominated social systems leads to increased harassment incidents which create safespacesforattackersbecausetheirvictimsremain silent.
Research on online silencing and gender inequality [18,19] demonstrates how societal expectations restrict women's ability to express themselves freely. Broad studies which use social and theoretical approaches to research [20-22] show that cultural norms and power imbalances together with institutional structures determine how cyber harassment becomes accepted and continues to exist. The understanding of these socio-cultural factors provides essential knowledge which enables the creation of effective educational programs and awarenesscampaignsandinterventionstrategiesthat focus on solving the main problems behind harassmentinsteadoftreatingitsvisibleeffects.
1.1.5
Cyber harassment creates effects that extend beyond digital platforms because it disrupts victims' psychological and social health. Researchers have discoveredthatpeoplewhoexperiencestresstogether with anxiety and psychological distress [20,21] will face difficulties in maintaining focus and they will experience sleep problems and decrease their time spentengaginginbothvirtualandreal-lifeactivities.
The long-term social effects of harassment emergewhensomevictimschoosetolimittheironline activities or they decide to stay away from particular social media platforms. Studies on user safety concernsandbehavioralchanges[17,18]indicatethat repeated exposure can lead to social withdrawal, reduced self-confidence, and even changes in professional or educational engagement. Research demonstratesthatcyberharassmentcausesemotional and social and mental health effects [21,22] because its impact creates multiple effects that continue to exist.
Cyber harassment occurs through established repeatingpatternswhichtakeplaceinstructuredtime intervals. Offenders usually select their victims through systematic approaches which result in ongoing abusive behavior. Studies on trolling and repeated harassment behavior [8,9] show that such actions are intentional, coordinated, and recurring ratherthanisolatedincidents.
The occurrence and severity of online harassment increase because of both environmental elements and external factors which include times when people use social media more and important global events and emergency situations. The research

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
shows that outside pressures and online patterns of behavior which users encounter lead to higher technology-enabled harassment rates [9,10] in these specificenvironments.
1.1.7 Legal and Protective Perspectives
Cyber harassment needs both strong legal systems and effective protective systems to achieve successful resolution. Research highlights the importance of laws, policies, and institutional safeguards [23-25] that aim to protect users, provide recoursetovictims,anddeterpotentialoffenders.
The existing legal system faces operational difficulties because people lack understanding of the law and because enforcement agencies operate weaklyandbecausepeoplefaceobstacleswhentrying to report violations. The research studies on legal frameworks and justice systems [25-27] demonstrate that successful implementation requires educational programsandplatform-basedsolutionstoenablelaws todeliveractualprotectionandjusticeforvictims.
1.2 Significance of the Study
Cyber harassment has become more common in society which requires complete research for its root causes and its development process and all its consequences. The research investigates victim responsepatternswhereonlineabuseoccurstoshow howharassmentdevelopsandprogressesandimpacts victimsthroughpsychologicalandsocialeffects.
The study uses victim experience findings and justice system research from [27,28] together with cyberbullying environment and institutional response research from [29,30] to create a complete understandingofthesubject.Theresearchestablishes safer digital environments which guide policy development and promote responsible online behavior to diminish the negative effects of cyber harassmentonindividualsandcommunities.
2. PROBLEM STATEMENT
The ability to communicate and connect through social media is available to users; however, the platform exposes many users (especially women) to online abuse. Today, we are able to implement reactive solutions that can help provide an efficient and user-focused solution for maintaining safety during communication in particular, when communicatingwithunknownindividuals.
KeyIssuesIdentified:
Many women have received inappropriate, creepy, and/orabusivemessagesviatheinternet.Harassment can be directed toward women through different methods; including minor inappropriate commentary aswellasseverelythreatening/stalkingbehaviors.
Experiences of online abuse can negatively affect women'semotionalwellbeingandoverallself-esteem.
Many women do not use social media or use social
media less than they would if they had not had negativeexperiencesonline.
The effectiveness of most current solutions; such as reporting and blocking; is that they only address the issueofabuseafterthefact.
There are many forms of indirect/subtle harassment that are hard for women to see and consequently demonstratethattheyoccurred.
Women have nowhere to go for assistance after receivingharmfulmessages.
There is currently no system for providing timely assistance to women who have received harmful messages.
Asolutionisrequiredthatpromotesandenhancesthe online safety of women in an effective and userfriendlymanner.
The main aim of this project is making a system thatoffersamoresecure,hassle-freeandharassmentfreeexperienceforwomenonsocial mediaplatforms. The system uses artificial intelligence for detecting, classifying, and hiding abusive or creepy messages in real time, giving users control over their digital interactions again. Specific Objectives of this project are:
ThefocusistobuildanAI-basedChromeextension which can monitor and analyse, in real time, Instagram DMs for messages that are potentially abusive,creepy,orharassing.Inthiscase,thistask would involve the multiclass classification of incoming messages into categories such as safe, creepy,mildharassment,mediumharassment,and severe harassment, based on linguistics and emotionaltone.
It provides total user privacy since all data is processed directly on the user's device, without sending any information to any server or cloud platforms.
To provide full control for the user over enabling and disabling the extension, thus being able to decide when the tool is working and how it treats therecognizedmessages.
Implement a strike system that monitors repeat offenders, so that whenever a user has been identified for multiple strikes, their messages wouldbeautomaticallydeletedorflagged,or both -creatingassuranceoflong-termsafetyforUSERS.
Let's implement action to reduce stress on the emotional and psychological levels for users being harassed online, and let them communicate freely andwithoutfear.
Encourage the further responsible and ethical use of AI in applications related to social safety, while making a case for continued research and innovationindigitalsafetytools. Withthesegoals,Guardianhopestoprotectusers

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
not only from harassment but to change the way technologyisusedinordertomakedigitalecosystems thatthriveonempathy.Thisisaprojectthatimagines aworldwheresocialmediacontinuestobeaplatform forstayingconnectedandcreative-notasourceoffear oranxiety.
Scope:
Thesystemisdesignedtodetectandfilterabusive or harassing text messages on the Instagram web interface.
It uses machine learning-based natural language processingtoclassifymessagesona severityscale rangingfromsafetosevereharassment.
All processes run locally at the user's device, thus providingdataprivacy.
Users have full control over when they want to enableordisablethistool.
This system tracks the behaviour of senders to protect users against repeated patterns of harassment.
Limitations:
The system currently accepts only text-based harassment;images,emojis,andvideoswhichmay containinappropriatecontentarenotanalysed.
A significant challenge that is presented is regardingthediversityoflanguage-themodelmay have a low degree of accuracy with slang, geographically specific phrases, and text that is purelynon-English.
Currently,theprojectisfocusedsolelyonthewebbased version of Instagram and does not include mobile applications or other social networking sites.
Real-time processing relies completely on the user's system capabilities, which I have found can impactitsspeedandresponsiveness.
The quality and balance of the training dataset determinetheaccuracyofthedetectionsystem;as online language evolves, it is a process that needs constantimprovement.
2.3 Applications
2.3.1 Real-world relevance
The Guardian Project has a particularly noteworthysignificanceintoday'sdigitalenvironment in which online communication has become commonplace. The sharp increase in users of social media has created an environment in which digital harassment, stalking and verbal abuse are prevalent, more so among women. Current reporting and moderation systems formal or informal are often either delayed or incapable of identifying subtle but no less harmful, manipulative, or repeated harassment. Guardian fills the gap by providing a smart, proactive and privacy-centric approach to
equippingusers withtoolsthatallowthemtoassume controloftheironlineexperience.
To put this in context, the extension provides immediate benefit for social media users that deal with unwanted or uncomfortable messages frequently-adifficultexperiencetonavigate,letalone secure mental peace and safety from. It reduces emotional exhaustion generated from continual exposuretotoxiccontentandenablesuserstoengage more fully in processes that require confident interactioninonlinespaces.
Lastly, Guardian tends to the individual user's safety while contributing to larger social designs for digital well-being, ethical online experience, and gender presence in technology inclusive products. Furthermore, it serves to make online environments safer, contributing to more women and marginalized individualsfearless,open,andfullparticipantsoftheir social,educational,andprofessionalnetworks.
Ultimately, Guardian is not simply a technology company offering an extension, it is a social design initiative formed out of a need for so many users to fosterarespectfulandharassment-freeonlineculture.
2.3.2 Potential Domains
While Guardian was conceptualized for the needs of Instagramdirectmessages,itsarchitectureanddesign principles are extensible to multiple platforms and domains.Itmightbeappliedinthefollowingareas:
Other Social Media Platforms:
The very same detection and filtering system can be applied to the integration of Facebook with Twitter (X), WhatsApp, Snapchat, or Telegram to safeguard users against abusive or harassing messages across platforms.
Email and Messaging Applications:
This capability could be integrated into email programs or work messaging systems, such as Gmail, Slack, or Microsoft Teams, and could also automatically filter messages to automatically delete spammessagesinaworkenvironment.
Online Gaming Communities:
Many (but especially women) experience toxic chats while playing multiplayer games. An AI similar to a guardian may be developed to oversee a game's chat systems and could automatically hide or censor inappropriate language or harassment, given prior recommendations.
Learning environments:
Moreover, educational environments may be utilized by this technology to help moderate discussion boards, chat rooms, virtual classrooms, and other learning environments and establish the expectations ofstudentsbeingrespectfulandpolitetooneanother.
Corporate/Customer Support Settings:
Employees in the roles of customer support often

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
endure rude or inappropriate language from customers. The results from Guardian can be used to filter out inappropriate communication in customer supportchatstomakeforabetterworkplace.
Public Forums and Comment Sections: It can also be extended to news websites, blogs, and communityforumstoautomaticallyhidetheharassing commentswhichwillhelpinreducingonlinetoxicity.
In all, the Guardian system shows wide-ranging versatility and can easily be extended from social media into any text-based communication platform requiring real-time, privacy-safe, intelligent text moderation. It constitutes a step forward in the integration of artificial intelligence with ethical computing to advance digital safety and user empowerment.
3.1
One big problem today? Harassment online has grown fast alongside social networks. Studies show techmakesiteasiertotargetwomen,withresearchon technology-facilitated violence against women [1,2] showinghowharmspreadsfurtherthanbefore.What stands out is how messages never stop - popping up hereandtherewithoutwarning.Beinghiddenbehind a screen lets people say things they might not in person. This constant pressure wears down mental well-beingmoredeeplyovertime.
Women facing cybercrime often encounter old biases dressed in new forms, as highlighted in [3,4] studies on patterns of cybercrime against women. Power imbalances shaped by gender play a big role behind these attacks. Online abuse tends to mirror how disrespect shows up in everyday life, just shifted intodigitalspaces.Researchspanningmanycountries points out that bullying travels easily between different kinds of apps, with global studies on cyberbullying trends [5] showing its spread across platforms.Itpopsupwherepeoplechatprivately,post publicly, even when they work together online. The patternstaysconsistent-nosinglesitecontainsit.
Looking at how people react when exposed to degradingmediashowsapattern-overtime,itmakes disrespect seem ordinary online, especially in [6] research on media-induced harassment and objectification. Studies tracking what targets feel during digital attacks reveal deeper wounds: unease creeps in, self-assurance fades, fear takes hold, participationdropsoff,whichisstronglysupportedby [7,8] findings on user experiences of cyber harassment.
When images reduce humans to props, behavior shifts without notice. Repeated scenes of exploitation quietly reshape expectations around interaction. Those caught in abusive exchanges often pull back,
voice dimming. What feels like background noise in entertainment may feed real distress far away from screens. Emotional weight piles up silently, altering whospeaksandwhostaysquiet.
When global crises hit, studies find more online attacks targeting women - linked to longer screen time, mixed with increased reliance on chat apps and virtualmeetings,particularlyobservedin[9]research on pandemic-related gender-based violence. Devices dominate daily life now, so risks sneak in quietly, which makes real-time monitoring essential. The complexitysurprisesmost: digital routines,emotional stress, and social patterns knot together, defying fast solutions.
Figuring out how online abuse changes over time helps create better tools to catch it. Studies of nasty onlinegroupsshowmeanactionsusuallygrowslowly, with escalation patterns in harassment communities [10] showing how behavior intensifies over time. Starting off with light mockery, things might shift toward focused bullying, danger signs appear later. If nobodystepsin,smalljabsturnintoorganizedcruelty.
Noticinghow people mean whattheysaymatters just as much as the words they type, which is emphasized in [11] intent detection research. A commentcanstingevenifitseemssmall -itspurpose shapes its impact. Over time, watching users interact reveals patterns nobody planned at first. When someone spends more hours scrolling through feeds, their chance of facing cruelty - or acting cruel - tends torise without warning, a trendsupported bystudies on [12] social media usage and cyberbullying behavior. What begins casually might shift into somethingheavierbyslowdegrees.
One key step forward involves sorting cyberbullyingbyhowseriousitis,asseeninseveritybased classification approaches [13], instead of just flagging it as present or absent. That shift brings clearer insight into digital harassment. Another angle looksatthe talk around eachmessage, because single linesoutofnowherecanmissthepointentirely,which aligns with [14] context-aware detection methods. Understanding what came before often changes everything.
How people act around each other shapes how often harassment happens. When others step in - or stayquiet-itcanmakeasituationworseorhelpstop it, according to [15] studies on bystander behavior. Lookingatwomen'sexperiences,ongoingharassment makes them more careful, say less, and join fewer conversations online - findings tied to [16] perceived safety and behavioral adaptation. Spotting clear insults matters, yet systems must also notice subtle actionsandcontextiftheyaretorespondwell.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
Lately, studies have turned attention to how varied languages and casual ways of talking shape onlinechats.Onsocial networks,people toss together differenttongues,shortcuts,made-upwords,symbols, plus looserulesforsentences - thismixtripsup most softwaretryingtokeeptrack.Meaningshidebetween lines, needing more than just word-for-word reading to catch them. Older methods struggle because they miss context, failing to label these messages right. Finding meaning beneath the words matters more now, so today's methods build models aware of surroundings, ignoring language barriers. Because people speak differently, systems must adaptworkingwellnomatterwhousesthem,wherethey're from,orhowtheyexpressthemselves.
Studies focusing on how safe women feel using internet platforms, particularly research on digital safetyconcerns[17],showjusthow harditis tokeep personal information protected. Without clear rules set by authorities, handling these attacks becomes eventougher.
When people face online abuse, they tend to step back - maybe by posting less or leaving apps altogether, as seen in [18] studies on user responses to harassment. Women especially pull away from speaking out in open forums after being targeted, whichishighlightedin[19]researchongender-based silencing. This pulling back feeds into deeper imbalances over time. Silence grows where voices shouldbeheard.
A heavy toll on the mind often follows cyberbullying,withresearchonmentalhealthimpacts [20] showing how digital attacks fuel stress, anxious thoughts, sadness, along with lasting inner wounds. Work centered on India's experience with online mistreatment, particularly regional studies on cyber harassment[21],addsdepth,uncoveringsocialnorms andbeliefsshapingwhogetstargeted,howit’sseen.
When it comes to online abuse aimed at specific genders, findings from studies on gender-targeted cybercrime [22,23] show rules are needed. Still, reports about how those rules work - or do notsuggest laws lag behind new ways people talk online, as discussed in [24,25] legal protection and enforcementresearch.Becauseofthisdelay,toolsthat watch and respond instantly might help close what's missing.
One reason more people want smart tools is how trickyonlineabusehasbecome.Becauselawsandtech must work together, machines might help back up new rules, with studies on integrating legal and technologicalsolutions[26]supportingthisapproach.
Feminist takes on cybercrime [27] unpack how online abuse is shaped by systemic forces, showing why tech must respond to gendered risks. From anotherangle,researchintouserperceptionsofharm and justice [28] reveals that responses should speak plainly - making outcomes clear matters just as much asfixingissues.
Early warning tools matter more when dealing withdigitalabuse,asfindingsonharassment,stalking, and blackmail [29] suggest the importance of early intervention. Instead of waiting, watching patterns helps - research into risk factors and institutional responses [30] shows steady oversight cuts harm fromonlinebullying.
Most new research points to one thing clearly: howwellasystemcanchangemattersjustasmuchas its starting design. When language shifts - slang spreads, shortcuts pop up, hidden meanings emergeold methods often fall short. Because online abuse rarely stays the same, tools stuck in place lose strength fast. A fresh update here or there can help, yet its value vanishes if the system refuses to adapt with each change. Sharpness comes not from single fixes, rather through steady tweaks that question old assumptions. Today’s method may fail tomorrow’s unseen shifts unless choices bend without breaking. Performance across forums, apps, or chats depends lessonsizeandmoreonreadinesstoevolve.
Even so, knowing more about online bullying hasn’t fixed all tech flaws yet. Most older tools just huntkeywordswithoutseeinghowwordsfittogether now. On top of that, those smart algorithms need hand-craftedtraits,makingthemstiffwhenfacedwith newtypesofdata.
Though deep learning has boosted how machines grasp context, it usually demands heavy computing power along with vast amounts of labeled data. Because of their design, several current systems struggle to operate in real time, limiting usefulness when quick responses matter. Running on cloud infrastructure introduces issues too - delays creep in, personal information may be exposed, and keeping datasafegrowsharder.
Most times, things stay just beyond grasp when there is no way to touch or turn a knob. Off-the-shelf gadgets often do their thing without asking, making changes nearly impossible. What stands clear is a growing demand - simple designs, built around people,workingfastwithoutdelay.
Beyond tech hurdles, how people interact with tools shapes how well bullying detectors work. When interfaces show clear reasons behind flags and let users adjust settings, confidence grows. Picture softwarethatratesthreatlevelsplainlywhilenudging gently instead of shouting alarms - people tend to

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
prefer that. Too many mistaken red flags wear down attention fast, causing folks to ignore warnings altogether.Smoothblendsintodailyapphabitsmatter just as much as correct spotting, pushing design toward calm reliability rather than perfect scores alone.
Still,aftermanystudies,somepiecesstaymissing whenspottingonlinebullying.Whatstandsout?Tools that catch abuse right as it happens just do not exist yet. These delays matter - most setups only respond afterharmoccursinsteadofstoppingitearly.
Few tools handle how badly someone might be targeted, stuck instead with yes-or-no labels. What shows up next could miss mockery or hidden slurs unlesssystemslearnnuancebetter.
Fears over private details living on distant servers keepgrowing-thispushesdemandforlocalhandling. People now watch closely where information goes, so tools that guard personal content matter more than ever.
Facing such issues head-on, the new Guardian setup uses a tweaked DistilBERT brain inside a web browser add-on for live warnings based on what's happening. Instead of sending data away, everything gets handled right on your device - keeping things private. By sorting abuse into different intensity buckets, it tells apart mild nudges from serious threats. Unlike older tools, this one moves quicker, grows easier, yet feels simpler to use without weighingusersdown.
4.1 Technologies, Tools and Data Selection
Guardianisbasedonahybridmixofstate-of-theart web technologies and the machine learning ones, which have been trained to work together flawlessly insidetheuser’sbrowser.Thefront-endofthesystem is being developed in the form of a Google Chrome extension, in which the latest Chrome Manifest Version3notonlyplaysamajorsecurityrolebutalso enhancesperformance.
This choice allows the extension to directly interact with the web-based interface of Instagram, reading the message content without alienating the user by requesting the permissions on the intrusive side.
The AI model is running in the local server designed with the FastAPI as the working backend systemandPythonastheprogramminglanguage.The FastAPI is responsible for the system's ability to promptly process messages as they come in, thereby minimizing latency and guaranteeing responsiveness thatisclosetoreal-time.
The AI model is being developed using a dataset thatconsistsofInstagramdirectmessageswhichhave
been processed with great care and anonymity. The dataset includes a wide variety of communication types.Labelingwasdonebyhumanstocategorizethe messages into safety categories from harmless to various levels of harassment and stalking. Data cleaning was a major component of the process and included the removal of duplicates, normalization of textual variations, and addressing of class imbalance toavoidbiasduringtraining.
AttheheartofGuardian'sAIskillsisaDistilBERT model, which has been tuned to perfection-an optimized and compact version of the original BERT transformer model. DistilBERT is able to handle a lot more of BERT's language understanding capabilities, but with fewer parameters that reduce the inference time.
The model was developed as a sequence classification neural network, and the annotated dataset provided it with a large variety of examples acrossmanyriskcategories.
HuggingFace’sTrainerAPIplayedacrucialroleinthe fine-tuning process by managing epochs, batch sizes, and learning rates in an efficient manner so as to increase the prediction accuracy. Regular validations on the subset of a validation dataset ensured the model’s strength and that it could generalize to new messages.
The method allows for very sophisticated detection, revealing even the most discreet linguistic signs like threats, stalking behaviours, unwanted attention and rude language that vary in severity, amongothers.
Data Collection and Labelling:
The training corpus was created in due course of the systematic extraction and anonymization of raw Instagram direct messages. Human annotators assigned several classification tags indicating the safetylevelofthemessageandthetypeofharassment over the message, which was then followed by preprocessing stages that included tokenization, normalization, and implementation of balancing techniques designed to ensure equal representation acrossthedifferentcategories.
TheDistilBERTtransformerunderwentfine-tuningon this dataset through the application of supervised learning paradigms. Model parameters were iteratively modified and validated by means of crossvalidationtechniquestoachievethebestclassification performancewhilepreventingoverfitting.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
Local AI Service Setup:
After the training, the distilled model was exported and incorporated into a Python FastAPI service that was running locally on the user’s computer. With this configuration, it is possible to make fast and private predictions without the need for external API calls or anyclouddependency.
Extension Development and Integration:
ThecontentscriptthatispartoftheChromeextension continually tracks the document object model (DOM) of the Instagram web interface to almost instantly capturethereceivedDMtext.
The extracted messages are sent through HTTP requests from the background service worker of the extension to the local FastAPI endpoint for the classificationprocess.
The classification results come back right away, leading to visual cues such as color-coded highlights thatshowtherisklevelsofthemessages.
A strike system counts the flagged messages for each sender. When four infractions are reached, the sender’s messages will be hidden automatically, and theuserwillreceiveanotificationalert.Thisfeatureis unique as it combines automated protection and user awarenessperfectly.
Data Persistence and User Interface:
All the metadata consisting of the strike counts and hidden users is stored using Chrome's Local Storage API. This allows the retention of data through the browser sessions without depending on external storage. The extension’s popup interface shows the users currently hidden senders and gives them the option to manually unhide contacts, thus offering completecontrolovermoderation.
Deployment, Testing, and Validation:
The whole system is installed by utilizing the Chrome developer mode to load the unpacked extension, which opens the door for quick iterations and easier debugging. The team went through a series of tests that were aimed at checking the system's responsiveness, classification accuracy, and usability while at the same time assuring the system's robustnessandfriendliness.
Guardian, the cutting-edge and privacy-respecting moderationtool,anditsusersaretogetherfightingthe battle against harassment and creepy messages in a waythatisultimatelyineffectiveandnon-translucent, relyingsolelyonthepowerfulyetsimplecombination of lightweight transformer models, local inference, andsmartUIintegration.
Guardian is designed in a modular system architecture comprising two major components: Frontend(ChromeExtension)andBackend(Local ML Server), which interact in real time to analyse and filter social mediaDMs. The textis pre-processedand passed to a fine-tuned DistilBERT model trained for multi-classharassmentdetection.
Themodelclassifiesmessagesintosixcategories ranging from safe content to different levels of harassment such as creepy flirt , mild harassment stalkingbehavioretc.
Based on the severity, the system either allows themessageorinstantlyalertstheuser.
Backend (Local ML Server)

Frontend (Chrome Extension)


Figure5.1:Twomodulesoftheproject:ChromeExtension, MLserver.
1. Social Media Web Interface
Thisisthefront-endplatformor,inotherwords,Social Media's website, where the user views and receives messages.
Since Chrome extensions have the ability to interface with webpages, the Guardian extension "listens" for newmessagesthatappearintheDOM.
Here is where the system starts - the content script is injectedintothispagetoreadmessages.
2. Content Script
This content script runs inside the context of the Instagramwebpage.
It continuously monitors the message area using DOM selectors.
Eachtimeanewmessageappears,it:
Extractsthetextcontent.
Sends it to the background service worker foranalysis.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

Figure5.2:DetectionofincomingtextsusingContent Script.
3. Background Service Worker
This is the central controller of the Chrome extension.
Itreceivesmessagesfromthecontentscript.
Foreverymessage:
ItsendsthetexttotheMLInferenceAPI(FastAPI backend).
Gets the predicted label, such as Safe, Creepy, Harassment.
Updatesstrikecountsforeachsender.

Figure5.3(a):Serverresponseas‘Safe’

Figure5.3(b):Serverresponseas‘Harassment’
AI backend is running locally at http://127.0.0.1:8000.
It hosts the fine-tuned DistilBERT model, trained ontheharassment/DMdataset.
Whenthebackgroundscriptsendsamessage:
APIpredictsthecategoryofthemessage.
Returnsalabelsuchas:
Safe
Creepyflirt
Mildharassment
Mediumharassment
Severeharassment
Oncetheclassificationisreceived:
The background worker sends the result back to thecontentscript.
Thecontentscriptvisuallyhighlightsorhidesthat messageintheInstagramUI.
Example: Safe→Nochange
Creepyflirt→Yellowhighlight Harassment→Redhighlightorblurredmessage
2. User Alerts & Strike Tracking
Each sender has a strike counter stored in chrome.storage.local.
When the user repeatedly sends toxic messages (forinstance,4strikes):
Itautomaticallyhidesthesender.
The user can unhide them at any time using the extensionpopup.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
Instagram Web

6.1 Output Results

Detected text
The Guardian Chrome Extension successfully integrates with the Instagram web interface and classifies DirectMessages(DMs)inrealtime.

Content Script

Highlight Message Background Service Worker

Sends to ML
Interface API

Receive Classification
Figure5.4:SystemArchitectureforMulticlassHarmful ContentClassification
5.2 Tools, Platforms and Libraries used
Table5.1:Tools,platforms,andLibrariesusedinproject.
COMPONENTS TOOLS/TECHNOLOGIES USED
Language Python(Backend),JavaScript (Frontend)
ML Model DistilBERT(HuggingFace Transformers)
Backend Framework FastAPI
Browser Extension Framework ChromeManifestV3
Dataset
Customlabeleddataset
Frontend Libraries HTML,CSS,JavaScript
Storage ChromeLocalStorage
Testing Platform Chrome(SocialMediaWeb Interface)
IDE VSCode

Figure6.1:Guardian:ChromeExtension
Observedoutcomes:
Safemessagesremainunaltered.
Creepyorflirtymessagesarehighlightedinyellow.
Harassmentortoxicmessagesarehighlightedinred.
Repeated offenders are automatically hidden, and the userisalertedthroughapopupnotification.

Figure6.2(a):Classificationofmessagesreceived.

Figure6.2(b):Classificationofmessages
When some harassing text/message is noticed, after 3 strikes,thewarningmessageispoppedup.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

Figure6.3Pop-upwarningmessage
Afterthepoppingupofthewarningmessagethat particular harassing message is highlighted with different colours according to the severity of the text.

Figure6.4:Messagehighlightedaccordingtoseverity
6.2 Classification model evaluation
6.3
Theoverallperformanceofthemodelonthetest dataset,consistingof1963samples,standsat0.78or 78%. The succeeding analysis provides a snapshot of theperformanceacross all sixclassesofoutput,using thestandardclassificationmetricsofPrecision,Recall, andF1-Score.
It has shown very good performance in classifyingthemostdistinctiveandcriticalcategories: Stalker Behaviour: The model performs highly reliablyintheidentificationofstalkerbehaviour,with anextremelyhighF1-scoreof0.96,precision0.97,and recall 0.95, hence robust, which indicates very effectivefeatureengineeringforthisclass.
The major "non-harmful" class performs very well and includes an F1-score of 0.90. Besides that, precisionandrecallequateat0.90,whichimpliesthat the model captures safe interactions quite well and themisclassificationofothertypesassafeisrare.
Another critical category that does well would include severe harassment, with an F1 of 0.76 (Precision:0.79,Recall:0.73).Precisionishighhere;if itidentifiessomethingassevereharassment,itwillbe correct.

7. CONCLUSION
7.1 Summary of Findings
Guardian, the project revealed some necessary information about the effect of online harassment on women and why options have not yet met the bar. As the study progressed, one fact became clear: women are repeatedly harassed on social media platforms. These types of attacks are not necessarily loud or obvious. But they’re little constant nudges that can makeapersonfeeluneasyorunsafe,or chippedaway at over time. The majority of the well-established methods of defense only act after damage has been inflicted. This emphasized the need for an exposurereducing,pre-damageintervention.
Testing Guardian on Instagram in real time proved that proactive protection is more than a dream it really is achievable. The program went beyond picking out blatantly abusive messages to discovering unsettling trends that often remain hidden in plain sight. That included what team membersdescribedasunwantedflirtingthatbecomes more aggressive, manipulative language, threats and other repeated creepy conduct. These exchanges can make women feel trapped, even when no single message is so awful that it should be flagged by established moderation. It demonstrated how contextual AI can solve this problem, rather than just blockingkeywords.
Anotherrealizationwehadwasthatuserscarea lotaboutprivacy.Foreventhemostprivacyconscious among us, there are lingering doubts about not only where your personal messages go during a sensitive exchange, but also who gets to look at them, who

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
storesthemandwhoownsthatinformation.Guardian was built with these issues in focus. Because all processing is done on the user’s device, no data is transmitted. Nothing is saved, shared or tracked by outsidesources.Thismadeitpossibleforuserstofeel safe, but not surveilled. That’s not something existing socialmediatoolshaveWhiletesting,werealizedthat control matteredtofeeling safeonline.Elementssuch asthestrikesystem,gentlewarnings,andthecapacity tobypassamoderationgaveuserstheimpressionthat theywereincontrol.
Rather than restricting how they communicate, Guardian gave users the ability to define their own limits. It was not about silencing other people; it was about arming users with the means to choose how and when they wanted to protect themselves from harm. The Guardian programme demonstrates a safer web is achievable. With good design and people-centred technology, social media can be a less frightening place. It can be restoredto whatit wasalways intended to bea place for connection, for creativity, and for meaningful expression where everyone, and especially women, canfeelprotectedandrespected.
7.2 How it meets the objective
TheobjectiveTherewasonetheprojectwaseasy to identify at the beginning. The purpose of it was to build a tool that empowers women to feel safer and more at ease sharing their lives on social media, particularlyInstagram.Lookingbackontheobjectives set out at the start, Guardian is delivering in spades. The second big aim was to identify abusive, disturbing,orharassingmessagesastheycomein,not after users have already seen them. By coupling an adapted DistilBERT model with real-time monitoring through a Chrome extension, the much lauded Guardian emerged out of the shadow. It scours the content of messages in a matter of seconds and indicates the potential level of risk before a user receives the message. The more than the usual blocked message checks was another important aim. Guardian labels messages into a range of categories, from safe to extreme harassment. It’s a bit more nuanced of an approach for users to understand the type of experiences they’re actually having rather than seeing everything under one umbrella. Privacy was a concern from the beginning, and this requirement is adequately fulfilled. Unlike conventional systems, Guardian does not transmit or leak the content of messages to any server All the processes, such as text processing, classification, and decision making, are performed locally on the user’s device.This keepseverythingunderlockandkey.
User control and autonomy was also important. ItissimpleforuserstotoggleGuardianonoroff,and to re-examine or unhide any messages that were
flagged, on their terms. It’s all with their permission, andtheirknowledge.
The two strikes policy was a success in dealing with long-term harassment. automatically hide that user’sfuturemessagesfromyouaswellandsendyou an alert. This eliminates the burden on users having tokeepan eye onorreport offenders whorepeatedly breaktherules.
Now emotional well-being was what motivated thisentireproject.Guardian enhancesthisbylimiting the amount of traumatic content being viewed and engagingindiscussionsinamoresecure atmosphere. As Guardian runs silently in the background, women will be able to enjoy handling their Instagram accountswithgreaterpeaceofmind,moreconfidence andfreedom.
8.1 What Improvements can be made
Guardian already provides robust protection from online abuse, but a few tweaks could make the systemmorepracticalanduser-friendly inthefuture.
One important factor that needs to be improved is the language support. At the moment the model performs best in English, but harassment is increasingly expressed in regional languages, slang, and mixed language communication, particularly on socialmedia.Withmultilingualtrainingdata,Guardian wouldbeabletohelpawiderrangeofusersbelonging toavarietyofcommunitiesandcultures.
Further enhancement in terms of classification accuracytoconsideremotionaltone,contextandprior conversation is also possible. This would help it identify even more subtle forms of harassment, sarcasm, threats dressed up as compliments, and patternsofabusethatmightrepeatovertime.
Further the UI can also be tweaked to allow more customization. For example, users could define their own sensitivity levels or choose content categories that are most relevant to them. A dashboard for flagged messages, strike counts, and prevalentharassmenttypescouldpotentiallyincrease userawarenessandcontrol.
Increased support for platforms is another useful direction. Guardian is for now only built for Instagram web messaging, but abuse takes place across many digital platforms. More such developments could also bring the real time protection support to WhatsApp Web, Facebook Messenger,Xandevenemailproviders.
Safetyescalationfeaturescouldbeaddedaswell. Forexample,alertingusersaheadoftimethattheyare responding to a known harasser, or displaying quick links to support services, mental health resources, or platform reporting tools in the event that the harassmentintensifies.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
Users who want to contribute anonymized data toimprovetheAImodelcouldoptintoa securecloud mode.Thiswouldenableclassificationperformanceto increase as a function of time and at the same time keepuserprivacyintact.
Guardian will be more necessary as digital communication changes. But online harassment is increasingly a rite of passage for many women. Traditional reporting systems are typically slow, stressful and reactive. Only when users are harmed can anything be done. Guardian overcomes that by taking action ahead of damage. Prevention is Now a BigPartofOnlineSafety.
Going forward, Guardian could be a game changer for the way social platforms treat user wellbeing. Features such as real-time message protection, AI-basedseveritydetection,andlong-termtrackingof harasser behaviour point to a model of social media thatcouldbetakenupgloballyby companies.Itcould help change policy and pave the way for new standardsinresponsibleplatformdesign.
In addition to individuals, Guardian could also serve communities, schools and organizations. Institutions interested in providing a safer online space for their members particularly younger women orindividualsnewtotheworldof social mediacould incorporate this instrument into their digital safety curriculums. This could enhance education and accountabilityinrespectful communication.
Kineticsalsosawpotentialforcollaboration with law enforcement and cyber safety officials. While Guardian keeps everything private by default, users couldchoosetosharerecordsofsevereharassmentin secure formats to help report and stop repeat offenders. This could make legal action more accessibleforthosewhoneedit.
1 Rajan, Benson. "Harassment and abuse of Indian women on dating apps: a narrative review of literature on technology-facilitated violence against women and dating app use."Humanities andSocialSciencesCommunications12.1(2025): 55.
2 Balabantaray,SubhraRajat,MausumiMishra,and Upananda Pani."a sociological study of cybercrimesagainst women inindia: deciphering the causes and evaluating the impact on the victims."international journal of asia-pacific studies19.1(2023).
3 Alauddin Middya, Bimal Mandal (2021). Cyberviolence Against Women: Exploring
Patterns of online Gender Based Harassment (2021)
4 Bhat, Rashid Manzoor, and Peer Amir Ahmad. "Social Media and the Cyber Crimes Against Women-A Study."Journal of Image Processing and Intelligent Remote Sensing (JIPIRS) ISSN(2022):2815-0953.
5 Negi Advocate, Dr Chitranjali. "An overview of worldwide cyberbullying and cyberviolence against women, teenagers, LGBTQ on social media: Facebook, Instagram, Telegram, WhatsApp, Snapchat, YouTube, LinkedIn and Twitter."Boston College International and ComparativeLawReview,Forthcoming(2023).
6 Galdi, Silvia, and Francesca Guizzo. "Mediainduced sexual harassment: The routes from sexually objectifying media to sexual harassment."SexRoles84.11(2021):645-669.
7 Burke WinkelmAn,Sloane, etal."Exploringcyber harassment among women who use social media."Universal journal of public health3.5 (2015):194.
8 Qian Zhang, Qinxuan Chen, Xinrong Gu (2024). Trolling, Cyberstalking, Body-shaming, Slutshaming – A Study on Online Abuses of Social Media(Zhang,10)
9 Mochamad Iqbal Jatmiko, Muh. Syukron, Yesi Mekarsari (2020).Covid-19,Harassment and Social Media: A Study of Gender - Based Violence Facilitated by Technology During the Pandemic (Jatmiko,2024)
10 Kejsi Take, Victoria Zhong, Chris Geeng, Emmi Bevensee, Damon McCoy, Rachel Greenstadt (2024). Stoking the Flames: Understanding Escalation in an Online Harassment Community (Take2024,23)
11 Abarna, Sheeba, et al. "Identification of cyber harassment and intention of target users of socialmedia platforms."Engineering applications ofartificialintelligence115(2022):105283.
12 Barlett, Christopher P., et al. "Social media use and cyberbullying perpetration: A longitudinal analysis."Violence and gender5.3 (2018): 191197.
13 Talpur, Bandeh Ali, and Declan O’Sullivan. "Cyberbullying severity detection: A machine learning approach."PloS one15.10 (2020): e0240924.
14 Ashraf, Noman, Arkaitz Zubiaga, and Alexander Gelbukh. "Abusive language detection in youtube comments leveraging replies as conversational context."PeerJComputerScience7(2021):e742.
15 Herry, Emily, and Kelly Lynn Mulvey. "Gender‐base cyberbullying: Understanding expected bystander behavior online."Journal of Social Issues79.4(2023):1210-1230.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
16 Sahu, Tamanna, and Ritu Raj. "The role of social media in shaping Women's paranoia about Harassment and Safety."International Journal of Interdisciplinary Approaches in Psychology3.5 (2025):1333-1344.
17 Nayayani, A. (2024). Women’s Safety in Digital Space. Indian Journal of Public Administration, 70(3),546-561
18 Chadha, Kalyani, et al. "Women’s responses to online harassment."International journal of communication14(2020):19-19.
19 Nadim, Marjan, and Audun Fladmoe. "Silencing women? Gender and online harassment."Social ScienceComputerReview39.2(2021):245-258.
20 Ghowrui,Ahana,etal."Effectsofcyberbullyingon women’smentalhealth."IJPR6.1(2024):25-29.
21 Lal, Disha, Udaya Kumar Giri, and Shrish Kumar Tiwari. "Virtual Vulnerability: Addressing Cyber Harassment against Women in India."DS Journal ofCyberSecurity2.3(2024):1-14.
22 Dar, Showkat Ahmad, and Dolly Nagrath. "Are Women a Soft Target for Cyber Crime in India."Journal of Information Technology and Computing3.1(2022):23-31.
23 B. Vijayalaxmi. (2020), Cybercrime Against WomeninIndia:ACriticalAnalysis.(B,2020)
24 Neha Gupta, Soyonika Gogoi (2025). Cyberbullying And Gender: An Analysis Of Legal Protections For Women On Social Media (Gupta, 2025)
25 Mythili, K. C., and K. Nagamani. "Safeguarding womenindigitalspaces:Legalresponsestocyber harassment and objectification social media."Development Policy Review43.5 (2025): e70039.
26 Nusrat Ali Rizvi (2025). The Role of Law in Combatting Gender-BasedViolence on Social MediaandOnlinePlatforms(Rizvi,2025)
27 Lazarus, Suleman, Mark Button, and Richard Kapend."Exploringthevalueoffeministtheoryin understanding digital crimes: Gender and cybercrime types."The Howard Journal of Crime andJustice61.3(2022):381-398.
28 GabrielGrill,JaneIM,SaritaSchoenebeck,Marilyn Iriarte (2023). Women’s Perspectives on Harm andJusticeafterOnlineHarassment(Grill,2023)
29 Orusa Karim, Sumbl Ahmad Khanday (2026). Cybercrime and Women-Online Harassment, Stalking, And Blackmail; A Sociological Analysis (Karim,2026)
30 Gallegos, Ada, et al. "Cyberbullying against women in digital environments examining manifestations, risk factors, and institutional responses."DiscoverPsychology(2026).