
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
ATHARVA WAGH1 , MOHD SAQLAIN2 , MOHD RAFE3, MUSAYYAB MULLA4 , Manasi Shimpi5
1Atharva wagh, Frontend developer (Team Leader)
2Mohd saqlain, Backend developer
3Mohd rafe, Database configurator
4Musayyab mulla, networking manager
5Manasi Shimpi Professor (Team Mentor), Dept. of computer, Abdul Razzak Kalsekar Polytechnic,Panvel, Maharashtra, India
Abstract - The Cyber Owl ecosystem is a privacy-centric, cross-platformparental monitoringsystemthatisintendedto monitor abusive language and explicit visual content in realtime. Toaddress theprivacyissues ofcloud-basedmoderation, thesystemusesan"IntelligenceLocalized"architecture,where allsensitiveaudioandscreendata is processedentirelywithin the Windows machine of the child. A stealthy background daemon captures system audio using loopback capture and screen frames, which are then passed through a highly optimizedCPU-friendlyedge machinelearningpipeline,which includes Speech-to-Text, a four-stageNLPclassifier,andCNNbased NudeNet. When sensitive content is detected, a concurrent Flask and SocketIO backend acts as a broker, relaying encrypted notifications to a remote Flutter-based Android dashboard. This is a comprehensive system that enables parents to receive real-time threat notifications, historical data, and remote device controls, all while maintaining robust child safety with data privacy.
Key Words: Cyber Owl, Content Moderation, Child Safety,Privacy-Preserving,EdgeComputing,Multimodal Detection, Real-Time Monitoring, Abuse Detection, Nudity Detection, Natural Language Processing , Computer Vision, Edge AI, Speech-to-Text, TF-IDF Vectorization, Support Vector Machines, Convolutional Neural Networks, Nude Net, Adhocratic Algorithm, Parental Controls, On-Device Inference
Theexponentialgrowthindigitalmeansofcommunication has revolutionized global connectivity but also placed vulnerable users, especially children, in serious online threatssuchascyberbullying,harassment,andexplicitvisual media. These online threats necessitate proactive and intelligent solutions that can detect abusive language and nudityinreal-timebeforetheyaffecttheuser.Thisresearch proposesCyberOwl,aprivacy-preservingecosystemforthe moderationofabusivelanguageandnudityinonlinemedia. The system operates entirely on consumer-grade edge hardware,providingarobustsolutionfordigitalparenting withoutcompromisinguserdataprivacy.
Traditional content moderation relies heavily on human reviewersorcloud-basedautomatedsystems,bothofwhich present fundamental limitations. Human moderation is inherentlyreactive,struggleswiththesheervolumeofdaily generatedcontent,andinflictsaseverepsychologicaltollon workers. Conversely, cloud-based AI solutions introduce significant privacy vulnerabilities and latency issues by continuouslytransmittingsensitiveusermedia toexternal corporate servers. These challenges coupled with growing parentalconcernsandstrictdataprotectionregulationslike GDPR and COPPA motivate the development of an edgecomputing framework that performs sophisticated threat detection locally, effectively balancing child safety with absolutedataprivacy.
The Cyber Owl framework operates through a highly optimized, three-tiered architecture built for edge computing.ThefirsttierisastealthWindows background daemon that captures system audio via loopback and samples screen frames, routing them through a localized machinelearningpipeline.Thismultimodalinferenceengine seamlessly combines a four-stage multilingual text classification cascade utilizing Adhocratic deterministic matching and TF-IDF vectorization with a CPU-optimized NudeNet convolutional neural network. When harmful contentisidentified,thesecondtieraconcurrentFlaskand SocketIO backend bridge brokers encrypted alerts and synchronizesdatabases.Finally,thesealertsarepushedto the third tier, a remote Flutter-based Android dashboard, enabling immediate parental oversight without ever transmittingtherawaudioorvisualmediafiles
While the system shows clear signs of effective real-time detection capabilities, it also exhibits certain practical limitations.Thetextmoderationsystemisbasedsolelyon speech-to-text transcriptions of system-captured audio

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
streamsanddoesnotincludedirectkeystrokeloggingortext inputmonitoringforencryptedtextmessageapplications. Additionally,inordertomaintainlowCPUutilizationrates, thesystem'svisualanomalydetectionsystemislimitedtoa samplingrateofonly5framespersecond,whichmayleadto missed anomalies in high-framerate full-motion video streams. Also, the system's architectural validation and performancebenchmarksarestrictlylimitedtomid-range Windows-basedoperatingsystemplatforms.
TheBagOfWordsModelisatextrepresentationmethodin whichthetextisrepresentedasanunorderedcollectionof words. This method loses grammar and word order but retainsfrequencyinformation.However,thismethodisnot efficientduetoitshighdimensionalityandfailuretoaccount for any semantic relationships between words. Term Frequency-InverseDocumentFrequencyisanextensionof theBagOfWordsModelinwhichthediscriminativepower of words in a document collection is used for weighting. Whenitcomestokeywordmatchingusingmultiplepatterns, the naive approach is to slide a window through the text. Thisisimpracticalduetoitsworst-casetimecomplexityof O(n*m*k).Ontheotherhand,theAho-Corasickalgorithm is a finite state automaton constructed from the set of patterns. This allows for matching all patterns in a single pass.

Fig -2.1:Aho-Corasickautomatonstructureshowingthe rootnodebranchingintodifferentcharacterstatesfor patternmatching.
Support Vector Machines (SVM) operate as supervised learningtechniquesthatusehyperplanestooptimizeclass separations.Duringbinaryclassification,theSVMalgorithm isdesignedtofindahyperplanethatmaximizesthemargin between two classes, using a regularization parameter to optimize between maximizing the margin and minimizing training error. Since SVM is not designed to provide class probabilities,Plattscalingisusedtofitasigmoidfunctionto
SVMoutput,providingclassprobabilitiesneededfornuanced thresholdingdecisions.
Convolutional Neural Networks (CNNs) serve as deep learning architectures specifically designed for processing grid-structureddata,makingthemthedominantapproach for visual content classification. The core mathematical operation in these networks is the convolution, wherein learnable filters slide across the input image to produce featuremapsthatdetectspecificpatternssuchasedgesor higher-leveltextures.

Fig -2.2: A structural diagram of a Convolutional Neural Networkshowingtheinputlayer,sequentialconvolutional layerswithfilters,poolinglayers,
Themultimodalfusionstrategiesdefinehowtocombinethe dataprocessedthroughseparateprocessingpipelines.Early fusion, also known as feature-level fusion, involves concatenating features from separate modalities before a classification step, allowingcross-modal interactions tobe learnedbutrequiringintricatefeatureengineeringtocope withheterogeneousinputdata.Ontheotherhand,latefusion, alsoknownasdecision-levelfusion,involvescombiningthe outputsofseparateclassifierstrainedonseparate,modalityspecificmodels.
For content moderation, where sensitivity and recall are prioritized, a late fusion decision engine employing an OR rule is particularly appropriate, resulting in an unsafe classificationifeitherthetextualorvisualmodalityexceedsa certainthreshold.VisualCheck:Theimageisofaharmless videogamemenu.(Score:Safe),Audio/TextCheck:Asevere threat is yelled in the voice chat. (Score: Unsafe), Final Decision:Safe(Visual)ORUnsafe(Text)=UNSAFE.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

Fig -2.3:Threefusionmethodframeworks:(a)Early fusionframework;(b)intermediatefusionframework; and(c)latefusionframework.
The system design strictly follows the basic principles of modularity,concurrency,gracefuldegradation,andprivacy bydesign.Modularityensuresthateverymodulefunctions in isolation for easy testing and updating. Concurrency is achievedthroughtheuseofseparatedaemonthreadsforthe audio and visual pipelines to take advantage of CPU utilization. A major constraint is that the analysis of any content must be performed entirely locally on the user's devicewithoutanytransmissionofanyaudio,video,text,or detection data to any external servers. Additionally, the systemissubjecttocertainnon-functionalrequirementsthat must be strictly satisfied. These include ensuring that the overall latency for the entire detection process does not exceed 500ms, CPU utilization is kept below 25%, and memoryutilizationdoesnotexceed500MBwithouttheuse ofGPUacceleration
Theoverallarchitectureisacleanlayeredone,optimizedfor separation of concerns. The Data Acquisition Layer is the entrypointofthesystem,anditisherethatthesystemaudio iscapturedthroughaloopbackrecordingmechanismviathe SoundCardlibraryandscreencapturesatarateof5frames persecondviatheMSSlibrary.
The next layer is the Processing Layer, where the actual detectionengineswillbeimplemented,transcribingaudio chunksviatheGoogleSpeechRecognitionlibraryandimages via the NudeNet detection model. Finally, the overall structure of the REST API endpoint will serve as a bridge between the backend business logic and the frontend presentationlayer.
Thesystemutilizesasolidrelationaldatabaseschemathatis gearedtowardslocaltelemetrystorage,userhandling,and local caching. Some of the major entities in the schema includetheUSERStableforauthentication,theCHILDREN table for account handling through foreign keys, and the DEVICES table for a list of registered hardware. All past detection occurrences are securely stored in the DETECTION_LOGStable,whichcontainsdetailssuchastime stamps,sourcemodality,labelapplied,andconfidencelevel. Another table, ALERTS, manages the status of alerts. A separate KEYWORDS table is used as the basic store for multilingual abusive words for the text classification pipeline.
Audio Pipeline and Speech-to-Text; The first step in the textualanalysisisthecaptureofthesystem’soutputaudio usingtheSoundCardlibrary.Thislibraryallowsforloopback recordingacrossplatforms.Theresponsetimeisreducedby segmentingthecapturedaudiointosequentialsegmentsof 2.5secondsbeforeitisprocessed.Thissegmentedaudiois then processed using the Google Speech Recognition API, which recognizes the spoken language in real-time and provides accurate transcriptions for multiple languages, includingEnglish,Hindi,andMandarin.
The transcribed text is then subjected to a four-stage classificationcascadeofconsiderablesophistication.Atthe firststage,AhoCorasickautomataareusedforidentifying explicitnuditykeywords.Thisproducesanimmediatematch withaconfidencescoreof0.99forexplicitkeywords.Atthe second stage, language-specific automata are used for matchingkeywordsrelatedtoabusewithacomprehensive lexicon containing 1,409 English keywords, 218 Hindi keywords,and250Chinesekeywordsataconfidencelevelof 0.95.
Thethirdstageusesspecializedpatternmatchingtodetect concatenated tokens. This stage helps resolve common errors in speech recognition systems whereby individual words are incorrectly combined. Thereafter, text that has gone through the first three stages is vectorized and classified by a Linear Support Vector Classifier with a thresholdof0.70.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
This section will discuss the mathematical underpinnings andimplementationspecificsoftheAho-Corasickalgorithm in relation to the first defensive layer and elucidate how high-severity explicit words such as “porn” and “nude” bypasscomputationallyexpensivemachinelearningmodels and result in a guaranteed confidence level of 0.99 in a timelymanner.
This section describes how the Aho-Corasick pattern matchingsystemisextendedtoincludealargedatabaseof 1,877 words. It also describes how it handles lower-tier abusive words in English, Hindi, and Chinese by giving a confidence level of 0.95 when it finds bullying or severe profanitywords.
The current subsection deals with a special problem encountered by engineers. Specifically, the speech-to-text enginestendtotreatrapidlyutteredwordsasasingleword (e.g.,thestring"shutup").Thissectionexplainsthespecial pattern-matching logic that has been developed for the purposeofavoidingsuchfalsenegatives.
4.1.4
Forthosetextsthatcannotbematchedbyasimplekeyword search,thisfinalstepservesasacatch-allmechanism.This section describeshowthevectorizationprocessusingTFIDF occurs, followed by a classification process using a Support Vector Machine and Platt Scaling to determine whetherthephrasemeetsa0.70probabilityrequirementfor abuse.
However, there are considerable constraints involved in implementingsuchcontentmoderationsystems,especially in terms of latency, throughput, as well as user privacy concerns. For instance, real-time content moderation requires efficient implementation to achieve sub-second detection latency, which often requires developers to use knowledgedistillationtocompactlargeteachermodelsinto more manageable student models. Moreover, there are considerable ethical as well as legal privacy concerns associatedwithtransmittingusercommunicationstocloud servers.
To mitigate the inherent privacy risks associated with transmitting potentially sensitive media to external
corporate servers, contemporary parental control applications often attempt to conduct content analysis entirelyon-device.

Fig -4.1:Aconceptualdiagramcomparingavulnerable cloud-processingarchitecture(datastreamingtoservers) withasecureon-deviceedge-processingarchitecture (dataremainingentirelyonthelocalmachine).
Whilesignificantadvancementshavebeenmadeinthefield of deep learning and the integration of multimodal approaches,thereareseveralresearchgapsthatthecurrent literaturehasfailedtoaddress.Oneofthemajorissueswith thecurrentacademicmodelsisthattheyare basedonthe premisethattheresearcherhastheluxuryofusingaGPU and other such computational tools. This is not the case whenthefocusisonthe efficientdesign requirementsfor theaveragehomecomputer.
Moreover,thereisalackofproperfocusonreal-timeaudiobasedabusedetection,andthereisalsoalackofsupportfor multilingual or code-mixed linguistic varieties. Also, most commercialparentalfilteringtoolsusecloudcomputingfor processing, which contradicts the need for privacypreserving filtering. This gives rise to the need for an integratedlocalCPU-optimizedmultimodalframework.
Another critical gap in the current literature lies in the effective synchronizationandfusionofmultipledatastreamsunderstrict hardware constraints. While heavy, cloud-based models can processaudio,text,andvisualinputssimultaneouslytounderstand complexinteractions suchassarcasm,rapidlyevolvinginternet slang, or contradictory audio-visual signals lightweight, CPUbound models often struggle with this contextual nuance. Furthermore,accuratelyaligningthesediversemodalitiesinrealtime

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Earlyformsofautomatedabusedetectionsystemsused"bad words"ora"blocklist/lexicon"approach.Whilethesewere simple,precise,yetlacking,thereweremanyproblemswith these approaches. First, there were many false positives, where "bad words" were found in harmless contexts, and manyfalsenegatives,whereattackerswouldusetrickssuch astypostoavoidthese"badwords."Tosolvetheseissues, researchersbegantousemachinelearningtechniques,such asfeatureengineeringandSupportVectorMachines(SVM) withcharactern-grams,makingthesystemmuchharderto evadebyattackers.
Theadventofdeeplearning,andmorespecificallysequencebaseddeeplearningnetworks,sawanenormousaccuracy improvement in NLP. CNNs and LSTMs eliminated the requirementforfeatureengineering bylearningthetext's deeprepresentationsdirectlyfromthedata.Thisnaturally led to the development of more accurate abuse detection transformerssuchasBERT.BERT'suseofbidirectionalpretraining resulted in unprecedented accuracy in abuse detection.However,thenumberofparametersinsuchdeep learningnetworksisextremelyhighandrequiressignificant GPUsupportforreal-timeexecution.
5.2.1
Earlymethodsforfindingexplicitcontentreliedheavilyon classical computer vision techniques such as skin color modelsandhandcraftedfeaturesthatattemptedtoidentify bodyparts.Thetheorybehinditwassimple:explicitcontent would be identifiable by its clear skin tones that could be easilyidentifiedandflagged.However,thesemethodswere simpleandlight-weightbutranintohardwalls.
Theyfailedtoaccountfortheentirespectrumofhumanskin tonesthatwerenotwellrepresentedinthesemodels,threw away too many innocent close-ups of faces due to misclassificationas“danger,”andlackedanyrealnotionof contexttoidentifysafeskintonesvs.explicitcontent.
TheadventoftheAlexNetbreakthroughcompletelychanged the face of the visual moderation world and marked the beginning of the meteoric rise of Convolutional Neural Networks(CNNs)fortheanalysisofimages.Thefirstdeep learning-based classifiers, such as Yahoo’s Open NSFW, providedasingleprobabilityscorethatreflectedtheoverall appropriateness of an image. Today, the best approaches havemovedbeyondthesimplisticsingle-labelclassification andintomoresophisticatedobjectdetectionpipelines.For instance,theNudeNettoolusessuchadvancedarchitectures todetectandlocatetheexactexposedanatomicalfeatures withpreciseboundingboxcoordinatesandconfidencelevels foreachfeaturetype.Thisallowsforextremelynuancedand flexiblefilteringbasedontheexacttypeofexposurerather than broad and sweeping prohibitions against entire categoriesofimages.
Theneedtoregulatethecontentasit’screated,inreal-time, putsverystrictcomputationalconstraintsonthedesignof thesystem.Inorderforthecontentdetectiontobeeffective inprotectingtheuser,thedetectionneedstobeperformed inamatterofmillisecondstoafewsecondsafterthecontent creation so that the problematic content can be stopped before the user even sees it. However, these strict latency requirementsputverystrictconstraintsonthecomplexityof the deep learning models that can be used. In order to achievereal-timecontentdetectionwithoutslowingdown the entire system, researchers have had to resort to very aggressiveoptimizations,likeknowledgedistillation,which compressesalarge,slowteachermodelintoasmall,efficient studentmodel.
Essentially, content moderation requires an in-depth dive intothepersonalconversationsofusers,whichinitselfisa majorethicalandlegaldilemmainthematterofprivacy.The traditionalcloud-basedsystemmakesthissituationworse byconstantlystreaminguseraudio,text,andimagedatato remotecorporateserversforprocessing.
5.3.2.1.
Thiscontinuoustransmissionofsensitivemultimodaldata rangingfromambientroomaudiotoprivatetextexchanges and screen captures creates multiple points of vulnerability.Oncedataleavesthelocaldevice,itbecomes

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
susceptibletointerceptionduringtransitandexposesusers to potential data breaches at the remote server level. Furthermore,theindefinitestorageandopaqueanalysisof this data by third-party entities raise significant legal complianceissues.
6.1.1
The major problem with the Cyber Owl text classification system lies in its architecture based on system audio loopback.Thesystem'spipelineisbasedsolelyonspeech-totext recognition of the captured chunks of the system's audio.Thismeansthatanythingtheusertypesmanuallyand silently,withouttheuseofspeech-to-textfunctionality,will neverbepickedupbytheNLPsystem.Thismeansthatifthe childisusinganend-to-end encryptedmessengerservice, chatting in a text-based game, or reading online forums without screen readers enabled, they will never be monitored by the system. This is because the system's designers wanted it to not become an invasive form of a keylogger and thus violate the "Intelligence Localized" principle.
6.1.2
However,dependingonaudiotranscriptionalsointroduces arangeofvulnerabilities,dependingonthelevelofaccuracy andreliabilityofthespeechrecognizer.Forexample,noise, cheapmics,multiplepeopletalkinginmultiplayerchat,and thick regional accentscanaffect the level oftranscription. Thesystemalsoattemptstomitigatetheeffectsofmerging errors with concatenated token resolution. However, the NLPcanonlyperformaswellasthedataitisgiven.Ifthe ASRenginemisinterpretsawordoraslurredword,thenthe Aho-Corasick automata and the TF-IDF models are essentiallybeingfedbaddata,whichofcourseincreasesthe likelihoodofamisseddetectionintheaudiochannel.
6.2.1
ToensurethattheaverageCPUusageremainsfirmlyunder 25%,andthememoryusagedoesn’texceed500MBwithout theuseofadedicatedGPU,thevisualdetectionpipelinewas heavilyconstricted.Theendresultwastheimpositionofa fixedframerateofprecisely5framespersecond.Whilethe current state of digital video and games maintains a minimumframerateof30-60FPS,theCyberOwlsystemis onlyconcernedwithasmallpercentageofthescreen’svisual data. There is a clear temporal gap between the time an
imageisdisplayedonthescreenandthetimetheNudeNet CNN is capable of detecting it: “A very explicit or harmful image,suchasaquicklyscrolledsocialmediafeedoravideo frame, could appear on the screen for a moment before passingthroughthe200mscapturewindowswithoutbeing detectedbytheNudeNetCNN.”
ThevisualpipelineworksonthedownsampledJPEGimages ofsize640*480individually,withoutreferencingtheimages beforeorafterthecurrentone.Itdoesn’tbenefitfromthe contextoftheimagesintermsoftime.TheNudeNetmodel identifies the exposed anatomical features in the grid of images,butitdoesn’tprocesstheimagesina sequence of framestocomprehendthecontentinalargersense.:
Byintentionallydiscardingrecurrentarchitecturesand3D convolutionstomaintainCPUefficiency,thevisualpipeline evaluateseachframeasanisolatedevent.Consequently,the trajectory,pacing,andevolvingnarrativeofavideostream areentirelylost.Whilethesystemreliablydetectsfeatures inafrozenmoment,itremainscompletelydetachedfromthe sequentialeventsthatdefinethemedia'strueintent.
Thislackoftemporalcontinuitycreatessignificantsemantic ambiguity, resultinginhigh false-positiveratesfor benign material. The context-free visual engine assigns the same threatleveltoaclinicalbiologydocumentaryasitwouldto explicit pornography,asbothgenerateidentical bounding boxsignalsattheframelevel.Withoutevaluatingsequential frames,theedge-optimizedsystemstrugglestodifferentiate educationalcontextfromharmfulcontent.
Thisresearch,whiledemonstratingrobustedge-computing capabilities for privacy-preserving content moderation, is bound by several inherent architectural and operational limitations that prioritize real-time efficiency over exhaustive data capture. Primarily, the natural language processingpipelineisstrictlyaudio-centric,relyingentirely onspeech-to-texttranscriptionscapturedviasystemaudio loopback;thisrenderstheframeworkfunctionallyblindto direct keystrokes and text-based communications within end-to-end encrypted messaging applications, while also introducing vulnerabilities related to automatic speech recognition (ASR) transcription errors caused by backgroundnoiseorheavyaccents.Furthermore,toadhere to strict CPU and memory constraints without requiring dedicatedGPUacceleration,thevisualdetectionpipelineis aggressivelylockedtoa5-frames-per-secondsamplingrate.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Thistemporalbottleneck,combinedwiththeConvolutional Neural Network's evaluation of isolated static frames without recurrent sequential context, creates blind spots where split-second explicit anomalies in high-framerate multimediamayevadedetection.Additionally,thesystem's currentprogrammaticimplementationisheavilycoupledto theWindowsoperatingsystemanddemandsabaselineof mid-range computational power (such as an Intel Core i7 with 16GB of RAM), effectively precluding native deploymentonmacOS,Linux,orhighlyresource-constrained microcomputerswithoutinducingseverethermalthrottling orinterfacelag.
8. References
1) Davidson,T.,Waring,W.,Pomarole,U.,&Mejova,Y. (2017).AutomatedHateSpeechDetectionandthe ProblemofOffensiveLanguage.Proceedingsofthe International AAAIConferenceon WebandSocial Media.
2) Waseem,Z.,&Hovy,D.(2016).HatefulSymbolsor HatefulPeople?PredictiveFeaturesforHateSpeech Detection. Proceedings of the NAACL Student ResearchWorkshop.
3) Badjatiya, P., Gupta, S., Gupta, M., & Varma, V. (2017).DeepLearningforHateSpeechDetectionin Tweets. Proceedings of the 26th International ConferenceonWorldWideWeb.
4) Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv preprintarXiv:1810.04805.
5) Nobata,C., Tetreault,J.,Thomas,A.,Mehdad, Y., & Chang, Y. (2016). Abusive Language Detection in Online User Content. Proceedings of the 25th InternationalConferenceonWorldWideWeb.
6) Krizhevsky,A.,Sutskever,I.,&Hinton,G.E.(2012). ImageNet Classification with Deep Convolutional NeuralNetworks.AdvancesinNeuralInformation ProcessingSystems.
7) Jones, M. J., & Rehg, J. M. (2002). Statistical Color Models with Application to Skin Detection. InternationalJournalofComputerVision.
8) Deselaers,T.,Gass,T.,Dreuw,P.,&Ney,H.(2008). Bag-of-Visual-Words Models for Adult Image ClassificationandFiltering.200819thInternational ConferenceonPatternRecognition.
9) Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You Only Look Once: Unified, Real-Time Object Detection. Proceedings of the IEEE
Conference on Computer Vision and Pattern Recognition.
10) Gomez, R., Gibert, J., Gomez, L., & Karatzas, D. (2020). Exploring Hate Speech Detection in Multimodal Publications. Proceedings of the IEEE/CVF Winter Conference on Applications of ComputerVision.
11) Aho,A.V.,&Corasick,M.J.(1975).EfficientString Matching: An Aid to Bibliographic Search. CommunicationsoftheACM.
12) Salton, G., & Buckley, C. (1988). Term-Weighting Approaches in Automatic Text Retrieval. InformationProcessing&Management.
13) Joachims, T. (1998). Text Categorization with Support Vector Machines: Learning with Many Relevant Features. European Conference on MachineLearning.
14) Platt, J. (1999). Probabilistic Outputs for Support VectorMachinesand ComparisonstoRegularized Likelihood Methods. Advances in Large Margin Classifiers.
15) Kiela,D.,Firooz,H.,Mohan,A.,Goswami,V.,Singh, A., ... & Joulin, A. (2020). The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes.AdvancesinNeuralInformationProcessing Systems.
16) Kumar, S., Spezzano, F., & Subrahmanian, V. S. (2019).Real-timeHateSpeechDetectionSystemfor Live Streaming. IEEE International Conference on DataMining.
17) Han, S., Mao, H., & Dally, W. J. (2015). Deep Compression:CompressingDeepNeuralNetworks with Pruning, Trained Quantization and Huffman Coding. International Conference on Learning Representations.
18) Lane,N.D.,Bhattacharya,S.,Georgiev,P.,Forlivesi, C.,Jiao,L.,...&Kawsar,F.(2016).DeepX:ASoftware AcceleratorforLow-Power
19) O’Keeffe, G. S., & Clarke-Pearson, K. (2011). The Impact of Social Media on Children, Adolescents, andFamilies.Pediatrics.