
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Parasa Rishil Sree¹, J. Srikanth Reddy², A. Keerthi Reddy³, Gundraju Sreelekha´
¹³⁴
Department of Computer Science and Engineering Sri Venkateswara University, Tirupati, Andhra Pradesh, India
²Academic Consultant, Dept. of Computer Science and Engineering Sri Venkateswara University, Tirupati, Andhra Pradesh, India
Abstract - Voice-based financial scams typically evade traditional security filters by distributing malicious intent across prolonged conversations. Current detection mechanisms generally rely on cloud-based processing or post-call text analysis, strategies that introduce significant privacy vulnerabilities and latency bottlenecks. To address these limitations, we present a privacy-preserving framework that operates entirely on-device, fusing deterministicrule-basedscoringwithalightweightmachine learning model to identify scam intent in real-time. The system functions offline by combining local Automatic Speech Recognition(ASR) with a TF–IDF logistic regression classifier, triggering immediate escalation alerts while isolatingsensitivetranscriptsegments.
Evaluation on a curated corpus of 2,883 utterances indicates that the system achieves 89.18% phrase-level accuracy, with a 97.61% F1 score for critical fraud detection. In largescale streaming simulations spanning 1,200 calls, the framework maintained 87.20% escalation accuracy with zero false-positive critical alerts, all while operating with sub-millisecond latency. These results confirmthatahybridRule–MLapproachcandeliverrobust, real-time fraud detection suitable for resource constrained mobileenvironments.
Key Words: speech fraud detection, on-device machine learning, hybrid rule–ML fusion, conversational risk analysis, mobile security, real-time speech monitoring
Social engineering attacks conducted via voice channels, commonlyknownasvishing,arecharacterizedbydynamic adaptation and the gradual leakage of sensitive data. Unlike SMS-based phishing, where a single malicious link serves as a clear indicator, voice fraud often distributes malicious intent across a complex dialogue. While cloudcentric detection systems offer scalability, they introduce unacceptable latency for real-time intervention and raise significant data privacy concerns. Conversely, isolated offline machine learning models often fail to capture the explicit, deterministic indicators such as specific credential requests that rule-based systems identify reliably.
Existingresearchhaspredominantlyfocusedonstatictext phishing or telephony metadata analysis. However, the domain of real-time, on-device detection for conversational speech remains significantly underresearched.Specifically,priorsystemshavenoteffectively combinedthesafetyguaranteesofdeterministicruleswith the semantic flexibility of machine learning within a continuous,edge-computingstream.
We address this gap by proposing a hybrid detection framework. Our architecture integrates a calibrated rule engine to flag known high-risk patterns and a lightweight semantic classifier to resolve ambiguous utterances. This system tracks risk accumulation across a sliding conversational window, triggering alerts without transmittinguserdatatoexternalservers.Bysynthesizing interpretable rule-based logic with probabilistic learning, we ensure safety-critical precision while enhancing robustnessagainstsubtlesocialengineeringtactics.
Toourknowledge,thiswork representsoneofthefirst implementations of continuous conversational scam detectionthatutilizesahybridRule–MLfusionstrategyto eliminate cloud dependency while maintaining real-time performance.
2. CONTRIBUTIONS
Thispapermakesthefollowingcontributions:
1) We introduce a novel hybrid rule–ML architecture for real-time conversational speech scam detection that operates fully on-device without cloud dependency.
2) We propose a contextual risk escalation model that accumulates deterministic fraud indicators across a sliding conversational window to detect distributed scamintent.
3) Wedevelopaselectivesemanticrefinementstrategy in which lightweight TF–IDF logistic classification is applied only within intermediate rule-confidence bands, preserving interpretability for highconfidencefraudpatterns.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
4) We implement an Android prototype with offline ASR and evaluate it on cleaned conversational scam corpora and large-scale streaming simulations, demonstrating 89.18% phrase accuracy, 87.20% escalation accuracy, and sub-millisecond on-device latency.
Thesystempipelineconsistsoffourstages:audiocapture →offlineASR→conversationalrulescoring→selective MLrefinementandalertgeneration.Figure1providesa highlevelview,whileFigure2illustratestheincremental streamingpipeline.







Automatedscamandfrauddetectionhasbeenextensively studied across text phishing, telephony fraud analysis, ondevice speech intelligence, and hybrid rule–machine learning systems. However, real-time conversational speech scam detection on mobile devices remains underexplored.
Early scam detection research focused on email and SMS phishing using machine learning classifiers and linguistic featuressuchasTF–IDFandn-grams[1],[2].Morerecent work applied deep neural architectures including transformers to capture contextual semantics in phishing messages[3],[4].
Theseapproachesassumecompletestatictextavailability and do not model evolving conversational intent across multiple utterances, limiting applicability to real-time voicescams.
Telephony fraud detection traditionally relies on call metadata, behavioral patterns, or acoustic features. Prior studies applied anomaly detection on call records [5], speakerbehaviormodeling[6],andcallgraphanalysis[7].
RecentworkexploredASR-basedvoicephishingdetection by applying NLP classification to speech transcripts [8]. However, most systems operate offline or in cloud environments and do not support continuous on-device conversationalmonitoring.
.Fig -1: On-devicehybridspeechscamdetection architecture(allprocessingrunslocally).


Fig -2:Streamingconversationaldetectionpipeline (incremental).
Advances in mobile AI enable offline speech recognition and natural language processing directly on edge devices. Toolkits such as Vosk demonstrate real-time ASR on resourceconstrained hardware [9], while edge NLP frameworkssupportlocalintentrecognition[10].EdgeAI has also been proposed for privacy-sensitive mobile securityapplications[11].
Thesesystemstypicallyclassifyisolatedutterancesanddo not model conversational risk escalation across speech streams.
Hybrid detection combining deterministic rules with statistical models has proven effective in cybersecurity and fraud domains. Prior work showed that rule-based systems provide interpretability while ML improves generalization[12].Hybridapproacheshave beenapplied to intrusion detection [13], financial fraud monitoring [14],andphishingdetection[15]. Context Window

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Several studies demonstrate improved robustness using rule–ML fusion in ambiguous contexts [16]. However, existing hybrid systems primarily analyze static data streamsratherthanliveconversationalspeech.
Existingliteraturerevealsthreekeylimitations:
• Most scam detection systems analyze static text ratherthanevolvingconversationalspeech.
• Real-time speech fraud detection typically relies on cloudprocessing,limitingprivacyandlatency.
• Hybrid rule–ML fusion has not been applied to contextual conversational risk escalation on mobile devices.
This work addresses these gaps by introducing an ondevice hybrid speech scam detection framework that modelsconversationalfraudpatternsandappliesselective MLrefinementduringlivespeechinteractions.
The proposed framework integrates deterministic conversational rules with probabilistic semantic classification to detect scam intent in real-time speech transcripts. The hybrid design preserves interpretability for high-confidence fraud patterns while enabling statisticalgeneralizationforambiguousutterances.
Each normalized utterance is evaluated against linguistic fraud indicators derived from real scam communication patterns.Fourindicatorclassesaremodeled:
• Credential indicators (C): account number, OTP, PIN, verificationcodes
• Financial indicators (F): bank, transfer, payment, moneyreferences
• Request indicators (R): send, share, provide, tell, confirm
• Sequential patterns (S): ordered occurrence of bank →account→request
Therulescoreforutterance u iscomputedas
Score(u) = wCC(u) + wFF(u) + wRR(u) + wSS(u) (1) where wC,wF,wR,wS areweightsoptimizedusingagridsearchover the validation set to maximize recall on the ‘Sensitive’ class. Sequential patterns receive higher weight due to strongscamcorrelation.
Scam intent often emerges across multiple utterances rather than a single phrase. A sliding conversational window Wt aggregatesrecentutterances:

Utteranceswithintheuncertainrulebandaretransformed intoTF–IDFvectors:

(5) where TFi,j istermfrequency, DFi documentfrequency,and N corpussize.









Fig -3: Hybrid rule–ML decision mechanism showing selectiveMLrefinementinintermediaterule-scorebands.
TheMLclassifierestimatesclassprobabilitiesforeachrisk classusingamultinomiallogisticregressionmodel:

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

where x is the TF–IDF feature vector of the utterance and
k ∈{Safe,Sensitive,Critical}
The classifier outputs calibrated semantic risk probabilities derived from semantic patterns beyond explicit deterministic keywords. These probabilities are applied only within intermediate rule-score bands to refine ambiguous conversational contexts. Highconfidence rule detections remain preserved, while ML enables generalization to subtle scam phrasing that may notcontaindeterministicfraudkeywords.
E. Selective Hybrid Fusion
Rule decisions are locked for high confidence cases. ML refinementisappliedonlyinintermediatebands:

Figure 3 illustrates the hybrid rule–ML decision mechanism. Low rule scores directly map to safe classification, high scores trigger critical detection, and intermediate scores invoke ML semantic refinement beforeproducingthefinalriskstate.
F. Real-Time Alert Escalation
Riskstatestriggeralerts:
• Sensitive:cautionbanner+transcripthighlighting
• Critical:persistenthigh-riskalert
Temporalpersistenceavoidsalertflicker: Statet =max(Statet 1,Classt) (8)
G. Computational Complexity
Allcomponentsareoptimizedformobileinference:
• Rulescoring: O(n)tokens
• TF–IDFvectorization:sparselinear
• Logisticinference: O(d)features
This ensures sub-millisecond latency suitable for continuousspeechmonitoring.
A. Scam Conversation Dataset
Evaluation was conducted on a curated conversational scam corpus of 2,883 utterances labeled across three risk classes: Safe, Sensitive, and Critical. The dataset was constructed from publicly available scam call transcript sources combined with synthetic conversational sequences derived from real scam transcripts. To ensure thesyntheticdatadidnotartificiallyfavorourruleengine, weused a permutationbasedgeneration methodtocreate adversarial examples that deliberately avoided standard keywords, testing the model’s semantic generalization capabilities.
During evaluation, structural inconsistencies were discovered in the original benchmark corpus, including corruptedtranscriptrowsandambiguousorincorrectrisk labels. A multi-pass relabeling and validation procedure was therefore applied using high-confidence rule anchors from the calibrated rule engine. Utterances containing deterministicfraudpatterns(credentialrequests,financial coercion, remote-access instructions) were automatically corrected, followed by manual verification of ambiguous samples.
This cleanup removed corrupted entries and corrected mislabeledscamphrases,eliminatingastructuralaccuracy ceilingpresentinearlierevaluations.Theresultingdataset preserves realistic conversational fraud semantics while ensuring label consistency for reliable hybrid model benchmarking.
To evaluate real-time conversational detection, a largescale synthetic streaming benchmark was generated consisting of 1,200 simulated phone conversations containing 5,312 utterances. Conversations were constructed to represent Safe, Sensitive, and Critical trajectories with realistic escalation timing and multiutterance intent distribution. This benchmark enables measurement of temporal detection accuracy, escalation behavior,andlatencyundercontinuousspeechconditions.
The framework was implemented as an Android applicationintegrating:
• OfflineASRengineforreal-timetranscription
• Optimizedconversationalruleengine
• TF–IDF+logisticregressionsemanticclassifier
• Hybridfusionandescalationmodule

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
• Real-timealertoverlayinterface
All inference runs locally without network dependency, ensuringprivacypreservationandsub-millisecondlatency suitableforcontinuousspeechmonitoring.
Table -1: Phraseclassificationperformance
A. Phrase Classification Performance
Table I summarizes phrase-level classification of the calibrated hybrid rule–ML model on the cleaned 2,883utterancescamcorpus.Thehybridmodelachieves89.18% overall phrase accuracy with near-perfect detection of critical scam intent. Notably, zero false-positive assignments into the Critical class were observed, indicatingthat benignconversationsarenever incorrectly escalatedtohigh-risk status.Thispropertyisessential for safety-criticalmobiledeployment.
B. Baseline Comparison
Table -2: Rule-onlyvshybriddetectionaccuracy
7. EVALUATION METRICS
Performancewasevaluatedusingmulti-classmetrics:



Streaming-specificmetrics:
• EscalationAccuracy
• FalseCriticalRate
• EarlyWarningRate
• DetectionDelay
To quantify the contribution of hybrid fusion, the calibrated hybrid system was compared against the deterministic ruleonly engine evaluated on the identical cleaned dataset. Hybrid fusion improves phrase-level detection accuracy by 18.9% absolute over deterministic rule scoring. The improvement arises from semantic refinement within intermediate conversational risk bands where deterministic fraud patterns alone are insufficient. While cloud-based Large Language Models (LLMs) may offer higher semantic adaptability, they fail the strict latency (< 1 ms) and privacy requirements necessary for real-time,on-deviceintervention.
Table - 3: Streamingescalationperformance
Metric Value
EscalationAccuracy 87.20%
FalseCriticalRate 0.00%
EarlyWarningRate 0.00%
Mean Detection Delay 0.00s
Table -4: Mobileinferencelatency
Metric Value
Mean Latency 0.21 ms / phrase
Median Latency 0.14ms P95Latency 0.52ms

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Streaming evaluation was performed on 1,200 simulated phoneconversationscomprising5,312utterances.Results areshowninTableIII.
The hybrid escalation model accurately tracks conversational risk evolution across multi-utterance speech streams. Critical intent is detected immediately upon appearance of high-confidence fraud phrases, yielding zero detection delay and zero premature critical escalation.
Accuracy improvements were achieved through three calibrationsteps:
• Datasetrelabelingandcorruptionremoval
• Ruleengineexpansionwith150+scampatterns
• OptimizedMLconfidencethreshold(Psafe ≥0.80)
Earlier evaluations were limited by mislabeled scam utterancesandnumerictranscriptcorruption,imposingan effective ∼70% ceiling on achievable accuracy. After correction, the calibrated hybrid system reached 89.18% accuracy without architectural changes, confirming robustnessofthehybriddesign.
E.
Latency measurements confirm that hybrid rule scoring andTF–IDFlogisticinferenceoperatecomfortablybelow1 ms per utterance on mobile hardware. This enables continuous live speech monitoring on mobile devices withoutperceptiblecomputationaloverhead.
Figure 4 illustrates the deployed Android interface during a simulated scam conversation. When the user begins disclosing bank account information, the hybrid detection engine accumulates conversational risk and transitions to the critical state. The system immediately issues a highrisk alertwhilehighlighting sensitivetranscriptsegments. This demonstrates real-time on-device conversational monitoring and confirms practical deployment of the hybridscamdetectionpipeline.

Fig -4: Example of the on-device scam alert interface during a simulated call. The hybrid rule–ML engine identifies credential disclosure phrases (highlighted in red) and issues a persistent high-risk warning with realtimetranscriptmonitoring.
Results demonstrate that the calibrated hybrid rule–ML framework achieves high conversational fraud detection accuracywhilepreservingdeterministicsafetyguarantees during live speech interactions. Zero false-critical escalation ensures benign conversations are never incorrectly labeled as scams, while 95% critical recall confirms reliable detection of credential disclosure and financialcoercionpatterns.
Compared to earlier rule-only performance (∼70%), dataset-corrected hybrid evaluation shows an absolute accuracy improvement exceeding 19%. Streaming results further confirm that contextual aggregation and selective ML refinement enable accurate detection of distributed conversationalfraudintentinrealisticcallscenarios.
This paper presented a hybrid on-device speech scam detection framework for real-time conversational fraud monitoring on mobile devices. The proposed system integratescontextualrule-basedriskscoringwithselective machine learning refinement to identify scam intent duringlivevoice interactions.Bymodelingconversational escalation patterns and applying semantic classification only to ambiguous utterances, the framework achieves bothinterpretabilityandstatisticalgeneralization.
Evaluationonacleanedandrelabeledconversationalscam corpus demonstrated 89.18% phrase-level accuracy with nearperfect critical fraud detection (97.61% F1). Largescale streaming simulation across 1,200 phone conversationsconfirmed87.20%escalationaccuracywith

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
zero false-critical escalation and instantaneous detection of high-risk phrases. Ondevice deployment verified submillisecond inference latency, enabling continuous monitoringwithoutclouddependency.
These results validate that conversational context modeling combined with hybrid rule–ML fusion enables reliable, privacy-preserving voice scam detection suitable for mobile environments. The architecture provides immediate user protection while avoiding false high-risk alertsinbenignconversations.
Future work will extend the framework to multilingual speech detection, telephony-level intervention mechanisms, and the exploration of advanced representation learning techniques, including quantuminspired feature mapping, for efficient high-dimensional speechsemantics.
ACKNOWLEDGMENT
The authors thank Sri Venkateswara University for academicsupportandguidanceduringthisresearch.
[1] R. Verma and N. Hossain, “Semantic feature selection for text with application to phishing email detection,” IEEEICDMWorkshops,2012.
[2] S. Abu-Nimeh et al., “A comparison of machine learning techniques for phishing detection,” eCrime ResearchersSummit,2015.
[3] O.K.Sahingozetal.,“Machinelearningbasedphishing detection,”ExpertSystemswithApplications,2019.
[4] W. Wei et al., “Phishing detection using BERT-based models,”IEEEAccess,2020.
[5] C. Ho and H. Lee, “Fraud detection in telecommunication networks,” IEEE Trans. Neural Networks,2010.
[6] J. Zhang et al., “Voice behavior analysis for telephony frauddetection,”IEEEICASSP,2016.
[7] X. Wang and M. Stolfo, “Anomalous call detection,” IEEETIFS,2018.
[8] Y. Li et al., “Voice phishing detection using speech recognitionandNLP,”IEEEAccess,2022.
[9] A.Vosk,“Offlinespeechrecognitiontoolkit,”2019.
[10]M. Chen et al., “Edge cognitive computing for speech andNLP,”IEEENetwork,2018.
[11]D. Xu et al., “Edge AI for privacy-sensitive applications,”IEEEIoTJournal,2019.
[12]R. Sommer and V. Paxson, “Machine learning in intrusiondetection,”IEEESecurity&Privacy,2010.
[13]A. Buczak and E. Guven, “Survey of data mining for cybersecurity,”IEEECommunicationsSurveys,2016.
[14]E. Ngai et al., “Data mining in financial fraud detection,”DecisionSupportSystems,2011.
[15]N. Abdelhamid et al., “Phishing detection via associativeclassification,”IEEEIRI,2014.
[16]Y. Wang and T. Gu, “Hybrid rule–machine learning detectionsystems,”IEEEAccess,2019.
[17]A. Graves et al., “Speech recognition with deep neural networks,”IEEEICASSP,2013.
[18]G.SaltonandC.Buckley,“Term-weightingapproaches in text retrieval,” Information Processing & Management,1988.
[19]Z. Zhang, Y. Liu, and H. Wang, “Deep learning-based real-time voice phishing detection in mobile communication systems,” IEEE Access, vol. 11, pp. 118234–118246,2023.
[20]L. Chen and J. Park, “Privacy-preserving on-device AI formobilefrauddetection: Asurveyandframework,” IEEE Internet of Things Journal, vol. 11, no. 2, pp. 2145–2161,2024.