
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
AN INTEGRATED FRAMEWORK FOR PLAGIARISM DETECTION
AND AI-GENERATED CONTENT ANALYSIS
Dr.B.Srinivasa Rao1 , Vanam Prashanth2 , Pasham Bugga Sai Ram3 , Syed Nabraas4
1 Professor, Department of Computer Science and Engineering 2,3,4 B.Tech Students, Department of Computer Science and Engineering
Teegala Krishna Reddy Engineering College , Telangana, India ***
Abstract The rapid expansion of digital academic resources and the widespread adoption of large language modelshaveintroducedcompoundingchallengestoacademic integrity.Traditionalplagiarismdetectionsystems,operating primarily through lexical comparison, continue to struggle against paraphrased content and algorithmically generated text. This paper introduces An Integrated Framework For Plagiarism and AI Content Analysis, a unified detection framework that combines string matching, n-gram analysis, TF-IDF cosine similarity, Word2Vec embeddings, and Sentence-BERT semantic comparison with AI-content detection through perplexity analysis, burstiness measurement,lexicaldiversityscoring,logisticregression,and transformer-basedclassification.Arankedsourceattribution mechanism further supports interpretability of results. Experimental evaluation using the PAN-PC-11 plagiarism benchmark corpus and PAN 2025 AI-content datasets demonstratesthatthehybridframeworkachieves95.2±0.8% detection accuracy for plagiarism and 93.0 ± 0.7% for AIgenerated content, outperforming each individual method. Statistical validation confirms these improvements are significant(p<0.05).Resultsindicatethatintegratinglexical, semantic, and probabilistic detection layers yields substantially more reliable academic integrity verification than any single technique.
Keywords Plagiarism detection; AI-generated text detection; semantic similarity; stylometric analysis; academic integrity; natural language processing; transformermodels;sourceattribution.
I. INTRODUCTION
Thedigitizationofacademicknowledgehashadadual effectonscholarlycommunication.Open-accessrepositories andonlinedatabaseshavedramaticallyloweredbarriersto informationretrieval,enablingresearchersandstudentsto engage with global scholarship in ways previously unimaginable.Atthesametime,thisaccessibilityhascreated fertilegroundforacademicmisconduct.Plagiarism inits evolving forms has become a persistent challenge for institutions worldwide, one that conventional detection systems are increasingly ill-equipped to address alone. Historically,plagiarismdetectionwasbuiltaroundsurfaceleveltextcomparison.AlgorithmsliketheLongestCommon Subsequence and Rabin-Karp fingerprinting worked well against verbatim reproduction. However, as awareness of detectiontoolshasgrown,sohavethestrategiesforevading them. Paraphrasing, synonym substitution, and structural
rewriting preserve intellectual substance while masking lexical overlap the very signal that traditional systems dependupon.Detectionapproachesthatrelyexclusivelyon word-level matching are effectively blind to these more nuancedformsofmisappropriation[1].
The situation has grown considerably more complex with the proliferation of large language models capable of generating academically coherent text on demand. A submission produced by a generative AI system may not borrowasinglesentencefromanyidentifiablesource,yetit representsanequallyseriousbreachofacademicintegrity. Existing plagiarism checkers offer little guidance in such scenariosbecausethegeneratedtextisstatisticallynovel.AI detection systems have emerged to fill this gap, but they typically function in complete isolation from source comparison engines, limiting their diagnostic utility. Integratedframeworkaddressesthisfragmentationdirectly. Rather than treating plagiarism detection and AI-content identification as separate problems, the proposed framework handles both within a unified analytical architecture,augmentedbyasourceattributionmodulethat rankscandidateorigindocumentscontributingtodetected similarities.Thegoalisasystemthatissimultaneouslymore comprehensive and more interpretable than any of its constituentcomponents.
II. RELATED WORK
A. Traditional Plagiarism Detection
Early plagiarism detection research concentrated on computationally efficient exact-match algorithms. Clough (2000) provided a foundational overview of string-based and fingerprinting approaches that underpinned the first generationofcommercialdetectionsystems.Potthastetal. (2010)introducedamorestructuredevaluationframework that exposed the practical limitations of purely lexical methods when applied to obfuscated or paraphrased submissions[1].Fingerprintingtechniques,whichgenerate compacthash-baseddocumentsignaturesfromoverlapping n-gram sequences, improved scalability considerably but remained equally vulnerable to synonym-level modifications. N-gram overlap measures combined with Jaccard similarity offered a modest improvement in robustness by tolerating partial phrase-level matches. Alzahranietal.(2012)conductedacomprehensiveanalysis of linguistic patterns in plagiarism and concluded that no singlelexicaltechniquewassufficientforthoroughdetection in realistic academic environments [4]. These findings

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
established a clear research motivation for multi-method approaches.
B. Semantic Similarity Methods
The introduction of vector space models and TF-IDF weighting represented a meaningful step toward contentlevel similarity analysis. By representing documents as weighted term vectors and measuring cosine distance between them, systems could begin to capture topical overlap even when exact phrases were absent. The subsequentdevelopmentofWord2Vec(Mikolovetal.,2013) marked a more substantial conceptual advance, encoding semantic relationships between words through distributionalco-occurrencepatternsinlargecorpora[8]. Transformer-basedmodelshavesincesetnewstandardsfor contextual language understanding. Devlin et al. (2019) introduced BERT, producing bidirectional contextual representationsthatcapturesentence-levelsemanticswith substantially greater fidelity [7]. Reimers and Gurevych (2019) adapted this architecture to produce efficient sentence-level embeddings through siamese training, yieldingSentence-BERT amodelparticularlywellsuited toscalablesemanticsimilaritycomputation[9].
C. AI-Generated Content Detection
Thedetectionofmachine-generatedtexthasattracted growingresearchattentionsincethereleaseofGPT-2and successormodels.Gehrmannetal.(2019)proposedGLTR,a detectiontoolbasedontop-ktokenprobabilityanalysisthat exploited language models' tendency to select statistically predictabletokens[10].Perplexity-basedmethodsoperate onarelatedprinciple:textgeneratedbyalanguagemodel typically exhibits lower perplexity under that model than human-written text of comparable length and complexity. Stylometric analysis offers a complementary perspective. Research has consistently shown that machine-generated text tends to exhibit lower sentence-length variability (reduced burstiness) and a narrower vocabulary range (lower Type-Token Ratio) compared with typical human academic prose. Abburi et al. (2023) demonstrated that ensembleclassifierscombiningmultiplestylometricsignals could substantially improve detection accuracy across diverse writing domains [12]. Wu et al. (2024) provide a thoroughsurveyofLLM-generatedtextdetectionmethods and identify key limitations in cross-model generalization [13].
D. Research Gap
Despiteprogressinbothdomains,existingsystemstreat plagiarism and AI-content detection as entirely separate problems.Integratedframeworkscapableofsimultaneously identifyingtextualreuse,semanticsimilarity,andAI-origin characteristics within a unified architecture remain rare. ThisgapistheprimarymotivationforthisFramework.
III. PROPOSED FRAMEWORK
A. System Architecture
Thisframeworkisstructuredasamodularmulti-stage analyticalpipeline.Documentssubmittedforanalysispass through a preprocessing layer, three parallel detection engines, and a final aggregation module that produces a unified evaluation report. The modular design allows individual components to be updated independently as detectionmethods evolve, without requiringarchitectural changestotheoverallsystem.

Fig. 1. System architecture overview showing parallel analytical modules.
The primary processing stages are: (1) Document Ingestion and Text Extraction, (2) Text Preprocessing, (3) Feature Engineering, (4) Plagiarism Detection Engine, (5) Semantic Similarity Analysis Module, (6) AI-Content Detection Module, (7) Source Attribution Engine, and (8) ResultAggregationandReportGeneration.
B. Preprocessing
Alldocumentsundergotokenization,casenormalization, stop-wordremoval,punctuationfiltering,andlemmatization before any similarity analysis is performed. These steps ensureconsistentfeaturerepresentationsandreducenoise indownstreamcomparisonsacrossalldetectionmodules.
C. Traditional Plagiarism Detection
The plagiarism detection engine implements three lexicalmethods.StringmatchingusingtheLongestCommon Subsequencealgorithmprovidesabaselineverbatimoverlap score:
Sim(X, Y) = LCS(X, Y) / max(|X|, |Y|)
where X and Y denote the two documents being compared. N-gram similarity extends this using Jaccard overlapacrossoverlappingwordsequences:
J(A, B) = |A ∩ B| / |A ∪B|
TF-IDF vector representations are computed for each documentusing:
TFIDF(t, d) = TF(t, d) × log(N / DF(t))

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
where TF(t, d) is the term frequency, DF(t) is the documentfrequencyoftermt,andNisthetotalcorpussize. Document similarity is then computed using cosine similarity: Cosine(A, B) = (A · B) / (||A|| ||B||). This combination captures topical similarity even when exact phrasematchesareabsent.
D. Semantic Similarity
Word2Vec embeddings represent words as dense vectorslearnedfromdistributionalco-occurrencepatterns. Thesimilaritybetweentwowordvectorsis:
Sim(v₁, v₂) = (v₁ · v₂) / (||v₁|| ||v₂||)
Document-level similarity is obtained by aggregating wordembeddingsacrossthefulltext.Forrichercontextual analysis,Sentence-BERTgeneratesfixed-lengthembeddings E_d and E_s for the query and source documents respectively,andsemanticsimilarityiscomputedas:
Sim_semantic = cosine(E_d, E_s)
This method substantially improves the detection of paraphrased plagiarism where surface wording has been altered but conceptual content remains substantially borrowed.
E. AI-Generated Content Detection
Perplexitymeasureshowpredictablyalanguagemodel assigns probabilities to a token sequence. Low values suggesttextwhosestructureconformscloselytolanguage modellearneddistributions:
Perplexity = 2^(-1/N * sum log2 P(wi))
Burstiness captures variability in sentence length: Burstiness=σ²/μ.Lexicaldiversityismeasuredusingthe Type-Token Ratio: TTR = Unique Words / Total Words. Lowervaluesonbothmetricsarecharacteristicofmachinegeneratedprose.Thesethreefeaturesarecombinedwitha logisticregressionclassifier:
P(AI|X) = 1 / (1 + e^⁻(ʷᵀX + b))
where X represents the stylometric feature vector. A transformer-basedsoftmaxclassifiersupplementsthiswith deep contextual features: Softmax(zᵢ) = e^ᵣᶦ / Σⱼ e^ᵣʲ. The outputsofbothclassifiersarefusedintoafinalAIprobability estimate.
F. Source Attribution
Candidate source documents are ranked by similarity contribution: Rank = argmax(SimilarityScore). When multiplesourcescontributetodetectedsimilarities,thefinal aggregatedscoreis:
FinalScore = α · Sᴄᴏₛᴵᵏᴇ + β · Sₛᴇᴹᵃᴺᴛᴵᴄ whereαandβareweightingcoefficientscontrollingthe contributionoflexicalandsemanticsimilarityrespectively. This ranked attribution provides reviewers with interpretableevidenceregardingthemostlikelyoriginsof suspiciouscontent.

IV. EVALUATION AND RESULTS
A. Dataset Description
The experimental evaluation of this Integrated Framework was conducted using a combination of established benchmark corpora and additional supplementary datasets to ensure both rigor and comparability with prior research. The primary dataset is thePAN-PC-11plagiarismcorpus,awidelyrecognizedgoldstandardbenchmarkdevelopedthroughthePANseriesof shared tasks in plagiarism detection research. The corpus containsamixtureofmanuallyandautomaticallygenerated plagiarism cases, including paraphrased, obfuscated, and structurally modified text, making it representative of the range of plagiarism types encountered in real academic environments[1].
To further validate the robustness of the proposed frameworkunderclassicaldetectionconditions,additional experimentswereconductedusingthePAN-PC-09dataset, an earlier PAN benchmark that continues to serve as a reference point for baseline comparisons in the literature [2]. For AI-generated content detection evaluation, recent benchmarksamplesdrawnfromthePAN2025sharedtask on generative plagiarism were incorporated, providing exposure to machine-generated documents produced by contemporarylargelanguagemodelsincludingLLaMAand Mistralvariants.Asubsetofpubliclyavailableacademictext from the Cornell arXiv repository was also included to supplementtheevaluationwithauthenticresearch-domain writing.
In total, the experimental dataset comprised approximately1,200textualdocumentsorganizedintothree primary categories: 500 authentic human-written documentsincludingacademicessays,researchsummaries, andtechnicalreports;400plagiarizeddocumentsspanning verbatim copies, paraphrased variants, and structurally rewritten documents; and 300 documents generated by largelanguagemodelsusingpromptsdesignedtosimulate academicwritingacrossscientificandtechnicaldomains.All documentsunderwentstandardizedpreprocessingincluding tokenization, case normalization, stop-word removal, and lemmatizationbeforefeatureextraction.

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
B. Dataset Splitting Strategy
The dataset was divided into training, validation, and testingsubsetsusingan80-10-10splittoensureunbiased performanceevaluation.Stratifiedsamplingwasappliedto preserveclassdistributionacrossoriginal,plagiarized,and AI-generated document categories, preventing skewed evaluation due to class imbalance. To further improve robustness, 5-fold cross-validation was performed on the trainingset.Ineachfold,themodelwastrainedon80%of thetrainingdataandvalidatedontheremaining20%.Final performance metrics were computed exclusively on the held-out test set, which was not used during training or modelselectiontoensureunbiasedestimation.
C. Training Procedure
The hybrid detection framework integrates multiple models trained independently and combined during inference. Traditional similarity methods such as string matching,n-gramoverlap,andTF-IDFcosinesimilaritydo not require training and were applied directly to preprocessed documents. For semantic similarity, the Word2Vec model was trained on the corpus using a continuous bag-of-words (CBOW) architecture. SentenceBERTembeddingswereobtainedusingapre-trainedmodel without additional fine-tuning to ensure generalization acrossunseendocumenttypes.
For AI-content detection, a supervised learning approach was adopted. Stylometric features including perplexity,burstiness,andType-TokenRatiowereextracted fromeachdocumentandusedtotrainalogisticregression classifier. Additionally, a transformer-based classifier was fine-tuned on labeled human-written and AI-generated samplesfromthetrainingpartition.Allmodelsweretrained andevaluatedunderidenticalpreprocessingconditionsto ensurefaircomparison.
D. Hyperparameter Configuration
Themodelswereconfiguredusingempiricallyselected hyperparametersvalidatedontheheld-outvalidationsetto balancedetectionperformanceandcomputationalefficiency. The key parameter settings are as follows: Word2Vec: vectorsize=300,windowsize=5,minimumwordcount=2, training epochs = 10; TF-IDF: n-gram range = (1, 2), maximumfeatures=10,000; Logistic Regression: solver= ‘liblinear’,regularization=L2,C=1.0; Sentence-BERT: pretrained model “all-MiniLM-L6-v2” producing 384dimensional sentence vectors; Transformer Classifier: learningrate=2×10⁻⁵,batchsize=16,fine-tuningepochs= 3. All models were tuned using validation data to prevent overfitting.
E. Evaluation Protocol
Performanceevaluationwasconductedusingstandard classification metrics including accuracy, precision, recall, and F1-score. All reported results represent the average performance across multiple runs with different random seeds to reduce variance due to random initialization. To
ensurestatisticalreliability,eachexperimentwasrepeated three times using seeds {42, 123, 456}, and the mean performance with standard deviation is reported. This approach reduces the likelihood of overestimating model performanceduetofavorabledatasplitsandbetterreflects thestabilityofeachmethodinpracticaldeployment.
F. Classification Metrics
Letting TP, TN, FP, and FN denote true positives, true negatives,false positives,andfalsenegativesrespectively, theevaluationmetricsaredefinedas:
Accuracy = (TP + TN) / (TP + TN + FP + FN)
Precision = TP / (TP + FP)
Recall = TP / (TP + FN)
F1 = 2 × (Precision × Recall) / (Precision + Recall)
G. ROC Analysis
ReceiverOperatingCharacteristicanalysisevaluatedAIdetection classifier performance across threshold values. The True Positive Rate (TPR = TP / (TP + FN)) and False PositiveRate(FPR=FP/(FP+TN))werecomputedforeach model. The Area Under the Curve (AUC) provides a threshold-independent summary of classification quality. HigherAUCvaluesindicatestrongerdiscriminationbetween human-written and AI-generated content, and the hybrid framework achieves significantly higher AUC than any individualdetectionmethodtested.
H. Reproducibility and Experimental Configuration
Toensurereproducibilityoftheproposedframework, all experiments were conducted using fixed model configurations and controlled random initialization. The implementation was carried out in Python 3.10 using standard machine learning and NLP libraries including Scikit-learn, Gensim, Sentence-Transformers, NLTK, and spaCy. Hardware comprised an Intel Core i7 12-core workstationwith32GBRAMandanNVIDIA RTXGPUfor transformerinferenceacceleration.
Toensureconsistentexperimentalresults,allstochastic processeswerecontrolledusingafixedrandomseed(seed= 42) for the primary run, with additional runs using seeds 123and456forstatisticalvalidation.Thisincludesdataset shuffling, model weight initialization, and training sample ordering.Evaluationmetricswerecomputedusingstandard functionsfromsklearn.metrics.Thecompleteexperimental pipeline,includingpreprocessing,featureextraction,model training, and evaluation, was executed under identical conditions for all methods to ensure fair comparison. The implementationcodeandexperimentalconfigurationswill be made publicly available to support reproducibility and furtherresearch.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
I. Performance Results
TABLE I Traditional
TableIdemonstratesaclearperformancegradient across detection methods. String matching achieves 72.3 ± 1.2% accuracy, adequate for verbatim plagiarismbutinsufficientagainstrephrasedcontent. The TF-IDF cosine approach reaches 83.4 ± 0.9% by capturing topical overlap independent of exact phrasing. Word2Vec achieves 88.2 ± 0.7% and Sentence-BERTreaches92.1±0.6%,withcontextual embeddings handling paraphrased content substantially better. The integrated framework achieves 95.2 ± 0.8% accuracy, demonstrating that combiningallthreeanalyticallayersyieldsresultsnone could achieve individually. The standard deviation
valuesconfirmconsistentperformanceacrossdifferent randomseeds.
Table II shows that individual stylometric features achievemoderateAIdetectionaccuracy:perplexityreaches 78.0±1.1%,burstiness75.2±1.3%,andTTR72.4±1.5%. No single stylometric feature is sufficient to reliably distinguish human-written from AI-generated text, particularlygivenstylisticimprovementsinrecentlanguage models. Logistic regression combining multiple features improves accuracy to 86.1 ± 0.9%, while the transformer classifierreaches90.3±0.8%.ThehybridAImodelachieves 93.0±0.7%accuracyandAUCof0.95.

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net
J. Ablation Study
Toevaluatethecontributionofeachcomponentwithin the proposed hybrid framework, an ablation study was conductedbysystematicallyremovingindividualmodules andmeasuringtheresultingperformancedegradation.The hybridframeworkconsistsofthreemajorcomponents:(1) lexical similaritymethods,(2)semanticsimilaritymodels, and(3)theAI-contentdetectionmodule.Eachcomponent was removed independently while keeping the rest of the systemunchanged.
The ablation results in Table IV demonstrate that the semantic similarity module contributes significantly to overallperformance,particularlyindetectingparaphrased content.Removingthismodulecausesa6.8percentagepoint dropinaccuracy.TheAI-detectionmoduleplaysanequally criticalrole:itsremovalcausesthelargestdegradation(7.6 percentagepoints),confirmingitsimportanceforidentifying machine-generated submissions. The performance degradation observed across all ablated configurations confirms that each component is essential for achieving optimal accuracy and that their integration is synergistic ratherthanmerelyadditive. TABLE III
K. Feature Importance Analysis
To better understand the contribution of individual stylometric features in AI-content detection, feature importance analysis was conducted on the logistic regression classifier trained for AI-generation probability estimation.Thelogisticregressionmodelcoefficientswere analyzedtoestimatetherelativeimportanceofeachfeature. Perplexity emerged as the most significant indicator, reflecting the distinctive probability distribution patterns characteristic of machine-generated text. Burstiness and Type-Token Ratio also contributed meaningfully by capturing structural regularity and lexical narrowness respectively.Thesefindingsconfirmthatcombiningmultiple stylometricsignalsprovidesamorerobustrepresentationof AI-generated content compared to relying on any single feature, which is consistent with the ablation results reportedabove.
L. Statistical Validation
Toensurethattheobservedperformanceimprovements arestatisticallysignificantandnotattributabletofavorable datasplitsorrandomvariation,multipleexperimentalruns wereconductedusingdifferentrandomseeds(42,123,456). Thereportedresultsrepresentthemeanperformancewith standarddeviationacrosstheseruns.Forthehybridmodel, accuracy was observed as 95.2 ± 0.8% across three independent runs, indicating stable and consistent performance.
Additionally, statistical significance testing was performedusingapairedt-testtocomparethehybrid framework against all baseline methods. The null hypothesis that the hybrid model and each baseline achieve equivalent performance was tested at a significancelevelofα=0.05.Theresultsconfirmthat the performance improvements of the hybrid frameworkarestatisticallysignificant(p<0.05)inall pairwise comparisons, demonstrating that the observed gains are not attributable to random variationintheexperimentalsetup.
M. Visualizations


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Figure 3 shows a right-skewed distribution of plagiarismsimilarityscores,withmostdocumentsinthe0–30% range, indicating largely original content. A smaller portionfallswithin40–60%,representingparaphrasedtext where meaning is retained but wording is changed. Documentsabove 60%are few but indicatestrongtextual reuse.The30%and60%thresholdsclassifysimilarityinto risk levels, highlighting the need for a hybrid approach combining lexical and semantic methods for effective detection.
Ultimately,thistieredclassificationallowseducators and investigators to prioritize high-risk cases while minimizing the manual review required for low-scoring, originalsubmissions.

4. AI probability scores for human-written and AIgenerated documents.
Figure4showstheAIprobabilitydistribution.Humanwritten documents cluster strongly below 0.3, while AIgenerated samples concentrate above 0.7. The relatively sparse intermediate region suggests the hybrid model providesconfidentclassificationsinmostcases.

Fig. 5. ROC curves comparing detection models with AUC values.
Figure 5 presents ROC curves for all AI detection classifiers. The hybrid framework model achieves AUC = 0.95, the highest among all evaluated methods. Its curve
mostcloselyapproachesthe ideal upper-leftcornerofthe ROCspace,reflectinghighsensitivitywithcomparativelylow false-positiverates.

Fig. 6. Confusion matrix heatmap for hybrid AI content detection.
Figure6presentstheconfusionmatrixforthehybridAI detection model. True positive and true negative counts substantially exceed false classifications. Notably, false negatives (AI-generated content misclassified as humanwritten)arefewerthanfalsepositives afavorablepattern giventhatmissedAIdetectionconstitutesthemoreserious integrityfailure.

Fig. 7. Processing time comparison across detection methods by document word count.
Figure 7 analyzes processing time as a function of documentsizeforeachdetectionmethod.Stringmatching processes documents fastest but at the cost of detection depth. Semantic embedding methods incur greater computational overhead due to transformer inference requirements. The hybrid framework's processing times remain within acceptable bounds for typical academic documentlengths.
Figure8providesadirectaccuracycomparisonacross all evaluated methods in bar chart form. The progressive improvement from traditional lexical approaches through semantic models to the integrated hybrid framework is

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
clearlyvisible,reinforcingthepaper'scentralargumentthat nosingledetectiontechniqueissufficient.

V. CONCLUSION
Thispaperhaspresentedanintegratedframeworkthat addressesthedualchallengeofplagiarismdetectionandAIgenerated content identification within a single analytical architecture. By combining string matching, n-gram similarity, TF-IDF cosine analysis, Word2Vec embeddings, and Sentence-BERT representations with stylometric AI detectionandmachinelearningclassifiers,theframework overcomes the fundamental limitations of any single detectionapproach.Arankedsourceattributionmechanism further enhances interpretability, and comprehensive reproducibilitydocumentationensuresthatresultscanbe independentlyverified.
Experimentalevaluationona1,200-documentdataset drawnfromthePAN-PC-11benchmarkcorpus,PAN-PC-09, PAN 2025 AI-content samples, and arXiv academic text confirmsthatthehybridframework achieves95.2±0.8% detection accuracy for plagiarism and 93.0 ± 0.7% for AIgenerated content identification, with an AUC of 0.95. Ablationanalysisconfirmstheessentialcontributionofeach component: removal of any individual module causes measurable and significant performance degradation. Statistical validation using paired t-tests confirms that improvementsoverallbaselinemethodsaresignificantatp <0.05.
Several limitations merit acknowledgment. The performanceofAIdetectioncomponentsissensitivetothe specificgenerativemodelsusedduringtraining;aslanguage models continue to improve, periodic retraining will be necessary to maintain accuracy. Transformer-based components impose non-trivial computational costs that mayaffectscalabilityforinstitutionsprocessingverylarge documentvolumes.Futureresearchshouldexploreefficient approximation strategies, cross-lingual extension of the framework,andadaptivedetectionmechanismscapableof evolvingalongsideemerginggenerativesystems.Integration with knowledge graph-based source attribution could further enhance identificationofconceptuallyderived but lexicallydissimilarcontent.
The development of reliable, interpretable, and computationallytractableacademicintegritysystemsisnot merely a technical challenge but a practical necessity for maintainingthecredibilityofscholarlycommunicationinan eraofubiquitousgenerativeAI.ThisFrameworkrepresents a meaningful contribution to this effort, and its modular, reproducibledesignprovidesastrongfoundationonwhich futureimprovementscanbesystematicallybuilt.
REFERENCES
[1] M. Potthast, B. Stein, A. Barron-Cedeno, and P. Rosso, "An evaluation framework for plagiarism detection," in Proc. COLING,2010.
[2]M.Potthastetal.,"Cross-languageplagiarismdetection,"Lang. Resour.Eval.,vol.45,no.1,pp.45–62,2011.
[3]D.Clough,"Plagiarisminnaturalandprogramminglanguages: Anoverviewofcurrenttoolsandtechnologies,"Res.Memo., Dept.Comput.Sci.,Univ.Sheffield,2000.
[4] A. Alzahrani, N. Salim, and A. Abraham, "Understanding plagiarismlinguisticpatterns,textualfeatures,anddetection methods," IEEE Trans. Syst. Man Cybern., vol. 42, no. 2, pp. 133–149,2012.
[5] M. A. El-Rashidy et al., "Reliable plagiarism detection system basedondeeplearning,"NeuralComput.Appl.,2022.
[6]M.Alvarez-Carmonaetal.,"Semantically-informeddistanceand similaritymeasuresforparaphraseplagiarismidentification," ExpertSyst.Appl.,2018.
[7] J. Devlin, M. Chang, K. Lee, and K. Toutanova, "BERT: Pretraining of deep bidirectional transformers for language understanding,"inProc.NAACL,2019.
[8]T.Mikolovetal.,"Efficientestimationofwordrepresentations invectorspace,"inProc.ICLR,2013.
[9] N. Reimers and I. Gurevych, "Sentence-BERT: Sentence embeddingsusingsiameseBERTnetworks,"inProc.EMNLP, 2019.
[10] S. Gehrmann, H. Strobelt, and A. Rush, "GLTR: Statistical detection and visualization of generated text," in Proc. ACL, 2019.
[11] K. Krishna et al., "Paraphrasing evades detectors of AIgeneratedtext,"inACLFindings,2023.
[12]H.Abburietal.,"Asimpleyetefficientensembleapproachfor AI-generatedtextdetection,"arXiv:2311.03084,2023.
[13] M. Wu et al., "A survey on LLM-generated text detection methods,"Comput.Linguist.,MITPress,2024.
[14] T. Kehkashan et al., "AI-generated text detection: A comprehensivereview,"Inf.Process.Manag.,2025.
[15] A. Amirzhanov et al., "A systematic survey of plagiarism detectionalgorithms,"Front.Comput.Sci.,2025.
[16] R. So, "Detection of AI-generated academic papers," Project Rachel,2024.
[17]P.Gosar,"Stylometricfingerprintingwithcontextualanomaly detection,"Preprints,2025.
[18]A.Kujur,"AcomparativeanalysisofAI-generatedandhumanwrittentext,"SSRN,2024.
[19] V. S. Sadasivan et al., "Can AI-generated text be reliably detected?"inProc.NeurIPSWorkshopMLSafety,2023.
[20] A. Najjar et al., "Leveraging explainable AI for LLM text attribution,"arXiv:2501.03212,2025.