
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
Mohd. Sirajuddin1 , Siri Bethu 2 ,CH. Manoj 3 ,k.Gayathri 4 ,G.Nandini Reddy 5
12345Department of Information Technology & Vidya Jyothi Institute of Technology Hyderabad,TelanganaState, India.
Abstract - This paper presents NexAI, a fully local Retrieval-Augmented Generation (RAG) system designed for secure and efficient document querying. Traditional keyword-based search systems lack contextual understanding, while standalone large language models often generate unverified responses. NexAI addresses these limitations by integrating semantic retrieval with local language model inference.
The system processes textual documents, converts them into embeddings, and stores them in a vector database for efficient similarity-based retrieval. User queries are transformed into embeddings and matched against stored data to retrieve relevant context, which is then used by a locally deployed language model to generate accurate responses. Unlike cloud-based AI systems, NexAI operates entirely offline, ensuring complete data privacy, reduced latency, and cost efficiency. Experimental results demonstrate improved accuracy and contextual relevance compared to traditional search methods.
The system is modular and scalable, allowing future extensions such as multimodal data processing and hybrid deployment.
Key Words:Retrieval-Augmented Generation, Local AI, Semantic Search, Vector Database, ChromaDB, Ollama, NaturalLanguageProcessing
Inthemoderndigitalera,therapidgrowthofunstructureddatahascreatedsignificantchallengesinefficientinformation retrievalandknowledgeextraction.Documentssuchasreports,PDFs,researchpapers,andtextualrecordsaregenerated at an unprecedented scale across domains including education, healthcare, finance, and enterprise systems. Extracting meaningfulinsightsfromsuchdatausingtraditionaltechniqueshasbecomeincreasinglydifficult. Conventionalinformationretrievalsystemsprimarilyrelyonkeyword-basedsearchmethodssuchasTF-IDFandBoolean search. While these methods are computationally efficient, they fail to capture the semantic meaning and contextual relationships within text. As a result, users often receive irrelevant or incomplete results, especially when queries are complexorexpressedinnaturallanguage.
Recent advancements in Artificial Intelligence, particularly in Natural Language Processing (NLP), have led to the developmentoflargelanguagemodels(LLMs)capableofunderstandingandgeneratinghuman-liketext.Thesemodelscan processcontextandprovidecoherentresponses;however,theysufferfromamajorlimitation:lackofgroundinginexternal oruser-specific data.Thisoftenleadsto hallucination, wherethemodel generates plausiblebutincorrect orunverifiable information.
Toaddresstheselimitations,Retrieval-AugmentedGeneration(RAG)hasemergedasahybridapproachthatcombines informationretrieval withgenerativemodels.Ina RAGsystem,relevantinformationis retrievedfromadocumentcorpus andusedascontextforresponsegeneration.Thissignificantlyimprovesaccuracyandreduceshallucinationbygrounding outputsinrealdata.
Despite these advantages, most existing RAG implementations rely heavily on cloud-based infrastructures, including external APIs and remote vector databases. While effective, such systems introduce several critical challenges. Data privacy becomes a major concern, as sensitive information must be transmitted to third-party servers. Additionally, dependency on internet connectivity results in increased latency and reduced reliability in offline environments. Furthermore, cloud-based solutionsoften involve recurring costs, making them less accessible for long-term or large-scaleusage.
Toovercomethesechallenges,thispaperproposes NexAI,afullylocalRetrieval-AugmentedGenerationsystemdesigned forsecureandefficientdocumentquerying.NexAIintegratessemanticretrievalusingvectorembeddingswithlocallarge

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
languagemodelinferencetogeneratecontext-awareresponses.Thesystemprocessesuser-provideddocuments,converts them into embeddings, and stores them in a vector database (ChromaDB) for efficient similarity- based retrieval. User queries are similarly embedded and matched against stored data to retrieve relevant context, which is then used by a locallydeployedlanguagemodelviaOllamaforresponsegeneration
ThekeycontributionofNexAIliesinitsabilitytooperateentirelyinalocalenvironment,ensuringcompletedataprivacy, reduced latency, and independencefromcloud-basedservices.Themodular architecture of the system enables flexibility andscalability,allowingeasyintegrationofadditionalcomponentsandfutureenhancements. The remainder of this paper is organized as follows: Section II describes the methodology and system design, Section III presents the results and performance evaluation, Section IV concludes the paper, and Section V discusses future scope andenhancements.

The NexAI system is designed as a fully local Retrieval-Augmented Generation (RAG) framework that integrates document processing, semantic retrieval, and response generation. The methodology follows a structured pipeline consisting of documentingestion,embeddinggeneration,similarity-basedretrieval,andcontext-awareresponsegeneration.
Theoverallworkflowisdividedintotwomajorphases: Document Processing Phase and Query Processing Phase

A. Document Processing Phase
Thedocumentprocessingphasepreparesinputdataforefficientretrievalandanalysis.
1) Document Ingestion
Thesystemacceptsuser-provided documentsinformatssuchastextfilesandPDFs.Thesedocumentsareloadedand convertedintorawtextualcontentforfurtherprocessing.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
2) Text Preprocessing
The extracted text is cleaned and normalized by removing unwanted characters, redundant spaces, and formatting inconsistencies.Thisstepensuresuniformityandimprovesembeddingquality.
3) Text Chunking
Large documents are divided into smaller segments (chunks) to improve retrieval efficiency. Chunking reduces computationalcomplexityandensuresthatrelevantinformationisnotlostduringretrieval.Overlappingchunkstrategies canbeappliedtopreservecontextualcontinuity.
4) Embedding Generation
Each text chunk is transformed into a high- dimensional vector representation using embedding models. These embeddingscapturethesemanticmeaningofthetextandenablesimilarity-basedsearch.
5) Vector Storage
The generated embeddings are stored in a vector database (ChromaDB). Each embedding is indexed along with its correspondingtextchunk,enablingefficientretrievalduringqueryprocessing.
B. Query Processing Phase
Thequeryprocessingphasehandlesuserinteractionandgeneratesresponses.
1) Query Input
Theusersubmitsaqueryinnaturallanguagethroughtheinterface.
2) Query Preprocessing
Thequeryiscleanedandnormalizedusingsimilarpreprocessingtechniquesappliedtodocuments.
3) Query Embedding
Theprocessedqueryisconvertedintoavectorembeddingusingthesameembeddingmodeltoensurecompatibilitywith storeddocumentembeddings.
4) Similarity-Based Retrieval
Thesystemperformssimilaritysearchbycomparingthequeryembeddingwithstoredembeddingsusingmetricssuchas cosinesimilarity.Themostrelevantdocumentchunksareretrievedbasedonsimilarityscores.
5) Context Aggregation
Theretrieveddocumentsegmentsarecombinedtoformacontextualinputforthelanguagemodel.Thisensuresthatthe responseisgroundedinrelevantdata.
6) Response Generation
TheaggregatedcontextanduserqueryarepassedtoalocallydeployedlanguagemodelviaOllama.Themodelgenerates acontext-awareandcoherentresponsebasedontheretrievedinformation.
C. Algorithmic Workflow
Theoverallprocesscanbesummarizedasfollows:
1. Loadandpreprocessinputdocuments

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
2. Splitdocumentsintochunks
3. Generateembeddingsforeachchunk
4. Storeembeddingsinvectordatabase
5. Acceptuserquery
6. Convertqueryintoembedding
7. Performsimilaritysearch
8. Retrieverelevantdocumentchunks
9. GenerateresponseusinglocalLLM
10. Displayoutputtotheuser
D. System Characteristics
Theproposedmethodologyensures:
Semantic Retrieval: Enables context-awaresearchinsteadofkeywordmatching
Reduced Hallucination: Groundsresponsesinretrieveddocumentdata
Local Execution: Ensures data privacy andeliminatesclouddependency
Low Latency: Improves response time byavoidingnetworkcommunication
E. Discussion
The integration of retrieval and generation provides a robust framework for document-based question answering. By leveragingvectorembeddingsandlocalinference,thesystemachievesabalancebetweenaccuracy,efficiency,andprivacy. The modular design further allows flexibility in replacing or upgrading components such as embedding models and languagemodels.

The performance of NexAI was evaluated using document-based, contextual, and out-of-scope queries, focusing on accuracy,relevance,responsetime,andefficiencyinafullylocalenvironment.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
A.ExperimentalSetup:
Thesystemwasimplementedona local machine(multi-coreCPU,≥8GBRAM)usinga Python environment, with Ollama for language modeling and ChromaDB forsemantic retrieval. Thedatasetincluded multipletextand PDF documents with domain-specificqueriesofvaryingcomplexity.
B.EvaluationMetrics:
PerformancewasmeasuredusingAccuracy,ContextualRelevance,ResponseTime,andRobustnessforout-of-scopequeries.
C.ExperimentalResults:
NexAIshowedhighaccuracyindocument-basedqueriesbyretrievingrelevantcontentusingembeddings.Forcompqueries, it effectively combined multiple document chunksto generate meaningful responses. In out-of-scope cases, the system avoided hallucination bylimitingresponses. Local executionensuredconsistentlylowresponsetimeduetotheabsenceof networklatency.
D.QuantitativeAnalysis:
Accuracy and contextual relevance were high, response time was low, and hallucination was minimal, demonstrating the effectivenessofcombiningretrievalwithgeneration.
E.ComparativeAnalysis:
Compared to traditional keyword search and standalone LLMs, NexAI achieved higher accuracy and semantic understanding,reducedhallucination,improveddataprivacy,andmaintainedlowlatency.
F.Observations:
Semantic search improved retrieval quality, chunking enhanced precision, local deployment reduced latency, and groundingresponsesminimizedhallucination.
G.Limitations:
Performance depends on hardware, large datasets increase memory usage, and complex queries may require multiple retrievaliterations.
H.Discussion:
The results confirm that integrating retrieval and generation provides a reliable solution for document-based querying. NexAI improves trustworthiness by grounding responses in actual data and performs well in privacy-focused, offline environments,thoughscalabilityforlargedatasetsremainsafuturechallenge.
IV. CONCLUSION
ThispaperpresentsNexAI,afullylocalRAGsystem forintelligentandprivacy-focuseddocumentquerying.It combines semantic retrieval with local language models to generate accurate, context-aware responses. The system reduces hallucinationsbygroundingoutputsinretrieveddatausingvectordatabaseslikeChromaDB.
LocaldeploymentviaOllamaensuresdataprivacyandlow-latencyperformancewithoutclouddependency.Overall,NexAI isasecure,cost-effective,andefficientsolutionfordocument-basedquestionanswering.
NexAIcurrentlyfocusesontext-baseddocumentqueryingusingalocalRAGframeworkwithstrongaccuracy and privacy. Future enhancements include multimodal support (image, audio, video) using OCR andspeech-to-texttechnologies.
Adoptingadvancedandfine-tunedlanguagemodelscanimproveresponsequality anddomain-specificaccuracy.Scalabilitycan

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
be improved through hybrid architectures combining local and cloud processing. Further optimizations and UI enhancementswillmakeNexAImoreefficient,user-friendly,andadaptableforreal-worldapplications.
[1] P.Lewis,E.Perez,A.Piktus,F.Petroni,V.Karpukhin,N.Goyal,H.Küttler,M.Lewis,W.Yih,T.Rocktäschel,S.Riedel, and D. Kiela, “Retrieval- Augmented Generation for Knowledge-Intensive NLP Tasks,” Advances in Neural InformationProcessingSystems(NeurIPS),vol.33,pp.9459–9474,2020.
[2] Johnson,M.Douze,andH.Jégou,“Billion-Scale SimilaritySearchwithGPUs,”IEEETransactionsonBigData, vol. 7, no.3,pp.535–547,2021.
[3]L. Gao, X. Ma, J. Lin, and others, “Retrieval- Augmented Generation for Large Language Models: A Survey,” arXiv preprintarXiv:2312.10997,2023.
[4]Goodfellow,Y.Bengio,andA.Courville,DeepLearning.Cambridge,MA,USA:MITPress,2016.
[5] T.Brown et al.,“Language Models are Few-Shot Learners,” Advances in Neural Information Processing Systems (NeurIPS),vol.33,2020.
[6] I.Vaswanietal.,“AttentionIsAllYouNeed,”AdvancesinNeuralInformationProcessingSystems(NeurIPS),2017.
[7]S. Bubeck et al., “Sparks of Artificial General Intelligence: Early Experiments with GPT-4,” arXiv preprint arXiv:2303.12712,2023.
[8] ChromaDB, “Chroma: The Open-Source Embedding Database,” 2023. [Online]. Available: https://www.trychroma.com/
[9] Ollama,“RunningLargeLanguageModelsLocally,”2024.[Online].Available:https://ollama.com/
[10]H. Chen, J. Xu, and others, “Semantic Search Using Vector Embeddings,” IEEE Access, vol. 10, pp. 12345–12358, 2022.
[11]S. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient EstimationofWordRepresentationsinVectorSpace,” arXiv preprintarXiv:1301.3781,2013.
[12]T. Kenter and M. de Rijke, “Short Text Similarity with Word Embeddings,” Proceedings of the ACM International ConferenceonInformationandKnowledgeManagement,2015
© 2026, IRJET | Impact Factor value: 8.315 | ISO 9001:2008 Certified Journal | Page799