
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
Shaik Riyaz 1 , Abhishek Arun Kumar 2 , S. Harshitha 3 , P. Varshita4
1ShaikRiyaz,SeniorAssistantProfessorDepartmentofComputerScienceandEngineering,GeethanjaliCollege ofEngineeringandTechnologyHyderabad,India
2AbhishekArunKumar,Student,DepartmentofComputerScienceandEngineering,GeethanjaliCollegeof EngineeringandTechnologyHyderabad,India
3S.Harshitha,Student,DepartmentofComputerScienceandEngineering,GeethanjaliCollegeofEngineering andTechnologyHyderabad,India
4P Varshita,Student,DepartmentofComputerScienceandEngineering,GeethanjaliCollegeofEngineeringand TechnologyHyderabad,India ***
Abstract - The introduction of Artificial Intelligence in the legal profession has created significant changes in terms of document processing, analyzing, and management. In this regard, the research will explore the concept of "AI Driven Legal Document Intelligence," concentrating on the use of AI in contract review, legal analysis, and legal document data extraction. Specifically, it is expected that through the use of AI, including the implementation of algorithms based on Natural Language Processing and MachineLearning,itwould be possible to considerably shorten the process of analyzing a large number of documents while enhancing the detection of risks and obligations. Moreover, the transition from conventional keyword searching to the use of semantic analysis and gaining will be analyzed based on the extent to which better understanding of the document’s contents is achieved.
Key Words: Artificial Intelligence, Legal Technology, Document Intelligence, Natural Language Processing, Machine Learning.
An AI-Driven Legal Document Intelligence System project entails creating an automated system that will be able to analyze legal documents including contracts, agreements, legislation, and case laws. In practice, experts have to manually go through complex legal documents to get meaningfulinformationoutofthem,buttheprocesstakesa lotoftimeandmaycontainerrorscommittedbyhumans.In theproposedproject,modernNaturalLanguageProcessing tools will be used for processing legal texts using transformers like LegalBERT, CaseLaw-BERT, and Longformer. Among other functions performed by this systemareclauseextraction,riskprediction,anddocument summarization, making it possible to provide quality services regarding legal document analysis. Additionally, methods like SHAP and LIME will be applied to interpret algorithmsofAIinaclause-by-clausemanner.Apartfrom that,awebapplicationwithadashboardwillalsobepartof the project where risk prediction, important clauses, and documentsummarieswillbepresentedvisually.
Thereareanumberofchallengesassociatedwithexisting analysis systems of legal documents. The major problem facedbymostofthesesystemsisthattheyonlyhaveaccess to limited data sets, and therefore they cannot be used in analyzingdifferentkindsoflegaldocumentslikecontracts, case law, and statutes. Besides, most of the artificial intelligence approaches in use today act as black boxes wheretheyproducepredictionbutfailtoexplainhowthe decisionwasarrivedat.Theselimitationsmake itdifficult for many users to adopt such systems since they are not transparent and accountable. Moreover, most of these systems do not integrate multiple sources of legal data, making the analysis incomplete and lacking in contextual understanding.Otherchallengesincludelackofapractical interfaceandfailuretobedeployableinareal-worldsetting. Asaresult,thereisaneedforanadvancedAIsystemthat analyzeslegaldocuments.
Significant advances have been achieved in the field of ArtificialIntelligence(AI)andNLP,whichmeansthatnow theanalysisoflegaldocumentsisbasednotonlyonkeyword search but also on semantics. Quite a few studies can be foundthatinvestigatemachinelearninganddeeplearning models aimed at clause extraction, finding liability, and ensuringcompliancewithregulationsamongahugebodyof legal texts. Thus, this literature review constitutes a good startingpointforcreatinganintelligentsystemthatisgoing to be used for contractual risk assessment. The research work Automated Legal Risk Assessment (2024) employs CNNandBERTmodelsthataremeanttoclassifyclausesand predict risks of litigation. The system performs extremely well in terms of risk prediction; however, it is designed exclusivelyforworkingwithonedocument,andthereisno featureoftranslationintootherlanguages.
The research article entitled "Transformer-based Clause AnalysisforIndianLaw(2025)"usesLEGAL-BERTasanaid inaccuratelyclassifyingcontractualduties.Itemphasizesthe importance of AI technology in differentiating between regular language and aggressive language within the indemnityclause.Despitethestudy'semphasisonaccuracy

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
inresults,itoperatesinablackboxmanner,andthereareno visualizationtechniqueslikeSHAPandLIMEtoexplainwhy certain clauses are identified as being aggressive. Furthermore,itdoesnotofferanyclear-cutguidelinesand leavestheinterpretationoffindingsuptothediscretionof thereader.Theothersignificantresearcharticlediscussedin the literature review is the "Hybrid LSTM & LLM based model for Contract Summarization (2024)". This article demonstratestheabilityofLLM'sstructureincapturingtext dependencies. However, there isno persistencelayerthat canbeusedforhandlingdataefficientlyinthedatabase.
Moreover, a review article published in Legal Analytics (2025)conductsacomparativeanalysisofdifferentmachine learning algorithms like SVM and Random Forests for contractauditing.WhileSVMisidentifiedasoneofthebest modelsinthisstudy,itoverlookstheinclusionofadvanced elements from Generative AI, risk dashboards, and predictionfeaturesnecessaryforreal-worldapplicationsin law.Clearly,allpreviousmodelsfocusonthedichotomyof Risk versus Non-risk cases and lack a holistic model for multi-document ingestion (PDF/Docx), clear-cut "AI Verdicts," and explainable AI. In this context, LexAI: Legal DocumentIntelligencePlatformaimstosolvetheseissuesby providingmultiplesources ofdata mining,AIVerdict,and dashboardsusingSHAP/LIMEapproach.
Atpresent,legal documents analysisandrisksmonitoring tools are usually standalone systems, which focus on a particular aspect of analysis, such as clause extraction or classificationinsteadofcombiningallthetoolswithinone unified analysis system. The first type of existing systems largelyusesconventionalmachinelearningandinitialdeep learningapproachestoanalyzethedatasetorthecaselaws academically.Despitetheusageofadvancedtransformers, for example, BERT, Legal-BERT, or CaseLaw-BERT, such systems can be defined as “single-purpose” solutions. In otherwords,theyperformwellwhenanalyzinglegalterms ordetectingcoreclausessuchasIndemnity,Confidentiality, or Termination. Nevertheless, the output is usually unstructuredandunprocessed,meaningthattheuserhasto interprettheresultsindependently.
Another type of the existing systems is represented by customized architectures such as Longformer that can be usedtoanalyzeverylengthydocumentsincludingjudicial transcripts and multi-page agreements. Even though such technologiesallowprocessingalargenumberoftokensin just seconds, the functions of providing plain-language summaries, sentiment analysis or even jurisdiction comparisonsaremissing.Asaresult,legalexpertshaveto spendalotoftimereadingandanalyzingenormousamounts of text before realizing what risks they have to face. Moreover,mostoftheexistingsolutionsoperateinisolation; forinstance,thesystemmaydeterminethatacertainclause
belongs to a "high-risk" category, but fails to draw final conclusionsandtomaketheconnectionbetweentheclause anditsinterpretation bycommonpeople.Thus,currently, there are no unified tools for combining heterogeneous perspectivesonvariouslegalaspectsintooneholisticview. Themajorityoftheavailableplatformsdonotsupportsuch important functionalities as real-time risk assessment or identificationofjurisdictionaldifferences.
Although there have been significant improvements in artificial intelligence technologies within the field of law, certain fundamental weaknesses remain within the methodologiesadoptedthusfar.Onemajorlimitationisthe use of a single source database, whereby only one of two sources,suchasthecompany’scontractagreementsorcase laws, is analyzed. The approach prevents the model from generalizingacrossthevariouslegaljurisdictionspresentin anactualbusinessenvironment.Furthermore,thefailureto integrate information from the various sources leads to incomplete comprehension of the context, whereby the modelunderstandstheliteralinterpretationofthecontract butfailstounderstanditsriskimplications.
Inaddition,thevastmajorityofsuchsystemsisbasedonthe “blackbox”concept.Theygiveaccurateresultsbutdonot provideanyexplanationforthesedecisions,whichleadsto poorreliabilityduetotheabsenceofrationalargumentsand justification (explainability gap). Furthermore, existing solutionsaremerelyscientifictoolsthatcannotbeputinto practice and used because of the lack of proper visual interfacesandintuitivedashboards.Nothavingsuchtools, users are unable to understand and interpret data easily. Lastly,modernsystemsfaceproblemsconcerningthelackof real-timeprocessingandactionableguidancebecausethey generateplain data ratherthanrecommendationssuch as “Safe to Proceed” or “Consult a Lawyer.” All these issues emphasize the importance of developing an all-around solutionliketheproposedLexAIsystem.
TheLexAI:LegalDocumentIntelligencePlatformrepresents an all-encompassing and end-to-end solution aimed at streamlininglegalanalysis.Theuseofcutting-edgenatural language processing techniques combined with a userfriendly design enables a clear-cut and actionable interpretationofanylegaldocument.
Thisistheinitialstagewherethesystemprocessesmultiple inputtypessuchasPDFs,DOCXfiles,andTXTdocuments. Thesystemappliesmachinelearningalgorithmstoextract information from the legal document and normalize its structure,regardlessofhowunstructuredandscanneditis.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
B.StructuralClauseSegmentationModule
In this phase, the system leverages semantic parsing techniquestodividelong-formtextsintoseparateclauses. This helps the system identify different parts of the text basedonstructuralcuessuchasheadings(Indemnity,Force Majeure).
C. SimpleLanguageTranslationModule
Thismoduleaimsattranslatingcomplicatedlegaljargonand language into understandable and easy language so that ownership rights and responsibilities can be easily comprehended by non-professional people and small businessowners.
D. RiskAssessment&ScoringModule
Itisbasicallytheanalyticbrainthatusesadomain-specific transformermodellikeLEGAL-BERTtoconductassessment and assign a numerical risk rating from zero (0.0) to one (1.0) based on the comparison between the text and enterpriseliabilitystandards.
E. VerdictEngine
Thedecision-makingbrainofthesystem,whichtakesallthe riskscoresinthedocumenttogiveaclearconclusionabout thedocumentasawholeintermsofitsrisksbeinglow,high, orcontactingalawyer.
F.ExplainableAI(XAI)InsightModule
To ensure trust among professionals, the interpretability features such as SHAP and LIME are incorporated. This allowsforidentificationoftheparticular"triggerwords"in theclausethatleadtotheredflagwarning.
G.InteractiveVisuals&DashboardModule
Thevisualizationmoduleensurestheuserhasaccesstothe high-fidelity interactive experience to explore the informationprovided.Theusercanutilizesuchfeaturesas risk distribution charts or side by side feed comparisons, whichhelptofilteroutandexplorespecificlegalissues.
H.Export&ReportingModuleinMultipleFormats
Thismoduleenablesthecreationofanexecutivesummary reportaswellastheformalanalysisreportineitherPDFor JSON format. The module is intended to be used in the enterprise setting where stakeholders will need to share theirfindingsorrecordversioninghistoryorrisks.
WorkflowSummary
As was mentioned above, the workflow starts with document ingestion and segmentation. Further, risk classification based on neural network takes place. These insightsareinterpretedbytheXAImodule,convertedinto human-readable format, and presented in visuals on the interactivedashboard.Finally,thedecisionisreached,and persistentdatarecorded.
Thisisasystematicapproachindevelopingmodelsthrough iterationandincrementationdesigns.First,legaldocuments includingcontracts,laws,andstatutesaregathered.These documents will be preprocessed by cleaning, tokenizing, normalizing,andsegmentingthempriortobeingfedintothe model.Legaldatasetswillbeemployedinordertotrainand test several transformer models, such as LegalBERT, CaseLaw-BERT,andLongformer.Modelsareevaluatedbased on accuracy, precision, recall, and F1-score, thereby determining which model is most efficient. The model includes an interpretability approach including SHAP and LIME. Finally, deployment will include uploading the documentontothewebapplicationviaframeworkssuchas Flask.
ForLexAImodel,itwasmadeinsuchawaythatapartfrom encouragingmodularity,itwouldalsopermitscalabilityand beabletohandlelegaltextualdataincomplexways.Ineach ofthelayers,thereareuniqueprocessesperformedbyeach layer, which collectively ensure the proper analysis of documents and identification of hidden legal threats. The systemintegratesvariousphasesintoonecoherentprocess justlikeinanywell-orientedjudicialanalysis.
TheDataLayerservesasthefoundationfortherepository withinthesystem,whereitcontrolstheprocessofingesting legal data in various forms. The Data Layer processes two main types of legal data: Contracts (Agreements, NonDisclosure Agreements (NDAs), and internal policies) and Legal Reference Material (Statutes, Case Law, and Regulations).Throughthisprocess,theDataLayerensures that there is enough context provided within the system whenassessingrisksassociatedwithdocuments.
It prepares the raw legal document for in-depth analysis throughvariousrefiningprocesses.Thefirstprocessinvolves Text Cleaning & Normalization which helps eliminate any unnecessary noise. After that comes Tokenization & Chunkingwhichwilldividelargedocumentsintosmalland semanticallyrelevantsegments.Annotation&Labelingfollow next where clauses along with possible risks are marked. Lastly,DatasetIntegrationtakesplacewheretheannotated segmentswillbeassembledintoasuitableformtobefedto theAImodels.
At the heart of the system lies the Modeling Layer, which leveragesstate-of-the-arttransformerssuchasLegal-BERT, Longformer, and CaseLaw-BERT. The layer analyzes the semanticsandcontextoftheinputdatatoenablethesystem todetectaggressionandclassifyclausesprecisely.Moreover,

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
it is capable of integrating knowledge graphs and graph neural networks (GNN) modules to discover associations between various legal entities. Each model is subjected to ModelEvaluation,whichusesprecision,recall,andF1-scores forvalidatingthepredictionsmadebythemodel.
4 ExplainabilityLayer(XAI)
Forachievinggreaterlegalclarity,theExplainabilityLayeris allabouttranslatingtheoutputofacomplicatedmodelinto meaningfulinsights.ItdoessousingtheSHAPAnalyzerto give explanations about the global reasoning behind the model,alongwiththeLIMEInterpretertogiveexplanations ataclauselevel.Thislayermakesitpossibleforlegalexperts to know the exact words and phrases like "unilateral termination"thatgeneratedahighriskscore.
5 PresentationLayer
PresentationLayerprovidesthefinaloutputinaneasy-tounderstand format. This layer contains the Visualization Dashboard, which generates interactive risk scores, summaries of clause impacts, and "Simple Terms." This interface is available through the Web Interface (Flaskbased).Userssuchaslawyersandcomplianceofficershave accesstoaneasy-to-useinterfacethroughthePresentation Layer.Analysisreportscanbeviewedandriskexplanation outputs can be exported by the users for making their decisionsatlast.

Module 1: Ingestion & Structural Normalization Ingests documents in multiple formats (PDF, DOCX, TXT). This module carries out text extraction, metadata acquisition (pages, word counts), and structural normalization to guaranteeconsistencyintheprocessingoflegaldocuments.
Module 2: Structural Decomposition & Semantic Parsing Applies regex pattern matching and natural language processing techniques to decompose raw texts into "Clauses."Thismodulerecognizeskeyheaders(forexample, Indemnity, Termination, IP Rights) and breaks down the documentintostructuralfragments.
Module 3: Neural Classification and Risk Analytics The primary machine learning pipeline with Legal-BERT and RoBERTamodel.Thismoduleconductsclausecategorization and evaluates numerical risk scores from past legal databasesandjurisdiction-specificlegalwording.
Module 4: Explainable AI (XAI) & Interpretability Layer Applies SHapley Additive exPlanations (SHAP) and Local InterpretableModel-agnosticExplanations(LIME)forpost hocinterpretationsofAIpredictions.Thismodulediscovers theunique"triggerwords"contributingtoriskscores.
Module 5: Generative Intelligence & Rationale Synthesis Applies the Google Gemini 2.0 Flash model. This module creates"SimpleTerms"plainlanguagesummariesandthe final"AIVerdict"byevaluatingthedistributionofriskscores acrossthedocument.
Module 6: Data Store & Concurrency Control Uses the SQLAlchemy 2.0 database schema implementation using SQLitewithWALmode.Thismoduletakescareofhandling concurrencyindatatransactionstoensurethatdocument analysisoutputsaresavedsafelywhilethereareconcurrent AIinferencerequestsmade.
Module 7: Interactive Dashboard UI Module The last presentationlayerthatisresponsibleforrenderingthefinal interactive risk chartsandIntelligenceFeed.UsesChart.js and Tailwind CSS frameworks for the final user interface layer.
A. LandingPage
Theinitialinterfaceoftheplatformcomesupwithaunique designthatguidestheusertothehomepage.Theinterface includesdynamicdatathatfocusesonthelatencybelowone secondandaccuracyaswellastheriskanalysis.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072

B. ArchitecturalComponents
This section displays the intelligent components of the platform,whichincludeSemanticMappingforanalyzingthe intention behind clauses, Compliance Guardrails for benchmarkingagainststandardssuchasGDPR/CCPA,and QuantumAnalyticsforpredictivemodelingwithregardto jurisdictionalhistory.

C. SecureDocumentIngestion
Thesystemboastsasafeandsimpleinterfaceforuploading multiplefileformatssuchasPDF,DOCX,andothers.Itisthe initial step within the structural decomposition and risk mappingprocessflow.

D. Real-TimeAnalysisPipeline
Whenadocumentisuploaded,anautomaticprocessscreen appears,showingtheAI’sprogressfromthestagewhereit parses the documents and extracts clauses to the stage whererisksareanalyzedandeventually,legalinsightsare generated.

The last dashboard gives a combined perspective on the healthstatusofthedocument.Thedashboardincludesthe boldverdictbytheAI(suchas“ProceedwithCaution”),the riskdistributioninthedoughnutchart,andanintelligence feed comparing the legal terms with their simpler equivalents.

F. IntelligenceFeed
Inthismodule,anin-depthanalysisofeachindividualpiece of a document can be performed. Clauses will be automaticallyclassifiedaccordingtovariouscategorieslike Payment,Benefits,Termination,etc.,andtheywillreceivea color-codedriskratingbadgeaswell.Thefeedhighlightsthe

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
"Simple Terms," where the legal obligations like EPF and Gratuitywillbeexplainedincontext.






7:ExtractedclausewithSimpleTermsTranslationand Riskscore

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 04 | Apr 2026 www.irjet.net p-ISSN: 2395-0072
For black box explanation, a card called Explainability is provided as part of the system. This employs SHAP techniquetographicallydepicthowspecificconcepts,such asLiabilityCapsandGoverningLaw,affecttheoverallrisk score quantitatively. Also, a LIME summary card gives an account of the predictions with high confidence using particularphrasinginthejurisdiction.

7. CONCLUSIONS
The project achieved its objectives in developing a legal document intelligence system that utilizes artificial intelligencetechnologytoautomaticallyanalyze,summarize, andassesstherisksinlegaldocuments,includingcontracts, agreements, and judgements. To ensure accurate identification of clauses and obligations within the documents,theprojectemployedtransformermodelssuch asLegalBERT,CaseLaw-BERT,andLongformer. Toenable transparencyindecisionmaking,explainableAItechniques like SHAP and LIME were integrated into the process to ensurecomprehensibleinterpretationattheclauselevel.A webdashboardwasdesignedtodisplaytheresults.
8. ACKNOWLEDGEMENT
TheauthorsthanktheDepartmentofComputerScienceand Engineering,fortheirguidanceandsupport.
9. REFERENCES
[1] M.Young,TheTechnicalWritersHandbook.MillValley, CA:UniversityScience,1989.
[2] R. Nicole, "Automated Risk Assessment in Legal Documents,"J.NameStand.Abbrev.,inpress.
[3] J. Smith, "Data Visualization in Legal Analytics," InternationalJournalofAI,vol.5,pp.22-29,2025.
[4] P.BrasandT.Henderson,"ExplainableAIintheLegal Domain:ChallengesandSolutions,"IEEEInternational ConferenceonBigData,pp.450-459,2022.
[5] G. S. Nelson, "Generative AI in Legal Research and Analysis: The Role of Large Language Models," North CarolinaJournalofLaw&Technology,vol.25,pp.112138,2023.
[6] J. J. Nay, "Natural Language Processing and Legal Intelligence,"ArtificialIntelligenceandLaw,vol.28,no. 1,pp.1-14,2019.
[7] I. Chalkidis, M. Fergadiotis, P. Malakasiotis, and I. Androutsopoulos,"LEGAL-BERT:TheMuppetsstraight outofLawSchool,"Proceedingsofthe2020Conference onEmpiricalMethodsinNaturalLanguageProcessing, pp.2898-2904,2020.
[8] M. Jha and K. Rao, "Explainable AI in the Indian Judiciary: Bridging the Gap between Algorithms and Justice,"ProceedingsoftheInternationalConferenceon Intelligent Systems and Knowledge Management (ISKM),pp.112-118,2024.
[9] P. Kumar and S. R. Dash, "AI-Based Legal Document Summarization: An Indian Perspective," Journal of Emerging Technologies and Innovative Research (JETIR),vol.11,no.3,pp.241-255,2024.
[10]S.Verma,"PredictiveRiskModeling,"J.LegalAnalytics, vol.9,no.2,pp.55-60,2023.