Skip to main content

Biomedical Abstract Simplification Using Large Language Models (LLMs) with Control Mechanism

Page 1


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

Biomedical Abstract Simplification Using Large Language Models (LLMs) with Control Mechanism

Thrishul1 , B. Bharath Kumar2 , K. Sai Karthik3 , Ch. Sai Kiran4, D.Kavitha5

1234Department of Information Technology, TKR College of Engineering and Technology, Telangana, India

5Assistant Professor, Department of Information Technology, TKR College of Engineering and Technology, Telangana, India

Abstract - Biomedicalresearcharticlesoftencontainhighly complex terminology and sentence structures, making them difficult to comprehend for non-expert readers, patients, and interdisciplinary researchers. This creates a significant accessibility gap between advanced medical research and its practicalunderstanding.Toaddressthischallenge,thispaper proposes an intelligent biomedical text simplification system based on a transformer-driven sequence-to-sequence architecture. The proposed system leverages a FLAN-T5 encoder–decoder model to convert complex biomedical abstracts into simplified, readable text while preserving core semantic meaning. The system incorporates preprocessing, tokenization, dense vector embeddings, contextual encoding, and beam search–based decoding to generate high-quality simplified outputs. Additionally, multiple simplification levels mild, medium, and strong are supported to cater to diverse user requirements. A Stream lit-based user interface enables real-time interaction and visualization of results. Experimental observations demonstrate that the proposed approacheffectivelyenhancesreadability withoutsignificant loss of informational content, making biomedical literature more accessible and user-friendly.

Key Words: Biomedical Text Simplification, Natural Language Processing, Transformer Models, FLAN-T5, Text Preprocessing, Beam Search Decoding, Stream lit Application

1. INTRODUCTION

1.1 Background and Motivation

Therapidgrowthofbiomedicalresearchhasresultedinan exponential increase in scientific publications, clinical reports, and healthcare documentation. While these resources are invaluable for medical professionals, they often contain complex terminology, dense sentence structures,anddomain-specificexpressionsthataredifficult to understand for non-expert readers, patients, and interdisciplinary researchers. This lack of accessibility createsasignificantbarrierbetweenbiomedicalknowledge anditseffectiveutilization.

Natural Language Processing (NLP) techniques have emerged as a powerful solution to bridge this gap by enabling automated understanding and transformation of textual data. Among these, text simplification has gained

considerable attention as it focuses on reducing linguistic complexitywhilepreservingtheoriginalsemanticmeaning. Inthe biomedical domain, effectivetextsimplification can improve knowledge dissemination, patient education, and cross-domain collaboration, making it a critical research area.

1.2 Problem Statement

DespiteadvancementsinNLP,simplifying biomedical text remainsachallengingtaskduetothepresenceofspecialized vocabulary,longcompoundsentences,andcontext-sensitive meanings.Traditionalrule-basedandstatisticalapproaches oftenfailtopreservecriticalmedicalinformationorproduce oversimplified outputs that distort the original intent. Additionally, many existing systems lack flexibility in controlling the level of simplification, limiting their applicabilitytodiverseusergroups.

Thereisaclearneedforanintelligentandadaptivesystem capable of simplifying biomedical abstracts while maintaining contextual accuracy, semantic integrity, and readability.Suchasystemshouldalsoprovideaninteractive interface to allow users to experiment with different simplificationlevelsinrealtime.

1.3 Objectives of the Proposed System

Theprimaryobjectiveoftheproposedsystemistodevelop an intelligent and reliable biomedical text simplification framework that can automatically transform complex biomedical abstracts into simplified and easily understandable text. The system aims to reduce linguistic complexitywhilepreservingtheoriginalsemanticmeaning andcriticalmedicalinformation,ensuringthatthesimplified output remains informative and contextually accurate. By focusing on biomedical abstracts, the proposed approach targetsahighlyspecializedandinformation-denseformof text that presents unique challenges in natural language processing.

Another key objective of this work is to leverage recent advancements in transformer-based architectures to improve contextual understanding and text generation quality.TheproposedsystemutilizesaFLAN-T5encoder–decoder model, which is specifically designed to handle sequence-to-sequence tasks efficiently. By exploiting its abilitytocapturelong-rangedependenciesandcontextual

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

relationships, the system aims to generate coherent and fluentsimplifiedtextthatalignscloselywiththeintentofthe originalabstract.

Thesystemalsoseekstoprovideflexibilitybyallowingusers to control the degree of simplification based on their individual requirements. Different user groups, such as medicalstudents,patients,andinterdisciplinaryresearchers, mayrequirevaryinglevelsofsimplification.Toaddressthis need,theproposedsystemsupportsmultiplesimplification levels, enabling users to choose outputs ranging from minimally simplified text to more aggressively simplified versions. This adaptability significantly enhances the practicalusabilityofthesystem.

In addition, the proposed work aims to bridge the gap betweenadvancednaturallanguageprocessingmodelsand real-worldapplicationsbyincorporatinganinteractivewebbasedinterface.TheintegrationofaStream lit-baseduser interface allows users to input biomedical abstracts and obtain simplified outputs in real time. This design choice emphasizes usability and accessibility, ensuring that the system can be easily adopted by users without technical expertise in machine learning or natural language processing.

1.4 Organization of the Paper

Thispaperisstructuredtoprovidea clearandsystematic presentationoftheproposedbiomedicaltextsimplification system. Following the introduction, the second section presentsa detailed literaturesurveythat reviews existing research and methodologies related to text simplification and biomedical natural language processing. This section highlights the strengths and limitations of current approachesandidentifiestheresearchgapsthat motivate theproposedwork.

The third section describes the overall architecture of the proposedsystem,detailingeachfunctionalmoduleinvolved inthetextsimplificationprocess.Thisincludesdiscussions ontextpreprocessing,tokenization,embeddinggeneration, transformer-basedencodinganddecoding,andbeamsearch optimization. The system architecture is explained with reference to the workflow diagram to provide a clear understandingofdataflowandcomponentinteractions.

2. Literature Survey

2.1 Existing Text Simplification Techniques

Textsimplificationhasbeenanactiveareaofresearchinthe fieldofnaturallanguageprocessing,aimingtoreducetextual complexity while preserving the original meaning. Early approachestotextsimplificationreliedheavilyonrule-based techniques,wherepredefinedlinguisticruleswereusedto replacecomplexwordsandrestructuresentences.Although thesemethodsprovidedsomelevelofinterpretability,they were highly dependent on handcrafted rules and lacked

scalability, particularly for domain-specific texts such as biomedicalliterature.

With the advancement of machinelearning, statistical and corpus-based methods were introduced to overcome the limitationsofrule-basedsystems.Theseapproachesutilized parallelcorporaconsistingofcomplexandsimplifiedtextto learn transformation patterns. While statistical methods demonstrated improved performance over rule-based techniques, they often struggled with contextual understandingandfailedtohandlelong-rangedependencies effectively,whicharecommoninbiomedicalabstracts.

Recent developments in deep learning have significantly transformed text simplification research. Neural network–based models, particularly sequence-to-sequence architectures,haveshownpromisingresultsbylearningendto-endmappingsbetweencomplexandsimplifiedtext.The introductionoftransformer-basedmodelsfurtherenhanced performance by enabling better contextual representation and parallel processing. These models have demonstrated superior fluency and semantic preservation compared to traditionalapproaches,makingthemsuitablecandidatesfor complexdomainssuchasbiomedicaltextsimplification.

2.1 Biomedical Text Simplification Approaches

Biomedicaltextsimplificationpresentsadditionalchallenges duetothepresenceofspecializedterminology,abbreviations, and domain-specific expressions. Several studies have exploredtheuseofdomain-adaptedwordembeddingsand medical ontologies to address these challenges. Ontologydrivenapproachesattempttoreplacecomplexmedicalterms with simpler equivalents using curated biomedical knowledgebases.However,suchmethodsareoftenlimited bytheavailabilityandcoverageofdomain-specificresources.

Neuralapproacheshavegainedpopularityinbiomedicaltext simplification due to their ability to learn contextual representations directly from data. Pretrained language models fine-tuned on biomedical corpora have shown improvedperformanceinhandlingmedicalterminologyand sentencestructure.Encoder–decodertransformermodels,in particular, have been effective in generating simplified biomedical text while maintaining semantic consistency. Despite these advancements, many existing models lack flexibilityinadjustingsimplificationintensitybasedonuser needs.

2.3 Limitations of Existing Systems

Although significant progress has been made in text simplification, existing systems still suffer from several limitations. Many approaches focus primarily on lexical simplification and fail to adequately address syntactic complexity, resulting in outputs that remain difficult to comprehend.Additionally,somemodelstendtooversimplify text, leading to the loss of critical biomedical information, whichisunacceptableinmedicalcontexts.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

2.4 Research Gap Identification

From the reviewed literature, it is evident that while transformer-basedmodelshaveimprovedthequalityoftext simplification,thereremainsagapindevelopingsystemsthat balance readability, semantic preservation, and user adaptability. Existing approaches often prioritize model performance without considering real-world usability or accessibility.Moreover,limitedattentionhasbeengivento providing multiple levels of simplification within a single unifiedframework.

Theidentifiedresearchgapliesindesigninganend-to-end biomedical text simplification system that combines advanced transformer architectures with flexible simplification control and an interactive user interface. Addressingthisgapcansignificantlyenhancetheaccessibility ofbiomedicalliteratureandsupportawiderrangeofusers, fromhealthcareprofessionalstonon-expertreaders.

3. Proposed System Architecture

3.1

Overall System Overview

The proposed biomedical text simplification system is designedasanend-to-endpipelinethattransformscomplex biomedicalabstractsintosimplifiedandreadabletextusinga transformer-based sequence-to-sequence model. The architecture follows a modular design approach, enabling efficientpreprocessing,contextualunderstanding,controlled simplification, and output generation. Each module in the system is responsible for a specific function, ensuring scalability,maintainability,andclarityindataflow.

Thesystemacceptsbiomedicalabstractsasinputthrougha web-based interface and processes them through multiple stages,includingtextpreprocessing,tokenization,embedding generation,transformer-basedencodinganddecoding,and finaloutputgeneration.TheintegrationofaFLAN-T5model atthecoreofthearchitectureenablesthesystemtocapture contextualsemanticsandgeneratecoherentsimplifiedtext. Additionally,beamsearchdecodingisemployedtoenhance output quality by selecting the most probable simplified sequences.

3.2 User Interface Module

The user interface module serves as the interaction layer betweentheuserandtheproposedsystem.Itisimplemented usingtheStreamlitframework,whichprovidesalightweight and responsive web-based environment for real-time text inputandoutputvisualization.Throughthisinterface,users cansubmitbiomedicalabstractsandselectthedesiredlevel ofsimplificationbasedontheircomprehensionneeds.

This module is designed to be intuitive and accessible, allowinguserswithminimaltechnicalbackgroundtointeract withthesystemeffectively.Byenablinginstantfeedbackand simplifiedtextdisplay,theinterfacebridgesthegapbetween

complexbackendprocessingandpracticalusability,thereby enhancingtheoveralluserexperience.

3.3 Text Preprocessing Module

Thetextpreprocessingmodulepreparestheinputbiomedical abstract for further processing by removing noise and ensuring consistency in text format. This stage involves operations such as lowercasing, removal of special characters, normalization of whitespace, and sentence segmentation. These preprocessing steps are critical for reducingvariabilityintheinputdataandimprovingmodel performance.

By standardizing the input text, the preprocessing module ensures that the transformer model receives clean and structured data. This contributes to better tokenization efficiency and more accurate contextual representations duringsubsequentstagesofprocessing.

-3:Systemarchitectureoftheblockchain-basedpeerto-peerenergytradingplatform

Fig

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

3.5 FLAN-T5 Encoder–Decoder Model

ThecoreoftheproposedsystemistheFLAN-T5encoder–decoderarchitecture,whichperformsthetaskofcontextual understandingandtextgeneration.Theencoderprocesses the input embeddings and generates a contextual representation that captures the meaning of the entire biomedical abstract.This representation isthenpassed to the decoder, which generates the simplified text in a sequentialmanner.

The use of FLAN-T5 enables the system to leverage instruction-tunedlearning,allowingittoadapteffectivelyto the task of text simplification. This architecture is particularly well-suited for sequence-to-sequence transformations,makingitanoptimalchoiceforsimplifying complexbiomedicallanguage.

3.6 Beam Search Optimization

Toimprovethequalityofthegeneratedsimplifiedtext,beam searchdecodingisemployedduringtheoutputgeneration phase. Unlike greedy decoding, beam search maintains multiple candidate sequences at each decoding step and selects the most probable sequence based on cumulative likelihoodscores.

This optimization technique helps reduce grammatical errorsandincoherentsentencestructuresintheoutput.By balancing exploration and exploitation during decoding, beam search enhances the fluency and readability of the simplifiedbiomedicaltext.

3.7 Output Generation Module

Thefinalmoduleofthesystemisresponsibleforgenerating anddisplayingthesimplifiedbiomedicalabstracttotheuser. Based on the selected simplification level, the decoder producesanoutputthatalignswiththedesiredcomplexity reductionwhilepreservingessentialinformation.

The generated text is then rendered on the Stream lit interface,allowinguserstoreviewandreusethesimplified content. Thismodulecompletesthe end-to-endworkflow, ensuring that the systemdeliversa seamlessand efficient textsimplificationexperience.

4. Methodology

4.1

System Workflow and Processing Pipeline

The methodology of the proposed biomedical text simplification system follows a structured and sequential workflow designed to ensure accurate transformation of complex biomedical abstracts into simplified text. The process begins with user input through the web-based interface, where the biomedical abstract is submitted for simplification.Oncereceived,theinputtextisforwardedto

the preprocessing stage to eliminate inconsistencies and preparethetextformodelingestion.

Afterpreprocessing,thestandardizedtextflowsthroughthe tokenization and embedding stages, where it is converted into numerical representations suitable for transformerbased processing. These representations are then passed throughtheencoder–decoderarchitecture,whichperforms contextualanalysisandgeneratessimplifiedtext.Thefinal outputisrefinedusingbeamsearchdecodinganddisplayed totheuser.Thispipelineensuressmoothdataflow,modular processing,andreliableoutputgeneration.

4.2 Text Preprocessing Strategy

Text preprocessing plays a critical role in enhancing the performance and reliability of the proposed system. Biomedical abstracts often contain irregular formatting, specialsymbols,andcomplexsentencestructuresthatcan negatively impact model performance if not handled properly.Toaddressthis,thepreprocessingstageperforms normalization operations such as lowercasing, removal of unnecessarysymbols,andwhitespacecorrection.

Sentence segmentation is also applied to improve the model’s ability to process long and information-dense abstracts.Bybreakingtheinputintomanageablelinguistic units, the preprocessing strategy ensures that the downstream transformer model can focus on semantic understandingratherthanstructuralinconsistencies.This stepsignificantlycontributestothequalityandcoherenceof thesimplifiedoutput.

4.3 Model Configuration and Training Strategy

The proposed system employs a FLAN-T5 transformer model configured in an encoder–decoder setup for sequence-to-sequence text simplification. The encoder is responsibleforcapturingcontextualrepresentationsofthe biomedical abstract, while the decoder generates the simplifiedversionbasedonthelearnedcontext.Themodel benefits from instruction tuning, enabling it to adapt effectivelytothetaskofbiomedicaltextsimplification.

Althoughthesystemprimarilyleveragesapretrainedmodel, task-specific configuration is applied to optimize performance. Parameters such as maximum input length, outputlength,anddecodingstrategyarecarefullyselectedto balance simplification quality and semantic preservation. Thisconfigurationensuresthatthemodelgeneratesfluent and meaningful simplified text without omitting critical biomedicalinformation.

4.4 Simplification Level Control Mechanism

Akeymethodologicalcontributionoftheproposedsystemis the incorporation of multiple levels of text simplification. Insteadofproducingasinglefixedoutput,thesystemallows userstoselectthedesiredsimplificationintensitybasedon

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

theircomprehensionneeds.Thiscontrolmechanismenables the generation of mildly simplified text that closely resemblestheoriginalabstract,aswellasmoreaggressively simplifiedtextfornon-expertusers.

The simplification level selection influences the decoding behaviorofthetransformermodel,particularlyduringbeam searchoptimization.Byadjustingdecodingparameters,the systemeffectivelycontrolssentencecomplexity,vocabulary choice, and output length. This adaptability enhances the usabilityandapplicabilityofthesystemacrossdiverseuser groups.

4.5 Implementation Environment

Theimplementationoftheproposedsystemiscarriedout using a modern and scalable software stack. The backend processing is implemented in Python, leveraging deep learninglibrariesformodelexecutionandtextprocessing. Thetransformermodelisintegratedintothesystemusing establishedNLPframeworksthatsupportefficientinference anddeployment.

The user interface is developed using Stream lit, enabling rapid prototyping and real-time interaction. The overall implementation environment is designed to support modular development, easy maintenance, and potential futureexpansion,ensuringthatthesystemcanbeextended withadditionalfeaturesorintegratedintolargerbiomedical informationplatforms.

5. Results and Discussion

5.1 Output Analysis of the Simplified Text

The performance of the proposed biomedical text simplification system is evaluated through qualitative analysisofthegeneratedsimplifiedoutputs.Thesystemis tested using multiple biomedical abstracts containing complex terminology and long sentence structures. The simplified outputs demonstrate a noticeable reduction in linguistic complexity while retaining the core semantic meaningoftheoriginaltext.Medicaltermsarepresentedin amorereadableform,andsentencestructuresaresimplified toimproveoverallcomprehension.

Thetransformer-basedFLAN-T5modeleffectivelycaptures contextual relationships within biomedical abstracts, enablingcoherentandfluenttextgeneration.Comparedto the original input, the simplified text exhibits improved readability,reducedsentencelength,andclearerexpression of ideas. These observations indicate that the proposed system successfully achieves its primary objective of enhancingaccessibilitywithoutcompromisinginformational value.

Inaddition,thegeneratedoutputsmaintainlogicalflowand contextualconsistencyacrosssentences,whichiscriticalin biomedical text interpretation. The simplification process

avoids abrupt sentence fragmentation and preserves the explanatory structure of the original abstract. This demonstrates the model’s ability to balance simplification with semantic continuity, making the output suitable for educationalandinformationalpurposes.

5.2 Impact of Simplification Levels

Oneofthekeystrengthsoftheproposedsystemisitsability togeneratesimplifiedtextatmultiplelevelsofcomplexity. When mild simplification is selected, the output remains closetotheoriginalabstractwhilereducingminorlinguistic complexities.Thislevelisparticularlysuitableforreaders with basic biomedical knowledge, such as undergraduate studentsorinterdisciplinaryresearchers.

Incontrast,mediumandstrongsimplificationlevelsproduce more concise and reader-friendly outputs by further reducing sentence complexity and substituting difficult terminology with simpler expressions. These levels are effectivefornon-expertusers,includingpatientsandgeneral readers,whomaylackfamiliaritywithbiomedicallanguage. Theavailabilityofmultiplesimplificationlevelssignificantly enhances the flexibility and practical usefulness of the system.

Furthermore, the differentiation between simplification levelsallowsthesystemtoadapttodiversecomprehension needswithoutrequiringseparatemodelsorpipelines.This unifiedapproachensuresconsistencyinoutputqualitywhile offering customization. Such adaptability is essential in biomedical communication, where the same content may needtobeinterpreteddifferentlydependingonthetarget audience.

5.3 Discussion on System Effectiveness

Theresultshighlighttheeffectivenessoftransformer-based architectures in handling complex biomedical text simplification tasks. The use of beam search decoding contributestoimprovedgrammaticalstructureandoutput coherence compared to greedy decoding approaches. Additionally, the integration of a user-friendly interface enablesreal-timeinteraction,makingthesystempractical foreverydayuse.

The modular design of the system further enhances its effectivenessbyallowingindividualcomponentstofunction independentlywhilecontributingtotheoverallworkflow. Eachstage,frompreprocessingtooutputgeneration,playsa role in maintaining text quality and readability. This structureddesignensuresthaterrorsorinconsistenciesat onestagedonotsignificantlydegradethefinaloutput.

However,thesystemalsofacescertainlimitations.Insome cases,highlyspecializedbiomedicaltermsmaystillappear inthesimplifiedoutputduetotheneedtopreservesemantic accuracy.Despitethis,theoverallperformanceofthesystem demonstratesastrongbalancebetweensimplificationand

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 02 | Feb 2026 www.irjet.net p-ISSN: 2395-0072

informationretention.Thesefindingsvalidatethesuitability oftheproposedapproachforimprovingtheaccessibilityof biomedicalliterature.

6. Conclusion and Future Scope

6.1

Conclusion

This paper presented an intelligent biomedical text simplificationsystembasedontransformer-drivenencoder–decoderarchitecture.ByleveragingtheFLAN-T5model,the proposedsystemeffectivelytransformscomplexbiomedical abstractsintosimplifiedandreadabletextwhilemaintaining semantic integrity. The modular design, combined with preprocessing, beam search optimization, and usercontrolledsimplificationlevels,ensures bothperformance andadaptability.

The qualitative results demonstrate that the system enhancesreadabilityandaccessibilityofbiomedicalcontent, makingitsuitableforawiderangeofusers.Theintegration of a Stream lit-based interface further strengthens the practical applicability of the system by enabling real-time interactionandvisualizationofsimplifiedoutputs.

6.2

Future Enhancements

Futureworkcanfocusonextendingthesystemtosupport full-lengthbiomedical articlesratherthanabstractsalone. Incorporating domain-specific medical ontologies and evaluationmetricssuchasreadabilityscorescouldfurther improveoutputquality.Additionally,multilingual support andintegrationwithhealthcareinformationsystemscould broadentheimpactandusabilityoftheproposedsolution.

REFERENCES

[1] Y.Zhao,X.Wang,andZ.Liu,“Neuraltextsimplification with semantic consistency,” Proceedings of the AAAI ConferenceonArtificial Intelligence,vol.33,no.1,pp. 738–745,2019.

[2] C.Scarton,G.Paetzold,andL.Specia,“Simplificationof scientific abstracts using BERT and reinforcement learning,” Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP),pp.3004–3016,2020.

[3] Y. Guo, Z. Jin, and Q. Zhang, “Lay summarization of biomedicaltextswithreadabilitycontrol,”Proceedings ofthe2021ConferenceoftheNorthAmericanChapter of the Association for Computational Linguistics (NAACL),pp.415–425,2021.

[4] B. Ondov, M. Berger, and A. Johnson, “Improving readabilityofbiomedicalliteratureusingGPT,”Journal ofBiomedicalInformatics,Elsevier,vol.132,pp.104123, 2022.

[5] Y.Guo,Z.Jin,andQ.Zhang,“PLABA:Adatasetforlayfriendlybiomedicalabstracts,”Proceedingsofthe2023 Conference on Computational Natural Language Learning(CoNLL),pp.210–220,2023.

Turn static files into dynamic content formats.

Create a flipbook
Biomedical Abstract Simplification Using Large Language Models (LLMs) with Control Mechanism by IRJET Journal - Issuu